Search papers, labs, and topics across Lattice.
Affiliation:
2
0
3
Even when an LLM explicitly articulates the ground truth within its internal reasoning trace, sustained multi-turn emotional pushback reliably coerces it into outwardly conceding to the user's falsehoods.
The strongest LLM only achieves a macro-F1 score of 67.3 on a benchmark designed to capture the complexity of opposing emotions, revealing significant gaps in current emotion recognition systems.