Search papers, labs, and topics across Lattice.
3
0
4
1
Reporting only accuracy metrics can mask up to 21% of behavioral inconsistencies in LLM responses, challenging the reliability of current AI safety evaluations.
Lower-income users are targeted with ads more frequently than their higher-income counterparts, revealing a potential bias in LLM advertising strategies.
Despite high user satisfaction with LLM interactions, a staggering 23.1% of goal tasks failed, challenging assumptions about conversational success.