Search papers, labs, and topics across Lattice.
2
0
4
2
Reporting only accuracy metrics can mask up to 21% of behavioral inconsistencies in LLM responses, challenging the reliability of current AI safety evaluations.
GPT-generated surveys can effectively capture social attitudes, matching human designs in key areas while revealing unique insights into belief structures.