Search papers, labs, and topics across Lattice.
3
0
5
0
Machine translation benchmarks have functionally saturated, but pairing human-authored failure cases with deterministic verification rules reveals critical multimodal blind spots that automated metrics consistently miss.
Social reasoning in language models is rooted in distinct training data, with targeted unlearning revealing its vulnerability to data removal.
Systematic gaps in AI evaluation reporting are exposed, revealing inconsistencies that hinder reliable comparisons across thousands of models and benchmarks.