Search papers, labs, and topics across Lattice.
1
0
2
11
Rephrasing benchmark problems can flip model answers, revealing that stronger LLMs are paradoxically more fragile to wording changes than weaker ones.