Search papers, labs, and topics across Lattice.
Affiliation:
2
0
2
0
Rephrasing benchmark problems can flip model answers, revealing that stronger LLMs are paradoxically more fragile to wording changes than weaker ones.
LLMs reveal distinct compositional skill gaps that traditional benchmarks fail to capture, highlighting the need for tailored interventions.