Search papers, labs, and topics across Lattice.
4
20
9
9
A single adversarial argument can reduce LLM accuracy to near zero, exposing a critical vulnerability in their belief systems.
Agents struggle to maintain planning accuracy in complex tool ecosystems, with GPT-5.4's performance plummeting from 51.90% to 11.36% under severe blocking conditions.
LMMs can't MacGyver their way out of a paper bag: they struggle to creatively repurpose objects in visually complex environments, revealing a critical gap in grounded reasoning beyond pattern recognition.
Forget specialized models: CoALM proves a single LLM can now master both multi-turn conversations *and* complex tool use, even outperforming GPT-4o.