Search papers, labs, and topics across Lattice.
5
0
7
13
Achieving trillion-parameter performance with just 35 billion parameters by scaling agent horizons reveals a new frontier in model efficiency.
Even the best LLMs struggle with Olympiad-level combinatorics, achieving only 65.4% on a benchmark designed to expose their reasoning limitations.
Current multimodal agents fail to consistently pass CAPTCHA tests, revealing fundamental limitations in their ability to replace humans in automated workflows.
Turn messy human expertise into neatly packaged, agent-usable skills with this automated system that distills heterogeneous traces into portable and correctable AI skills.
Forget retraining: this Red-Blue game hardens AI systems against jailbreaks and CVEs by teaching defensive principles without parameter updates.