Search papers, labs, and topics across Lattice.
5
0
9
11
MOPD achieves superior capability integration in LLMs by distilling knowledge from multiple RL teachers without losing performance, setting a new standard for post-training methods.
See how ideas like "democracy" or "freedom" have subtly shifted their meaning across different news sources and time periods, all within a single, comparable framework.
Current autonomous agent benchmarks miss nearly half of safety violations and over 10% of robustness failures because they only check final outputs, a problem Claw-Eval directly addresses.
1.58-bit LLMs are surprisingly more resilient to sparsity than their full-precision counterparts, opening new avenues for extreme compression.
LLMs trained with a novel "second-order rollout" that generates critiques in addition to responses learn more effectively from the same data, unlocking better reasoning.