Search papers, labs, and topics across Lattice.
8
0
15
25
T1, a Mixture-of-Experts model of 122B total trained with reinforcement learning, operating a real shell in a cloud sandbox for up to 300+ tool-call turns per task, rewarded by executing each task's own verifier by executing each task's own verifier.
Recursive task synthesis not only slashes generation costs to $0.05 per task but also produces increasingly complex challenges that boost model performance by up to 10 points on key benchmarks.
Staleness-Adaptive Trust Regions reshape update geometry in asynchronous reinforcement learning, achieving record performance while controlling for high-staleness updates.
Behavior localization is revolutionized, enabling developers to seamlessly connect high-level modification requests to specific code locations in complex AI harnesses.
Self-distillation in a verifiable environment enables web agents to achieve competitive performance without reliance on external teacher models.
HiLS-Attention achieves over 64x context length extrapolation with 90% retrieval accuracy, outperforming traditional full attention mechanisms.
FlashMemory-DeepSeek-V4 slashes GPU memory usage by over 90% for ultra-long contexts while enhancing model accuracy.
Learnable critics that evaluate the model's own GUI grounding proposals, rather than relying on static geometric heuristics, unlock substantial gains in accuracy.