Search papers, labs, and topics across Lattice.
5
0
9
5
T1, a Mixture-of-Experts model of 122B total trained with reinforcement learning, operating a real shell in a cloud sandbox for up to 300+ tool-call turns per task, rewarded by executing each task's own verifier by executing each task's own verifier.
Recursive task synthesis not only slashes generation costs to $0.05 per task but also produces increasingly complex challenges that boost model performance by up to 10 points on key benchmarks.
Staleness-Adaptive Trust Regions reshape update geometry in asynchronous reinforcement learning, achieving record performance while controlling for high-staleness updates.
Behavior localization is revolutionized, enabling developers to seamlessly connect high-level modification requests to specific code locations in complex AI harnesses.
Agents struggle with long-horizon tasks, achieving only a 15.2% success rate even with advanced models, highlighting a critical gap in current AI capabilities.