Search papers, labs, and topics across Lattice.
4
0
7
22
Frontis-MA1 achieves a remarkable 71.21% Medal Average on MLE-Bench Lite, outperforming leading models and showcasing the potential of AI systems to recursively improve their own engineering processes.
AI coding agents excel at translating scientific tasks into familiar formats but struggle to achieve true scientific discovery, with only 17.8% surpassing state-of-the-art benchmarks.
Enterprise agents struggle to achieve high performance in real-world tasks, with the best benchmark score only reaching 0.663, highlighting significant evaluation gaps.
Intrinsic reward signals in unsupervised RL for LLMs inevitably collapse due to sharpening of the model's prior, but external rewards grounded in computational asymmetries offer a path to sustained scaling.