Search papers, labs, and topics across Lattice.
3
0
5
CalibForge reveals that adversarial calibration can dramatically enhance the effectiveness of training data for terminal agents, leading to unprecedented performance improvements on standard benchmarks.
Fine-tuning on DeNovoSWE catapults LLM performance in generating entire software repositories, achieving nearly an 8x improvement on a challenging benchmark.
Autonomous ML research agents achieve significantly better long-horizon performance by maintaining durable state through a shared workspace, suggesting that orchestration and memory are more critical than raw reasoning power.