Search papers, labs, and topics across Lattice.
3
0
8
11
OneDayAgent achieves a groundbreaking 0.821 score on long-horizon tasks, proving that a single harness can effectively manage execution across diverse LLM backends without tuning.
Forget manual data curation – now LLMs can autonomously engineer training data that boosts student model performance by over 57%.
LLMs can slash token usage by 70% and boost reasoning accuracy by 14.8% in long-horizon tasks simply by learning when to remember and forget intermediate thoughts.