Search papers, labs, and topics across Lattice.
3
0
8
9
OneDayAgent achieves a groundbreaking 0.821 score on long-horizon tasks, proving that a single harness can effectively manage execution across diverse LLM backends without tuning.
LLMs are surprisingly bad at automating the creation of executable visual workflows from natural language, highlighting a significant gap in their ability to translate intent into reliable, deployable code.
LLMs can slash token usage by 70% and boost reasoning accuracy by 14.8% in long-horizon tasks simply by learning when to remember and forget intermediate thoughts.