Search papers, labs, and topics across Lattice.
4
0
6
7
Black-box RL can boost agent performance by nearly 15 points on complex tasks, revealing a new frontier for scalable optimization.
Autonomous evolution of agent harnesses can yield significant performance improvements, yet struggles in specific task environments reveal critical limitations.
Building agents that can reliably automate complex, multi-step workflows over local files and tools just got a whole lot easier.
Today's code-generating AI falls apart when faced with real-world software engineering tasks that demand cross-repository reasoning and external knowledge, achieving less than 45% success on the new BeyondSWE benchmark.