Search papers, labs, and topics across Lattice.
3
0
4
5
Static environments can be transformed on-the-fly to better suit agent learning, resulting in up to a 9.0-point performance boost with fewer execution steps.
FinanceHarness not only automates financial deep research but also reveals that even advanced LLMs struggle with specialized financial tasks, scoring below 40% on rigorous benchmarks.
Autonomous research agents are surprisingly unreliable, with existing systems hallucinating references 21% of the time and failing to align methods with code as often as 80%, but a new "Chain-of-Evidence" approach can fix this.