Search papers, labs, and topics across Lattice.
3
0
6
2
Living-Harness enables agents to learn from past failures dynamically, leading to substantial performance improvements in interactive tasks.
LLMs struggle with financial reasoning under real-world conditions, revealing critical flaws in their ability to handle complex, long-horizon tasks.
MLLMs are often overconfident, but a new confidence-driven training and test-time scaling approach can boost accuracy by 8.8% across benchmarks.