Search papers, labs, and topics across Lattice.
3
0
4
9
EarlyEval can cut evaluation costs by up to 44% without sacrificing accuracy, revolutionizing how we assess LLM agents.
Current AI coding agents struggle with large-scale refactoring tasks, achieving only a 41.2% success rate on a newly curated benchmark designed to challenge their capabilities.
Pruning tool outputs directly within the agent leads to a remarkable 39% reduction in token usage without sacrificing performance.