Search papers, labs, and topics across Lattice.
2
0
2
2
EarlyEval can cut evaluation costs by up to 44% without sacrificing accuracy, revolutionizing how we assess LLM agents.
Benchmark scores for coding agents may mislead progress assessments, with only 39% of GSO tasks passing validity checks across machines.