Search papers, labs, and topics across Lattice.
8
0
5
11
EarlyEval can cut evaluation costs by up to 44% without sacrificing accuracy, revolutionizing how we assess LLM agents.
Parallel reasoning in LLM agents can cut decoding time by up to 43% while maintaining performance, reshaping agent efficiency.
Brevis compresses tensor data by synthesizing a DSL program, achieving over 30% smaller archives than leading general-purpose compressors while ensuring bit-exact reconstruction.
As SE agents become mainstream, the shift to evaluation-driven development reveals new challenges that could redefine engineering workflows.
LLM agents can restore compatibility in over 60% of outdated repositories, but their effectiveness varies dramatically based on the complexity of the required changes.
Benchmark scores for coding agents may mislead progress assessments, with only 39% of GSO tasks passing validity checks across machines.
Execution in LLM-based program repair is often a costly default that yields minimal benefits, suggesting a need for a strategic reevaluation of its use.
LLMs can generate code 55% faster by executing code *while* generating it, challenging the traditional generate-then-execute paradigm.