Search papers, labs, and topics across Lattice.
4
0
4
11
Benchmark scores for coding agents may mislead progress assessments, with only 39% of GSO tasks passing validity checks across machines.
LLM agents can restore compatibility in over 60% of outdated repositories, but their effectiveness varies dramatically based on the complexity of the required changes.
Execution in LLM-based program repair is often a costly default that yields minimal benefits, suggesting a need for a strategic reevaluation of its use.
LLMs can generate code 55% faster by executing code *while* generating it, challenging the traditional generate-then-execute paradigm.