Search papers, labs, and topics across Lattice.
3
0
5
13
Current AI coding agents struggle with large-scale refactoring tasks, achieving only a 41.2% success rate on a newly curated benchmark designed to challenge their capabilities.
Stop feeding your LLM-based bug reproduction tools irrelevant code: iCoRe's correlation-aware retrieval boosts test generation accuracy by up to 31.7%.
LLMs struggle to translate code into formal specifications, as evidenced by their poor performance on the new Model-Bench benchmark, revealing a critical gap in their ability to support formal verification.