Search papers, labs, and topics across Lattice.
4
0
3
3
Agent-involved code reviews speed up decision-making but compromise on quality, challenging the assumption that faster reviews equate to better outcomes.
ML evaluation harnesses, the unsung heroes of model development, are plagued by surprisingly mundane software engineering issues like missing documentation and unimplemented features, hindering reliable assessment.
AI code review agents may scale defect screening, but their suggestions are adopted less often and, when adopted, can actually *worsen* code quality, underscoring the critical need for human oversight.
Turns out, almost all AI agent tool descriptions are "smelly," and while fixing them improves performance, it also introduces a tricky efficiency trade-off that can be solved by carefully choosing which components to include.