Search papers, labs, and topics across Lattice.
University College London
3
0
4
As SE agents become mainstream, the shift to evaluation-driven development reveals new challenges that could redefine engineering workflows.
Today's agents are surprisingly bad at real-world terminal tasks, with even frontier models failing nearly 40% of the time on everyday workflows.
AI agents and humans exhibit over 10 distinct repair behaviors when performing urgent hot fixes, suggesting opportunities for targeted human-automation collaboration.