Search papers, labs, and topics across Lattice.
Affiliation:
1
0
2
Frontier LLM agents still struggle to judge their own completion: while older models lose tracking accuracy mid-task, the newest generation reliably monitors intermediate execution only to freeze conservatively at the finish line.