Search papers, labs, and topics across Lattice.
Stony Brook University
1
0
3
Matched execution scores can hide up to 64.3 points of command-path failure, revealing a critical gap in evaluating LLM coding agents.