Search papers, labs, and topics across Lattice.
2
0
3
14
Long-horizon evaluations reveal that agents suffer from compounding errors, but without proper baseline comparisons, we can't fully grasp the reasons behind their failures.
Real-world coding tasks, reverse-engineered from actual commits and scenarios, make Tencent WorkBuddy Bench a game-changer in contamination-resistant evaluation for coding agents.