Search papers, labs, and topics across Lattice.
3
0
3
7
Long-horizon evaluations reveal that agents suffer from compounding errors, but without proper baseline comparisons, we can't fully grasp the reasons behind their failures.
Real-world coding tasks, reverse-engineered from actual commits and scenarios, make Tencent WorkBuddy Bench a game-changer in contamination-resistant evaluation for coding agents.
A Qwen3-8B model, trained with a new SFT+RLAIF recipe on a challenging new benchmark, SWE-QA-Pro, beats GPT-4o in repository-level code understanding.