Search papers, labs, and topics across Lattice.
2
1
6
10
Matched execution scores can hide up to 64.3 points of command-path failure, revealing a critical gap in evaluating LLM coding agents.
Reasoning-based safety guardrails, once thought to be a strong defense against jailbreaks, crumble with just a few strategically placed tokens.