Search papers, labs, and topics across Lattice.
4
0
6
5
Static scanners fail against adaptive evasions, but a new behavior-centric auditor can detect 97% of malicious skills with minimal false positives.
Coding agents guess their way through underspecified instructions, leading to alarming action-boundary violations that challenge the notion of safe autonomy.
Forget task-specific overfitting: training coding agents on atomic skills unlocks surprisingly broad generalization to complex software engineering tasks.
LLMs struggle to translate code into formal specifications, as evidenced by their poor performance on the new Model-Bench benchmark, revealing a critical gap in their ability to support formal verification.