Search papers, labs, and topics across Lattice.
Affiliation:
6
0
8
5
This work proposes Collision Mesh Poisoning (CMP), the first poisoning attack against robotic manipulation delivered through the 3D asset supply chain, and evaluates several defenses and shows that they are insufficient to defend against CMP, highlighting the need for new defenses.
Current AI coding agents struggle with large-scale refactoring tasks, achieving only a 41.2% success rate on a newly curated benchmark designed to challenge their capabilities.
Static scanners fail against adaptive evasions, but a new behavior-centric auditor can detect 97% of malicious skills with minimal false positives.
Coding agents guess their way through underspecified instructions, leading to alarming action-boundary violations that challenge the notion of safe autonomy.
Forget task-specific overfitting: training coding agents on atomic skills unlocks surprisingly broad generalization to complex software engineering tasks.
LLMs struggle to translate code into formal specifications, as evidenced by their poor performance on the new Model-Bench benchmark, revealing a critical gap in their ability to support formal verification.