Search papers, labs, and topics across Lattice.
5
0
6
Model-assisted CBRN planning shows promise, but significant uplift is limited to the radiological domain, challenging assumptions about AI's broad applicability in high-stakes scenarios.
No defense against prompt injection can achieve both high security and high fidelity, exposing a hidden cost that could compromise critical tasks in LLM applications.
Current agents are alarmingly susceptible to skill-based attacks, with success rates reaching over 86%, exposing a critical vulnerability in AI safety.
Semantic watermarks, embedded via AMR, survive paraphrasing attacks that obliterate token-level watermarks.
Current red-teaming efforts miss the forest for the trees: ARES reveals that safety failures often stem from a systemic breakdown between the LLM *and* the reward model, not just the LLM itself.