Search papers, labs, and topics across Lattice.
3
0
5
3
A composite attack can breach a supposedly robust self-check defense, achieving up to 67% success where individual methods fail.
A guard-agnostic amplifier improves safety classifier performance but exposes a troubling trade-off between attack success and benign refusal rates.
LLM safety filters, which rely on semantic pattern matching, can be bypassed at scale by encoding harmful prompts as coherent mathematical problems, revealing a fundamental vulnerability.