Search papers, labs, and topics across Lattice.
2
0
2
Pairing decoy images with encoded jailbreak prompts can reduce attack success rates by up to 73 percentage points, revealing a surprising interaction in VLM defense mechanisms.
A guard-agnostic amplifier improves safety classifier performance but exposes a troubling trade-off between attack success and benign refusal rates.