Search papers, labs, and topics across Lattice.
3
0
4
Guardrails in AI systems may claim safety benefits that are largely illusory, with their actual effectiveness hinging on the specifics of input encoding and channel dynamics.
Pairing decoy images with encoded jailbreak prompts can reduce attack success rates by up to 73 percentage points, revealing a surprising interaction in VLM defense mechanisms.
A guard-agnostic amplifier improves safety classifier performance but exposes a troubling trade-off between attack success and benign refusal rates.