Search papers, labs, and topics across Lattice.
3
0
3
0
Instruction-dense visual jailbreaks can covertly embed harmful instructions in images, bypassing safety measures that protect against direct text generation.
GhostPrompt achieves over 30% higher attack success rates while slashing computation time by nearly 70%, transforming how adversarial prompts can be utilized across diverse images.
LLMs can inherently recognize policy violations, and PVDetector exploits this to achieve unprecedented detection accuracy against prompt injection attacks.