Search papers, labs, and topics across Lattice.
Affiliation:
3
1
7
0
Subtle prompt changes can destabilize LLMs significantly, but four key factors can mitigate this sensitivity by targeting low-order interactions.
Current multimodal agents fail to consistently pass CAPTCHA tests, revealing fundamental limitations in their ability to replace humans in automated workflows.
VLMs don't fail to *recognize* harmful intent when jailbroken; instead, visual inputs *shift* their internal representations into a distinct "jailbreak state," opening a new avenue for defense.