Search papers, labs, and topics across Lattice.
This paper introduces ToxScreen, a benchmark designed to assess the ability of defenders to detect backdoored triggers in large language models (LLMs) under realistic conditions, such as having no access to training data or a trusted reference model. The study reveals that while gradient-based prompt optimization is ineffective for trigger recovery, a token look-up method based on attack success rates can successfully identify triggers when backdoors are present. The findings highlight a distinct operational mechanism for backdoors compared to jailbreaks, providing valuable insights for developing detection strategies in high-stakes AI applications.
A novel detection approach reveals that while traditional methods fail, a simple token look-up can effectively recover backdoor triggers in LLMs.
As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time. We ask whether a defender can recover such a trigger under realistic affordances, namely white-box access to the weights and knowledge of the behavior of concern, but no training data, no trusted reference model, no knowledge of the trigger, and no certainty that the model is poisoned. To evaluate whether a defender can recover such a trigger under realistic settings, we release ToxScreen, a benchmark of roughly 800 backdoored models spanning attack objectives, trigger mechanisms, poisoning rates, model scales, and backdoor training mechanisms. We also assert that the backdoors are high-quality: they achieve high attack success rates, generalize to unseen harmful inputs, and preserve clean-task performance. Scoring recovery of the planted trigger, we find that gradient-based prompt optimization fails in recovery, whereas a token look-up that ranks candidates by attack-success rate recovers the trigger wherever the backdoor is effective. To understand this more, we study the relationship between attack behaviors and the weights of an LLM. We find a phenomenon whereby backdoors operate via different mechanistic strategies than jailbreaks, allowing defenders to filter jailbreaks. Finally, no method reliably surfaces every backdoor, but a broadly jailbreakable model is itself anomalous, a useful signal even when the exact trigger is not recovered. We release all models and evaluation code