Search papers, labs, and topics across Lattice.
This paper investigates the vulnerability of retrieval-augmented generation (RAG) systems to poisoning attacks, where adversarial documents can manipulate outputs. The authors identify a phenomenon called Attention Collapse, where attention entropy decreases in attacked generations, contrasting with the dispersed attention seen in benign cases. They introduce D-SCAN, a lightweight detection framework that leverages these attention dynamics to effectively identify such attacks, even when the final output remains unchanged.
Attention Collapse reveals a hidden vulnerability in RAG systems, enabling the detection of poisoning attacks that traditional methods fail to catch.
Retrieval-augmented generation (RAG) is indispensable for enhancing large language models. However, RAGs are increasingly susceptible to poisoning attacks, in which adversarial documents are injected to manipulate generator outputs. Previous methods rely on output-side signals such as perplexity and consistency checks to detect such attacks. Nevertheless, our analysis reveals that deliberate attacks often induce false confidence, where poisoned outputs exhibit even lower perplexity than benign ones, rendering uncertainty-based detection ineffective. To address this challenge, we explore the internal dynamics of the generator and identify a distinctive signature termed Attention Collapse. Unlike the dispersed attention in benign generations, attacked generations exhibit a decrease in entropy as attention concentrates on poisoned documents. Building on these findings, we propose D-SCAN (Document-level Signal Collapse Analysis), a lightweight detection framework that monitors attention dynamics to identify attacked generations. Extensive experiments on multiple attack benchmarks demonstrate the effectiveness of our method. Moreover, D-SCAN can detect attacks even when they fail to alter the final answer. Code is available at https://github.com/yingtaoren/D-Scan.git.