Search papers, labs, and topics across Lattice.
This paper introduces the Groundhog Bit-Flip Attack (GBFA), a novel Denial-of-Wallet availability attack targeting Mixture-of-Experts (MoE) large language models (LLMs) by exploiting the correlation between specific experts and tokens. By flipping routing-layer bits, the authors demonstrate a significant increase in token usage鈥攁veraging 5912% inflation鈥攚hile maintaining semantic fidelity across various LLM tasks. The findings expose a critical vulnerability in MoE architectures, underscoring the potential for adversarial manipulation in scalable LLMs.
Flipping just a few bits can inflate LLM output by over 5900%, revealing a critical vulnerability in Mixture-of-Experts architectures.
Mixture-of-Experts (MoE) architectures enable scalable and efficient large language models (LLMs) by selectively activating expert sub-networks through a routing mechanism. However, this adaptive design introduces a new attack surface: specific experts become disproportionately correlated with certain tokens (e.g., end-of-sequence), allowing adversaries to manipulate model behavior via lightweight perturbations. In this work, we present \textbf{Groundhog Bit-Flip Attack (GBFA)}, the first bit-flip-based \textit{ Denial-of-Wallet availability attack} against MoE-based LLMs. By identifying and flipping routing-layer bits associated with related expert activations, we demonstrate that GBFA substantially extends the decoding token usage across three different LLM modes: conversational, reasoning, and agentic tasks, while largely preserving semantic fidelity. Across four main real-world MoE-based LLMs, manually deactivating on average fewer than \textbf{4 experts} drives average output inflation to $\mathbf{5912\%}$, with the majority of test samples reaching max tokens. These results reveal a robustness vulnerability of MoE architectures to bit flip, and highlight the potential of GBFA as an availability attack against LLMs.