Search papers, labs, and topics across Lattice.
This paper addresses the critical issue of backdoor generalization in machine learning, specifically how a backdoor learned from a training-time trigger can activate with an inference-time trigger that was not present during training. The authors introduce Lilith, a novel black-box framework that effectively creates a compact target-side vulnerability using a single training anchor and constructs a family of inference triggers that maintain the necessary representation geometry. Experimental results demonstrate that Lilith achieves high success rates in family-wise attacks with minimal impact on benign utility, revealing a significant blind spot in current backdoor evaluation methods.
Backdoor attacks can generalize across trigger families, with Lilith achieving high success rates while preserving benign model performance.
Machine-learning services increasingly rely on public data, third-party providers, and outsourced training, creating opportunities for data-poisoning attacks that implant persistent malicious behavior while preserving benign utility. However, existing backdoor studies largely evaluate exact trigger reuse, training-exposed trigger diversity, or variations along predefined transformation axes. They therefore leave a critical blind spot: whether a backdoor learned from one training-time trigger can generalize to an inference-time trigger family absent from victim training. We formulate this problem as backdoor generalization under training--inference trigger shift and introduce Lilith, a black-box anchor-to-family framework. Using only disjoint surrogate resources, Lilith first induces a compact target-side vulnerability with a single training anchor, then constructs a bounded inference-only family that preserves the anchor-induced representation geometry. We characterize this mechanism through anchor clearance and family reach, deriving sufficient conditions for family-wise target preservation under local regularity and bounded surrogate--victim discrepancy. Experiments across datasets, architectures, poisoning rates, and defenses show that Lilith achieves high family-wise attack success with limited utility degradation and a small trigger generalization gap. Additional analyses show that family activation depends on representation alignment rather than the proposal mechanism, exposing a broader threat overlooked by exact-trigger evaluation.