Search papers, labs, and topics across Lattice.
This paper introduces STAIN-FL, a stealthy targeted backdoor attack framework for federated video anomaly detection that leverages contextual triggers found in surveillance environments. By manipulating anomaly-to-benign labels and applying gradient masking, STAIN-FL maintains high clean accuracy while effectively inducing misclassifications for triggered anomalies. Evaluations on the UCF-Crime dataset reveal that sparse attacks can achieve significant backdoor accuracy (up to 56.7%) with minimal impact on overall model performance, underscoring the persistent risk of such attacks in real-world applications.
Sparse backdoor attacks can misclassify over half of triggered anomalies while keeping clean accuracy losses below 2%, revealing a stealthy threat in federated learning systems.
Federated video anomaly detection trains model collaboratively without sharing raw surveillance footage, but limited server-side visibility lets compromised clients to inject backdoor via malicious updates. This paper introduces STAIN-FL, a stealthy targeted backdoor attack injection framework that uses naturally occurring surveillance conditions, including low-light scenes, indoor settings, and crowd density, as contextual triggers. STAIN-FL combines anomaly-to-benign label \textit{manipulation} with gradient masking over least-updated coordinates to preserve clean accuracy while inducing trigger-conditioned misclassification. We evaluate STAIN-FL on \texttt{UCF-Crime} using 1024-dimensional I3D features in a non-IID four-client multi-agency setting, comparing FedAvg and FedProx under sparse and continuous attacks. Results show that sparse attacks have low-detectability, operationally significant attacks rather than high-intensity attacks: they keep the mean clean-accuracy drop below $2\%$, yet still misclassify more than half of triggered anomalies at peak backdoor accuracy under FedAvg ($56.7\%$) and FedProx ($54.2\%$). Under FedAvg, the sparse backdoor remains above the $25\%$ backdoor-accuracy threshold for an average of $336$ post-attack rounds, highlighting the persistence risk of contextually triggered attacks in surveillance systems.