Search papers, labs, and topics across Lattice.
This study challenges the face-centric approach to presentation attack detection (PAD) by introducing TPO, a novel dataset of non-facial objects (tomatoes, potatoes, and onions) to train PAD models. The authors demonstrate that a detector trained on TPO achieves a remarkable average AUC of 92.70% across multiple face PAD benchmarks, outperforming models trained on synthetic faces and competing with those trained on real faces. Their findings suggest that PAD representations can be effectively learned without facial content, which could lead to advancements in privacy-preserving detection methods.
Transferable presentation attack representations can be learned without relying on facial content, achieving high performance across standard benchmarks.
Face presentation attack detection (PAD) is traditionally formulated as a face-specific problem, although many of the visual artifacts introduced by print, replay, and recapture processes are not inherently tied to facial appearance. In this work, we investigate whether transferable PAD representations can be learned without using faces during downstream PAD training. To this end, we introduce TPO, a controlled face-free presentation attack dataset consisting of bona fide, print, and replay recordings of, almost randomly chosen, tomatoes, potatoes, and onions acquired under protocols that closely mirror conventional face PAD datasets. Using a foundation-model-based PAD architecture, we demonstrate that a detector trained on TPO achieves an average AUC of 92.70% across four standard cross-dataset face PAD benchmarks, outperforming training on synthetic faces and remaining competitive with models trained on real face datasets. Conversely, models trained on face PAD datasets transfer consistently above chance to TPO, suggesting that the learned representations capture characteristics of the presentation process rather than object semantics. Furthermore, incorporating TPO into conventional face PAD training consistently improves cross-dataset performance under fixed optimization budgets, indicating that face-free data provides complementary information rather than simply additional training samples. Finally, representation and frequency analyses provide further evidence that transferable PAD representations cannot be explained by a single spectral artifact but instead encode richer presentation cues shared across object categories. Together, these results provide empirical evidence that transferable presentation attack representations can be learned independently of facial content, opening new opportunities for privacy-preserving and identity-independent PAD development.