Search papers, labs, and topics across Lattice.
This paper introduces DINO-VPT, a novel vision-only framework for face anti-spoofing that utilizes hierarchical visual prompt tuning to address both physical and digital spoofing threats. By employing a Prompt Routing Network (PRN) to dynamically inject prompts based on input features, DINO-VPT effectively disentangles various spoofing artifacts, eliminating the need for complex multimodal fusion. Evaluations on the UniAttackData benchmark reveal that DINO-VPT outperforms existing state-of-the-art Vision-Language Model (VLM) approaches, demonstrating the efficacy of a streamlined architecture in achieving high accuracy in unified face anti-spoofing tasks.
A lightweight vision-only framework achieves superior face anti-spoofing accuracy without the complexity of multimodal fusion.
With the increasing diversity of spoofing attacks, there is a growing demand for unified Face Anti-Spoofing (FAS) models capable of detecting both physical and digital threats. While existing Vision-Language Models (VLMs) demonstrate high generalization in this context, they heavily rely on complex multimodal fusion and external text encoders. In this paper, we propose DINO-VPT, a lightweight, vision-only framework leveraging hierarchical visual prompt tuning. By dynamically injecting prompts conditioned on input features via a Prompt Routing Network (PRN), our method effectively disentangles diverse spoofing artifacts without requiring multimodal fusion. Evaluations on the UniAttackData benchmark demonstrate that DINO-VPT achieves higher accuracy than state-of-the-art VLM-based methods. Our results indicate that a properly structured vision-only architecture can achieve state-of-the-art performance in unified FAS without the need for multimodal supervision.