Search papers, labs, and topics across Lattice.
This paper evaluates the effectiveness of large audio language models (LALMs) for spoofing-aware speaker verification (SASV), contrasting their performance with traditional modular ASV-countermeasure (CM) systems. While LALMs initially performed poorly in a zero-shot setting, task-specific adaptation significantly improved their capabilities, demonstrating that they can achieve competitive SASV performance through various optimization strategies. The findings suggest that LALMs can serve as a robust and auditable foundation for unified SASV, highlighting their potential advantages over conventional cascaded approaches.
Task-specific adaptation transforms LALMs from near-chance performers to competitive players in spoofing-aware speaker verification.
Recent advances in text-to-speech and voice cloning make high-quality spoofing inexpensive and scalable, threatening voice authentication systems, especially automatic speaker verification (ASV). Existing defenses mainly address this threat through binary countermeasures (CMs) for deepfake detection or spoofing-aware speaker verification (SASV), where current systems are dominated by modular ASV-CM fusion and cascaded pipelines. Although large audio language models (LALMs) have shown promise on related audio tasks, including CM and ASV, their use for SASV remains unexplored, despite their capacity to produce natural-language rationales for auditing and robustness beyond discriminative predictions. This work systematically evaluates LALMs for SASV against conventional pipelines under zero-shot prompting, supervised adaptation, reasoning-oriented training, and reinforcement-learning-based optimization. Our results show that pretrained LALMs are near chance in the zero-shot setting, confirming that they are not natively suited to SASV, but that task-specific adaptation closes this gap. We further find that competitive SASV performance can be achieved through several distinct routes. These findings position LALMs as a promising and auditable foundation for unified SASV, while clarifying where conventional cascade systems still lead.