Search papers, labs, and topics across Lattice.
This study explores the impact of multimodal aggregation on automatic speaker verification (ASV) systems, revealing that as multiple anonymized utterances are processed, the performance of speaker verification improves significantly. By incorporating prosodic and linguistic cues alongside acoustic data, the researchers demonstrate that multimodal systems achieve better accuracy than their unimodal counterparts. Notably, even with just five anonymized utterances, the combination of audio and text reduces the equal error rate (EER) by over 15%, highlighting vulnerabilities in speaker anonymization techniques.
Even anonymized speech reveals over 15% more speaker information when combining audio and text, challenging current privacy assumptions.
Most automatic speaker verification (ASV) systems operate on individual utterances, despite real-world interactions typically consisting of multiple utterances. As speech accumulates, increasingly rich speaker information becomes available through acoustic, prosodic, and linguistic cues, potentially challenging speaker anonymization methods that primarily target vocal characteristics. We investigate ASV in a multi-utterance, multimodal setting and examine whether aggregating information across anonymized speech impacts privacy. We first study audio-only aggregation across multiple anonymized utterances and observe consistent performance improvements as more speech becomes available. We then incorporate prosodic and linguistic information, showing that multimodal systems outperform unimodal approaches. Finally, we compare aggregation strategies and find that frame-level aggregation yields the lowest EERs. Even with only five anonymized utterances, combining audio and text reduces EER by over 15% relative to audio-only aggregation, demonstrating that substantial speaker-discriminative information remains accessible despite anonymization.