Search papers, labs, and topics across Lattice.
This study introduces the first multilingual spoken hallucination detection benchmark, featuring 12,013 samples across English, Russian, and Kazakh, with controlled hallucinations of varying types and severity. The authors evaluate fine-tuned multilingual encoders and multimodal decoder models in both transcript-based and direct audio processing settings, revealing that transcript-based detection consistently outperforms audio processing, particularly in the presence of ASR errors. Notably, the research shows that detectors trained on synthetic data can effectively transfer to real-world fake news detection, achieving a macro-F1 score of 0.82-0.88 on original text.
Transcript-based detection outperforms direct audio processing in identifying spoken hallucinations, revealing a critical gap in current methodologies.
While text-based hallucination detection has been extensively studied, spoken hallucination detection remains largely unexplored, particularly for low-resource languages. We present the first multilingual spoken hallucination benchmark comprising 12,013 news samples across English, Russian, and Kazakh with controlled hallucinations of three types and three severity levels. Samples comprise original articles and aligned hallucinated counterparts in text and audio. We complement the synthetic corpus with 290 fact-checked fake news items collected natively in Russian (225) and Kazakh (65), translated into the other language and rendered through the same TTS-ASR pipeline. We assess fine-tuned multilingual encoders and, in zero-shot in-context settings, multimodal decoder models on transcript-based versus direct audio processing. Transcript-based detection generally outperforms direct audio processing, with binary-task degradation for strong encoders tracking per-language ASR error. On real-world fakes, synthetic-trained detectors transfer strongly (macro-F1 0.82-0.88 on original text), while Russian provenance analysis reveals both veracity-related and model-dependent machine-style signals, quantifying a key confound in synthetic hallucination benchmarks.