Search papers, labs, and topics across Lattice.
This paper introduces MAD2, a benchmark for multimodal claim verification in spoken dialogues, comprising 1,000 dialogues with over 3,368 claims. The study reveals that incorporating dialogue context significantly enhances verification accuracy, particularly in live-moderation scenarios, while also demonstrating that the conversational structure is more critical for verification than the framing of misinformation. Key findings indicate that using preceding context can achieve performance comparable to offline methods, highlighting the importance of context in spoken dialogue verification.
Conversational structure, not just misinformation framing, is the key to improving spoken claim verification accuracy.
Every day, millions absorb claims from podcasts and streams that no fact-checker ever sees. Spoken misinformation is built through conversation, where credibility comes not from facts alone but from how claims are framed, reinforced, or left unchallenged across turns. Yet fact-checking has focused on isolated text, leaving dialogue audio under-studied. We introduce MAD2, a new Multi-turn Audio Dialogues benchmark for spoken claim verification, containing 1,000 two-speaker dialogues with 3,368 check-worthy claims and approximately 10 hours of audio, and propose calibrated multimodal fusion of a context-aware audio encoder and a dialogue-aware text model. Across settings, adding dialogue context improves verification, but the gains depend on scenario type. Using only preceding context often matches offline performance, supporting live-moderation settings, and audio contributes most when transcript-based models are destabilized by additional context. Overall, conversational structure matters more for verification than misinformation framing.