Search papers, labs, and topics across Lattice.
The authors benchmark end-to-end Audio Language Models (ALMs) on unseparated, mixed-speaker interview audio of children who stutter across semantic summarization and speech entailment tasks. Testing this regime probes whether audio foundation models can bypass brittle ASR cascades to parse non-normative pediatric acoustics, linguistic disfluencies, and multi-speaker attribution simultaneously. While ALMs successfully capture high-level semantic gist, their reasoning fidelity degrades sharply and suffers from adult-speaker leakage as speech disfluency increases.
Audio language models can grasp the broad gist of stuttered child speech, but their reasoning completely collapses and leaks multi-speaker context as disfluency rates rise.
Child speech differs from adult speech in acoustics, prosody, and linguistic structures. Speech disfluencies (such as repetitions) further challenge automatic understanding. While Audio Language Models (ALMs) show strong semantic reasoning from speech audio, their ability to reason about disfluent child speech in mixed-speaker settings remains unexplored. We investigate this through two tasks: child-focused semantic summarization and speech entailment. Experiments use recordings of children who stutter in mixed speaker interviews without explicit speaker separation. Models are instruction-guided to focus on the child, preserve clinically relevant disfluencies, and avoid adult-speech leakage. Evaluation combines LLM-based judges and reference-based metrics, anchored by transcript-oracle baselines to isolate errors. Results show that while ALMs extract high-level meaning from stuttered speech, reasoning degrades significantly with increased