Search papers, labs, and topics across Lattice.
This paper addresses the challenge of maintaining context adherence in multi-round spoken dialogue systems by identifying a critical gap between latent context awareness and active adherence during decoding. The authors introduce a novel Context-Aware Decoding (CAD) approach that utilizes internal attention mechanisms to enhance the influence of relevant historical utterances during inference. Evaluations on the Audio MultiChallenge benchmark reveal significant improvements in Semantic Memory and Self Coherence, underscoring the effectiveness of the proposed method in achieving context-faithful dialogue generation.
Bridging the gap between context awareness and adherence could redefine how spoken dialogue systems maintain coherence across conversations.
Despite the success of end-to-end (E2E) spoken dialogue systems, maintaining strict context adherence in multi-round conversations remains a challenge. While prior works attribute these failures to models forgetting dialogue history, we highlight an equally critical but overlooked bottleneck: a gap between latent context awareness and active adherence. Although models internally recognize relevant past utterances, strong parametric priors often overshadow these signals during decoding. To bridge this gap, we propose an audio-adapted Context-Aware Decoding (CAD) approach. By leveraging internal attention mechanisms to isolate key historical rounds, our approach contrasts output distributions with and without this key context during inference, directly amplifying multimodal contextual signals. Evaluations on the Audio MultiChallenge benchmark demonstrate significant improvements in Semantic Memory and Self Coherence subtasks, successfully enforcing strict, context-faithful adherence.