Search papers, labs, and topics across Lattice.
JoyAI-Talker is a full-duplex speech dialogue system designed to enhance empathetic interactions in voice agents by integrating a modular Thinker-Talker architecture and a unified speech-text joint training pipeline. This approach effectively mitigates cognitive degradation, preserving the model's reasoning and logical capabilities while enabling expressive speech synthesis through a text-controllable generation paradigm. Evaluations indicate that JoyAI-Talker achieves a high response rate of 0.88 during user interruptions, showcasing its potential for fluid and natural speech dialogue in real-world applications.
Achieving a response rate of 0.88 in full-duplex interactions, JoyAI-Talker redefines how empathetic voice agents can engage in natural conversations even amidst interruptions.
We present JoyAI-Talker, a full-duplex speech dialogue system that delivers robust foundation model capabilities while empowering empathetic interaction and voice agent intelligence. JoyAI-Talker adopts a modular Thinker-Talker architecture and further implements a unified speech-text joint training pipeline to mitigate the common "cognitive degradation" bottleneck, thereby largely preserving the model's core textual reasoning, STEM, and logical capabilities while extending them to speech-based interaction. For expressive speech synthesis, the Talker module employs a text-controllable generation paradigm that enables natural-language instructions to flexibly control vocal attributes and localized paralinguistic events, such as laughter and sighs, supporting more expressive and fine-grained speech responses. To enhance conversational empathy, we introduce the Persona-Adaptive Empathetic Response (PAER) framework. PAER employs a hierarchical cognitive pipeline to extract non-verbal speaker cues, such as gender, age, and emotional state, from raw input audio, incorporate them into the Thinker's CoT reasoning, and generate context-adaptive responses that align semantically appropriate text with fine-grained control over utterance-level expressiveness and localized paralinguistic events, including sighs, speaking rate, and volume. We further integrate Joy-Duplex, a state-driven, plug-and-play full-duplex framework that functions as an efficient gating engine for real-time turn control. Extensive evaluations show that JoyAI-Talker achieves highly competitive performance on foundational T2T and S2T benchmarks. In full-duplex evaluation, the system reaches a high response rate of 0.88 under user interruptions while maintaining an extremely low false-trigger rate under background speech, demonstrating its readiness for fluid and natural speech dialogue.