Search papers, labs, and topics across Lattice.
Speech recognition, text-to-speech, audio generation, music AI, and spoken language understanding.
#16 of 24
4
X-AuT, a progressive framework that selects layer combinations through short behavioral probes and restores the pruned model through representation alignment, cross-scale distillation, scheduled student-policy supervision, and LoRA finetuning, is introduced.
TART is the first framework to generate guitar tablature with both fingering and expressive technique annotations directly from guitar audio, and is the first framework to generate guitar tablature with both fingering and expressive technique annotations directly from guitar audio.
Techniques for improving speech recognition-based transcript creation in multiple languages from videos to better train these automated tools are presented.
The central finding is that aggregate WER hides code switching behavior, and the best system by WER (an ASR model) is statistically indistinguishable from a leading audio LM on WER, yet the audio LM is significantly better on every switch localized metric.
A continuous-time mechanism for duration-aware acoustic modelling in TTS using neural controlled differential equations (CDEs) is proposed and CDEs are position as a promising design space for continuous-time and duration-aware style-sensitive TTS.
This paper describes the system for Task~2 of the second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge, which adapt Qwen3-Omni-30B-A3B-Instruct with a segment-evidence-aware data and post-training pipeline and obtains 90.92% accuracy on the final official evaluation set.
These results validate the deployed navigation and inspection closed loop, while HROS provides an extensible software foundation for memory-augmented, voice-aware, and continuously improvable embodied inspection agents.
CEM-TUDASR, a computationally efficient unsupervised Transformer-based super-resolution framework for WCE image enhancement without paired low-resolution (LR) and high-resolution (HR) training data, is proposed.
ToxicRAG is presented, a one-document-per-target attack that expresses misinformation as a coherent knowledge-update narrative that matches or exceeds the strongest evaluated baseline in every combination of dataset--model combinations.
Experiments show that ZipCodec substantially outperforms existing streaming codecs at comparable bitrates in both reconstruction and downstream tasks, while operating at a significantly lower frame rate.
A lightweight Text Risk Score (TRS) is proposed, which estimates synthesis risk from interpretable text features without manual annotation or model training and shows a positive correlation with content errors and duration abnormalities, demonstrating its usefulness as a low-cost pre-synthesis indicator for complex-text risk diagnosis in low-resource multilingual TTS.
It is demonstrated that articulatory representations created through articulatory inversion can be used as an interpretable basis for accent comparison and that optimal transport provides a framework for accent comparison across arbitrary recording types.
This paper presents SurgicalRoomAgent, a voice-interactive multi-agent system for smart operating rooms based on large language models (LLMs), which achieves natural language understanding, device control, intraoperative recording, and surgical report generation through a layered architecture comprising a voice interaction pipeline and an agent core.
Results show that scaling consistently improves performance, while multitask learning acts as a beneficial regularizer primarily for smaller-capacity models.
The Eloquence team's approach to Task 2 of the 2nd MLC-SLM challenge at Interspeech 2026, which involves multilingual Multiple-Choice Question Answering (MCQA) across 21 languages, is detailed, which involves multilingual Multiple-Choice Question Answering across 21 languages.
A replay-regularized, attack-aware curriculum that steps exposure based on measured attack influence is proposed that shows improved overall robustness and reduced attack-level imbalance compared with standard multi-attack training.
A unified post-training framework that equips pretrained text-to-speech models with natural-language control over segment-level emotion and duration is proposed, highlighting post-training as a practical approach to extending existing speech synthesis models.
A prompt extension framework for TUSS is proposed that incorporates downstream task information into the input prompts and switches the loss function according to the given prompt during training, enabling outputs with different signal characteristics at inference time.
Human affect is fundamentally relational: vocal expressive coupling operates across speakers at sub-second timescales, exposing the severe limitations of modeling individual emotional states in isolation.
When privacy rules forbid raw audio access, general LLMs falter on sparse ASR transcription mistakes—making two-stage, detector-gated span correction the far more effective paradigm.
Weakly-supervised LLM annotations can bridge the phonemic data gap in low-resource languages, slashing Filipino sentence-level G2P error rates from nearly 20% down to 0.54% while successfully disambiguating prosodic homographs.
Even top-performing voice MLLMs break down when forced to talk science, routinely failing to verbalize symbolic notation, decode technical jargon, and adapt across progressive multi-turn interactions.
Tracking text progress directly from latent speech tokens rather than rendered audio cuts alignment error by 88% while slashing compute overhead by 20x.
Multilingual speech foundations leave massive performance on the table: specialized monolingual Whisper models beat Whisper-large-v3 across 77 of 102 languages while slashing character error rates by nearly 3x.
Leading audio LLMs suffer up to a 41-percentage-point performance drop when prompted in native Southeast Asian languages compared to English, exposing severe blind spots in acoustic temporal grounding and paralinguistics.