Search papers, labs, and topics across Lattice.
80 papers published across 5 labs.
This study introduces StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework.
TART is the first framework to generate guitar tablature with both fingering and expressive technique annotations directly from guitar audio, and is the first framework to generate guitar tablature with both fingering and expressive technique annotations directly from guitar audio.
Techniques for improving speech recognition-based transcript creation in multiple languages from videos to better train these automated tools are presented.
The central finding is that aggregate WER hides code switching behavior, and the best system by WER (an ASR model) is statistically indistinguishable from a leading audio LM on WER, yet the audio LM is significantly better on every switch localized metric.
This study introduces StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework.
TART is the first framework to generate guitar tablature with both fingering and expressive technique annotations directly from guitar audio, and is the first framework to generate guitar tablature with both fingering and expressive technique annotations directly from guitar audio.
Techniques for improving speech recognition-based transcript creation in multiple languages from videos to better train these automated tools are presented.
The central finding is that aggregate WER hides code switching behavior, and the best system by WER (an ASR model) is statistically indistinguishable from a leading audio LM on WER, yet the audio LM is significantly better on every switch localized metric.
A continuous-time mechanism for duration-aware acoustic modelling in TTS using neural controlled differential equations (CDEs) is proposed and CDEs are position as a promising design space for continuous-time and duration-aware style-sensitive TTS.
This paper describes the system for Task~2 of the second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge, which adapt Qwen3-Omni-30B-A3B-Instruct with a segment-evidence-aware data and post-training pipeline and obtains 90.92% accuracy on the final official evaluation set.
These results validate the deployed navigation and inspection closed loop, while HROS provides an extensible software foundation for memory-augmented, voice-aware, and continuously improvable embodied inspection agents.
CEM-TUDASR, a computationally efficient unsupervised Transformer-based super-resolution framework for WCE image enhancement without paired low-resolution (LR) and high-resolution (HR) training data, is proposed.
ToxicRAG is presented, a one-document-per-target attack that expresses misinformation as a coherent knowledge-update narrative that matches or exceeds the strongest evaluated baseline in every combination of dataset--model combinations.
Experiments show that ZipCodec substantially outperforms existing streaming codecs at comparable bitrates in both reconstruction and downstream tasks, while operating at a significantly lower frame rate.
A lightweight Text Risk Score (TRS) is proposed, which estimates synthesis risk from interpretable text features without manual annotation or model training and shows a positive correlation with content errors and duration abnormalities, demonstrating its usefulness as a low-cost pre-synthesis indicator for complex-text risk diagnosis in low-resource multilingual TTS.
It is demonstrated that articulatory representations created through articulatory inversion can be used as an interpretable basis for accent comparison and that optimal transport provides a framework for accent comparison across arbitrary recording types.
This paper presents SurgicalRoomAgent, a voice-interactive multi-agent system for smart operating rooms based on large language models (LLMs), which achieves natural language understanding, device control, intraoperative recording, and surgical report generation through a layered architecture comprising a voice interaction pipeline and an agent core.
Results show that scaling consistently improves performance, while multitask learning acts as a beneficial regularizer primarily for smaller-capacity models.
The Eloquence team's approach to Task 2 of the 2nd MLC-SLM challenge at Interspeech 2026, which involves multilingual Multiple-Choice Question Answering (MCQA) across 21 languages, is detailed, which involves multilingual Multiple-Choice Question Answering across 21 languages.
A replay-regularized, attack-aware curriculum that steps exposure based on measured attack influence is proposed that shows improved overall robustness and reduced attack-level imbalance compared with standard multi-attack training.
A unified post-training framework that equips pretrained text-to-speech models with natural-language control over segment-level emotion and duration is proposed, highlighting post-training as a practical approach to extending existing speech synthesis models.
A prompt extension framework for TUSS is proposed that incorporates downstream task information into the input prompts and switches the loss function according to the given prompt during training, enabling outputs with different signal characteristics at inference time.
X-AuT, a progressive framework that selects layer combinations through short behavioral probes and restores the pruned model through representation alignment, cross-scale distillation, scheduled student-policy supervision, and LoRA finetuning, is introduced.
Human affect is fundamentally relational: vocal expressive coupling operates across speakers at sub-second timescales, exposing the severe limitations of modeling individual emotional states in isolation.
When privacy rules forbid raw audio access, general LLMs falter on sparse ASR transcription mistakes—making two-stage, detector-gated span correction the far more effective paradigm.
Weakly-supervised LLM annotations can bridge the phonemic data gap in low-resource languages, slashing Filipino sentence-level G2P error rates from nearly 20% down to 0.54% while successfully disambiguating prosodic homographs.
Even top-performing voice MLLMs break down when forced to talk science, routinely failing to verbalize symbolic notation, decode technical jargon, and adapt across progressive multi-turn interactions.
Tracking text progress directly from latent speech tokens rather than rendered audio cuts alignment error by 88% while slashing compute overhead by 20x.
Multilingual speech foundations leave massive performance on the table: specialized monolingual Whisper models beat Whisper-large-v3 across 77 of 102 languages while slashing character error rates by nearly 3x.
Leading audio LLMs suffer up to a 41-percentage-point performance drop when prompted in native Southeast Asian languages compared to English, exposing severe blind spots in acoustic temporal grounding and paralinguistics.
This work introduces Brain2Semantics2Text, a method that reconstructs text through an intermediate semantic embedding space and describes the core principles of the approach, its implementation, and the strategies used to mitigate the challenges of learning a reliable neural-to-semantic mapping.
Full-duplex multimodal interaction and deep agentic reasoning no longer require separate systems: Gander delivers native barge-in capabilities and streaming video-audio understanding by decoupling reflex-level generation from deliberate task execution.
X2Streaming-ASR slashes commit latency to as low as 27 ms while achieving superior recognition accuracy, setting a new benchmark for streaming ASR systems.
An auditable decision record that carries four aligned cues into a late calibration step: a passive detector score, a conditional keyed-probe score on a marked derivative, retrieval support, and a speaker-profile margin, together with explicit disagreement coordinates is answered.
Standard metrics routinely misclassify valid real-time summarization as translation failure, but structuring LLM evaluation around deterministic MQM error penalties recovers human system rankings with up to 0.707 Kendall concordance.
This work presents TontaubeV1, a model that preserves natural prosody while enabling streaming from a single consumer GPU, and is designed primarily for English and German, with additional multilingual support.
This study conducted a study with a professional percussionist, co-developing a gesture mapping toolkit and recording a ten-track album, which supported the development of a continuous gesture recognition method for percussive mapping and surfaced insights into the design process.
Natural-language instructions can now handle both zero-shot speech synthesis and surgical acoustic editing within a single unified model, operating at 4-step distilled inference speeds without classifier-free guidance.
Simply cutting off a voice assistant mid-sentence with generic audio is enough to break its safety guardrails, spiking jailbreak success rates by up to 39 percentage points.
In-the-wild audio deepfakes break standard forensic assumptions: 21 of 29 temporal coherence metrics invert their discriminative direction outside training distributions, with markers like entropy flipping signs entirely between synthetic speech and music.
Some Irish traditional tunes are 29% more likely to be learned based on their social and melodic features, revealing a surprising mechanism behind cultural selection and diversity.
Achieving an unprecedented equal error rate of 0.88% in audio deepfake detection, this model sets a new benchmark for robustness against sophisticated voice synthesis attacks.
Correctly aligning audio descriptions with representations can boost classification performance by nearly 5 points, revealing the critical role of semantic refinement in audio tasks.
Every speech-to-speech model misgenders speakers based on content, not voice, with misgendering rates soaring to 90% when voice and content clash.
The first systematic characterization of explicit paralinguistic control in a TASTE based model is provided, which establishes TASTE based modeling as a practical route toward full-duplex systems while identifying natural conversation robustness, speech generation latency, and feature general paralinguistic control as open challenges.
A data-centric stabilization recipe that jointly improves coverage and fidelity through three mechanisms: controllable instruction diversification for systematic expansion, LLM-based drift filtering for quality control, and attribute-aligned supervision that grounds prosody control in acoustic perturbations is proposed.
Synchronized burst pulses in dolphin communication reveal complex social dynamics that standard analysis methods overlook.
AV-STE is proposed, a modular streaming audio-visual front-end that restores corrupted semantic speech tokens from noisy audio and lip video before they reach the speech LLM, and remains entirely frozen, preserving its pretrained conversational capabilities.
Conversational Voice, a pipeline that converts real two-speaker excerpts into three complementary training-data artifacts, is presented, a pipeline that converts real two-speaker excerpts into three complementary training-data artifacts.
A novel HOA coding scheme based on the explicit and blind estimation of the Relative Spatial Room Impulse Response (ReSRIR), using a beamformed version of the HOA signal as a reference signal to derive an efficient parametric representation for immersive audio coding is proposed.
Language orthogonalization transforms self-supervised speech models, enabling them to detect Parkinson's disease across languages without being misled by language identity.
Audio language models can grasp the broad gist of stuttered child speech, but their reasoning completely collapses and leaks multi-speaker context as disfluency rates rise.
Physical "Red/Black" air-gapping can no longer scale to dynamic multi-domain environments, forcing critical infrastructure to determine whether software-defined isolation can match the provable guarantees of dedicated hardware.
Expensive bidirectional cross-attention is unnecessary for speech emotion recognition: routing acoustic cues unidirectionally into text representations slashes parameter count by 60% while outperforming heavier state-of-the-art transformers.
Clinical behavioral screening can reach a 0.59 F1 without ever exposing raw video or audio of minors, but moving from coarse screening to granular psychometric item prediction hits an immediate performance wall across current multimodal architectures.