Search papers, labs, and topics across Lattice.
84 papers published across 7 labs.
British and American audio descriptions differ significantly in language and narrative style, impacting how visually impaired viewers experience films.
CASA not only sets a new benchmark in automatic speaking assessment but also clarifies how acoustic and content features interact, paving the way for more interpretable AI assessments.
HybridSB-MoE achieves superior speech enhancement by intelligently fusing spectral and waveform models, outperforming traditional methods in both efficiency and quality.
Achieving phoneme conversion in under 0.15 ms could revolutionize real-time Thai text-to-speech applications.
Layer selection for speech-based PD detection is more about the dataset than the model architecture, revealing a critical flaw in current approaches.
CASA not only sets a new benchmark in automatic speaking assessment but also clarifies how acoustic and content features interact, paving the way for more interpretable AI assessments.
HybridSB-MoE achieves superior speech enhancement by intelligently fusing spectral and waveform models, outperforming traditional methods in both efficiency and quality.
Achieving phoneme conversion in under 0.15 ms could revolutionize real-time Thai text-to-speech applications.
Layer selection for speech-based PD detection is more about the dataset than the model architecture, revealing a critical flaw in current approaches.
Misalignment in speculative decoding can lead to a staggering 21-frame error in ASR, but innovative tracking methods can significantly boost efficiency and accuracy.
VoxAudio revolutionizes vocalized audio synthesis by embedding intelligible speech seamlessly within complex soundscapes, outperforming traditional methods that compromise on clarity and control.
Prioritizing lexically rich narratives can accelerate ASR transcription quality improvements by overcoming the cold start problem in language documentation.
A single-layer speech enhancement model outperforms naive architectures and achieves competitive quality with a significant speedup through progressive knowledge distillation.
Leveraging dialogue as a conditioning signal significantly enhances video-to-music generation, outperforming existing models with a novel dataset that ensures reproducibility.
Transcript-free voice cloning across 14 languages achieves unprecedented fidelity and intelligibility, setting a new standard for zero-shot TTS systems.
Luna-TTS achieves unprecedented low error rates in real-time speech generation, outperforming leading commercial systems while enabling advanced features like zero-shot voice cloning and emotional control.
Real-time, context-aware music generation can transform in-vehicle experiences by adapting soundtracks to evolving driving conditions.
Heavy speaker overlap drastically hinders speech recognition accuracy in smart glasses, revealing critical limitations in current audio-language models.
Text-to-music models may appear controllable, but a closer look reveals that much of their output is simply a reflection of their training data rather than genuine instruction-following.
Natural-language critiques can transform how we evaluate and optimize song generation models, leading to more human-aligned outputs.
Achieving a staggering reduction in Word Error Rate from 12.15% to 2.79%, MiDashengLM-Gen sets a new standard for text-to-audio generation.
Dialect recognition in ASR can be significantly improved without sacrificing Mandarin accuracy, thanks to a novel self-distillation approach.
Achieving speech intelligibility that consistently surpasses ground-truth recordings, Phoenix TTS redefines the boundaries of zero-shot TTS and voice conversion.
CookVoice achieves high-quality, style-controllable voice generation with a fraction of the parameters, outperforming larger models in efficiency and flexibility.
Out-of-domain generalization in speech encoders hinges on their proximity to unseen TTS embeddings rather than their distance from natural speech.
Continuous non-autoregressive models outperform discrete methods in speech enhancement, revealing a critical shift in paradigm effectiveness.
Deep learning models can dramatically improve the accuracy of Relative Transfer Matrix estimation, outperforming traditional methods in noisy environments.
Family-specific audio tagging can now achieve unprecedented accuracy in noisy, naturalistic environments, bridging the gap in infant-centered audio understanding.
Achieving seamless audio-visual identity swapping in talking videos while preserving original dynamics could revolutionize content creation and personalization.
ASR-roundtrip evaluation can miss nearly half of the critical reading errors in Chinese news TTS, revealing a significant gap in current assessment methods.
DINO-A reveals that smaller patch sizes in Vision Transformers consistently yield better audio representation quality, challenging assumptions about model architecture in audio tasks.
Achieving a WER of 23.44% for Burmese medical ASR, this work sets a new benchmark that challenges the capabilities of larger models.
Trust in AI interviewers hinges on perceived stakes and the subtleties of conversational grounding, not just their technical capabilities.
A lightweight model for music AVQA achieves 96% accuracy by leveraging frozen audio encoders, challenging the notion that larger models are always better.
A unified framework that not only generates but also interprets piano performances reveals critical insights across all skill levels, challenging traditional assessment methods.
Spontaneous vocal guidance in drone teleoperation reveals a three-phase structure that could revolutionize how we design adaptive control systems for voice-operated robots.
Korean pop music from the 1960s to 1980s is perceived as lagging behind US trends by four to five years, but this gap narrows significantly in the 1990s and beyond.
Unsupervised learning of pitch-contour tokens reveals hidden structures in Korean traditional music, aligning with expert categories and enhancing analysis.
Gesture generation systems struggle to match the expressive quality of human motion, with top submissions lagging far behind motion-capture benchmarks.
Whispered speech recognition just got a major upgrade, slashing hallucination rates from over 25% to just 4.5% with a new self-supervised uncertainty learning approach.
Learning audio effects without dry references not only boosts performance but also aligns better with real-world recording conditions.
E2E speech models are vulnerable to a stealthy DoS attack that can drastically increase their output length and resource consumption without altering the original input.
Jointly summarizing and translating long-form spoken content could revolutionize how we handle multilingual information processing.
MazzikaAI transforms real-time Arabic maqam accompaniment by seamlessly integrating expert musical knowledge with generative AI, achieving unprecedented responsiveness and microtonal fidelity.
Real-time turn-taking detection can be achieved with unprecedented accuracy and low latency using a novel dual-head modeling approach.
DonorRank reveals that the right selection of donor languages can drastically enhance zero-shot ASR performance in low-resource contexts, challenging existing heuristics.
Fragile gains in low-resource ASR performance are exposed, with standard methods outperforming complex alternatives in Garhwali speech recognition.
Ex-Omni-2D generates visually coherent dialogue responses that seamlessly integrate text, speech, and video, all while avoiding the need for extensive multi-modal training data.
Even the best voice agents struggle with real-world conversational tasks, scoring below 50% in effective assistance across diverse scenarios.
A lightweight model that predicts children's age alongside phonemes can outperform larger models, revolutionizing phoneme recognition in children's speech.
Current TTS evaluators miss the mark, with MOS predictors focusing solely on sound quality and Audio-LLMs struggling to generalize across speech dimensions.
British and American audio descriptions differ significantly in language and narrative style, impacting how visually impaired viewers experience films.
Unifying sign language translation and production reveals a novel approach that significantly improves motion accuracy while retaining competitive translation performance.
Achieving high-fidelity audio synthesis on FPGAs could redefine standards in professional audio applications.
This low-cost, customizable handpan interface not only democratizes music creation but also enhances learning with integrated visual aids.
Denoising bioacoustic signals using ridge-guided training synthesis can dramatically enhance the clarity of vocalizations, leading to better classification outcomes in noisy environments.
PhonoQ-derived features boost phonological classification accuracy, revealing intricate speech patterns that traditional models miss.
Fine-grained audio captioning just got a major upgrade—AudioMap achieves state-of-the-art results by redefining how we reward temporal accuracy and descriptive richness in audio events.
EER may mislead researchers about the effectiveness of voice anonymization, while privacy-ZEBRA offers a more accurate lens on information leakage.
BiTSE achieves unprecedented improvements in speech extraction fidelity by effectively leveraging spatial and temporal cues in noisy environments.
Diarization-derived features can rival complex speech embeddings in predicting language proficiency, making automated assessments more accessible and efficient.
AMIE (Video) outperformed human physicians in critical clinical tasks, signaling a leap toward AI that can effectively engage in complex medical consultations.
Environmental audio manipulation is easier to detect than synthetic speech, but existing detectors fail to handle both effectively, revealing critical gaps in current audio forensics.
State-of-the-art audio description systems struggle to match expert human performance, revealing significant gaps in current methodologies.
ST-Omni-R1 not only excels in sound-event recognition but also sets a new standard for spatial audio reasoning, achieving nearly double the accuracy of existing models.
Retrieval-guided initialization in RAG-Audio nearly matches the performance of direct retrieval in brain-to-audio tasks, revolutionizing how we reconstruct audio from neural signals.
Leading models fall short in emotional intelligence, with EmoS achieving 83.8% accuracy and nearing human-level performance.
Single-corpus evaluations can obscure up to 80 points of macro-F1 variability, revealing the hidden pitfalls of current infant cry analysis methods.
A2I-Set enables AudioCanvas to outperform existing methods, achieving unprecedented visual expressiveness and alignment in audio-to-image generation.
AIxSpeed lets learners consume audio content 30% faster while enhancing comprehension and satisfaction, reshaping how we approach learning efficiency.
Explicitly planning musical structure with MusicLayout allows for unprecedented control and inspection in text-to-music generation, transforming how we create and refine compositions.
Inaudible low-frequency signals can cripple LALM performance by up to 67%, revealing a hidden vulnerability in audio processing systems.
Unified audio generation just got a major upgrade—SonicWeave's innovative routing mechanism boosts compositional quality and expert specialization, outperforming traditional models.
DAVE achieves superior speech separation performance even in the presence of unreliable visual inputs, thanks to its innovative decoupled framework and extensive training dataset.
Differentiable SoundFont proxies can revolutionize MIDI velocity estimation, achieving superior performance across instruments by focusing on acoustic dynamics rather than raw waveforms.
A training-free dynamic clustering method significantly improves long speech separation performance, especially in sparse scenarios with unknown speaker counts.
Generalizing DoA estimation across diverse microphone arrays could revolutionize audio processing in mobile and dynamic environments.
Achieving a top rank in the MeViS-Audio challenge, this system showcases the power of agreement-based segmentation in accurately identifying objects described by spoken expressions.
A-PACK reveals that deferring audio pruning can lead to a 78% reduction in prefill costs while boosting performance in omni-modal LLMs.
Gestures alone can predict referential intent, revealing their critical role in multimodal dialogue even when speech is ambiguous.
Emotion-sensitive neurons in LALMs are language-specific, but pooling cross-lingual evidence reveals powerful, transferable Multilingual Emotion Neurons that enhance affective control.
Achieving high-fidelity speech synthesis with a 40.8% reduction in real-time processing time could revolutionize interactive voice applications.
High speech overlap isn't the primary challenge in cocktail-party scenarios; innovative audio-visual strategies and large language models can cut recognition errors by 57%.
BAMU achieves a remarkable MOS improvement in speech quality by dynamically allocating quantization resources based on frame complexity, outperforming traditional fixed-depth codecs.
Traditional acoustic localization methods can fail dramatically in complex environments, but this new approach cuts worst-case errors while keeping median accuracy intact.
Achieving over 90% performance retention with a staggering 20x KV cache compression could redefine efficiency in long-context audio inference.
The interplay between representation and modeling strategies can redefine how we optimize audio generative systems for context and uncertainty.
FullDiT not only outperforms leading commercial music generators but also redefines how we approach music rendering by leveraging full-context generation techniques.
LSAR transforms audio retrieval by mapping audio directly into a sparse lexical space, enabling efficient and interpretable retrieval without the pitfalls of transcription.