Search papers, labs, and topics across Lattice.
80 papers published across 2 labs.
Unvoiced speech regions boost deepfake detection accuracy by nearly 50%, outperforming traditional full-audio approaches.
A training-free self-embedding strategy can significantly enhance the detection of partial deepfake speech, even when only small segments are manipulated.
Conversational avatars no longer require latency-heavy cascades: generating speech and full-body gesture end-to-end from shared LLM hidden states runs 5.4× faster than two-pass pipelines while matching cascaded teacher fidelity.
Unsupervised speech models can be dramatically improved for accented speech, achieving a 23.6% boost in performance with minimal adaptation.
Critical adjacency from expert critics can significantly enhance music recommendation systems, achieving a notable AUC of 0.767 in cold-start scenarios.
Conversational avatars no longer require latency-heavy cascades: generating speech and full-body gesture end-to-end from shared LLM hidden states runs 5.4× faster than two-pass pipelines while matching cascaded teacher fidelity.
Unsupervised speech models can be dramatically improved for accented speech, achieving a 23.6% boost in performance with minimal adaptation.
Critical adjacency from expert critics can significantly enhance music recommendation systems, achieving a notable AUC of 0.767 in cold-start scenarios.
Transcript-based shortcuts in dialogue models lead to a staggering drop in accuracy, revealing a critical flaw in current evaluation methods.
XTTSv2 can anonymize voices while maintaining high speech quality and intelligibility, achieving near-optimal privacy without any language-specific training.
Multimodal models struggle with consistency, showing significant judgment discrepancies between text and speech inputs, especially in Arabic contexts.
Authentic telemedicine conversations in Bengali reveal that existing models can significantly improve their performance on clinical reasoning tasks with the right data.
Cross-lingual word-to-speech mappings can be effectively learned from visual grounding without the need for transcriptions or extensive model training.
Homophones in English are not just phonetically identical; they reveal significant differences in pronunciation that reflect their meanings.
PFGS can reduce word error rates by up to 19.3% compared to random selection, highlighting the critical role of phoneme frequency in TTS augmentation for ASR.
Emotion recognition in visual intelligence is heavily skewed towards linguistic cues, revealing a critical gap in purely visual affect recognition capabilities.
A striking task-dependent robustness gap reveals that while ASR thrives on direct audio retrieval, AQA falters due to bottlenecks in mediated information access.
Traces of a forgotten language can enhance re-learning speed by 14%, challenging the notion of critical periods as purely maturational.
High-performing cough-based TB classifiers fail to generalize across datasets, revealing that data collection artifacts overshadow disease-related signals.
A runtime monitoring system that achieves 85% accuracy in detecting air traffic control procedural violations could significantly enhance aviation safety protocols.
The Dissonance Spectrum outperforms traditional music representations, achieving the highest mean scores in open-ended music question answering and emotion recognition tasks.
Whisper's adaptation for Baniwa ASR reveals that multilingual models can effectively bridge the gap for low-resource languages, achieving competitive error rates.
Encoder models hold their ground against generative LLMs in ASR evaluation, but the latter enhance interpretability and hypothesis selection.
Leveraging speech act information can dramatically enhance derailment forecasting accuracy, especially in low-data environments.
ACF-Net outperforms existing methods in fine-grained visual categorization by effectively handling the complexities of asymmetric audio-visual inputs.
A training-free self-embedding strategy can significantly enhance the detection of partial deepfake speech, even when only small segments are manipulated.
VoiceMem achieves a 30-point accuracy boost over traditional memory systems while delivering real-time, emotionally aware interactions without added latency.
Existing ASR systems falter in recognizing code-switched speech, with the best performing system still achieving a staggering 35.93% word error rate.
Training voice agents in an audio-native environment can more than double their task success rates and enhance efficiency without external dependencies.
Current generative models struggle with temporal drift and responsiveness in streaming audio-video generation, as revealed by the new StreamAV-Bench benchmark.
Adaptive contrastive decoding boosts audio-visual speech recognition performance, striking the right balance between noise resilience and prediction fidelity.
Sound object identification can eliminate howling in conferencing setups by intelligently managing audio playback, challenging conventional echo cancellation methods.
CSAVocoder achieves real-time spatial audio generation with enhanced fidelity by effectively integrating dynamic spatial cues, outperforming traditional vocoders.
LALMs face significant challenges in audio comprehension, especially in extracting relevant facts from lengthy signals, with performance deteriorating as audio duration increases.
Achieving superior acoustic echo control with a model that operates at just 2% of the computational cost of its more complex counterpart is a game-changer for real-time applications.
AI-generated sounds can be distinguished from real ones with high accuracy by analyzing decay-region group delay, revealing a critical forensic cue in audio forensics.
Embedding a neural codec representation enables full recovery of manipulated audio segments, a breakthrough for content integrity in speech recordings.
GAN-based NDDF not only outperforms traditional methods in reverberant conditions but also simplifies directivity pattern estimation, revolutionizing spatial sound capture.
Enhanced diagnostic feedback for L2 Mandarin learners reduces mispronunciation errors by over 23%, transforming how pronunciation is taught and assessed.
Models trained on LAION-BVD achieve state-of-the-art performance in multimodal tasks, showcasing the dataset's potential to redefine video understanding.
LLMs can effectively assess voice-agent interactions, but their reliability hinges on specific metrics and evaluation setups, revealing a nuanced landscape for automated judgment.
Transcript-based detection outperforms direct audio processing in identifying spoken hallucinations, revealing a critical gap in current methodologies.
The nuanced interplay between pitch and duration in Debussy's music reveals that broader plateaus in pitch may mislead segmenting efforts, complicating traditional music analysis methods.
Arbitrary shapes can now serve as dynamic waveform generators, transforming how we synthesize sound from geometric forms.
Synchronizing multiple musical outputs on a single immutable timeline could revolutionize how composers manage heterogeneous audio formats.
Achieving state-of-the-art performance in speech decoding with a dataset that features 80 hours of deep, within-subject MEG data sets a new standard for neural data quality and quantity.
DAPF models excel in dementia detection but struggle to provide trustworthy explanations, revealing a gap between performance and interpretability.
TurnBench reveals that human-like turn-taking in dialogue remains elusive for AI, with systems struggling to balance accuracy and false positives in interruption detection.
Scammers engage older targets with 15% more conversational turns, yet their requests for sensitive information remain unchanged, revealing a surprising consistency in their tactics.
Task-disentangled LoRA enables seamless integration of audio-visual tasks, outperforming both unified and task-specific models in multi-modal learning.
AudioLens-R1 redefines audio clustering by allowing models to adaptively organize speech based on user-specified perspectives, achieving unprecedented accuracy improvements.
FireRedAudio achieves state-of-the-art performance in audio understanding and generation by leveraging decoupled continuous representations, setting a new standard for unified audio-language models.
Forgetting is drastically reduced in audio classification tasks, with SPECTRA outperforming traditional methods by leveraging subspace structure for feature replay.
Fine-grained spatio-temporal reasoning in audio-language models can dramatically enhance the understanding of complex multi-event audio sequences.
Watermarking can expose hidden vulnerabilities in audio deepfake detection systems, leading to significant performance drops that vary by dataset.
Transforming long-form audio meeting comprehension, the GRGA model leverages graph-based planning to overcome acoustic loss and memory challenges, achieving superior QA performance.
Automating clinical note generation from speech could drastically reduce healthcare workers' documentation time while preserving essential patient information.
The Qwen3-Omni model can reconstruct complex narratives from garbled audio inputs, revealing a hidden layer of reasoning that challenges our understanding of audio language processing.
REDnet can effectively separate audio from an unknown number of speakers using variable microphone setups, achieving unprecedented accuracy in challenging scenarios.
Improved spatial audio generation in real-world speech scenes achieves unprecedented fidelity and accuracy using a visually guided framework.
Unvoiced speech regions boost deepfake detection accuracy by nearly 50%, outperforming traditional full-audio approaches.
ADEPS can encode spatial audio with zero-shot flexibility across any microphone array, outperforming traditional methods in fidelity and quality.
Achieving expressive TTS now hinges on effectively optimizing non-verbal vocalizations, with design choices impacting NV fidelity more than previously understood.
Achieving robust multi-species birdsong classification on low-power hardware, PolyChirp can monitor up to 10 species simultaneously while conserving energy for an entire breeding season.
Despite strong downstream performance, SLMs struggle with instruction-following due to weak alignment between speech and text representations, revealing a critical gap in their training methodology.
DF-MoE achieves superior deepfake detection performance by harnessing a diverse set of multimodal features, setting a new standard in the field.
WnW reduces GPU memory usage to 20% of audio tokens without sacrificing accuracy, challenging the limitations of existing KV cache methods in long-form speech processing.
Leveraging MLLMs without any training, this framework achieves competitive audio-guided video segmentation, showcasing the power of foundation models in practical applications.
Post-training with LoRA can boost accuracy in some models while hindering others, revealing the nuanced interplay between architecture and adaptation in audio-dependent tasks.
Amplitude modifiers can ensure Lipschitz continuity in neural networks, enabling robust audio signal recovery with guaranteed convergence.
DiaScriber achieves unprecedented accuracy in multi-speaker scenarios, overcoming the challenges of overlapping speech and rapid transitions.
Achieving a 40% reduction in character error rate, this syllable-level UASR framework unlocks new potential for low-resource language recognition without costly phoneme resources.
ASR errors can dramatically worsen the performance of advanced retrieval-augmented generation systems, with error amplification reaching up to 67%.
Achieving synchronized audio-visual generation and trajectory control, EchoWM sets a new standard for immersive media experiences in generative models.
Multilingual and multimodal audio understanding is critically under-evaluated, with EXAM$^2$ revealing up to 21.7% performance gaps in current models.
Real-world audio-visual speech enhancement struggles, with baseline models achieving only -4.069 dB SI-SDR in challenging overlapping scenarios.
Winning systems in the AT-ADD challenge achieved over 90% accuracy in detecting all types of audio deepfakes, showcasing the potential of innovative detection strategies.
EmoTra-TTS achieves up to 87% improvement in emotional transition quality, redefining how TTS systems can convey nuanced human emotions in speech.
Achieving top performance in long-horizon audio-visual generation, JoyAI-Echo-1.5 sets a new standard for coherent storytelling and interactive environments.
Speech2MaskTrack reveals how structured speech constraints can dramatically improve object tracking accuracy in video segmentation tasks, even in the absence of acoustic cues.
Watermarking TTS outputs without retraining or quality loss is now feasible, thanks to a novel method that exploits noise correlations.
Achieving over 15% improvement in speech decoding accuracy across subjects reveals a breakthrough in the generalizability of brain-computer interfaces.
Current LALMs can identify basic audio content but often fail to diagnose and reason about degradation, revealing a significant blind spot in their capabilities.
Articulatory feature supervision can dramatically enhance phoneme recognition in non-canonical speech, revealing interpretable error patterns that align with phonological structures.
Pruning-based correction can cut speaker leakage in multi-talker ASR by up to 29%, transforming transcript reliability in noisy settings.