Search papers, labs, and topics across Lattice.
62 papers published across 2 labs.
Unvoiced speech regions boost deepfake detection accuracy by nearly 50%, outperforming traditional full-audio approaches.
Self-embedding steganography can transform clean speech into a resilient defense against partial deepfake manipulation without requiring any training.
VoiceMem's dual-brain architecture boosts conversational accuracy and emotional engagement, setting a new standard for real-time interaction in speech language models.
Traces of a forgotten language can enhance re-learning speed by 14%, challenging the notion of critical periods as purely maturational.
High-performing cough-based TB models fail to generalize across datasets, revealing that data collection artifacts overshadow disease-related signals.
VoiceMem's dual-brain architecture boosts conversational accuracy and emotional engagement, setting a new standard for real-time interaction in speech language models.
Traces of a forgotten language can enhance re-learning speed by 14%, challenging the notion of critical periods as purely maturational.
High-performing cough-based TB models fail to generalize across datasets, revealing that data collection artifacts overshadow disease-related signals.
A runtime monitoring framework for air traffic control procedures achieves an impressive F1 score of 0.85, highlighting critical procedural violations that could prevent future accidents.
The Dissonance Spectrum reveals intricate frequency interactions that conventional models overlook, leading to superior performance in music understanding tasks.
Whisper's adaptation for Baniwa ASR reveals that multilingual models can effectively bridge the gap for low-resource languages, achieving competitive error rates.
Encoder models hold their ground against generative LLMs in ASR evaluation, but the latter enhance interpretability and hypothesis selection.
Leveraging speech act information can dramatically enhance derailment forecasting accuracy, especially in low-data environments.
Howling in conferencing systems can be dramatically reduced by identifying sound objects and muting channels, rather than relying on traditional echo cancellation methods.
ACF-Net outperforms existing methods in fine-grained visual categorization by effectively handling the complexities of asymmetric audio-visual inputs.
Embedding a neural codec representation allows for the recovery of manipulated audio segments, a breakthrough for content integrity in speech recordings.
Self-embedding steganography can transform clean speech into a resilient defense against partial deepfake manipulation without requiring any training.
By decomposing phonemes into low-level phonological components, this framework offers L2 Mandarin learners unprecedented clarity in pronunciation feedback.
CSAVocoder achieves real-time spatial audio generation with improved fidelity by effectively leveraging dynamic pose information and inter-channel cues.
Achieving 98% computational savings while improving AEC performance reveals the untapped potential of knowledge distillation in audio processing.
Models trained on LAION-BVD achieve state-of-the-art performance in multimodal tasks, showcasing the dataset's potential to redefine video understanding.
LLMs can effectively assess voice-agent interactions, but their reliability hinges on specific metrics and evaluation setups, revealing a nuanced landscape for automated judgment.
Transcript-based detection outperforms direct audio processing in identifying spoken hallucinations, revealing a critical gap in current methodologies.
The nuanced interplay between pitch and duration in Debussy's music reveals that broader plateaus in pitch may mislead segmenting efforts, complicating traditional music analysis methods.
Arbitrary shapes can now serve as dynamic waveform generators, transforming how we synthesize sound from geometric forms.
Synchronizing multiple musical outputs on a single immutable timeline could revolutionize how composers manage heterogeneous audio formats.
Achieving state-of-the-art performance in speech decoding with a dataset that offers unprecedented depth and breadth in MEG recordings could revolutionize non-invasive brain-computer interfaces.
DAPF models excel in dementia detection but struggle to provide trustworthy explanations, revealing a gap between performance and interpretability.
TurnBench reveals that human-like turn-taking in dialogue remains elusive for AI, with systems struggling to balance accuracy and false positives in interruption detection.
Scammers engage older targets with 15% more conversational turns, yet their requests for sensitive information remain unchanged, revealing a surprising consistency in their tactics.
Task-disentangled LoRA enables seamless integration of audio-visual tasks, outperforming both unified and task-specific models in multi-modal learning.
AudioLens-R1 redefines audio clustering by allowing models to adaptively organize speech based on user-specified perspectives, achieving unprecedented accuracy improvements.
FireRedAudio achieves state-of-the-art performance in audio understanding and generation by leveraging decoupled continuous representations, setting a new standard for unified audio-language models.
Forgetting is drastically reduced in audio classification tasks, with SPECTRA outperforming traditional methods by leveraging subspace structure for feature replay.
Fine-grained spatio-temporal reasoning in audio-language models can dramatically enhance the understanding of complex multi-event audio sequences.
Watermarking can expose hidden vulnerabilities in audio deepfake detection systems, leading to significant performance drops that vary by dataset.
Transforming long-form audio meeting comprehension, the GRGA model leverages graph-based planning to overcome acoustic loss and memory challenges, achieving superior QA performance.
Automating clinical note generation from speech could drastically reduce healthcare workers' documentation time while preserving essential patient information.
The Qwen3-Omni model can reconstruct complex narratives from garbled audio inputs, revealing a hidden layer of reasoning that challenges our understanding of audio language processing.
REDnet can effectively separate audio from an unknown number of speakers using variable microphone setups, achieving unprecedented accuracy in challenging scenarios.
Improved spatial audio generation in real-world speech scenes achieves unprecedented fidelity and accuracy using a visually guided framework.
Unvoiced speech regions boost deepfake detection accuracy by nearly 50%, outperforming traditional full-audio approaches.
ADEPS can encode spatial audio with zero-shot flexibility across any microphone array, outperforming traditional methods in fidelity and quality.
Achieving expressive TTS now hinges on effectively optimizing non-verbal vocalizations, with design choices impacting NV fidelity more than previously understood.
Achieving robust multi-species birdsong classification on low-power hardware, PolyChirp can monitor up to 10 species simultaneously while conserving energy for an entire breeding season.
Despite strong downstream performance, SLMs struggle with instruction-following due to weak alignment between speech and text representations, revealing a critical gap in their training methodology.
DF-MoE achieves superior deepfake detection performance by harnessing a diverse set of multimodal features, setting a new standard in the field.
WnW reduces GPU memory usage to 20% of audio tokens without sacrificing accuracy, challenging the limitations of existing KV cache methods in long-form speech processing.
JoyAI-Echo-1.5 not only excels in generating coherent long-form narratives but also sets a new benchmark for interactive world modeling with its innovative memory and geometric control techniques.
Leveraging MLLMs without any training, this framework achieves competitive audio-guided video segmentation, showcasing the power of foundation models in practical applications.
Post-training with LoRA can boost accuracy in some models while hindering others, revealing the nuanced interplay between architecture and adaptation in audio-dependent tasks.
Amplitude modifiers can ensure Lipschitz continuity in neural networks, enabling robust audio signal recovery with guaranteed convergence.
DiaScriber achieves unprecedented accuracy in multi-speaker scenarios, overcoming the challenges of overlapping speech and rapid transitions.
Achieving a 40% reduction in character error rate, this syllable-level UASR framework unlocks new potential for low-resource language recognition without costly phoneme resources.
ASR errors can dramatically worsen the performance of advanced retrieval-augmented generation systems, with error amplification reaching up to 67%.
Achieving synchronized audio-visual generation and trajectory control, EchoWM sets a new standard for immersive media experiences in generative models.
Multilingual and multimodal audio understanding is critically under-evaluated, with EXAM$^2$ revealing up to 21.7% performance gaps in current models.
Real-world audio-visual speech enhancement struggles, with baseline models achieving only -4.069 dB SI-SDR in challenging overlapping scenarios.
Winning systems in the AT-ADD challenge achieved over 90% accuracy in detecting all types of audio deepfakes, showcasing the potential of innovative detection strategies.
EmoTra-TTS achieves up to 87% improvement in emotional transition quality, redefining how TTS systems can convey nuanced human emotions in speech.
Speech2MaskTrack reveals how structured speech constraints can dramatically improve object tracking accuracy in video segmentation tasks, even in the absence of acoustic cues.
Watermarking TTS outputs without retraining or quality loss is now feasible, thanks to a novel method that exploits noise correlations.
Achieving over 15% improvement in speech decoding accuracy across subjects reveals a breakthrough in the generalizability of brain-computer interfaces.
Current LALMs can identify basic audio content but often fail to diagnose and reason about degradation, revealing a significant blind spot in their capabilities.
Articulatory feature supervision can dramatically enhance phoneme recognition in non-canonical speech, revealing interpretable error patterns that align with phonological structures.
Pruning-based correction can cut speaker leakage in multi-talker ASR by up to 29%, transforming transcript reliability in noisy settings.
TLive-Omni achieves superior live-commerce understanding by seamlessly integrating multi-modal inputs and optimizing for real-time response quality.