Search papers, labs, and topics across Lattice.
67 papers published across 2 labs.
TLive-Omni achieves superior live-commerce understanding by seamlessly integrating multi-modal inputs and optimizing for real-time response quality.
Human listeners misclassify genuine audio as fake 77% of the time, revealing a critical vulnerability in our ability to discern deepfake speech.
Achieving over 99% accuracy in speech deepfake source tracing while ensuring interpretability by design could redefine standards in forensic audio analysis.
Iterative proxy correction boosts sentiment analysis accuracy by refining initial proxies and adapting to incomplete multimodal inputs.
$TCP_\alpha$ guarantees complete separation of confidence scores for correct and incorrect predictions, transforming how we trust model outputs in music information retrieval.
TLive-Omni achieves superior live-commerce understanding by seamlessly integrating multi-modal inputs and optimizing for real-time response quality.
Human listeners misclassify genuine audio as fake 77% of the time, revealing a critical vulnerability in our ability to discern deepfake speech.
Achieving over 99% accuracy in speech deepfake source tracing while ensuring interpretability by design could redefine standards in forensic audio analysis.
Iterative proxy correction boosts sentiment analysis accuracy by refining initial proxies and adapting to incomplete multimodal inputs.
$TCP_\alpha$ guarantees complete separation of confidence scores for correct and incorrect predictions, transforming how we trust model outputs in music information retrieval.
Clinicians can now receive both an early warning and precise lead time for AECOPD exacerbations, thanks to a novel two-stage model that leverages continuous ventilator data.
Enhancing AI clones with listening behaviors can dramatically elevate user perceptions of authenticity and engagement.
Prosody can boost decision-making accuracy in dialogue systems by nearly 25% when it conveys user concerns that words alone cannot express.
Children showed no stress reduction from lower-pitched robot voices, challenging assumptions about voice pitch in child-robot interactions.
Achieving a 19.4% reduction in Mel Distance while using 45% fewer parameters, ear-VAE2 redefines the standards for high-fidelity music reconstruction.
Achieving real-time EEG auditory attention decoding with a power-efficient ASIC could revolutionize hearing assistance for cochlear implant users in noisy settings.
A unified music identification system can achieve robust performance across track and version identification tasks with just 10 seconds of audio input.
Participants transformed their dance experiences by leveraging a sound-based device that creates immersive soundscapes from movement, enhancing spatial awareness and improvisation.
A 14M parameter model that outperforms larger transformers while being 12 times smaller, reshaping the efficiency landscape of audio-visual processing.
NAPE achieves state-of-the-art performance in audio representation learning by simplifying the pre-training process to a single autoregressive prediction task.
High-performing ASR models can produce accurate transcripts even from contradictory audio, revealing a troubling disconnect between benchmark scores and real-world performance.
X2Streaming-TTS achieves true token-level synthesis with a median time to first audio token of just 15.8 ms, outperforming traditional pseudo-streaming models in quality and responsiveness.
Creators prefer steering adaptive lyrics through explicit controls, reshaping their creative process and audience engagement.
Fine-tuning with semi-hard negatives can significantly enhance the accuracy of sound effect retrieval through vocal imitation.
Resynthesizing audio from coarse tokens can be dramatically improved by leveraging the geometric structure of the RVQ layer, leading to better fidelity than traditional methods.
Fine-tuning Whisper models for multilingual medical ASR reveals that the best performance hinges on the adaptation strategy, with surprising shifts in internal representations based on language context.
Fine-tuning outperforms zero-shot inference, but the real game-changer is the use of synthetic data to elevate performance in culturally specific tasks.
Achieving centimeter-level precision in drummer motion synthesis from audio could revolutionize character animation in music-driven applications.
Extracting music-theoretic features can significantly enhance style classification accuracy and interpretability in melody analysis.
Uncertainty in data can be powerfully communicated through sound, with ominous tones evoking anxiety and clear tones promoting calmness.
Whisper's fine-tuning can reduce Mizo ASR error rates to as low as 7.22%, showcasing its potential for low-resource languages.
Semantically enriched speech representations in FireRedTTS3 lead to unprecedented stability and fidelity in voice cloning and editing tasks.
Models that leverage acoustic features can dramatically enhance the detection of nuanced emotional states in speech, outperforming traditional text-based approaches.
Deepfake speech detection may achieve sub-1% error rates in controlled settings, but real-world performance falters dramatically due to unforeseen challenges.
DynaForcing not only recovers dynamic motion in avatars but also improves visual quality, resolving a critical trade-off in real-time streaming applications.
Automated data curation and imbalance-aware training strategies significantly enhance LALMs' performance on culturally diverse folk music, yet deep musical understanding remains elusive.
Pop music culture analysis is critically limited by a lack of multi-modal approaches, risking the validity of findings derived from isolated data sources.
The DX7-GNN achieves unprecedented audio-to-preset retrieval accuracy by mimicking FM signal flow, outperforming traditional methods even with fewer parameters.
Achieving near-perfect anomaly detection with neuromorphic processors could revolutionize low-power monitoring in industrial settings.
A unified DNN-based approach reveals richer acoustic scene descriptions while maintaining competitive performance in Voice Activity Detection.
Achieving over 90% accuracy in real-time target speaker identification could revolutionize hearing aid technology by enabling selective amplification with minimal latency.
Geographic information can unlock the identification of bat species that automated systems typically overlook, increasing classification coverage significantly.
Emotion-sensitive neurons in multimodal models reveal shared mechanisms for recognizing emotions across speech and faces, with implications for enhancing emotion recognition capabilities.
AnyTalk enables seamless 3D speech animations for any character without the tedious rigging or animation data typically required.
MGSI reveals that integrating audio and visual sentiment cues at multiple temporal scales can dramatically enhance sentiment analysis accuracy in LLMs.
Depth integration in audio-visual segmentation leads to over 10% performance gains, revealing a critical yet overlooked modality in multimodal perception.
No existing speech retrieval method can robustly handle the full spectrum of user intents, revealing a critical gap in current technologies.
Targeted structural probes reveal that some neural audio watermarks can be erased with a single attack, while others remain impervious, highlighting a critical divide in watermarking effectiveness.
Noise reduction in hearing aids just got smarter—two new losses preserve spatial cues better than ever before.
GRPO-trained LALMs can now transform unstructured audio into structured media with unprecedented accuracy, achieving a 49-point F1 score leap.
Normalizing reference and hypothesis representations can drastically reduce error rates in multilingual ASR, but challenges remain with unconverted characters and scoring pipeline limitations.
Capturing the nuances of microtonal vocal music, this system reveals how computational methods can preserve complex musical traditions with unprecedented accuracy.
Neural network-driven predictions can enhance active noise control, achieving superior suppression of non-stationary speech signals.
Synthetic HRTFs can achieve localization performance on par with measured HRTFs, challenging assumptions about the necessity of physical measurements in spatial audio applications.
Semantic speech tokens can be refined to enhance intelligibility and consistency across speakers, with significant implications for voice conversion and TTS applications.
Compositional zero-shot singing-and-dancing generation is now possible, allowing for unprecedented flexibility in video creation from audio and visual prompts.
Achieving 94.7% accuracy with a feature extractor that eliminates multipliers could revolutionize low-power keyword spotting on edge devices.
Fault detection in I2S transport signals can be significantly improved by integrating structural and payload information through sonification, rather than relying solely on oversampling techniques.
Leveraging repeated EEG responses across sessions, this method achieves remarkable session-invariance and reduces character error rates in EEG-to-speech decoding.
DiffM2A achieves superior Ambisonic encoding fidelity, maintaining performance across diverse microphone configurations and boundary conditions.
Transforming audio captioning from a passive task into an adaptive, evidence-driven process could redefine how we approach fine-grained audio understanding.
Enhancement systems can significantly alter ASR outcomes, but the best choice varies by task and context, challenging the notion of a one-size-fits-all solution.
Cached LLM probability retrieval can improve ASR performance without the heavy lifting of retraining, outperforming traditional methods in a majority of tested scenarios.
DuplexGen achieves more authentic conversational dynamics by allowing timing to emerge organically rather than relying on rigid, handcrafted rules.
CineDub achieves unprecedented accuracy in multi-speaker dialogue dubbing from uncropped videos, setting a new standard for video dubbing technology.
Null tokens can serve as a powerful diagnostic tool for hallucination in ASR and NMT, revealing a trade-off between suppressing fabrication and maintaining valid outputs.
Iterative refinement in TTS can bridge the gap between low-resource data conditions and high-quality expressive synthesis, achieving near-supervised performance with minimal labeled data.
Realigning the Mimi codec reveals a precise mapping of semantic tokens to phonetic realizations, challenging previous assumptions about their representation.
CLARA's innovative clip-level approach reveals that fine-grained analysis can dramatically enhance hateful video detection accuracy, outperforming traditional methods.
A single linear layer transforms T2AV models into effective voice-cloning systems, achieving record-breaking speaker similarity while slashing inference time by ~30x.
ARENA uncovers vulnerabilities in large audio-language models that traditional text-based red-teaming methods miss, achieving near-perfect safety metrics across multiple systems.
PRISM achieves a remarkable 12.94 percentage point improvement in noisy audio classification without any additional training, redefining adaptation strategies for ATMs.