Search papers, labs, and topics across Lattice.
87 papers published across 6 labs.
A novel temperature-aware ANC method achieves superior noise reduction by dynamically adapting to real-time changes in engine noise and lining temperature.
Intonation and tone in Mandarin are not just separate features; they dynamically interact to shape meaning and speaker intent in profound ways.
Always Alert can detect distress signals with minimal disruption, achieving high accuracy while users go about their daily lives.
A wearable respiratory sensor can accurately distinguish stress from relaxation with 88% accuracy, making it a game-changer for non-invasive stress monitoring.
KVAE tokenizers achieve superior generative performance, setting a new benchmark for multimodal models in audio, image, and video synthesis.
A wearable respiratory sensor can accurately distinguish stress from relaxation with 88% accuracy, making it a game-changer for non-invasive stress monitoring.
KVAE tokenizers achieve superior generative performance, setting a new benchmark for multimodal models in audio, image, and video synthesis.
LILAC ensures that re-encoding audio streams returns the original tokens, solving a critical flaw in existing neural audio codecs.
Concept recoverability in AI grading systems varies significantly by architecture, revealing hidden biases that could undermine assessment fairness.
ECHO achieves a remarkable 94.9% tool-execution pass rate while maintaining patient data privacy, setting a new standard for locally-deployable health assistants.
Context biasing methods can slash biased word error rates by up to 88%, outperforming speech LLMs in challenging ASR scenarios.
AI music systems are not just homogenizing sounds; they are reshaping the very landscape of musical value and recognition.
Noise-aware residual correction boosts the realism of autoregressive audio-visual generation, tackling issues of identity drift and desynchronization head-on.
Achieving real-time audio-video generation at 27.12 FPS, Vorch-Streamer tackles the dual challenges of exposure bias and causal speech generation in long-form content.
Replacing standard residual connections with manifold-constrained hyper-connections boosts speaker recognition performance across multiple architectures.
A novel unified framework for electric guitar tone transfer and removal achieves superior performance by disentangling content and tone representations, reshaping how we approach guitar tone modeling.
Sequential blending of audio stems could revolutionize automatic music mixing, outperforming traditional methods by leveraging human-like processing strategies.
Current speech deepfake detection systems falter dramatically against emotionally expressive attacks, with performance dropping to near-random levels on the new AffectDF benchmark.
A single model can now seamlessly handle over 10 diverse audio-visual tasks without the need for task-specific architectures.
FormBharo shows that rule-based controls can enable smaller, cost-effective models to excel in form completion, even when faced with challenging real-world speech inputs.
A new dataset and model achieve a staggering 4.98% symbol error rate for classical music, setting a high bar for audio-to-score transcription in popular music.
ASR systems are not just failing technically; they perpetuate colonial hierarchies that silence marginalized voices, necessitating a radical rethinking of how we design these technologies.
Selective convergence in self-supervised learning allows models to prioritize salient audio-visual cues, leading to superior performance in multi-source localization without manual annotations.
Erratic beat tracking outputs can be effectively mitigated by a novel masked diffusion approach that models multiple plausible interpretations.
Fine-grained evaluation reveals that current text-to-audio models fail to preserve speech content and control audio attributes effectively.
Breaking the curse of multilinguality, this framework boosts low-resource speech performance while preserving high-resource capabilities, all with minimal data requirements.
MERaLiON-GR outperforms existing models in gender recognition across multiple Southeast Asian languages, showcasing the power of specialized speech models.
EmpaAva outperforms traditional chatbots by delivering real-time, emotionally aware interactions through a photorealistic 3D avatar.
AV-MSF achieves state-of-the-art impact sound rendering with minimal data, revolutionizing how we model acoustic properties of objects.
Models trained on the new DiVers dataset show remarkable resilience to noisy and acoustically diverse inputs, outperforming traditional benchmarks.
SFC redefines semantic understanding in spoken language tasks, achieving superior accuracy and adaptability in open-domain contexts.
Intonation and tone in Mandarin are not just separate features; they dynamically interact to shape meaning and speaker intent in profound ways.
Leveraging temporal differences can dramatically enhance video-to-audio generation quality, outperforming even dedicated multimodal representations.
Lip articulation accuracy improves dramatically when phoneme information is integrated into audio-driven rendering, reducing closure violations and enhancing realism.
Achieving high-quality music mixing while allowing for nuanced stylistic control could redefine how producers approach automated mixing.
Fine-tuning large audio-language models for emotion recognition can be drastically improved by leveraging hyperbolic geometry, leading to better performance on class-imbalanced datasets.
Diff-Symbo achieves unprecedented quality and diversity in text-controlled music generation, outperforming leading models by leveraging a novel latent diffusion framework.
The most effective playback similarity metric, CLEWS, is also the least expensive to implement, revolutionizing evaluation strategies in music transcription.
Chord recovery accuracy jumps from 18% to 54% with minimal supervision, showcasing a new frontier in music co-creation agents.
Standard music transformers lose equivariance as they scale, but the Equivariant Music Transformer recovers this crucial property, enhancing generative performance.
Voice input significantly hampers LLM accuracy due to structural transcription issues, while keyboard input is surprisingly more resilient.
Switching music tokenization can halve the Frechet Music Distance, proving that representation trumps model size in text-to-music generation.
Extracting implicit music styles from audio can dramatically enhance the quality and fidelity of symbolic music generation.
Multi-task supervision in V2N allows for unprecedented accuracy in piano transcription, setting new benchmarks in onset, offset, and velocity prediction.
H2S achieves a remarkable 48.54 mAP in audio-visual instance segmentation, setting a new benchmark in the field.
Calliphony transforms brush strokes into dynamic control signals for generative music, merging visual art with live performance in unprecedented ways.
Reward-based fine-tuning of a discrete diffusion model significantly enhances synthesizer inversion performance, outperforming traditional methods in audio matching tasks.
CLASVS achieves a remarkable 46.2% reduction in editing errors while preserving melody and singer identity, revolutionizing lyric editing in singing voice synthesis.
MeloCodec achieves superior singing voice representation by leveraging melodic priors, enabling precise pitch control without compromising timbre quality.
Achieving a 23.99% reduction in speaker verification errors in noisy classroom conditions could revolutionize AI applications in education.
Achieving over 82% output correctness, this new benchmark and model redefine the standards for audio-visual target speaker extraction by effectively integrating visual cues.
DAIEN-TTS achieves unprecedented control over speech synthesis by disentangling speaker characteristics from environmental factors, resulting in highly natural and contextually relevant audio outputs.
InvFlowFD achieves reference-free music quality assessment, aligning closely with human perception while eliminating the need for background data.
Always Alert can detect distress signals with minimal disruption, achieving high accuracy while users go about their daily lives.
ALPO achieves significant improvements in emotional expressiveness by decoupling text and speech learning signals, revealing the potential for more nuanced spoken dialogue systems.
Transfer learning in avian bioacoustics reveals that weak supervision and negative transfer are critical challenges, with significant implications for biodiversity monitoring accuracy.
Achieving 99% accuracy in recognizing overlapping sounds could revolutionize how robots interact in real-world scenarios, especially in emergency situations.
Despite advancements, AI-generated sound effects still struggle with temporal synchronization in complex scenarios, revealing a critical gap in current methodologies.
Simple vector addition in latent spaces can rival complex generative models for audio restoration, challenging the need for intricate architectures.
Language-Specialized Multi-Teacher On-Policy Distillation outperforms traditional RL methods, revealing a new pathway for enhancing multilingual ASR performance.
MSA-EchoLite achieves near state-of-the-art performance in acoustic echo cancellation with a fraction of the computational cost, redefining efficiency in lightweight AEC systems.
GROW achieves a 22.7% reduction in word error rate while accelerating training by 2.9x, redefining efficiency in TTS reinforcement learning.
Taste-sound correspondences in AI-generated music vary significantly across cultures, with response-style biases obscuring deeper perceptual differences.
SwanTale achieves unprecedented expressiveness in multi-speaker audio generation, outperforming existing models in both instruct and zero-shot tasks.
Emotional speech synthesis proves to be the toughest challenge for TTS systems, with performance varying dramatically across different speech domains.
The Gemini-3.1-Pro-Preview model outperforms traditional classifiers in sound-source identification, but even top models confidently produce incorrect answers 92-100% of the time.
Homebot redefines home automation by enabling personalized, hands-free AI interactions that respect user privacy and context.
LOUDAR adapts to unknown audio distortions on-the-fly, achieving superior restoration results without the need for paired training data.
Disordered speech alters high-level representations in ASR models, revealing that phoneme identity is recoverable only in the upper layers, which has significant implications for model adaptation strategies.
Structural design of the corpus, not its size, dictates what attributes a contrastive audio embedding can effectively encode.
EchoCache achieves a remarkable 2.46x speedup in audio-driven video generation without sacrificing quality or coherence.
Even the best audio-video generators struggle with basic acoustic principles, revealing a critical gap in their performance.
P-MUSE achieves unprecedented flexibility in music synthesis by allowing users to generate and edit music with or without aligned MIDI inputs.
A compact MEMS microphone array can generate directional ultrasonic signals, challenging conventional designs in acoustic transmission.
Disfluencies are not just noise; they carry crucial meaning that, when ignored, significantly degrades translation quality.
Detection systems can achieve perfect accuracy on participant-generated deepfakes, yet the generation side reveals a troubling evasion rate that poses deployment risks.
SpeechAgent-R's ability to adaptively coordinate skills and tools leads to a remarkable 15-point performance boost in complex audio reasoning tasks.
Achieving a median pick-level precision of 0.990, this workflow transforms dense acoustic data into actionable insights for fin whale monitoring.
SAGE reduces switching latency to 2.04 seconds while achieving superior target speaker extraction performance in the presence of neural noise and attention switching.
Touch-responsive paintings can now create immersive soundscapes, transforming solitary art into a collaborative auditory experience.
SGAD achieves a breakthrough in EEG-based auditory attention decoding, enhancing accuracy and stability while slashing response latency.
Sonification is not just a method but a transcategorical practice that maintains its identity across diverse contexts, reshaping how we understand its role in science and art.
UAF's innovative uncertainty weighting mechanism boosts classification accuracy by over 15% compared to traditional methods, revealing the critical role of representation confidence in animal vocalization analysis.
Few-shot prompting with chain-of-thought reasoning outperforms other strategies, achieving the best alignment with teacher assessments in music analysis scoring.
A novel temperature-aware ANC method achieves superior noise reduction by dynamically adapting to real-time changes in engine noise and lining temperature.
Self-evolving rubric rewards can dramatically enhance audio reasoning in models, outperforming traditional methods by adapting to the model's evolving capabilities.
Achieving a response rate of 0.88 in full-duplex interactions, JoyAI-Talker redefines how empathetic voice agents can engage in natural conversations even amidst interruptions.
Latent Softmax achieves up to 17.5% lower phoneme error rates in multilingual ASR by intelligently modeling tonal distinctions without sacrificing cross-lingual sharing.
Generative drum demixing not only enhances transcription accuracy but also produces editable audio stems, revolutionizing how we approach drum source separation.
Emotional TTS can be dramatically enhanced by personalizing expression based on individual and cultural perception, leading to more authentic interactions.
FATE not only retains temporal information but also encodes synchronization in a reusable embedding space, outperforming existing models in both semantic and temporal tasks.
Mapping MEG retrieval weights to cortical sources uncovers which speech features are critical for accurate perception, revealing that narrative context significantly boosts information recoverability.