Search papers, labs, and topics across Lattice.
95 papers published across 3 labs.
Standard word error rates artificially inflate speech BCI performance by ignoring unmodeled language, masking a critical capability trade-off that an open-vocabulary information-theoretic metric resolves while improving decoding accuracy by up to 16.3%.
Interrupted voice agents often hallucinate what they have already said because text generation outpaces audio playback—a failure mode solved by forcing the model to listen to its own real-time voice output.
Text-AB achieves a breakthrough in voice dubbing and dialogue synthesis by eliminating alignment constraints, resulting in unprecedented quality and efficiency.
Closed-architecture audio processors may be powerful, but they obscure the rich narratives and complexities of their internal processes, risking a loss of artistic and technical insight.
The resulting 82M-parameter model, Wayu-Paxa-TTS-Edge, enables on-device Thai TTS without reference audio and achieves the lowest pause-placement error and intra-word pause rates among the three systems.
Interrupted voice agents often hallucinate what they have already said because text generation outpaces audio playback—a failure mode solved by forcing the model to listen to its own real-time voice output.
Text-AB achieves a breakthrough in voice dubbing and dialogue synthesis by eliminating alignment constraints, resulting in unprecedented quality and efficiency.
Closed-architecture audio processors may be powerful, but they obscure the rich narratives and complexities of their internal processes, risking a loss of artistic and technical insight.
The resulting 82M-parameter model, Wayu-Paxa-TTS-Edge, enables on-device Thai TTS without reference audio and achieves the lowest pause-placement error and intra-word pause rates among the three systems.
Localized chord corruption can produce over 2.88 times the output change compared to traditional ACR replay, challenging assumptions about chord recognition in music generation systems.
Achieving superior RIR reconstruction and speech quality at just 375 bps challenges the effectiveness of conventional audio codecs in immersive audio applications.
Incorporating severity diversity in ASR training can slash error rates by over 60%, leveling the playing field for individuals with cleft lip and palate.
Integrating semantic understanding with acoustic fidelity can dramatically elevate the accuracy of speech quality assessments, challenging the dominance of self-supervised learning models.
StreamWSR achieves high-quality speech super-resolution with zero-look-ahead streaming, all in a compact model under 10 million parameters.
Continuous NAC representations can significantly boost speech quality and intelligibility in enhancement tasks, outperforming traditional discrete methods.
Test-time adaptation can dramatically enhance speech quality under mismatched acoustic conditions without requiring labeled data.
ToolDF achieves a remarkable 14.39-point boost in macro-F1 scores over traditional methods while providing interpretable evidence for mixed-authenticity audio deepfake detection.
Real Human-Human dialogue data can transform turn-taking in dialogue systems, achieving better proficiency without sacrificing semantic quality.
Neurophysiological data leakage poses a greater risk in the Musical Metaverse than previously recognized, necessitating a rethink of security protocols in real-time environments.
Users' perceived quality varies significantly with playback speed, and CAPQ-FAST reveals how to optimize this experience across different media types.
Achieving 2.8x compression in on-device speech tokenizers without sacrificing accuracy could revolutionize mobile AI applications.
PACodec slashes bitrate by 30% while preserving decoding quality, thanks to its innovative parallel quantization approach.
Traditional single-channel separation methods fall short of optimal performance, but this new geometric approach reveals a pathway to significantly reduce estimation error by addressing phase discrepancies.
StrixAE achieves unprecedented audio enhancement performance by integrating a multimodal language model with a tailored reinforcement learning approach, setting a new benchmark in real-world audio restoration.
Dual time-frequency spectral representations can dramatically enhance the quality of degraded music recordings, outperforming traditional methods in both objective metrics and listener satisfaction.
Achieving accurate direction estimation of acoustic intensity from 50 Hz to 20 kHz using tightly arranged cardioid microphone arrays could revolutionize sound localization in noisy environments.
VocalCap achieves unprecedented traceability in audio capture, ensuring that every recording is backed by rigorous evidence of its integrity and quality.
Systems that excel in dialect identification often leverage unique acoustic features, while ASR performance hinges on data normalization strategies.
Practical non-invasive BCIs cannot require days of user calibration, and this benchmark forces models to tackle word-level MEG decoding under strict 10-to-40-minute subject adaptation budgets.
Joint audio-video diffusion models hallucinate not from diffuse attention noise, but because bidirectional audio-video cross-attention systematically overrides text conditioning in favor of learned canonical priors.
Real-time speaker-attributed ASR is now feasible with VibeVoice-ASR-Streaming, achieving unmatched accuracy while processing speech on-the-fly.
Pseudo-triplet construction enables TTS systems to generate nuanced voice modifications that respond to performance directions while maintaining speaker identity.
Loneliness in older adults can be detected through a powerful combination of speech patterns and vocal characteristics, revealing critical insights into their emotional states.
Real-world speech recordings with overlapping noise events are now timestamped and categorized, filling a crucial gap in audio datasets.
OV capture in noisy environments can be significantly improved by leveraging the unique low-pass characteristics of bone-conducted vibrations in earbuds.
A 7B model can boost source-grounded speech planning accuracy from 68.4% to 91.9% by addressing citation legality before synthesis.
Achieving 0% detectable speech while preserving 85% activity recognition accuracy could revolutionize privacy in acoustic monitoring for elderly care.
Human evaluations show SonicCaps captions are not only more descriptive but also significantly enhance audio retrieval performance compared to existing datasets.
Attention-projection adapters can dramatically improve dysarthric ASR performance, but simpler models like LoRA often outperform more complex variants in practical applications.
Two-stage mixing systems can dramatically enhance audio quality, but their effectiveness hinges on proper task decomposition and model selection.
CRAW achieves state-of-the-art resilience against neural audio transformations while keeping audio quality intact, a game-changer for combating synthetic audio fraud.
Joint audio-video models often stay in sync with each other while drifting entirely from the script; explicitly routing text guidance onto the shared temporal axis cuts shot boundary error by 96% down to 42 milliseconds.
Standard word error rates artificially inflate speech BCI performance by ignoring unmodeled language, masking a critical capability trade-off that an open-vocabulary information-theoretic metric resolves while improving decoding accuracy by up to 16.3%.
Missing-note reconstruction can be tackled as a constrained MAP problem, enabling precise melodic completions even with significant data loss.
Reducing cpWER in multi-talker speech recognition by over 1% demonstrates that soft speaker posteriors can effectively enhance model performance without the pitfalls of hard segmentation.
Enhanced speech clarity in noisy environments could revolutionize human-robot interactions in public spaces.
Despite advances in large audio language models, the best performer in temporal audio grounding only achieves a mere 31.2 mIoU, revealing critical limitations in current capabilities.
TUTTI achieves unprecedented A2S performance by leveraging a synthetic dataset, outperforming traditional methods and enabling robust cross-instrument adaptability.
MADS achieves superior audio classification performance while halving the dimensionality compared to conventional spectral summaries, redefining efficiency in audio representation.
Imperceptible fingerprints prove to be far more reliable for speech deepfake attribution, even as perceptible ones fluctuate with emotional content and model changes.
ABSE-NET achieves superior speech enhancement in open-fit hearing aids without requiring in-ear microphones, tackling acoustic leakage head-on.
Sound extraction accuracy improves dramatically when leveraging the hierarchical relationships of an ontology, allowing for flexible querying across various sound categories.
AVERT achieves a new state-of-the-art in spoken dialogue state tracking by leveraging audio verification to correct persistent ASR errors, outperforming conventional text-based methods.
Achieving a 20x increase in transaction success rates for financial speech recognition in Nepali with just 300 training examples reveals the power of domain adaptation in low-resource settings.
Fine-tuned TTS models expose severe privacy risks, with membership inference attacks achieving speaker-level AUCs nearing perfect accuracy.
Achieving precise timing control in audio-visual generation without model retraining could revolutionize how we synchronize speech and visuals in AI applications.
Human judgments on musical similarity reveal surprising discrepancies with computational metrics, challenging the reliability of current AI models in creative domains.
Audio language models encode speaking style effectively but lose critical paralinguistic information before making predictions, revealing a significant gap in their capabilities.
Uncertainty-aware methods can significantly enhance speaker verification performance by reducing EER and improving separation through a unified approach that propagates uncertainty throughout the verification pipeline.
A two-stage detection and classification system boosts killer whale monitoring accuracy and speed, crucial for the conservation of endangered populations.
A single-tower speech codec can outperform complex dual-tower architectures, achieving low-bitrate efficiency without sacrificing semantic quality.
Code-switched phrases can now be spoken with their native accents without any training, thanks to a novel localized guidance approach that identifies phrase boundaries in real-time.
Achieving speaker anonymization with provable privacy guarantees while maintaining high utility could redefine standards in secure voice communication.
U-PAST achieves state-of-the-art speech enhancement performance with a fraction of the parameters, outperforming larger models in challenging acoustic conditions.
Achieving synchronized audio-video generation at 2K resolution with a compact 7B model could revolutionize content creation and accessibility in multimedia applications.
ASR systems show systematic performance disparities linked to the linguistic distance of speakers' first languages, revealing hidden biases in their design.
Despite advanced models, cry-reason classification struggles due to label limitations, not capacity, revealing critical insights into dataset quality.
Reducing audio tokens by 75% with stride-k subsampling preserves performance while slashing computational costs—no retraining required.
Context-Aware Interleaved Batching cuts Word Error Rate while preserving context, revolutionizing how we handle speech transcription in real-time applications.
EviBound not only boosts diagnostic accuracy but also ensures that mental health screenings are grounded in the appropriate evidentiary context, eliminating unsupported claims.
Lightweight classifiers can outperform larger models in handling voice assistant fallbacks, transforming user interactions from failures to opportunities.
Phonological structure leaves measurable traces in vocal music, enabling accurate language classification from unaccompanied singing.
Negative sentiment in parliamentary speech not only alters pitch and intensity but also reveals surprising cross-linguistic gender dynamics in filled pauses.
Latent adversarial refinement can enhance vocal-accompaniment separation quality while slashing sampling costs by leveraging a compact latent space.
A novel framework that combines audio-grounded verification with a rich dataset leads to significant advancements in long-paragraph audio captioning.
Transcribing guitar audio to tablature just got smarter—Noise2Fret integrates musical constraints to boost accuracy and efficiency.
Bridging the gap between audio embeddings and LLMs, this method achieves a remarkable 16.2% boost in deepfake voice detection accuracy on unseen domains.
Sound sources in the Arctic can be misidentified due to a phenomenon that renders them nearly undetectable at certain depths, complicating acoustic monitoring efforts.
Encoding expert music mixing conventions into language models can significantly elevate the quality of automatic music upmixing.
Achieving robust ASR with a front-end that reduces word error rates while using fewer than 1 million parameters and minimal computational overhead is a game changer for speech processing efficiency.
VIBE achieves superior instruction adherence and controllability in music generation from video, setting a new standard for semantic alignment in multimodal tasks.
LCAR slashes hallucination failures by over 50% in LLM-based ASR systems without any additional training or external models.
Clean audio can be weaponized as a backdoor trigger in speech enhancement models, achieving near-perfect attack success without altering the input.
Shifting the defense from visual to audio, this method effectively thwarts identity impersonation in 3D talking face generation while preserving visual quality.
Cleaner speech datasets may boost in-domain performance but can undermine generalization, revealing a hidden cost in Alzheimer's detection models.
XVAE-WMT achieves superior sound separation without requiring paired clean recordings, setting a new standard for interpretability in biomedical signal processing.
SE can enhance audio quality but may drastically mislead LLMs, with intent classification errors more than doubling in some cases.
Achieving 99.77% accuracy in neuromorphic speech recognition while slashing energy costs could redefine efficiency standards in AI processing.
MusGU+ reveals that existing generative music frameworks fall short in helping musicians effectively evaluate and adopt AI tools for their creative processes.
Subjective rewards in speech generation are not interchangeable, revealing critical nuances in how RL aligns with human listener preferences.
Combining contrastive learning with JEPA yields a model that enhances music representation without increasing complexity, outperforming traditional methods in key tasks.
Unigram entropy emerges as a key predictor of codec performance, revealing stark differences in how various architectures handle noise conditions.
Multi-emotion TTS can now achieve unprecedented accuracy and expressiveness, outperforming leading models in both trajectory and blending tasks.
Spectrally adaptive loss functions can drastically improve high-frequency speech clarity in streaming applications, addressing a critical gap in current enhancement techniques.
Weakly supervised TST can achieve substantial accuracy improvements by leveraging rhythmic regularities, even without precise onset annotations.
PhysWave achieves spatially consistent audio generation by combining physics-based priors with a unified control framework, revolutionizing text-to-FOA applications.
Text-conditioned generative music models outperform their audio-conditioned counterparts in faithfully following emotional cues, revealing critical insights into genre-specific emotion-following dynamics.
TEMPO not only achieves superior timestamping accuracy in audio-language models but also integrates innovative techniques like atomic timestamp tokens and reinforcement learning for enhanced performance.
Current audio-language models struggle with temporal music grounding, but targeted training can lead to significant performance improvements.