Search papers, labs, and topics across Lattice.
21
3
14
11
EduPanel achieves human-level reliability in evaluating teaching videos while enhancing scoring accuracy and preserving expert oversight.
Instruction-level feedback from audio-aware LLMs can drastically enhance the accuracy of multi-event audio generation, bridging a critical gap in current models.
ORCA not only boosts performance by 26.4 points but also restores critical speaker identity cues that traditional models overlook.
Taiwanese Mandarin TTS systems can achieve a 63.9% reduction in word error rates by using a context-adapted tokenizer and language model.
SR-FD reduces word error rates by over a third, transforming the landscape of intelligibility in few-step TTS synthesis.
Timestamp drift in ASR can be corrected with minimal parameter updates, achieving near-perfect alignment without sacrificing model performance.
Leveraging lecture context can boost technical term recognition in Mandarin ASR systems by over 15% without sacrificing overall accuracy.
LatentASR transforms frozen ASR models into adaptive systems that intelligently allocate compute resources, achieving significant WER reductions on challenging inputs without the need for extensive retraining.
CAAD achieves an 8% performance boost in speech language models while slashing inference latency and linguistic bias.
Targeting only the gaps in information, GDP-RAG achieves unprecedented accuracy in multi-hop question answering while slashing computational costs.
Bridging the gap between verbal and non-verbal vocalizations, this approach slashes speaker verification errors by over 40% while preserving speech accuracy.
Instruction-based steering can redirect attention in LALMs to acoustically relevant regions, achieving over 60% overlap with ground-truth sound event locations without any training.
Widely used emotion embedding similarity metrics for speech generation are more sensitive to speaker and linguistic features than actual emotion, rendering them unreliable for evaluating emotional expressiveness.
Semantic-level uncertainty estimation methods significantly enhance the reliability of audio-aware language models, outperforming traditional approaches in critical reasoning tasks.
Bridging the gap between audio reconstruction and language modeling objectives yields neural audio codecs that are both more acoustically faithful and linguistically predictable.
Speech-to-speech translation can now convey laughter and tears with human-like fidelity, thanks to a surprisingly data-efficient approach leveraging LoRA experts.
Systematic biases in LALMs can be triggered by subtle cues like gender and accent, revealing a complex landscape of fairness that traditional benchmarks miss.
Real-world speech disfluencies trip up even the most advanced full-duplex voice agents, exposing critical gaps in self-correction and multi-step reasoning abilities.
High-frequency details, often discarded, are actually crucial for spotting singing voice deepfakes, enabling significantly better detection.
Audio watermarks can now survive neural resynthesis, thanks to a latent space embedding technique that resists semantic compression by modern audio codecs.
Overcome LALM's struggles with localized dialectal prosody: a new Taiwanese audio-text dataset and fine-tuning strategy boosts accuracy by 6.5% on the TAU Benchmark.