Search papers, labs, and topics across Lattice.
National Taiwan University
12
3
10
12
AMRD achieves superior performance in speech emotion recognition by leveraging the strengths of multiple teachers while capturing relational structures that traditional methods miss.
EduPanel achieves human-level reliability in evaluating teaching videos while enhancing scoring accuracy and preserving expert oversight.
ORCA not only boosts performance by 26.4 points but also restores critical speaker identity cues that traditional models overlook.
Bridging the gap between verbal and non-verbal vocalizations, this approach slashes speaker verification errors by over 40% while preserving speech accuracy.
Widely used emotion embedding similarity metrics for speech generation are more sensitive to speaker and linguistic features than actual emotion, rendering them unreliable for evaluating emotional expressiveness.
Systematic biases in LALMs can be triggered by subtle cues like gender and accent, revealing a complex landscape of fairness that traditional benchmarks miss.
Speech-to-speech translation can now convey laughter and tears with human-like fidelity, thanks to a surprisingly data-efficient approach leveraging LoRA experts.
Text-only LLMs already contain surprisingly diverse levels of auditory knowledge, and this pre-existing knowledge strongly predicts their performance when adapted for audio-language tasks.
Speech quality assessment is skewed: male listeners consistently give higher scores than female listeners, and standard MOS models learn and perpetuate this bias.
Contrastive Decoding's power-up for audio language models hinges on fixing specific error types, like uncertainty and audio absence, but don't expect it to magically fix flawed reasoning.
Audio watermarks can now survive neural resynthesis, thanks to a latent space embedding technique that resists semantic compression by modern audio codecs.
Overcome LALM's struggles with localized dialectal prosody: a new Taiwanese audio-text dataset and fine-tuning strategy boosts accuracy by 6.5% on the TAU Benchmark.