Search papers, labs, and topics across Lattice.
6
3
9
5
Experiments conducted on the UASpeech and TORGO corpora suggest that Conformer models trained using the DSI tokens outperform the comparable baseline HuBERT discrete/continuous features by statistically significant WER reductions.
Loop Memory Attention enables models to revisit earlier computations, leading to a 2.2% accuracy boost in MathQA tasks compared to fixed loop depths.
Zero-shot adaptation can reduce word error rates by over 0.6% while achieving nearly 10 times faster real-time processing for elderly speech recognition.
Speaker adaptive training combined with confidence-guided pseudo-labeling leads to substantial gains in elderly speech recognition accuracy, outperforming conventional methods.
Achieving up to 27.73% absolute WER reduction without any training or additional data challenges the conventional reliance on fine-tuning for model compression.
Current Full-Duplex Speech Language Models stumble in multi-round conversations, struggling to maintain consistent performance across turns and various evaluation dimensions.