Search papers, labs, and topics across Lattice.
1
0
3
2
Explicitly modeling visemes in AV-HuBERT slashes word error rates by over 50% in noisy environments, proving that targeted visual guidance can dramatically improve speech recognition robustness.