Search papers, labs, and topics across Lattice.
3
0
4
8
NAPE achieves state-of-the-art performance in audio representation learning by simplifying the pre-training process to a single autoregressive prediction task.
Noise doesn't stand a chance against VIB-AVSR, which boosts LLM-based audio-visual speech recognition performance by integrating Variational Information Bottleneck layers for enhanced robustness.
Despite the intuition that noisy environments should make models rely more on visual cues, AVSR models stubbornly cling to audio, even when it's heavily degraded.