Search papers, labs, and topics across Lattice.
4
0
6
6
NAPE achieves state-of-the-art performance in audio representation learning by simplifying the pre-training process to a single autoregressive prediction task.
Noise doesn't stand a chance against VIB-AVSR, which boosts LLM-based audio-visual speech recognition performance by integrating Variational Information Bottleneck layers for enhanced robustness.
MambAdapter achieves superior performance in audio and speech tasks while drastically cutting down on computational resources.
Despite the intuition that noisy environments should make models rely more on visual cues, AVSR models stubbornly cling to audio, even when it's heavily degraded.