Search papers, labs, and topics across Lattice.
4
0
6
General-purpose self-supervised audio representations can outperform specialized supervised models, reshaping the landscape of audio understanding in ALLMs.
Video-LLMs fail to effectively learn and apply skills from long video memories, revealing a fundamental gap in their capabilities.
Reinforcement learning enables video-LLMs to re-watch and refine answers without the costly overhead of chain-of-thought training, achieving better performance with less computation.
UBD reveals that traditional accuracy metrics can mask critical sample-level discrepancies in model behavior, leading to a more nuanced understanding of data contamination effects.