Search papers, labs, and topics across Lattice.
7
0
9
4
AV-Flamingo outperforms existing models on complex audio-visual tasks, revealing that size isn't everything when it comes to reasoning capabilities.
LALMs can boost their temporal reasoning accuracy by 3.2% simply by better redistributing attention across audio tokens rather than relying on textual cues.
Video-derived geometric supervision enables a 33% reduction in collisions and a 150% increase in navigation success for VLAs, reshaping the landscape of obstacle-aware navigation.
Audio-language models can now reason about 30-minute-long audio clips with timestamp-grounded intermediate steps, unlocking a new level of fine-grained understanding.
Steering vectors work primarily by nudging the output value (OV) circuit in attention, not by re-weighting attention scores, and can be drastically sparsified without losing effectiveness.
AVLLMs may "hear" at intermediate layers, but they largely ignore audio cues in favor of vision when generating text, revealing a fundamental modality bias.
LALMs can be easily tricked into "hearing" things that aren't there, with success rates as high as 95% on targeted attacks.