NVIDIAClarifaiK-frameMar 12, 2026arXiv:2603.12254

Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing

Baifeng Shi, Stephanie Fu, Long Lian, Hanrong Ye, David Eigen, D. Eigen, Aaron Reite, Aaron A. Reite, Boyi Li, Jan Kautz, Song Han, David M. Chan, Pavlo Molchanov, Trevor Darrell, Hongxu Yin

AI Summary

AutoGaze, a novel module, is introduced to address the computational bottleneck of processing long, high-resolution videos in MLLMs by selectively removing redundant patches before they are processed by ViTs or MLLMs. Trained with next-token prediction and reinforcement learning, AutoGaze autoregressively identifies and retains only the most informative multi-scale patches necessary for video reconstruction within a defined error threshold. This approach achieves a 4x-100x reduction in visual tokens and up to 19x speedup, enabling MLLMs to process 1K-frame 4K videos and achieve state-of-the-art results on video benchmarks, including a 4.5% improvement over prior MLLMs on the newly introduced HLVid benchmark.

Key Contribution

MLLMs can now handle 4K videos up to 100x faster thanks to AutoGaze, which selectively attends to only the most informative patches.

Abstract

Multi-modal large language models (MLLMs) have advanced general-purpose video understanding but struggle with long, high-resolution videos -- they process every pixel equally in their vision transformers (ViTs) or LLMs despite significant spatiotemporal redundancy. We introduce AutoGaze, a lightweight module that removes redundant patches before processed by a ViT or an MLLM. Trained with next-token prediction and reinforcement learning, AutoGaze autoregressively selects a minimal set of multi-scale patches that can reconstruct the video within a user-specified error threshold, eliminating redundancy while preserving information. Empirically, AutoGaze reduces visual tokens by 4x-100x and accelerates ViTs and MLLMs by up to 19x, enabling scaling MLLMs to 1K-frame 4K-resolution videos and achieving superior results on video benchmarks (e.g., 67.0% on VideoMME). Furthermore, we introduce HLVid: the first high-resolution, long-form video QA benchmark with 5-minute 4K-resolution videos, where an MLLM scaled with AutoGaze improves over the baseline by 10.1% and outperforms the previous best MLLM by 4.5%. Project page: https://autogaze.github.io/.

Architecture Design (Transformers, SSMs, MoE)Computer Vision Multimodal Models

Citation Metrics

Citations0

Influential citations0

References110

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing

Related Papers