Search papers, labs, and topics across Lattice.
This paper introduces EcoFrame, a novel framework for adaptive visual evidence scheduling that enhances long-video understanding by optimizing frame selection based on the inference feedback from vision-language models (VLMs). Unlike traditional static methods, EcoFrame employs entropy-gated budget scheduling to dynamically adjust the frame budget based on output uncertainty, while utilizing attention-guided candidate proposals to focus on informative regions. Experimental results show that EcoFrame significantly improves the accuracy-efficiency trade-off, achieving a 64.4% accuracy on Qwen2.5-VL with a 1.85x speedup compared to existing methods like BOLT and AKS.
EcoFrame achieves a remarkable 13.5x inference speedup while maintaining accuracy, revolutionizing how we approach long-video understanding in VLMs.
Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of frames selected as sparse visual evidence. Existing relevance-based methods rely on static one-shot selection with fixed frame budgets and candidate pools, while agent-based schedulers achieve adaptivity through costly multi-round reasoning and interactive search. We propose EcoFrame, a training-free framework for low-overhead query-adaptive visual evidence scheduling. EcoFrame leverages the VLM's inference feedback to determine when to increase the frame budget and where to search for additional candidate evidence. Specifically, entropy-gated budget scheduling uses output uncertainty to stop early when the current evidence is sufficient or progressively expand the frame budget otherwise. Meanwhile, attention-guided candidate proposal converts frame-level attention into a temporal prior, enabling dense local search in informative regions while preserving global coverage when attention is diffuse. Experiments on Video-MME, LongVideoBench, and MLVU demonstrate that EcoFrame achieves a better accuracy--efficiency trade-off across multiple VLM backbones. On Qwen2.5-VL, EcoFrame achieves an average accuracy of 64.4, surpassing BOLT at 63.5, while providing a $1.85\times$ speedup over AKS and BOLT. Compared with the agent-based A.I.R., EcoFrame maintains comparable accuracy with up to a $13.5\times$ inference speedup. Code will be available at https://github.com/AK-DREAM/EcoFrame.