Search papers, labs, and topics across Lattice.
Affiliation:
1
0
3
0
Halving frame resolution costs virtually nothing in long-video MLLMs, but reinvesting those saved visual tokens to double temporal frame counts yields an immediate 2–3 point accuracy gain—all while a 30-year-old sparse approximation algorithm matches state-of-the-art custom selectors.