Search papers, labs, and topics across Lattice.
Harbin Institute of Technology (Shenzhen)
1
0
3
14
MLLMs do not need full-depth forward passes on every incoming frame: indexing video streams solely within early transformer layers slashes prefill latency by 52x without degrading downstream comprehension.