Search papers, labs, and topics across Lattice.
Affiliation:
1
0
3
4
MLLMs do not need full-depth forward passes on every incoming frame: indexing video streams solely within early transformer layers slashes prefill latency by 52x without degrading downstream comprehension.