Search papers, labs, and topics across Lattice.
2
0
6
6
An optimized open-source pipeline-parallel runtime is built that co-designs scheduling and parallelism through a JCT-aware scheduling layer and pipeline-integrated multi-token prediction.
Visual token dominance is the hidden culprit behind LVLM inference inefficiency, and this paper dissects the problem to reveal how to navigate the fidelity-efficiency tradeoff.