Search papers, labs, and topics across Lattice.
Renmin University of China
4
0
8
AVOC achieves a remarkable 4.9-point accuracy boost over the next best model in long-form audio-video comprehension, redefining efficiency in multimodal understanding.
LLMs can be sped up by over 2x without sacrificing accuracy, by compressing the input and predicting multiple output tokens at once using a unified framework.
By jointly training a keyframe sampler with an MLLM, MSJoE achieves state-of-the-art accuracy in long-form video understanding while significantly reducing computational cost.
Stop paying for verbose overthinking: BFS-PO slashes LRM output length while simultaneously boosting accuracy.