Search papers, labs, and topics across Lattice.
3
0
6
0
Achieving similar performance to larger models with significantly less data and faster inference speeds could redefine efficiency benchmarks in foundation models.
AVOC achieves a remarkable 4.9-point accuracy boost over the next best model in long-form audio-video comprehension, redefining efficiency in multimodal understanding.
By jointly training a keyframe sampler with an MLLM, MSJoE achieves state-of-the-art accuracy in long-form video understanding while significantly reducing computational cost.