Search papers, labs, and topics across Lattice.
Peking University, AI Laboratory
3
0
5
TRACE achieves over 50% accuracy in long-video question answering by ensuring every answer is grounded in comprehensive visual evidence, outperforming traditional methods at a fraction of the frame cost.
LLaVA-OV-2's codec-stream tokenization lets it crush existing video-language models, especially in tasks requiring fine-grained temporal understanding of high-frequency motion.
Ditching text-based chain-of-thought unlocks better audio-visual reasoning by interleaving textual steps with a unified latent space that preserves dense sensory information.