Search papers, labs, and topics across Lattice.
7
0
11
9
Achieving state-of-the-art performance in document parsing, NaviDC-OCR tackles geometric distortions and structural reasoning challenges that plague existing models.
Retrieval reasoning that learns from failures can dramatically boost the accuracy of multimodal retrieval systems.
Skip the costly full training runs: this new metric accurately predicts face recognition dataset quality using only lightweight proxy models.
Uniformly quantizing the entire diffusion action head of VLAs to W4A4 is not only possible, but can match or exceed FP16 performance, defying conventional wisdom and slashing memory footprint by 71%.
LLaVA-OV-2's codec-stream tokenization lets it crush existing video-language models, especially in tasks requiring fine-grained temporal understanding of high-frequency motion.
Forget generic retrieval signals – UniDoc-RL uses reinforcement learning to teach LVLMs how to actively perceive and reason about visual information, yielding a 17.7% performance boost.
Turns out, skipping the boring parts of a video (like static backgrounds) makes your vision AI both faster and smarter, beating state-of-the-art models with less data.