Search papers, labs, and topics across Lattice.
6
0
7
3
StreamOPD achieves near teacher-level performance in streaming video understanding without relying on memory or retrieval, reshaping the landscape of post-training techniques.
Mage-VL slashes visual token usage by over 75% while enhancing real-time multimodal performance, outperforming larger models in video understanding.
Mage-Flow achieves high-resolution image generation and editing in under a second on a single GPU, challenging the notion that larger models are always necessary for quality.
LLaVA-OV-2's codec-stream tokenization lets it crush existing video-language models, especially in tasks requiring fine-grained temporal understanding of high-frequency motion.
RL fine-tuning LMMs for tool use can collapse structural formats due to strong pretrained tool priors, but a surprisingly simple fix of targeted format rewards and frame-budget randomization can restore stability and boost performance.
Today's visual generation models are often evaluated on the wrong things, leading to inflated performance claims that mask critical failures in spatial reasoning, temporal consistency, and causal understanding.