Search papers, labs, and topics across Lattice.
5
0
8
10
Current video generation models struggle with visual reasoning, achieving only 51% accuracy on a new benchmark designed to probe their capabilities.
StreamOPD achieves near teacher-level performance in streaming video understanding without relying on memory or retrieval, reshaping the landscape of post-training techniques.
RL fine-tuning LMMs for tool use can collapse structural formats due to strong pretrained tool priors, but a surprisingly simple fix of targeted format rewards and frame-budget randomization can restore stability and boost performance.
Stop letting SFT ruin your LMMs: PRISM uses on-policy distillation to realign your model *before* RL, boosting performance by up to 6%.
Today's visual generation models are often evaluated on the wrong things, leading to inflated performance claims that mask critical failures in spatial reasoning, temporal consistency, and causal understanding.