Search papers, labs, and topics across Lattice.
3
0
4
3
Current video generation models struggle with visual reasoning, achieving only 51% accuracy on a new benchmark designed to probe their capabilities.
MLLMs excel at reasoning over diagrams but falter in parsing and editing them, revealing a significant gap that needs addressing.
Achieving better video distillation quality isn't just about precision; it's about ensuring broad mode coverage during training.