Search papers, labs, and topics across Lattice.
3
0
6
1
Mage-Flow achieves high-resolution image generation and editing in under a second on a single GPU, challenging the notion that larger models are always necessary for quality.
Transforming multimodal resources into executable agent skills boosts performance by nearly 12 percentage points, showcasing the power of diverse learning materials.
Unified multimodal models often *hurt* performance on multimodal understanding tasks, except for spatial reasoning, visual illusions, and multi-round reasoning, challenging the assumption that generation universally improves understanding.