Search papers, labs, and topics across Lattice.
This paper introduces a capability-driven data infrastructure that integrates capability-specific supervision with a curriculum scheduling approach to enhance generalist image generation. By organizing heterogeneous supervision based on the dependencies among generative capabilities, the authors curate a vast dataset comprising 440 million images and train multimodal diffusion models of 3B and 6B parameters. The results demonstrate significant improvements in visual coverage, rendering versatility, and effective transfer across various generative tasks, underscoring the importance of a holistic data design in advancing image generation technologies.
A novel capability-driven data infrastructure enables multimodal models to achieve unprecedented versatility and transferability in image generation tasks.
Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. A central challenge is not only how to curate each task-specific corpus, but also how to organize heterogeneous supervision according to the dependencies among generative capabilities. We present a \textbf{capability-driven data infrastructure} that couples capability-specific supervision construction with capability-aligned curriculum scheduling. Its three specialized yet interoperable data engines build complementary relational supervision for text-image grounding, inter-image transformation, and image-knowledge association, while caption experts align T2I and editing supervision across tasks and granularities. A multi-stage curriculum jointly evolves task composition, visual-concept distribution, data quality, and image resolution along the dependency order of capability acquisition, with capability-aware evaluation closing the loop through targeted retrieval, expert construction, and gap-aware resampling. At scale, the framework curates a 440M-image T2I corpus, 120M editing pairs, and over 27M image-entity pairs. With this infrastructure, we train multimodal diffusion models at two scales from scratch, with 3B and 6B sizes respectively. We conduct quantitative evaluation on CPI-Bench, along with qualitative evaluations across diverse text-to-image and editing scenarios. Experimental results present broad visual coverage, versatile rendering, and effective transfer across generative capabilities.