Search papers, labs, and topics across Lattice.
AI Laboratory, Fudan University
4
0
6
HPSD enables TI2V models to internalize high-quality visual cues, resulting in a remarkable boost in text-to-video performance while simultaneously enhancing image-to-video generation.
Intern-S2-Preview-397B not only excels in multimodal scientific reasoning but also enhances biological instruction performance without altering its foundational architecture.
Forget textual rules and coarse embeddings: a multimodal reward model that directly compares rendered visuals unlocks significant gains in vision-to-code RL.
Diffusion models can now reason their way through complex spatial tasks with near-perfect accuracy, thanks to a new framework that unlocks chain-of-thought reasoning within the latent space.