Search papers, labs, and topics across Lattice.
3
0
5
7
Intern-S2-Preview-397B not only excels in multimodal scientific reasoning but also enhances biological instruction performance without altering its foundational architecture.
Visual Pretraining outperforms text-only methods, revealing that rich visual cues can enhance language model performance in ways previously underestimated.
Current MLLMs still struggle to connect the dots between images and text when they're interleaved, highlighting a critical gap in real-world multimodal understanding.