Search papers, labs, and topics across Lattice.
13
0
10
7
Intern-S2-Preview-397B not only excels in multimodal scientific reasoning but also enhances biological instruction performance without altering its foundational architecture.
Current LLMs only achieve 27.3% accuracy in reasoning about scientific lineage, revealing a critical gap in their compositional capabilities.
Large-scale structured academic visual data can transform image generation from mere aesthetic appeal to verifiable knowledge-grounded creation.
Spatial intelligence in MLLMs can be dramatically enhanced without any architectural modifications or retraining, thanks to a novel collaborative cognitive mapping approach.
SEGA3D achieves an impressive 8.3 mIoU improvement over previous methods, redefining the standards for 3D vision-language segmentation.
Current video MLLMs struggle to grasp fleeting visual events, with top models barely surpassing 39% accuracy on critical momentary tasks.
LLM-powered agents can now produce surprisingly strong photographs in complex 3D environments, suggesting a path towards embodied AI with aesthetic awareness.
Visual degradations can cripple the spatial reasoning abilities of even state-of-the-art MLLMs, but targeted finetuning can restore—and even surpass—human-level performance.
LVLMs can achieve SOTA visual reasoning by learning to "see" in a way that optimizes for reasoning, even if it means deviating from strict geometric accuracy.
Small, open-source LLMs can now outperform larger, closed-source models in complex industrial design tasks by learning to orchestrate CAD/CAE tools within a reinforcement learning framework.
Current image editing models stumble when domain-specific knowledge is required, as revealed by a new benchmark spanning disciplines from natural science to social science.
A 4B-parameter model, InternVL-U, outperforms 14B-parameter models in multimodal generation and editing, proving that size isn't everything.
Sports expose surprising limitations in VLMs' spatial reasoning, as current models struggle to generalize from existing benchmarks despite fine-tuning gains on a new, large-scale dataset.