Search papers, labs, and topics across Lattice.
15
2
11
4
InstanceControl achieves superior image generation fidelity and control without the burden of instance labeling, revolutionizing how we handle complex scenes.
ILLUME-X achieves unprecedented quality in free-form interleaved text-image generation, setting a new benchmark for multimodal models.
Rebinding visual cache positions can boost multimodal reasoning accuracy by 5% while slashing computational costs dramatically.
LabVLA achieves unprecedented success rates in executing complex laboratory protocols, outperforming all existing models in both familiar and novel settings.
Single-view RGB input can revolutionize how robots perceive and manipulate transparent objects, achieving reliable grasping without complex depth sensing.
MLLMs can achieve 10% gains on multimodal reasoning benchmarks by using ground-truth anchored data curation and scaffold-stripping to avoid cognitive drift during self-evolution.
End-to-end joint optimization of planning and execution in an image restoration agent unlocks significantly improved performance compared to independently trained tools and all-in-one models.
Forget expensive human feedback loops: a VLM-powered reward function can efficiently align image editing diffusion models with human preferences.
Coordinating embodied multi-agent systems doesn't require end-to-end training; instead, offload planning to a VLM in simulation and transfer back to the real world for execution.
Current image editing models, even closed-source ones, still fall short on complex and creative instruction-based tasks, as revealed by a new interpretable QA-based evaluation framework.
Foundation models can be tamed to reconstruct realistic 4D interactions between hands and articulated objects from a single RGB video, even without pre-scanning or multi-view data.
RL's inherent resilience to catastrophic forgetting can be harnessed to improve continual learning in GUI agents, outperforming SFT alone.
Achieve spatially precise image edits in complex scenes by explicitly reasoning about object positions in text *before* visual grounding.
Achieve more realistic and coherent 4D scene representations by modeling motion within the SE(3) Lie group, outperforming NeRF-based methods.
Imagine AI scientists that not only reason but also autonomously conduct experiments in the real world – that's the promise of Intelligent Science Laboratories.