Search papers, labs, and topics across Lattice.
5
0
7
6
Despite advances in MLLMs, they still struggle with dynamic reasoning, falling far short of human capabilities in interpreting continuous visual cues.
TopoGPT achieves a remarkable leap in lane topology reasoning, producing geometrically consistent lane graphs that outperform existing methods by substantial margins.
Pruning 90% of visual tokens without sacrificing performance could revolutionize the efficiency of 3D scene understanding in multimodal models.
Latent reasoning can beat explicit Chain-of-Thought – but only if you force it to learn causal dynamics via a visual world model, not just language.
Training a single point cloud encoder across diverse 3D domains not only improves perception but also unlocks emergent behaviors and enhances robotic manipulation and spatial reasoning.