Search papers, labs, and topics across Lattice.
10
0
9
7
Achieving state-of-the-art performance in mobile manipulation hinges on aligning temporal granularity and action space, revealing critical insights into effective world-action modeling.
Current interactive world models fall short, with none passing the rigorous tests of WorldRoamBench designed to assess long-horizon stability across action, vision, physics, and memory.
Generating realistic 3D environments from satellite imagery in under 10 minutes could revolutionize how we visualize and interact with our planet.
End-to-end document transcription is now a viable alternative to brittle pipelines: ABot-OCR achieves state-of-the-art results by directly converting page images to clean Markdown.
Closing the sim-to-real gap in vision-language navigation requires benchmarks grounded in realistic 3D reconstructions, not just generated scenes.
Autonomous agents struggle to retain instructions when burdened with retrieving information from the open web, exposing a critical retrieval-reasoning trade-off.
VLMs often fail at spatial reasoning because they either ignore visual cues or exhibit unstable reasoning, but a novel process-shaping framework can fix this.
Network jitter in cloud-based robot control can be overcome by converting temporal lag into spatial pose offsets, restoring the VLA's original geometric intent without fine-tuning.
Ditch discrete waypoints: VLA models can now generate smooth, physically plausible robot trajectories by directly regressing continuous action functions.
By learning to project actions onto a low-dimensional manifold, ABot-M0 achieves faster and more stable robotic control policies compared to directly predicting actions in the full high-dimensional space.