Search papers, labs, and topics across Lattice.
3
0
5
5
This work proposes MobileVLA-R1 2.0, an RL-enhanced VLA framework that explicitly couples structured embodied reasoning with executable mobile robot control, and introduces a reasoning-conditioned action decoder that maps multimodal reasoning representations to task-level action targets, which are subsequently translated into embodiment-specific commands by robot controllers.
Transforming pre-training data with reasoning annotations boosts model performance by over 2 percentage points without altering the training objective.
Geometry-consistency awareness in video spatial reasoning can lead to a 12.6-point performance boost over current state-of-the-art models.