Search papers, labs, and topics across Lattice.
Affiliation:
6
0
7
10
This work introduces a comprehensive evaluation framework organized around four complementary dimensions of physical world reasoning, and constructs a diverse set of novel tasks that require models to integrate complementary information across modalities.
This work proposes MobileVLA-R1 2.0, an RL-enhanced VLA framework that explicitly couples structured embodied reasoning with executable mobile robot control, and introduces a reasoning-conditioned action decoder that maps multimodal reasoning representations to task-level action targets, which are subsequently translated into embodiment-specific commands by robot controllers.
A novel hierarchical framework enables autonomous vehicles to anticipate long-term traffic dynamics while making real-time decisions, achieving superior driving performance.
Streaming video MLLMs perform far better when historical context is actively internalized into evolving latent tokens rather than queried as passive external visual buffers.
Bridging the action-sufficiency gap in robotic manipulation, GIFT achieves up to 12.6 points improvement over existing models by integrating structured intermediate features.
A principled framework for General World Models reveals the limitations of current systems and the architectural requirements for future progress.