Search papers, labs, and topics across Lattice.
This paper introduces HyWorldVLA, a hybrid Vision-Language-Action model that integrates pixel-level supervision with latent representation learning to enhance robustness in autonomous driving scenarios. By employing a two-phase training approach鈥攑re-training with video latent prediction and co-fine-tuning with latent feature prediction鈥攖he model effectively balances the trade-offs between interpretability and noise sensitivity. Experimental results on NAVSIM v1 and v2 benchmarks reveal that HyWorldVLA surpasses existing pixel-based and latent-based models, setting a new standard for evaluating noise robustness in world modeling for autonomous driving.
HyWorldVLA achieves unprecedented noise robustness in autonomous driving by seamlessly integrating pixel-level grounding with latent representation learning.
Vision-Language-Action (VLA) models augmented with world modeling represent a promising paradigm for end-to-end autonomous driving. While pixel-level future prediction enables fine-grained spatiotemporal reasoning, it compromises robustness in noisy driving scenarios. Conversely, latent-based world models alleviate this sensitivity but often incur limited interpretability and representational degradation due to absent pixel-level grounding. To reconcile this trade-off, we propose HyWorldVLA, a hybrid world-VLA framework that unifies pixel-level supervision and latent representation learning. In the pre-training stage, HyWorldVLA predicts video latents encoded by a pre-trained video VAE, while simultaneously reconstructing video frames to provide precise pixel-level grounding. During the subsequent co-fine-tuning phase, the model exclusively predicts latent features, which are fed into an action expert to generate trajectories. Extensive experiments on NAVSIM v1 and v2 benchmarks demonstrate that HyWorldVLA significantly outperforms both pixel-based and latent-based world model baselines. Notably, we present the first comprehensive qualitative and quantitative analysis of world model noise robustness in autonomous driving, establishing a new benchmark for evaluating future architectures.