Search papers, labs, and topics across Lattice.
This paper introduces Infrastructure-centric World Models (I-WM), which leverage a roadside perspective to enhance autonomous driving by integrating long-term behavioral data from fixed sensors with the spatial breadth of vehicle-borne sensors. By proposing a dual-layer architecture and a phased sensor strategy, the authors demonstrate how I-WM can improve generative scene understanding, predictive dynamics, and collaborative communication in vehicle-to-everything (V2X) systems. The key finding is that this complementary approach not only enriches the understanding of traffic patterns but also anticipates safety-critical events, potentially transforming roadside perception in autonomous driving systems.
Infrastructure-centric world models could revolutionize roadside perception by combining long-term data with real-time vehicle insights, enhancing safety and efficiency in autonomous driving.
World models, generative AI systems that simulate how environments evolve, are transforming autonomous driving, yet all existing approaches adopt an ego-vehicle perspective, leaving the infrastructure viewpoint unexplored. We argue that infrastructure-centric world models offer a fundamentally complementary capability: the bird's-eye, multi-sensor, persistent viewpoint that roadside systems uniquely possess. Central to our thesis is a spatio-temporal complementarity: fixed roadside sensors excel at temporal depth, accumulating long-term behavioral distributions including rare safety-critical events, while vehicle-borne sensors excel at spatial breadth, sampling diverse scenes across large road networks. This paper presents a vision for Infrastructure-centric World Models (I-WM) in three phases: (I) generative scene understanding with quality-aware uncertainty propagation, (II) physics-informed predictive dynamics with multi-agent counterfactual reasoning, and (III) collaborative world models for V2X communication via latent space alignment. We propose a dual-layer architecture, annotation-free perception as a multi-modal data engine feeding end-to-end generative world models, with a phased sensor strategy from LiDAR through 4D radar and signal phase data to event cameras. We establish a taxonomy of driving world model paradigms, position I-WM relative to LeCun's JEPA, Li Fei-Fei's spatial intelligence, and VLA architectures, and introduce Infrastructure VLA (I-VLA) as a novel unification of roadside perception, language commands, and traffic control actions. Our vision builds upon existing multi-LiDAR pipelines and identifies open-source foundations for each phase, providing a path toward infrastructure that understands and anticipates traffic.