Search papers, labs, and topics across Lattice.
This paper introduces SafeGen, a goal-conditioned diffusion framework designed to generate safety-critical scenarios for vision-language model-based autonomous driving (VLMAD) systems. By framing scenario generation as a diffusion process guided by catastrophic end-states, the method effectively captures realistic human-vehicle interactions that traditional simulator-based approaches struggle to model. Experimental results show that SafeGen enhances the Judge Overall Score by an average of 24.25% across three VLMADs, indicating significant improvements in understanding and decision-making capabilities in real-world driving contexts.
SafeGen boosts VLMAD performance by over 24% in safety-critical scenario generation, bridging the sim-to-real gap that has long plagued autonomous driving systems.
VLMs are increasingly deployed in AD systems, creating an urgent need for rigorous safety evaluation under rare yet safety-critical scenarios. Among these, interactions with vulnerable road users represent a major source of real-world failures. However, existing safety-critical scenario generation methods predominantly rely on simulator-based pipelines, which suffer from a substantial sim-to-real gap and often fail to capture realistic, diverse, and unforeseen human-vehicle interaction dynamics. We present SafeGen, a goal-conditioned diffusion framework for safety-critical scenario generation in VLMADs. Our key insight is to formulate scenario generation as a goal-conditioned diffusion process, where a predefined catastrophic end-state serves as a strong supervisory signal, guiding the generation of temporally coherent video trajectories that naturally evolve toward safety-critical outcomes. Building on this formulation, we introduce Context Grounded End State Reasoning, which leverages VLMs to analyze benign driving contexts and infer latent vulnerabilities in human-vehicle interactions, producing structured end-state specifications that induce high-risk scenarios. Conditioned on these targets, we further propose End State Conditioned Video Evolution, which grounds semantic threats into physically plausible visual dynamics. Specifically, we instantiate high-risk agents within the scene via depth-aware geometric projection, followed by boundary-conditioned diffusion to generate intermediate frames with consistent motion patterns and temporal coherence. Extensive experiments across 3 VLMADs demonstrate that SafeGen increases the Judge Overall Score, a metric using a VLM judge to evaluate VLMADs'understanding and decision-making, by 24.25% on average compared to SoTA baselines. Furthermore, fine-tuning a VLMAD improves performance in real-world driving scenes by an average of 15.9%.