Search papers, labs, and topics across Lattice.
This paper investigates the limitations of driving world models as counterfactual simulators by highlighting a fundamental mismatch between direct action-conditioned predictions and factual driving outcomes. The authors formalize this gap using a causal framework and demonstrate that direct predictions fail to accurately represent counterfactual scenarios, as they do not account for the factual continuation of observed episodes. By introducing a simple, training-free pipeline that incorporates observed evidence into counterfactual views, they show significant improvements in prediction accuracy across two world models.
Direct predictions in driving world models can lead to substantial inaccuracies in counterfactual scenarios, revealing a critical gap in current methodologies.
Driving world models are often interpreted as counterfactual simulators for observed driving episodes: given a factual driving log, they are asked what would have happened under an alternative ego action. In this paper, we identify a fundamental mismatch between this goal and direct action-conditioned prediction. The direct prediction uses the shared history and the alternative action but not the factual continuation observed after that history. It can therefore generate a plausible future without preserving what actually happened in this episode. We formalize this gap using the causal recipe of abduction, action, and prediction and study it in a setting with a short time horizon, where the alternative ego action does not alter how surrounding agents evolve. To make the gap measurable, we construct a controlled simulation benchmark with factual outcomes and matched counterfactual outcomes. Across two representative world models, direct predictions fail to match the counterfactual ground truth, supporting our analysis. As a constructive check of this analysis, we introduce a deliberately simple, training-free pipeline that moves observed evidence into the counterfactual view and lets the frozen model complete what remains unknown. Even this simple construction raises the overall recovered fraction substantially and reduces perceptual distance to the matched counterfactual on both models. We hope this work draws attention to this gap and motivates better counterfactual prediction methods for driving world models.