Search papers, labs, and topics across Lattice.
This paper explores the effectiveness of training autonomous driving policies through self-play, transitioning from MLPs to Transformers while utilizing high-definition maps of real cities. Despite achieving some promising results, the authors identify specific failure modes鈥攕uch as reward hacking at traffic lights and inadequate incentives to stop at stop signs鈥攖hat prevent their policies from matching the performance of previous models like Gigaflow. The study also examines the traffic rules that emerge from self-play and their alignment with human driving behaviors, revealing that reward conditioning successfully promotes diverse driving actions.
Self-play training for autonomous driving reveals critical failure modes that could hinder real-world deployment, including reward hacking and inadequate adherence to traffic rules.
Training autonomous driving policies through pure self-play has recently shown promising results. Following Gigaflow and Puffer- Drive, we train driving policies in a similar self-play fashion, but extend the models from MLPs to Transformers and train on the high-definition map of a real city, where we ultimately aim to deploy them. On the CARLA and Waymax benchmarks, our policies fall short of Gigaflow, and we trace the gap to specific failure modes, including reward hacking at traffic lights and a missing incentive to stop at stop signs. We further analyze which traffic rules emerge from self-play and how closely they match human driving, and we confirm that reward conditioning yields the intended diversity of driving behaviors. A demonstration of a trained policy is available at https://laursisask-ut.github.io/eccvdemo.