Search papers, labs, and topics across Lattice.
This survey examines the evolution of end-to-end autonomous driving systems from simple camera-to-control models to sophisticated planning-oriented architectures that leverage structured representations and trajectory-level outputs. It emphasizes the importance of learned, supervised, and evaluated intermediate representations for ensuring safe and compliant driving, while categorizing existing methods based on input representation, planning output, supervision signal, and evaluation protocol. The authors highlight the need for consistent benchmarking to interpret architectural advancements and identify critical challenges such as uncertainty-aware planning and reproducible evaluation.
Modern autonomous driving systems hinge on learned representations that prioritize safety and compliance, not just raw performance metrics.
End-to-end autonomous driving has evolved from camera-to-control regression toward planning-oriented systems that use structured representations, trajectory-level outputs, and increasingly realistic evaluation protocols. This survey reviews this transition across behavior cloning, conditional imitation learning, privileged distillation, BEV and vectorized planning, unified perception-prediction-planning architectures, world-model-based planners, and vision-language-action systems. We argue that the key distinction in modern end-to-end driving is not whether intermediate representations are used, but whether they are learned, supervised, and evaluated to support safe, feasible, and route-compliant planning. To organize the literature, we synthesize existing methods along four axes: input representation, planning output, supervision signal, and evaluation protocol. We further examine the benchmark shift from open-loop trajectory matching to closed-loop simulation, non-reactive real-log evaluation, long-tail testing, and human-preference-aware metrics. Our analysis highlights that architectural progress is difficult to interpret without benchmark-consistent evaluation, and that displacement-based open-loop metrics alone provide limited evidence for safe and human-aligned driving. We conclude with open challenges in uncertainty-aware planning, learner-expert mismatch, runtime safety assurance, language-action grounding, world-model validation, and reproducible benchmarking.