Search papers, labs, and topics across Lattice.
Affiliation:, University of T眉bingen
2
0
1
26
This work proposes OPTED (on-policy fine-tuning for end-to-end driving) which decouples reinforcement learning from the post-training of the end-to-end policy: a privileged teacher is trained using RL on vectorized inputs (HD-map and bounding boxes) which provides supervision to the pre-trained student during closed-loop post-training.
A single diffusion step can match the performance of traditional deterministic models in dynamic environments, revealing new avenues for efficient online planning.