Search papers, labs, and topics across Lattice.
3
0
4
Human trajectory logs are no longer the performance ceiling for autonomous driving: closed-loop reinforcement learning paired with distilled foundation models outperforms human demonstration baselines across major open and closed-loop benchmarks.
Combining RLVR and OPD through SAF not only prevents entropy collapse but also boosts performance across multiple benchmarks, revealing the potential for more effective reinforcement learning strategies.