Search papers, labs, and topics across Lattice.
Affiliation:
5
0
7
6
A purely post-hoc and tuning-agnostic weight rectification framework that achieves Parameter Space Orthogonality, which is the necessary and sufficient condition for preserving historical performance to the first order is introduced.
A reward-compatible bounded mixing mechanism for $\gamma\mathrm{OPD}$ that balances verifiable outcome feedback with the discounted OPD advantage to move beyond purely teacher-dependent optimization is developed.
Unleashing pretrained video diffusion models for autonomous driving yields surprisingly strong planning performance, suggesting a new path beyond vision-language models.
End-to-end driving models are surprisingly bad at using navigation, but a new framework shows how to inject it for SOTA results.
A mere 0.01% of tokens can destabilize LLM reinforcement learning, but masking their gradient updates unlocks significant performance gains.