Search papers, labs, and topics across Lattice.
Affiliation:
4
0
5
4
A purely post-hoc and tuning-agnostic weight rectification framework that achieves Parameter Space Orthogonality, which is the necessary and sufficient condition for preserving historical performance to the first order is introduced.
A reward-compatible bounded mixing mechanism for $\gamma\mathrm{OPD}$ that balances verifiable outcome feedback with the discounted OPD advantage to move beyond purely teacher-dependent optimization is developed.
FastDSAC stabilizes training and boosts performance in robotic locomotion by effectively managing exploration and policy plasticity through innovative action constraints.
Achieve real-time autonomous driving policy generation with a new flow-matching RL algorithm that slashes inference latency without sacrificing performance.