Search papers, labs, and topics across Lattice.
Affiliation:
9
0
8
A purely post-hoc and tuning-agnostic weight rectification framework that achieves Parameter Space Orthogonality, which is the necessary and sufficient condition for preserving historical performance to the first order is introduced.
A reward-compatible bounded mixing mechanism for $\gamma\mathrm{OPD}$ that balances verifiable outcome feedback with the discounted OPD advantage to move beyond purely teacher-dependent optimization is developed.
NeuralParker achieves superior parking performance by retaining essential route context in complex environments, outperforming traditional planners.
RADAR achieves superior optimization performance by decoupling momentum estimation from update geometry, leading to consistent improvements over traditional adaptive optimizers.
The Cramér-geometric Bellman operator reveals a unique fixed point that could transform how we approach evaluation errors in distributional reinforcement learning.
Post-training on synthesized safety-critical scenarios can dramatically enhance the reliability of autonomous driving systems, reducing failures in rare but critical events.
Achieve real-time autonomous driving policy generation with a new flow-matching RL algorithm that slashes inference latency without sacrificing performance.
MLLMs that ace simple traffic rules still struggle when multiple rules interact, especially when they conflict, revealing a critical gap in their ability to handle real-world driving complexity.
A mere 0.01% of tokens can destabilize LLM reinforcement learning, but masking their gradient updates unlocks significant performance gains.