Search papers, labs, and topics across Lattice.
Beihang University
7
0
11
WDL-OPD boosts MATH500 accuracy from 0.630 to 0.685, showcasing a powerful new approach to stabilizing on-policy distillation.
SPOT redefines on-policy distillation by ensuring that probing decisions directly enhance downstream reasoning performance, not just teacher alignment.
Trajectory anchoring bias in VLA models can be mitigated by transforming future trajectory decisions into verifiable selections, leading to more reliable reasoning in autonomous driving.
Continuous and consistent robotic actions can be achieved without additional network parameters, revolutionizing how robots interpret and execute complex tasks.
Overcome the prohibitive cost of ground-truth labels in reinforcement learning by actively acquiring labels for only the most valuable samples, leading to stable training and improved performance even with limited annotation budgets.
RLHF can be made more stable and effective by explicitly verifying and reinforcing policy improvements against a historical baseline, rather than relying solely on instantaneous reward signals.
Forget noisy, biased LLM evaluators: CDRRM distills preference insights into compact rubrics, letting a frozen judge model leapfrog fully fine-tuned baselines with just 3k training samples.