Search papers, labs, and topics across Lattice.
Beihang University, Zhongguancun Laboratory
5
0
10
Reranking at the sentence level can dramatically enhance the quality of synthesized scientific responses by revealing critical contextual relationships often overlooked by traditional methods.
WDL-OPD boosts MATH500 accuracy from 0.630 to 0.685, showcasing a powerful new approach to stabilizing on-policy distillation.
Trajectory anchoring bias in VLA models can be mitigated by transforming future trajectory decisions into verifiable selections, leading to more reliable reasoning in autonomous driving.
Coordinating UAVs and ground vehicles for on-demand edge services doesn't have to sacrifice privacy or efficiency: LOSA's look-ahead mechanism achieves both.
RLHF can be made more stable and effective by explicitly verifying and reinforcing policy improvements against a historical baseline, rather than relying solely on instantaneous reward signals.