Search papers, labs, and topics across Lattice.
Affiliation:
2
0
3
The real driver of on-policy distillation gains in multi-turn agents is counterfactual rollback at the first fatal mistake鈥攁n insight leveraged here to achieve double-digit RLVR gains by internalizing rollback dynamics directly into gradient updates without resetting the environment.
Quadruped robots can now learn diverse skills and adapt to complex terrains without expert datasets, thanks to a novel keyframe-guided self-imitation learning framework.