Search papers, labs, and topics across Lattice.
Affiliation:
2
0
0
The real driver of on-policy distillation gains in multi-turn agents is counterfactual rollback at the first fatal mistake鈥攁n insight leveraged here to achieve double-digit RLVR gains by internalizing rollback dynamics directly into gradient updates without resetting the environment.
Frozen autonomous policies can survive out-of-distribution dynamics without retraining by turning streaming conformal prediction residuals directly into adaptive safety projection margins.