Search papers, labs, and topics across Lattice.
Affiliation:
2
0
3
0
A two-stage OPD-then-RL approach outperforms traditional methods by leveraging the strengths of both on-policy distillation and reinforcement learning without the interference seen in joint optimization.
Stochastic eviction can boost reasoning model accuracy by protecting critical value states while achieving significant memory savings.