Search papers, labs, and topics across Lattice.
Affiliation:
2
0
3
2
A two-stage OPD-then-RL approach outperforms traditional methods by leveraging the strengths of both on-policy distillation and reinforcement learning without the interference seen in joint optimization.
Smaller language models can efficiently replace larger ones in rubric-based reinforcement learning, achieving competitive performance with significantly reduced computational costs.