Search papers, labs, and topics across Lattice.
1
0
2
4
A reward-compatible bounded mixing mechanism for $\gamma\mathrm{OPD}$ that balances verifiable outcome feedback with the discounted OPD advantage to move beyond purely teacher-dependent optimization is developed.