Search papers, labs, and topics across Lattice.
Affiliation:
2
1
2
5
Mixed-policy reinforcement learning can enable language models to absorb knowledge more effectively than traditional supervised fine-tuning, especially in complex reasoning scenarios.
RLHF's implicit preference aggregation can now be explicitly controlled via differentiable loss functions corresponding to different voting rules, enabling principled trade-offs between axiomatic guarantees and optimization stability.