Search papers, labs, and topics across Lattice.
1
0
2
Tracking both task advantage and actual reward gains can drastically improve the efficiency of multi-task reinforcement learning for LLMs.