Search papers, labs, and topics across Lattice.
2
0
2
Tracking both task advantage and actual reward gains can drastically improve the efficiency of multi-task reinforcement learning for LLMs.
The Relative Surprisal Index reveals that the interplay between token probability and entropy is crucial for optimizing reinforcement learning in language models, leading to substantial performance gains.