Search papers, labs, and topics across Lattice.
Shanghai Jiao Tong University
2
0
2
Training LLMs with tokens selected by gradient magnitude rather than just entropy boosts reasoning performance across diverse tasks.
The Relative Surprisal Index reveals that the interplay between token probability and entropy is crucial for optimizing reinforcement learning in language models, leading to substantial performance gains.