Search papers, labs, and topics across Lattice.
2
0
4
20
Training updates that improve performance in LLMs can actually degrade inference quality鈥攗nless you use the new Monotonic Inference Policy Update framework.
Plasticity loss in RL isn't just about forgetting; it's about vanishing gradients, and a simple sample re-weighting can bring back the learning.