Search papers, labs, and topics across Lattice.
2
0
3
0
Adapting supervision weights based on the evolution of divergence histories boosts reasoning performance in language models without extra computational overhead.
Merging RL experts effectively requires balancing sharp, informative signals with stable, dispersed components, a challenge that ResMerge addresses with innovative spectral techniques.