Search papers, labs, and topics across Lattice.
Affiliation:
3
0
5
0
Merging existing expert capabilities can yield significant performance boosts, but choosing the right fusion method can make all the difference in multi-domain reinforcement learning.
Most RLVR datasets are just remixes of a few originals, and this paper shows how to trace them back to their source, revealing widespread data contamination.
Pass-rate-1 prompts got you down? Composition-RL boosts LLM reasoning by automatically composing multiple problems into new verifiable questions, making better use of your existing data.