Search papers, labs, and topics across Lattice.
2
0
5
0
Sharing rollout feedback across related samples can significantly boost exploration efficiency in RLVR, leading to better reasoning capabilities in large language models.
HarnessEval-W transforms world model evaluation from mere scoring to a transparent reasoning process that mirrors human judgment.