Search papers, labs, and topics across Lattice.
2
0
3
OPDVR transforms the landscape of model distillation by ensuring that only correct trajectories enhance learning, leading to significant performance gains on reasoning tasks.
Decomposing complex reasoning problems into verifiable subproblems unlocks significant performance gains in LLM reasoning, especially on hard problems previously stuck in gradient dead zones.