Search papers, labs, and topics across Lattice.
3
0
5
OPDVR transforms the landscape of model distillation by ensuring that only correct trajectories enhance learning, leading to significant performance gains on reasoning tasks.
Decomposing complex reasoning problems into verifiable subproblems unlocks significant performance gains in LLM reasoning, especially on hard problems previously stuck in gradient dead zones.
Fine-tuning LLMs on datasets filtered at the token level, rather than the sentence level, can boost performance by up to 13.7%.