Search papers, labs, and topics across Lattice.
2
0
5
2
OPD$^2$ not only boosts multilingual math reasoning but also narrows the performance gap between English and Korean models, revealing the hidden potential of language-specific training signals.
Capturing the delta between teacher and base models can revolutionize on-policy distillation, leading to superior reasoning capabilities in LLMs with just a brief post-training phase.