Search papers, labs, and topics across Lattice.
NAVER AI Lab, KAIST
3
0
7
29
OPD$^2$ not only boosts multilingual math reasoning but also narrows the performance gap between English and Korean models, revealing the hidden potential of language-specific training signals.
Retaining hazardous knowledge while selectively refusing dangerous queries could redefine safety strategies for large language models.
Capturing the delta between teacher and base models can revolutionize on-policy distillation, leading to superior reasoning capabilities in LLMs with just a brief post-training phase.