Search papers, labs, and topics across Lattice.
2
0
4
1
OPD excels at transferring reasoning skills over specific answers, revealing a nuanced relationship between teacher-student origins that can complicate multi-teacher setups.
LLMs can learn to reason *worse* from seemingly better training data: models trained on CoT data with lower loss can generalize poorly due to inheriting inefficient, divergent reasoning patterns.