Search papers, labs, and topics across Lattice.
2
0
5
11
Skill-SP not only pushes the performance ceiling of LLMs but also transforms initially misaligned models into high-performing agents through dynamic skill evolution.
LLMs can learn to reason *worse* from seemingly better training data: models trained on CoT data with lower loss can generalize poorly due to inheriting inefficient, divergent reasoning patterns.