Search papers, labs, and topics across Lattice.
Affiliation:
1
0
Increasing non-thinking supervision can actually hinder the accuracy of thinking mode in large language models, revealing a critical trade-off in training dynamics.