Search papers, labs, and topics across Lattice.
3
0
5
3
Over 57% of training samples suffer from suppressed criteria, but a new self-distillation method can reverse this trend and enhance performance in RL tasks.
OPRD closes the performance gap between student and teacher models while training 1.44x faster and using 54% less memory than traditional methods.
RLVR models exhibit "Early Correctness Coherence" under noisy supervision, suggesting a surprising opportunity for self-correction via dynamic label refinement.