Search papers, labs, and topics across Lattice.
1
0
2
Pretraining loss is a deceptive selection metric: at 30B MoE scale, downstream SFT performance is governed not by benchmark scores, but by the checkpoint's solution density under local weight perturbations.