Search papers, labs, and topics across Lattice.
3
0
5
Pretraining loss is a deceptive selection metric: at 30B MoE scale, downstream SFT performance is governed not by benchmark scores, but by the checkpoint's solution density under local weight perturbations.
Overparameterization isn't just a quirk of deep learning; it's provably *necessary* for stable, robust classification, even for discontinuous functions.
1-bit quantization, powered by k-means, can surprisingly outperform higher-bit integer quantization in generative tasks under a fixed memory budget.