Search papers, labs, and topics across Lattice.
2
0
5
2
Optimal self-distillation can significantly enhance generative model performance by intelligently mixing teacher and true velocity signals, correcting both under- and over-regularization issues.
DRA outputs are surprisingly variable, with inference and early-stage decisions being the biggest culprits, but structured outputs and ensemble querying can significantly reduce this stochasticity.