Search papers, labs, and topics across Lattice.
3
0
4
4
Retaining more copies of low-frequency content while aggressively pruning high-frequency duplicates can significantly boost model performance during pretraining.
LLMs can now be benchmarked for their ability to prepare training data, revealing that a new evaluation metric outperforms traditional methods in predicting downstream utility.
DataFlex makes data-centric LLM training dramatically easier, unifying disparate methods for data selection, mixing, and reweighting into a single, efficient, and reproducible framework.