Search papers, labs, and topics across Lattice.
3
0
6
Optimal hyperparameter choices for supervised fine-tuning can vary dramatically with model scale and architecture, challenging conventional wisdom in AI training practices.
Knowledge retention in language models is heavily influenced by training data diversity, with broad data reducing the recitation-to-use gap from 27.4 to 5.4 points.
Stop benchmarking algorithm discovery on the same old saturated datasets: DiscoGen offers millions of fresh, configurable tasks to truly test your ADA.