Search papers, labs, and topics across Lattice.
3
0
4
7
Token effectiveness varies dramatically with model size and data strategies, revealing that classic compute-optimal methods are often misguided in real-world applications.
Data synergy can either amplify or diminish model performance, revealing that the right dataset combinations are crucial for optimal language model training.
Larger models learn more not just because of increased capacity, but because they experience less interference during training, allowing them to retain rare and complex tasks that smaller models forget.