Search papers, labs, and topics across Lattice.
5
0
8
0
Higher epiplexity in training data can significantly boost model performance in unseen tasks, revealing a new avenue for data-driven generalization strategies.
Requential coding can compress billion-parameter models to sizes orders of magnitude smaller than traditional methods, revealing the hidden structure in datasets.
Emergent capabilities in transformer models arise stochastically, with larger models gaining critical skills earlier due to their ability to learn sparse attention patterns more effectively.
LLMs can ace physics exams, but put them in a universe with slightly different laws and their scientific reasoning skills fall apart.
Language models can leverage their own generative abilities to nearly eliminate forgetting during finetuning, challenging the need for stored exemplars.