Search papers, labs, and topics across Lattice.
Department of Computer Engineering
2
0
4
0
Attention-free models can outperform transformers in language generation, especially at smaller dataset scales, challenging the dominance of attention mechanisms in NLP.
Forget tokenization and attention: this tiny 733K parameter model classifies text directly from raw bytes using frequency-domain processing, outperforming larger tokenized models.