Search papers, labs, and topics across Lattice.
This paper advances the theoretical understanding of Transformers by establishing preliminary sample complexity bounds for learning C-RASP constructions, which are tailored to enhance the expressivity of attention-based models. The authors address a critical gap in existing research by shifting focus from expressivity to learnability, providing insights into the conditions under which Transformers can effectively learn specific tasks. Key findings suggest that the proposed bounds can inform the design of more efficient training protocols for large language models, thereby optimizing their performance on diverse tasks.
Transforming our understanding of Transformers, this work reveals that learnability may be as crucial as expressivity in optimizing large language models.
A theoretical understanding of Transformers is crucial to better understand the capacities and limitations of large language models (LLMs). There is much work analyzing the expressivity of attention-based models. By proposing handcrafted weights or using computational complexity arguments, a large amount of past theoretical works have sought to characterize which tasks are and which are not in the hypothesis class of Transformer models. However, little work investigates the learnability of such solutions. In this work, we make progress towards this goal. Inspired by recent loss landscape analysis work, we propose preliminary sample complexity bounds for learning C-RASP constructions with Transformers.