Search papers, labs, and topics across Lattice.
This study introduces TELLME, a novel method that enhances continual pre-training (CPT) for large language models by integrating Test-Enhanced Learning (TEL) principles, which utilize quizzes during training to boost efficiency. The approach addresses the challenges of acquiring large-scale domain-specific datasets and reduces computational costs while promoting better knowledge retention. Experimental results show that TELLME achieves up to a 23.6% performance improvement in the financial domain and a 9.8% increase in long-term memory retention compared to existing methods.
TELLME boosts language model performance by over 23% while slashing the need for extensive domain-specific datasets through innovative quiz-based training.
Continual pre-training (CPT) has been widely adopted as a method for domain adaptation in large language models. However, CPT has consistently been accompanied by challenges, such as the difficulty of acquiring large-scale domain-specific datasets and high computational costs. In this study, we propose a novel method called Test-Enhanced Learning for Language Model Enrichment (TELLME) to alleviate these issues. TELLME leverages the TestEnhanced Learning (TEL) principle, whereby the model's training efficiency is improved using quizzes during training. It integrates this principle with CPT, thereby promoting efficient domain-specific knowledge acquisition and long-term memory retention. Experimental results demonstrate that TELLME outperforms existing methods by up to 23.6% in the financial domain and achieves a 9.8% improvement in long-term memory retention.