Search papers, labs, and topics across Lattice.
The ChineseBabyLM Challenge invites researchers to develop data-efficient language models for Chinese using only 100 million tokens, focusing on three evaluation tracks: natural language understanding (NLU), cognitive alignment, and Hanzi knowledge. This initiative aims to foster innovation in training methodologies and model architectures tailored to the Chinese language, addressing the unique linguistic and cognitive challenges it presents. By establishing a competitive framework, the challenge seeks to advance the state of language modeling in Chinese NLP and enhance the cognitive plausibility of these models.
Training language models with just 100 million tokens could redefine efficiency benchmarks in Chinese NLP.
This paper describes the first ChineseBabyLM challenge, which will be held in the 2026 NLPCC conference. The challenge calls for researchers to train language models from scratch with 100 million Chinese tokens and evaluates the models on 3 tracks of tasks: NLU, cognitive alignment and Hanzi knowledge. There is no restriction on tokenizer, model architecture and the number of training epochs. Details of the challenge can be found in https://chinese-babylm.github.io/.