Search papers, labs, and topics across Lattice.
The ChinaVoices Challenge 2026 establishes standardized tasks and evaluation conditions for processing Chinese dialects, focusing on multi-dialect identification and automatic speech recognition (ASR) across 16 dialect categories. Utilizing approximately 320 hours of speech data, the challenge features two tracks鈥攔estricted and open data鈥攚hile analyzing submissions from 28 teams, with most systems surpassing the baseline performance. Notably, the findings reveal a correlation between dialect identification accuracy and ASR error rates, offering insights into effective strategies for dialect-discriminative acoustic representation and data normalization in speech processing systems.
Systems that excel in dialect identification often leverage unique acoustic features, while ASR performance hinges on data normalization strategies.
This paper summarizes the ChinaVoices Challenge 2026, which aims to establish unified task definitions and evaluation conditions for Chinese dialect speech processing and to advance multi-dialect identification and automatic speech recognition. The challenge covers 16 dialect categories and defines two tasks: Chinese Multi-Dialect Identification and Chinese Multi-Dialect Automatic Speech Recognition (ASR). It uses approximately 320 hours of speech across the Reference Set, Open Evaluation Set, and Hidden Evaluation Set. The two tasks use the same evaluation audio, and each includes restricted-data and open-data tracks. We describe the task settings, data, evaluation metrics, and Qwen3-ASR-1.7B baseline, and analyze the leaderboard results and submitted systems. In total, 28 teams submit results, 17 provide system reports, and systems from 15 teams pass the compliance review and are included in the analysis. Most eligible systems outperform the baseline, and the official top-three order remains unchanged on the Hidden Evaluation Set for both tasks. Dialect-level results show that categories with higher identification accuracy generally have lower ASR error rates, although the tasks assess related but distinct capabilities. Leading identification systems commonly exploit dialect-discriminative acoustic representations, whereas leading ASR systems emphasize data normalization, augmentation, and auxiliary CTC objectives. These results provide practical guidance for developing and evaluating Chinese multi-dialect speech processing systems.