Search papers, labs, and topics across Lattice.
This paper introduces CN-NewsTTS Bench v0.1, a comprehensive benchmark designed to evaluate the pronunciation accuracy of Chinese news text-to-speech (TTS) systems when processing complex written forms without relying on user-side rules or manual edits. The benchmark includes a development set of 200 records and a public test set of 800 records, along with 992 auto-evaluable targets and an automatic scoring system, allowing for systematic assessment across multiple TTS products. Key findings reveal that while the top-performing TTS system achieves a strict accuracy of 0.879, many others struggle, with several systems scoring below 0.60, highlighting significant variability in performance across products.
TTS systems struggle with complex Chinese news text, with top models achieving only 87.9% accuracy while many lag below 60%.
Chinese news text contains dense written forms such as scores, hyphenated model names, ranges, unit symbols, percentages, English abbreviations, and mixed Chinese-Latin-digit names. These forms are frequent in real listening workflows, and a text-to-speech (TTS) system can preserve the written string while changing the spoken meaning. We introduce CN-NewsTTS Bench v0.1, an open target-level benchmark for evaluating whether Chinese news TTS products pronounce such targets correctly from raw text, without user-side rules, LLM rewriting, SSML hints, or manual edits. The release contains a 200-record development set, an 800-record public test set, 992 public auto-evaluable targets, fixed transcripts from a three-ASR ensemble, an automatic target scorer, and initial results for seven product TTS systems. We additionally report ASR-route diagnostics, ASR-subset ablations, category-level results, confidence intervals, and provider configuration metadata. The best system reaches 0.879 strict accuracy, while several systems remain below 0.60.