Search papers, labs, and topics across Lattice.
The paper introduces KUTED, a new English-to-Central Kurdish speech-to-text translation (S2TT) dataset comprising 91,000 sentence pairs derived from TED talks. They found that orthographic variations in Central Kurdish significantly degrade translation performance. To mitigate this, they propose and implement a text standardization approach, achieving a 3.0 BLEU improvement on the FLEURS benchmark using a fine-tuned Seamless model.
Orthographic inconsistencies in low-resource languages like Central Kurdish can cripple S2TT performance, but systematic text standardization offers a surprisingly effective remedy.
We present KUTED, a speech-to-text translation (S2TT) dataset for Central Kurdish, derived from TED and TEDx talks. The corpus comprises 91,000 sentence pairs, including 170 hours of English audio, 1.65 million English tokens, and 1.40 million Central Kurdish tokens. We evaluate KUTED on the S2TT task and find that orthographic variation significantly degrades Kurdish translation performance, producing nonstandard outputs. To address this, we propose a systematic text standardization approach that yields substantial performance gains and more consistent translations. On a test set separated from TED talks, a fine-tuned Seamless model achieves 15.18 BLEU, and we improve Seamless baseline by 3.0 BLEU on the FLEURS benchmark. We also train a Transformer model from scratch and evaluate a cascaded system that combines Seamless (ASR) with NLLB (MT).