Search papers, labs, and topics across Lattice.
This paper introduces a non-autoregressive method for diacritic restoration in Arabic speech transcripts using Connectionist Temporal Classification (CTC) with hard constraints during decoding. By constructing a character-level diacritization lattice from undiacritized transcripts, the approach effectively narrows down hypotheses to valid diacritized forms. The proposed method achieves statistically significant reductions in diacritic error rates compared to a more complex multi-modal baseline, highlighting its efficiency and performance benefits in handling Arabic phonological distinctions.
Efficient diacritic restoration in Arabic speech can be achieved with a non-autoregressive method that significantly reduces error rates while simplifying the decoding process.
In this work, we address diacritic restoration for Arabic speech transcripts. Most speech data are undiacritized, limiting the ability of modeling fine-grained phonological distinctions. The speech modality has recently been explored as a way to complement text-based diacritic restoration efforts. We propose an efficient non-autoregressive approach for speech-to-text diacritization based on Connectionist Temporal Classification (CTC). Our method incorporates hard constraints during decoding by constructing a character-level diacritization lattice from an undiacritized transcript and restricting hypotheses to valid diacritized realizations. We evaluate on Classical Arabic and Modern Standard Arabic test sets (namely, ArVoice and ClArTTS) against a more computationally-complex multi-modal diacritic restoration baseline, and show statistically significant reductions in diacritic error rates in both, demonstrating that the proposed approach offers both performance and efficiency gains.