BolognaMar 17, 2026arXiv:2603.16258

Is Semi-Automatic Transcription Useful in Corpus Creation? Preliminary Considerations on the KIParla Corpus

Martina Simonotti, Eleonora Zucchini, Silvia Ballarè, Caterina Mauri

AI Summary

This paper investigates the impact of ASR-assisted transcription on the creation of the KIParla corpus of spoken Italian. Eleven transcribers, with varying experience levels, produced both manual and ASR-assisted transcriptions of the same audio segments from three conversation types. The study uses statistical modeling, word-level alignment, and annotation-based metrics to analyze transcription speed and accuracy, revealing that ASR assistance can speed up transcription but doesn't guarantee accuracy improvements, with results varying based on workflow, conversation type, and transcriber experience.

Key Contribution

ASR-assisted transcription doesn't automatically improve accuracy in corpus creation, and its effectiveness hinges on factors like workflow design and transcriber expertise.

Abstract

This paper analyses the implementation of Automatic Speech Recognition (ASR) into the transcription workflow of the KIParla corpus, a resource of spoken Italian. Through a two-phase experiment, 11 expert and novice transcribers produced both manual and ASR-assisted transcriptions of identical audio segments across three different types of conversation, which were subsequently analyzed through a combination of statistical modeling, word-level alignment and a series of annotation-based metrics. Results show that ASR-assisted workflows can increase transcription speed but do not consistently improve overall accuracy, with effects depending on multiple factors such as workflow configuration, conversation type and annotator experience. Analyses combining alignment-based metrics, descriptive statistics and statistical modeling provide a systematic framework to monitor transcription behavior across annotators and workflows. Despite limitations, ASR-assisted transcription, potentially supported by task-specific fine-tuning, could be integrated into the KIParla transcription workflow to accelerate corpus creation without compromising transcription quality.

Data Curation & Synthetic Data Natural Language Processing Speech & Audio

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Is Semi-Automatic Transcription Useful in Corpus Creation? Preliminary Considerations on the KIParla Corpus

Related Papers