Mar 16, 2026arXiv:2603.15034

Interpretable Predictability-Based AI Text Detection: A Replication Study

AI Summary

This paper replicates and extends a system for detecting machine-generated text, originally used in the AuTexTification 2023 shared task, by incorporating newer multilingual language models and 26 document-level stylometric features. The authors replaced GPT-2 with Qwen and mGPT for probabilistic features and used mDeBERTa-v3-base for contextual representations in both English and Spanish. Results demonstrate that the added stylometric features improve performance, and a multilingual configuration can match or exceed language-specific models, while also highlighting the importance of clear documentation for reproducibility.

Key Contribution

Stylometric features, combined with modern multilingual language models, significantly boost the performance of machine-generated text detection, often surpassing language-specific models.

Abstract

This paper replicates and extends the system used in the AuTexTification 2023 shared task for authorship attribution of machine-generated texts. First, we tried to reproduce the original results. Exact replication was not possible because of differences in data splits, model availability, and implementation details. Next, we tested newer multilingual language models and added 26 document-level stylometric features. We also applied SHAP analysis to examine which features influence the model's decisions. We replaced the original GPT-2 models with newer generative models such as Qwen and mGPT for computing probabilistic features. For contextual representations, we used mDeBERTa-v3-base and applied the same configuration to both English and Spanish. This allowed us to use one shared configuration for Subtask 1 and Subtask 2. Our experiments show that the additional stylometric features improve performance in both tasks and both languages. The multilingual configuration achieves the results that are comparable to or better than language-specific models. The study also shows that clear documentation is important for reliable replication and fair comparison of systems.

Interpretability & Mechanistic Interp Natural Language Processing Open-Source Models & Weights

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Interpretable Predictability-Based AI Text Detection: A Replication Study

Related Papers