Search papers, labs, and topics across Lattice.
This paper introduces a comprehensive evaluation framework for assessing the fairness of watermarking schemes in large language models across multiple languages, addressing the limitations of existing evaluations that focus solely on English. By employing a multi-faceted approach that includes empirical detection thresholds, a companion measurement for calibration failures, and diverse quality metrics, the authors reveal significant disparities in watermarking effectiveness across typological language families. The findings indicate that these fairness gaps are rooted in structural language properties, challenging the assumption that evaluation results are consistent across different languages.
Cross-lingual fairness gaps in language model watermarking are not just language-specific but are fundamentally tied to the structural properties of language families.
Watermarking schemes for large language model output are evaluated almost exclusively on English text using each scheme's detection threshold and a narrow set of quality measurements. Multilingual deployment exposes evaluation-design choices that are inconsequential on English but determine conclusions cross-lingually. We propose an evaluation framework with four components: detection thresholds calibrated empirically per deployment context, a threshold-independent companion measurement that distinguishes calibration failures from detection failures, three disjoint quality measurement paradigms (distributional, paired-semantic, and reference-perplexity), and a generalized-entropy decomposition of cross-language disparity over a typological family partition. Applied to six watermarking schemes, three open-weight generators, eleven languages spanning four scripts and eight typological families, and both base and instruction-tuned regimes, the framework reveals failure modes that single-language single-paradigm evaluation cannot surface. Across detection and quality, observed disparity is predominantly between-family on the typological partition, indicating that cross-lingual fairness gaps in watermarking are structural to language properties rather than idiosyncratic to particular languages.