TurinMar 29, 2026arXiv:2603.27768

TailNLG: A Multilingual Benchmark Addressing Verbalization of Long-Tail Entities

Lia Draetta, Michael Oliverio, Virginia Ramón-Ferrer, Pier Felice Balestrucci, Flaviana Corallo, Carlos Badenes-Olmedo, Alessandro Mazzei, Marco Antonio Stranisci, Rossana Damiano

AI Summary

The paper introduces TailNLG, a new multilingual benchmark (English, Italian, Spanish) derived from Wikidata, designed to evaluate Data-to-Text generation models' ability to verbalize long-tail entities. Experiments with LLMs in zero-shot settings reveal a consistent bias against rare entities, evidenced by lower embedding-based scores and higher model uncertainty. The study also demonstrates that existing evaluation metrics fail to consistently capture these biases, underscoring the need for improved evaluation frameworks.

Key Contribution

LLMs struggle to verbalize rare entities, exhibiting lower performance and higher uncertainty compared to common entities, even in multilingual settings.

Abstract

The automatic verbalization of structured knowledge is a key task for making knowledge graphs accessible to non-expert users and supporting retrieval-augmented generation systems. Although recent advances in Data-to-Text generation have improved multilingual coverage, little attention has been paid to potential biases in the verbalization of rare entities, frequently known as long-tail entities. In this work, we present the first systematic study of long-tail entities in Data-to-Text generation. We introduce TailNLG, a new multilingual benchmark in English, Italian, and Spanish, built from Wikidata and covering entities with varying levels of popularity. We evaluate three different families of large language models in zero-shot settings and compare their performance on rare versus common entities, as well as against the established WebNLG benchmark. Our results reveal a consistent bias against long-tail entities: embedding-based scores are lower, and model uncertainty is higher for rare entities. We further show that the impact of long-tail entities varies across models and languages, and that existing evaluation metrics do not consistently capture these differences, highlighting the need for more reliable evaluation frameworks.

Data Curation & Synthetic Data Eval Frameworks & Benchmarks Natural Language Processing

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

TailNLG: A Multilingual Benchmark Addressing Verbalization of Long-Tail Entities

Related Papers