Search papers, labs, and topics across Lattice.
This paper introduces AtlasNLP, a comprehensive resource that catalogs over 13,000 NLP datasets with a focus on country-level representation, addressing the critical lack of geographic metadata in existing datasets. The analysis reveals significant disparities in dataset coverage across different countries and tasks, highlighting that language diversity does not equate to geographic inclusivity. These insights underscore the need for improved documentation practices in NLP to ensure equitable representation and inform data collection strategies and AI policy.
Dataset representation in NLP is alarmingly uneven, with geographic blind spots that could skew AI development and policy.
Understanding which countries are represented in NLP datasets is essential for identifying gaps, targeting data collection, measuring progress, and informing AI policy. However, geographic metadata is very rarely available, and country-level representation is often hidden behind broad language-level claims. We introduce AtlasNLP, a country-aware atlas of over 13,000 NLP dataset records across normalized NLP task categories, tracking both the populations represented and where datasets are produced. AtlasNLP includes AtlasNLP-Gold, a human-curated reference set, and AtlasNLP-Core, an ACL-derived large-scale collection. Using this resource, we show that (1) dataset coverage is highly uneven across countries and tasks; (2) dataset production and representation are geographically asymmetric; and (3) language coverage does not imply geographic representation. These findings reveal blind spots in current dataset documentation practices and motivate more explicit geographic metadata for country-aware NLP evaluation.