Search papers, labs, and topics across Lattice.
This paper redefines the characterization of Persian in natural language processing (NLP) from a low-resource language to an annotation-scarce language, emphasizing the ecological factors affecting its NLP resources. Through a review of 34 Persian text resources and three quantitative analyses, the authors reveal that while Persian has a significant web presence, the availability of annotated data varies greatly across different tasks and domains. The findings highlight that the scarcity of annotations is not due to an overall lack of data but rather an uneven distribution and accessibility of resources, which has critical implications for developing NLP applications in Persian.
Persian isn't just low-resource; it's annotation-scarce, revealing a complex landscape of NLP resource availability that challenges conventional assumptions.
Persian (Farsi) is often described as a low-resource language in natural language processing, but that label collapses distinct shortages into a single category. This paper argues that Persian is more precisely described as annotation-scarce, provided that the term is understood as a property of its NLP resource ecology rather than an intrinsic property of the language. The review covers 34 representative Persian text resources available by July 2026 and adds three quantitative cross-checks. First, independent web measurements place Persian among roughly the twenty most visible content languages: W3Techs reports Persian on about 0.9% of websites with a known content language, while Common Crawl CC-MAIN-2026-30 identifies Persian as the primary language of 0.7039% of HTML pages. Second, a selective speech review shows a long resource trajectory from FARSDAT to recent corpora containing hundreds or thousands of hours of speech. Third, a matched Persian-English comparison normalizes task-specific annotation volumes by relative Common Crawl web presence. The resulting ratios vary sharply: Persian syntax and news NER are comparatively dense, whereas natural-language inference falls below the web-proportional baseline. The evidence therefore does not support a simple claim that Persian is globally deficient in labeled volume. Instead, annotation scarcity is expressed through uneven task and domain coverage, incompatible schemes, access and documentation friction, and limited supervision for specialist domains, preference data, and varieties beyond standard Iranian Persian.