Search papers, labs, and topics across Lattice.
This paper introduces IslamicTurathBench (ISTB), a comprehensive multi-task dataset designed to evaluate large language models (LLMs) on classical Islamic scholarship, addressing the lack of high-quality annotated resources in this domain. The dataset comprises 3,465 question-answer items derived from 35 authoritative works spanning over 12 centuries and covers seven key fields of Islamic Studies, allowing for nuanced assessment of model performance across varying scholarly demands and task formats. Key findings include baseline performance metrics from ten LLMs, providing a foundation for future research and development in applying AI to religious and cultural contexts.
LLMs struggle with the complexities of Islamic scholarship, and the newly introduced ISTB reveals significant gaps in their performance across different levels of scholarly demand.
Large language models (LLMs) are increasingly used for question answering, education, and research, including in religious and cultural domains where answers depend on specialised source traditions. Yet in Islamic Studies, key concepts, methods, and debates preserved in the authoritative scholarly tradition, known as turath, lack high-quality annotated resources. We introduce IslamicTurathBench (ISTB), a multi-task, multi-discipline dataset for evaluating LLMs on classical Islamic scholarship. Developed and reviewed by domain experts, ISTB contains 3,465 question-answer items drawn from 35 recognised source works spanning more than 12 centuries of scholarship across seven key fields of Islamic Studies. To enable comprehensive profiling of model capabilities, ISTB is structured along two axes: scholarly demand (Beginner, Intermediate, and Advanced) and task format (multiple-choice questions, passage-based comprehension, and open-ended knowledge questions). ISTB includes aggregated scores from a scholarly human reference panel and zero-shot baselines from ten systems. The dataset supports reproducible evaluation of language-model behaviour across source works, disciplines, scholarly-demand levels, and question formats in a historically layered scholarly domain.