Search papers, labs, and topics across Lattice.
This paper introduces L3Cube-IndicQuest v2, a comprehensive multilingual benchmark designed to assess the factual knowledge of large language models (LLMs) specifically in the context of Indian languages. The benchmark features 3,471 English question-answer pairs across nine domains, generated through a hybrid approach that combines LLM-based question generation with human verification, and is translated into 19 Indic languages, resulting in a dataset of 69,420 pairs. Evaluation of six LLMs reveals that the leading commercial model significantly outperforms others, including the Indic-specialized Sarvam 30B, across all languages tested, indicating the effectiveness of the benchmark and the performance disparities among models.
The leading commercial LLM outperforms open-weight models by a significant margin, revealing stark disparities in factual knowledge across languages.
We present L3Cube-IndicQuest v2, a large-scale gold-standard multilingual question-answering benchmark for evaluating the India-specific factual knowledge of Large Language Models (LLMs). The benchmark comprises 3,471 curriculum-grounded English question--answer pairs spanning nine domains, curated from educational curricula, competitive examination materials, and domain-specific reference books. We introduce a practical hybrid construction strategy that combines context-grounded LLM-based question generation and validation with semantic deduplication and human verification, enabling scalable creation of benchmark data while preserving annotation quality. The benchmark is translated into 19 Indic languages, yielding a publicly released multilingual dataset of 69,420 question--answer pairs across 20 languages. We evaluate six LLMs under three protocols: LLM-as-a-judge and two deterministic lexical criteria, exact-substring and word-overlap matching. All three produce almost the same model ranking, showing that the results do not depend on the choice of judge. The frontier commercial model leads by a wide margin, and among open-weight models Gemma4 31B outperforms the Indic-specialised Sarvam 30B in every evaluated Indic language.