Search papers, labs, and topics across Lattice.
This study evaluates the performance of large language models (LLMs) on trivia-style questions through a newly introduced multilingual benchmark, TriviaRoomQA, which encompasses 3,300 parallel multiple-choice questions across six European languages. The findings reveal that while LLMs excel in knowledge-intensive areas like history and mathematics, they struggle significantly with everyday cultural topics, indicating a notable knowledge gap. Additionally, the performance of these models varies by language, highlighting the complexity of factual knowledge access across different linguistic contexts.
LLMs may ace academic trivia but falter on everyday cultural knowledge, revealing a critical gap in their understanding.
Quiz rooms, trivia nights, and quiz shows challenge human knowledge across a wide range of topics, from canonical facts to everyday culture. In this paper, we examine whether large language models (LLMs) can perform competitively in such settings, using quiz-style questions to test them on both common and niche topics. We introduce TriviaRoomQA, a multilingual benchmark designed to evaluate everyday, culturally grounded, and long-tail knowledge across 288 topics. The benchmark contains 3,300 parallel multiple-choice questions in six European languages and additional 5,340 French-only questions for a more fine-grained case study. We evaluate 30 open-weight LLMs from European, Asian, and North American providers, covering models from 7 to 70B parameters. We find that models are strong on knowledge-intensive topics such as history, geography, and mathematics, but substantially weaker on everyday popular-culture topics such as celebrities, music, movies, and news. Moreover, model performance varies across languages even for the same underlying questions, suggesting that access to factual knowledge is not always language-independent. In sum, our dataset and experiments demonstrate an important knowledge gap which is not captured by existing academic-based saturated benchmarks.