Search papers, labs, and topics across Lattice.
This study investigates the performance of 13 large language models on the ProverbIT benchmark, a new Italian dataset designed to evaluate their ability to complete and select proverbs. While models excelled in completing proverbs, they struggled significantly in multiple-choice tasks, particularly when correct answers were not provided, revealing a critical gap in their reasoning capabilities. The analysis indicates that these models often default to literal synonyms and fail to recognize the absence of correct options, underscoring their reliance on memorization over genuine semantic understanding of culturally embedded language.
LLMs may ace proverb completion but falter dramatically when faced with multiple-choice questions, exposing a troubling reliance on memorized patterns over true comprehension.
Large Language Models (LLMs) have transformed computational linguistics and achieved remarkable performance across numerous natural language processing tasks, yet significant gaps persist in understanding how these systems process culturally embedded linguistic expressions. This paper introduces ProverbIT, a novel Italian benchmark comprising 100 multiple-choice questions designed to evaluate LLMs'ability to complete Italian proverbs. We assess 13 frontier models, including Large Reasoning Models (LRMs) and traditional LLMs, across three tasks: proverb completion, multiple-choice selection with correct answers, and multiple-choice selection without correct answers. Our evaluation reveals surprising results: while nearly all models demonstrate knowledge of the proverbs through successful completion tasks, performance drops dramatically when transitioning to multiple-choice formats without correct answers, with even state-of-the-art reasoning models showing substantial degradation. Through detailed Chain-of-Thought analysis of two LRMs, we uncover that models exhibit a strong bias toward selecting literal synonyms and frequently mention correct proverb endings during reasoning without successfully identifying their absence from the given options. These findings suggest that current LLMs rely heavily on memorized patterns rather than deeper semantic understanding of culturally grounded expressions, highlighting important limitations in their reasoning capabilities for figurative language comprehension.