Search papers, labs, and topics across Lattice.
This paper introduces CultureTalk-ID, a novel dialogue-based benchmark designed to evaluate cultural commonsense in Indonesian and its local languages through 4,496 authentic dialogues across 11 languages. By focusing on the dialogic context, the benchmark addresses the limitations of existing models that rely on isolated prompts, thereby capturing the nuances of cultural communication. The key findings reveal that LLMs struggle with culturally grounded reasoning and translation tasks, highlighting significant gaps in their understanding of local cultural contexts.
LLMs falter in culturally nuanced dialogue, revealing critical gaps in their understanding of Indonesian cultural commonsense.
Culture is lived through conversation, yet existing Indonesian cultural commonsense benchmarks evaluate LLMs on short and isolated prompts, stripping away the dialogic context in which cultural nuances actually surface. We introduce CultureTalk-ID, the first dialogue-based benchmark for cultural commonsense in Indonesian and its local languages, comprising 4,496 culturally grounded dialogues across 11 languages and 13 culturally salient topics, curated through a multi-stage human pipeline with native speakers to ensure authenticity. CultureTalk-ID introduces three complementary tasks, namely dialogue-based multiple-choice cultural commonsense reasoning, culturally faithful machine translation, and language steering, which jointly probe whether LLMs can understand, transfer, and generate culturally grounded language.