MonashMar 2, 2026arXiv:2603.01952

LiveCultureBench: a Multi-Agent, Multi-Cultural Benchmark for Large Language Models in Dynamic Social Simulations

Viet-Thanh Pham, Lizhen Qu, Thuy-Trang Vu, Gholamreza Haffari, Dinh Phung

AI Summary

The paper introduces LiveCultureBench, a novel benchmark for evaluating LLM agents in dynamic, multi-cultural social simulations, focusing on both task completion and adherence to socio-cultural norms. The benchmark uses a simulated town with diverse synthetic residents and an LLM-based verifier to generate structured judgments on norm violations and task progress. Experiments using LiveCultureBench reveal insights into cross-cultural robustness, the trade-off between effectiveness and norm sensitivity, and the reliability of LLM-based evaluation compared to human oversight.

Key Contribution

LLMs struggle to balance task completion with cultural norms in dynamic social simulations, revealing critical gaps in their cross-cultural robustness and highlighting the need for human oversight in automated benchmarking.

Abstract

Large language models (LLMs) are increasingly deployed as autonomous agents, yet evaluations focus primarily on task success rather than cultural appropriateness or evaluator reliability. We introduce LiveCultureBench, a multi-cultural, dynamic benchmark that embeds LLMs as agents in a simulated town and evaluates them on both task completion and adherence to socio-cultural norms. The simulation models a small city as a location graph with synthetic residents having diverse demographic and cultural profiles. Each episode assigns one resident a daily goal while others provide social context. An LLM-based verifier generates structured judgments on norm violations and task progress, which we aggregate into metrics capturing task-norm trade-offs and verifier uncertainty. Using LiveCultureBench across models and cultural profiles, we study (i) cross-cultural robustness of LLM agents, (ii) how they balance effectiveness against norm sensitivity, and (iii) when LLM-as-a-judge evaluation is reliable for automated benchmarking versus when human oversight is needed.

Constitutional AI & AI Ethics Eval Frameworks & Benchmarks Tool Use & Agents

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

LiveCultureBench: a Multi-Agent, Multi-Cultural Benchmark for Large Language Models in Dynamic Social Simulations

Related Papers