Search papers, labs, and topics across Lattice.
This paper introduces WSE-bench, a benchmarking framework designed to evaluate large language models (LLMs) on their storytelling capabilities within evolving world simulations. The study reveals that while larger models enhance sustained narrative generation, they do not guarantee improvements in canonical coherence or meaningful story development, highlighting a complex relationship between these storytelling dimensions. The findings indicate that sustained generation, coherence, and meaningful development are distinct and can sometimes conflict, challenging the assumption that enhancing one aspect will benefit the others.
The empirical analysis reveals that larger models may excel in generating narratives but often fail to maintain coherence and depth, exposing a critical trade-off in LLM storytelling capabilities.
Large language models can write fluent stories, but open-ended storytelling requires more than local fluency. In evolving world simulations and AI-native games, models must preserve facts, relationships, causal dependencies, and character states as the world changes. We introduce WSE-bench, a process benchmark that separately evaluates sustained generation, canonical coherence, and meaningful development in dynamic LLM storytelling. Generation Coverage records the proportion of planned narrative steps produced; Consistency tracks when canon breaks; and Richness measures how meaningfully branching, player-shaped trajectories develop. Across frontier models, Consistency and Richness do not form a smooth trade-off: their empirical Pareto frontier is non-concave, with several non-dominated intermediate configurations that no positive linear weighting can select. Added structure can enrich trajectories, but it does not uniformly improve coherence and may shorten them. Model scale chiefly improves sustained generation, without producing reliable gains in canonical coherence or meaningful development. These results show that sustained generation, canonical coherence, and meaningful development are distinct and sometimes competing capacities. WSE-bench makes those dynamics visible by extending narrative evaluation from finished stories to the processes that create them.