Search papers, labs, and topics across Lattice.
This paper introduces a human-grounded framework to measure the distributional breadth of content generated by large language models (LLMs) compared to human-written texts. By employing two novel metrics, LLM Coverage (LLM-Cov) and In-Boundary Rate (IBR), the authors quantify how LLMs produce plausible yet narrow outputs that often lack the stylistic and thematic diversity characteristic of human writing. The findings reveal that LLMs tend to concentrate their outputs around the central themes of human responses, highlighting a significant gap in cultural reach between machine-generated and human-generated narratives.
LLMs may generate plausible narratives, but they fall short in capturing the rich diversity and stylistic irregularities found in human writing, revealing a critical gap in their cultural reach.
When a large language model (LLM) writes Harry Potter fanfiction, it reliably produces fundamental elements of the Hogwarts universe, such as recognizable places and characters. Human-written Harry Potter fanfictions, however, typically include these fundamentals and much more, incorporating stylistically irregular content and relationship-diverse plotlines. This gap between LLM and human writing has been noted across a variety of domains. LLMs tend to produce"average"writing, while human writing contains more diverse content that covers a broader distribution. Existing work has shown the existence of this distributional"gap", but no work has proposed a systematic way to measure it. Our paper proposes a human-grounded framework that uses the empirical distribution of human writing on a topic to measure the distributional breadth of LLM-generated content on that same topic. We propose two metrics, LLM Coverage (LLM-Cov) and In-Boundary Rate (IBR), that separate the plausibility of LLM content from its distributional breadth. Across ideation and narrative tasks, we find that current LLMs produce plausible but narrow content that concentrates near the center of the human response space. Our framework can enable researchers to better assess the distributional breadth of LLM-authored content, which we term its"cultural reach".