Search papers, labs, and topics across Lattice.
This study analyzes the formal variation in novels generated by large language models (LLMs) compared to human-written texts, focusing on sentence structure, readability, and other stylistic measures. By contrasting outputs from GPT-5.5 and Qwen3-14B across different styles with human-written novels, the research reveals that AI-generated novels exhibit significantly less variability in sentence structure and other formal features. The findings indicate a phenomenon termed "variance overclosure," where a collection of AI-generated works shows a constrained formal range despite individual pieces resembling human fiction stylistically.
AI-generated novels show a striking lack of formal diversity, with repeated generations compressing sentence structure far more than human authors do.
While large language models can generate entire novels, there is little information about the level of formal variation in their output over many generations. Rather than asking whether individual passages can be identified as AI-generated, this study asks whether repeated AI generation can produce the same range of diversity which is found across human corpora. This paper contrasts six corpora based on generation source and target style: twenty novels generated using GPT-5.5 Thinking in a nineteenth-century British realist style, twenty novels generated using Qwen3-14B in a nineteenth-century British realist style, twenty novels generated using each of these models in a contemporary zero style, 205 nineteenth-century human-written British novels, and sixty-five contemporary human-written Zero-Style novels. At the document level, the research includes MATTR-500, Shannon entropy, average sentence length, readability, and punctuation rate measurements. The most robust and reliable result is compression of sentence structure. Repeated generations produce novels that vary far less from one another in sentence structure than human novels do. Compression is also present in the measures of readability, punctuation, and sentence length variability within novels. Lexical measures tend to be similarly compressed, with the exception of Qwen Zero-Style MATTR. Despite having distinct mean stylistic profiles, GPT and Qwen lack a stable pattern of cross-measure correlation. This article therefore distinguishes between variance overclosure, which represents a limited formal range between novels, and a more specific phenomenon of correlational overclosure. This means that an individual AI-generated novel may resemble human fiction stylistically, while a collection of AI-generated novels occupies a much narrower formal range.