Search papers, labs, and topics across Lattice.
This study investigates how periodic subject changes in text generation influence perceived novelty and connection in base language models, revealing that introducing a new subject every few hundred tokens significantly enhances judged surprise and connection. The research employs a controlled evaluation protocol across 24 conditions and three base models, demonstrating that interruptions lead to a 1.2 to 1.4 point increase in surprise and a 0.8 point increase in connection compared to habituation alone. Notably, the findings indicate that while certain connective prompts hinder performance, the simple act of interruption can substantially multiply valid candidate heuristics in specific tasks without compromising quality.
Injecting new subjects into a language model's output can boost perceived novelty by over 1.4 points, transforming how we understand model engagement.
Where does the novelty a base language model produces with no task come from, and what can an LLM judge of a long stream actually see? We dismantle a cognitively inspired generation loop over 24 conditions on three base models. Most of its effect lives in one operation: a new subject injected every few hundred tokens (an interruption) into a stream whose literal repetition is damped (habituation). We judge windows of generated text only, with the premise as the unit (n=10) and a judge measured for repeatability, against a second judge family and against human readers. Under that protocol the interruption raises judged surprise by 1.2 to 1.4 points and connection by 0.8 over habituation alone. A connective that asks for continuity hurts; a bare paragraph break adds nothing detectable on fresh text; a reset context does at least as well as a kept one; and a pre-registered replication on new premises confirms the primary contrast. Three things the window judge could not see changed the first version of this study, and we think they are of general use. The judge scores the experimenter's injected sentence as the model's own. A fixed rotation of injected sentences makes the model replay its earlier segments from beyond the judge's horizon, and the judge scores the replay as surprise and connection (65-80% of post-interruption windows at periods 150-300). And the local gains do not compose: no arm produces an integrated document. The salience monitor, the in-loop judge, memory across interruptions and a judge-gated Review run with a gate that opens add nothing. On a problem with a verifier (online bin packing), the interruption multiplies valid, distinct candidate heuristics three- to fourfold without raising the quality of the best. We report an evaluation protocol for long generation and a controlled characterization of a simple intervention, not a mechanism of creativity.