Search papers, labs, and topics across Lattice.
The authors investigated the low diversity of LLM-generated stories, finding that a small set of tokens related to names, settings, and professions (e.g., "Elias," "lighthouse," "clockmaker") dominate the output across four current models. These tokens are not prevalent in pre-training data or published literature but are likely overrepresented in preference datasets used for alignment. This highlights how even small, biased datasets used in alignment can disproportionately influence LLM outputs, leading to a lack of diversity.
LLMs are churning out eerily similar "lighthouse" stories, revealing that alignment data can have a surprisingly outsized impact on generation diversity, even more so than pre-training data.
LLM-generated stories are a popular use case, but they show very low variability. We sample 20,000 total stories from four current models using five prompts. We find that 11 words occur in 88.3% of generated stories, with little difference between models. These words include names (Elias, Mara, Elara), settings (lighthouses), and professions (clockmaker, librarian). These tokens do not often occur in published literature nor pre-training data, but they are found in preference data that is likely to have been used by all current models. Surprisingly, these "lighthouse" stories are infrequent when compared with the average post-training story, much of which contains references to copyrighted characters or adult content. This result demonstrates the potentially disproportionate impact of small datasets combined with powerful alignment algorithms.