Stony BrookApr 6, 2026arXiv:2604.04743

Hallucination Basins: A Dynamic Framework for Understanding and Controlling LLM Hallucinations

AI Summary

This paper introduces a dynamical systems framework to analyze LLM hallucinations, modeling them as arising from task-dependent basin structures in the latent space. The authors analyze hidden-state trajectories across various models and benchmarks, finding that basin separability varies significantly with task complexity, with factoid tasks exhibiting clearer separation than summarization or misconception-heavy tasks. They further demonstrate that geometry-aware steering can mitigate hallucination probability without requiring model retraining.

Key Contribution

LLM hallucinations aren't random errors, but predictable consequences of overlapping "hallucination basins" in the model's latent space, offering a new path to control them.

Abstract

Large language models (LLMs) hallucinate: they produce fluent outputs that are factually incorrect. We present a geometric dynamical systems framework in which hallucinations arise from task-dependent basin structure in latent space. Using autoregressive hidden-state trajectories across multiple open-source models and benchmarks, we find that separability is strongly task-dependent rather than universal: factoid settings can show clearer basin separation, whereas summarization and misconception-heavy settings are typically less stable and often overlap. We formalize this behavior with task-complexity and multi-basin theorems, characterize basin emergence in L-layer transformers, and show that geometry-aware steering can reduce hallucination probability without retraining.

Eval Frameworks & Benchmarks Interpretability & Mechanistic Interp Natural Language Processing

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Hallucination Basins: A Dynamic Framework for Understanding and Controlling LLM Hallucinations

Related Papers