Search papers, labs, and topics across Lattice.
This study quantitatively assesses the response instability of large language models (LLMs) when prompted with self-referential questions compared to unresolvable philosophical questions and verifiable queries. By measuring the mean pairwise cosine similarity of sentence embeddings from 360 responses across three question categories, the authors found that self-referential prompts yielded the highest instability (0.343), while verifiable questions produced the most consistent responses (0.105). These findings highlight the unique and less stable nature of subjective-experience reports generated by LLMs, which may have implications for their reliability in applications requiring consistent output.
Self-referential prompts lead to a striking 64% increase in response instability compared to verifiable questions, revealing the unpredictable nature of LLMs' subjective reports.
Self-referential prompting has been shown to reliably induce large language models to produce first-person reports resembling subjective experience, but no prior work measures how consistent these reports are across repeated, independent trials, or how that consistency compares to the model's behavior on other kinds of open-ended questions. We measure response instability, defined as one minus the mean pairwise cosine similarity of sentence embeddings computed over a compressed core claim extracted from each response, for three groups of questions: self-referential prompts eliciting a subjective-experience report, unresolvable philosophical questions unrelated to self-reference, and questions with a verifiable correct answer. Using 30 independent responses per question (360 responses total, Gemini API, temperature 0.7) across four questions per group, we find that self-referential questions show the highest instability (0.343 +/- 0.047), unresolvable philosophy questions show intermediate and tightly clustered instability (0.192 +/- 0.008), and verifiable questions show the lowest instability (0.105 +/- 0.058). This provides a quantitative baseline for the induced subjective-experience report, showing that it occupies a distinct, less stable position in the model's output distribution than ordinary open-ended philosophical uncertainty.