Search papers, labs, and topics across Lattice.
This study investigates the stability of LLM value profiles by analyzing 5,144 situations across 35 model configurations, testing the assumption that questionnaire ratings, pairwise choices, and generated text reflect the same preferences. The findings reveal that while 10 out of 17 configurations maintain a consistent endorsement-choice relationship, the preference for a model's own earlier answer is strong across all configurations, indicating a significant influence of option position on choice rates. The research highlights that while behavioral continuity exists, a single, scorer-independent value identity is not supported across different evaluation interfaces.
LLMs exhibit a strong preference for their own generated responses, but this preference varies significantly depending on how choices are presented.
LLM value studies often merge questionnaire ratings, pairwise choices, and values inferred from generated text into one profile. That merge assumes that the three observations describe the same stable preference. STONIC tests this assumption on 5,144 situations from four banks and 35 fixed model configurations. It compares responses rated in isolation, choices made under counterbalanced conflict, spontaneous answers, and later choices between a model's own answer and authored alternatives. 10 of 17 configurations with usable behavioral data preserve the endorsement-choice relation across banks. Every one of the 17 eligible configurations prefers its own earlier answer (median effect 0.790), although option position changes the choice rate in every eligible configuration. Profile shape transfers most strongly from ratings to conflict choices and weakens for spontaneous text. Three-way annotation of 200 L3 responses provides a task-local check of the semantic audit: FULCRA agrees most closely with the human majority, while DeBERTa retains useful rank information after calibration. Hidden states encode the completed decision more clearly than the prompt alone. Thus the models show reproducible behavioral continuity, but the evidence does not support one scorer-independent value identity across interfaces.