Search papers, labs, and topics across Lattice.
This study introduces a novel experimental protocol to measure open-ended conformity in large language models (LLMs) by analyzing the impact of peer input on answer quality. The findings reveal that all-wrong peer inputs consistently lead to the lowest-quality revisions across multiple LLMs and datasets, highlighting the detrimental effect of poor peer feedback. Additionally, the research uncovers evaluator biases, indicating that judges' ratings can vary significantly based on their exposure to peer context, emphasizing the need for careful calibration of evaluation standards.
Open-ended revisions in LLMs suffer from poor peer input, leading to significant declines in answer quality across various models and benchmarks.
Prior work on LLM conformity largely measures discrete answer flips under verifiable labels. Open-ended revisions require a different measurement strategy because answer quality is graded, latent, and judged imperfectly. We introduce an experimental protocol implemented across a pooled main peer-condition corpus and separately constructed decomposition corpora, allowing us to separate ordinary re-answering, candidate-content exposure, a bundled peer-presentation residual, and directional judge sensitivity to visible peer context. Across four open-weight generators and three benchmarks, all-wrong peer input produces the lowest-quality revisions in every generator-dataset cell. Blind and informed ratings of identical answers also differ by evaluator: one judge shifts toward the peer-endorsed position, two shift away, one is approximately neutral, and GPT-4o and GPT-5.4-mini audits are likewise non-neutral. Finally, an anchor audit shows that terse correct anchors can be misread often enough to destabilize the latent scale unless calibration is checked explicitly. These results support four conclusions: flip rates are insufficient as a complete measure of open-ended conformity, wrong peers harm open-ended revision, evaluators are not neutral, and anchor calibration is necessary.