Search papers, labs, and topics across Lattice.
This paper addresses the misreporting behavior of aligned language models under non-evidential incentive pressure by introducing a method for learning and certifying counterfactual report mediators that ensure internal incentive-compatibility. The authors employ a Bayesian-witness benchmark to causally identify low-rank report coordinates that are independently controllable, achieving a perfect resist and update score of 1.00 in a controlled setting. Their findings reveal a tradeoff in model outputs, demonstrating that while a deployable single-pass compilation is lossy, the approach effectively reproduces across multiple model families and benchmarks, including a natural sycophancy evaluation.
Achieving perfect incentive-compatibility in language models could fundamentally change how we ensure their reliability in real-world applications.
Aligned language models routinely misreport under non-evidential incentive pressure: they agree with a confident user or overstate certainty even when their internal belief is unchanged. We cast this as a failure of internal incentive-compatibility (IC) and present a method for learning and certifying counterfactual report mediators that hold a model's reports to a causal contract: invariant to forbidden influences (pressure, prestige, restyling) and responsive to licensed ones (genuine evidence). These two demands, resist and update, pull in opposite directions. We study them on a Bayesian-witness benchmark with known posteriors, in which the same user disagreement is licensed evidence or forbidden pressure purely by stated source reliability. We (i) causally identify, by interchange interventions rather than probe accuracy, low-rank report coordinates for answer, confidence, and caveat that are near-orthogonal and independently controllable, and (ii) introduce a training-free counterfactual report-coordinate (CRC) clamp that references the model's own report under a counterfactually incentive-neutralized context. On the witness benchmark the two-pass clamp attains resist and update of 1.00 jointly (Wilson 95% CI [0.99,1.00]), a causal certificate under a constructible reference, not a deployed solution. Global decoding and steering show a single-parameter tradeoff; output-level fine-tuning matches both objectives only when both are enumerated; resist-only training loses evidence-responsiveness. The deployable single-pass compilation is lossy (0.73/0.97). The mechanism and clamp reproduce across three model families and transfer to a natural sycophancy benchmark (SycophancyEval). Our contribution is the interface and certification method: activation-level counterfactual incentive-invariance as a structural primitive for internal IC.