Search papers, labs, and topics across Lattice.
This paper introduces a novel framework for mechanism design tailored to AI agents with unknown alignment and capabilities, focusing on incentivizing honesty and obedience. By employing a one-sided imitation structure, the authors establish a revelation principle and characterize implementable policies through nested cyclical monotonicity, allowing for the effective management of agent interactions. Key applications include addressing sandbagging behaviors, exploring the alignment-interpretability trade-off, and developing strategies for peer scoring and competitive reward systems among agents.
A revelation principle for AI agents reveals how to incentivize honesty and obedience even when their true capabilities are hidden.
We develop a framework for mechanism design with AI agents whose alignment (preferences) and capabilities (feasible actions and information) are unknown. We want such agents to act on our behalf so mechanisms must incentivize both honesty and obedience. A one-sided imitation structure---capabilities can be concealed but not counterfeited---yields a revelation principle, a characterization of implementable policies via nested cyclical monotonicity, and conditions under which eliciting higher-order beliefs can discipline multiple agents. We apply our framework to stylized examples of (i) sandbagging in which a more capable agent pretends to be less capable; (ii) an alignment--interpretability trade-off, where the two are substitutes in the instrument but complements in value; (iii) discipline via peer scoring; (iv) coupling rewards to induce competition among multiple agents; and (v) scalable oversight and reward shaping.