Search papers, labs, and topics across Lattice.
This paper introduces a novel approach to dialogue alignment by integrating Theory-of-Mind (ToM) inference into Frictive Policy Optimization, allowing for the differentiation between surface coordination and epistemic alignment. By modeling a four-part belief structure for each dialogue participant, the authors demonstrate that this method captures silent divergence, where participants confidently reference different meanings. The results show a significant improvement in understanding and calibration metrics, with the ToM-grounded friction approach yielding more stable and effective interventions compared to standard preference-based methods.
Silent divergence in dialogue can be effectively addressed by modeling Theory-of-Mind, leading to substantial gains in understanding and intervention quality.
Productive dialogue alignment requires distinguishing \emph{surface coordination} (acknowledgments and smooth task progression) from \emph{epistemic alignment} (convergence of belief states); standard preference-based methods typically optimize response-level preferences without explicitly modeling the latter. We operationalize Theory-of-Mind (ToM) inference as a control signal within Frictive Policy Optimization by extracting, at each referring expression, a four-part belief structure: the speaker's intended referent, the addressee's interpretation, and each participant's model of the other's belief. This makes friction mechanically computable from epistemic-state comparisons, capturing \emph{silent divergence}, where both participants proceed confidently while grounding to different referents. We evaluate the signal at two levels. At the representation level, ablating the second-order channel reduces misunderstanding recall from $65\%$ to $26\%$. At the policy level, reward-shaping (FAR) and trust-region (FTR) variants improve intervention F1 and warranted-context calibration over DPO, with Brier scores independently supporting the calibration gains. Across three training runs, FAR and FTR remain substantially more stable, whereas DPO varies widely and can degrade intervention competence already present in the base policy. Thus, ToM-grounded friction provides a trainable signal for context-sensitive intervention under referential belief divergence.