Search papers, labs, and topics across Lattice.
This paper introduces Reward-Optimized Probe-and-Respond (RO-PnR), a decision framework designed to effectively intervene in multi-turn dialogues addressing health misinformation. By weighing the costs and benefits of probing for user information against providing immediate corrections, RO-PnR adapts to user heterogeneity in health literacy and belief commitment. Experimental results demonstrate that RO-PnR achieves superior cost-adjusted utility, requiring 30% fewer turns compared to traditional always-probe methods across multiple datasets and models.
Probing for user information can be costly, but RO-PnR shows that strategic questioning can significantly enhance the effectiveness of health misinformation interventions.
Correcting health misinformation in dialogue requires more than producing a factual rebuttal: users differ in what they know, what they believe, and what they need to hear, so an effective intervention often depends on first asking the right clarifying question. Yet existing methods either respond immediately or probe indiscriminately, treating clarification as either unnecessary or always beneficial. We propose Reward-Optimized Probe-and-Respond (RO-PnR), a framework that learns when asking is worth its cost. At each turn, RO-PnR chooses between probing for more information and committing to a final correction, guided by a turn-level reward that weighs the expected gain from probing against its interaction cost. To capture how user heterogeneity affects probing value, we model each simulated user with a latent state along health literacy and belief commitment. Experiments show that RO-PnR achieves the highest cost-adjusted utility across three health-misinformation datasets and three base models, using 30% fewer turns than always-probe baselines.