Search papers, labs, and topics across Lattice.
This study evaluates the effectiveness of contextualized counterspeech generated by AI in mitigating online toxicity, contrasting it with generic approaches that fail to consider user context. By integrating various forms of contextual information and fine-tuning techniques, the authors conducted a mixed-design crowdsourcing experiment to assess the persuasiveness of these personalized interventions. The findings reveal that while personalization can enhance perceived adequacy and persuasiveness, not all contextualization strategies yield positive results, highlighting the need for careful design in counterspeech systems.
Contextualized AI-generated counterspeech can significantly enhance persuasiveness, but not all personalization strategies are equally effective.
AI-generated counterspeech offers a scalable and effective strategy to mitigate online toxicity by promoting more constructive dialogue. Yet, existing approaches adopt a generic, one-size-fits-all paradigm, overlooking the conversational context and characteristics of the targeted users. Here, we propose and evaluate multiple strategies for generating contextualized counterspeech that is adapted to the moderation setting and personalized to the moderated user. In detail, we explore a range of configurations that integrate different forms of contextual information and fine-tuning techniques. We conduct a comprehensive evaluation combining quantitative indicators with a pre-registered, mixed-design crowdsourcing experiment. To ensure robustness, we implement algorithmic measures of counterspeech quality based on ROUGE, BLEU, and BERTScore, observing overall consistent results across metrics. Furthermore, we analyze which characteristics of both the generated counterspeech and the moderated toxic message most strongly influence perceived persuasiveness, yielding insights into how contextualized interventions can be made more effective. Our findings show that personalization can be effective, but not uniformly so. Lightweight strategies combining conversational context and user history improve perceived adequacy and persuasiveness, whereas several other contextualization strategies degrade human-perceived counterspeech quality. Taken together, these results provide actionable directions for developing more personalized, effective, and responsible counterspeech systems, ultimately advancing human-AI collaboration in online content moderation.