Search papers, labs, and topics across Lattice.
This paper introduces FIRE, a multi-agent reasoning framework that categorizes hate speech into five distinct types and generates tailored counterspeech strategies accordingly. By leveraging a novel dataset, FactualCS, which includes over 4,700 annotated instances, the framework significantly improves factual and category-specific accuracy by approximately 12% and 11%, respectively, while also reducing toxicity by about 11%. Comprehensive evaluations indicate that FIRE outperforms existing methods, highlighting the importance of understanding the nuances of hate speech for effective counterspeech generation.
Tailoring counterspeech to specific hate speech categories can enhance effectiveness and reduce toxicity, achieving significant improvements over traditional methods.
Counterspeech effectively neutralizes the impact of online hate. Although prior work explores automated counterspeech generation, it largely emphasizes stylistic control while treating hate speech as homogeneous, overlooking that distinct forms of abuse require fundamentally different counterspeech strategies. To address this gap, we introduce FIRE (Factuality Informed Multi-Agent Reasoning Framework) that first decomposes hate speech into one of the five distinct categories (misinformation, stereotype, conspiracy, dehumanizing, non-factual), and then maps it to a targeted counterspeech style. To facilitate FIRE, we curate FactualCS, a novel dataset of $4,784$ instances that provides the annotations regarding hate categories, reasoning traces, and evidence mappings, which are critical elements for grounded generation that are missing in prior work. A comprehensive evaluation across $28$ baseline configurations demonstrates that FIRE significantly surpasses existing methods, despite using compact agents ($<$2B). FIRE achieves a $\sim$ $12 \%$ and $\sim$ $11 \%$ improvements in factual and category-specific accuracy respectively, while simultaneously reducing toxicity by $\sim$ $11 \%$ relative to the strongest baselines. Further human evaluation confirms that responses generated by FIRE are significantly preferred over the strongest baselines, underscoring its effectiveness for real-world deployment. These findings show that decomposing the underlying intent of hate speech is essential for generating safe, effective, and contextually precise counterspeech.