Search papers, labs, and topics across Lattice.
This study compares span-guided and unguided detoxification methods for harmful content, revealing a nuanced trade-off in human preferences based on the severity of the harmfulness. Span-guided approaches are preferred when they effectively localize edits without altering the original intent, while unguided methods are favored for broader rewrites that achieve greater overall harm mitigation. The findings highlight the need for tailored evaluation protocols that separately assess the sufficiency of harm mitigation and the preservation of meaning, rather than relying solely on aggregate scores.
Span-guided detoxification may preserve intent but can leave harmful messages intact, while unguided methods risk over-modification鈥攈ighlighting a critical trade-off in content moderation strategies.
Span-guided rewriting aims to preserve meaning by localizing edits to annotated harmful spans, but the same constraint can leave harmful intent insufficiently mitigated. We present a controlled exploratory comparison of span-guided and unguided detoxification on a mixed-source English evaluation set comprising manually curated inputs and HateXplain test items. We conduct a dense blinded human evaluation under a fixed single-generator setting. Human preferences reveal a trade-off rather than a uniformly superior rewriting strategy. Span-guided outputs are favored when localized editing preserves the original stance and avoids unnecessary modification, whereas unguided outputs are favored when broader rewriting achieves more complete mitigation. This contrast varies substantially across the study-defined strata: the two strategies are competitive in the strong stratum, while unguided rewriting is clearly preferred in the mild stratum. Rationale annotations trace this difference to complementary failure risks: residual harm after localized editing and over-modification after broader rewriting. We treat automatic evaluation as a diagnostic rather than a substitute for human judgment. Toxicity-similarity scalarizations, a multi-generator analysis, and two general-purpose LLM judges reproduce parts of the aggregate tendency but do not yield an analogous stratified contrast. These setting-specific findings do not establish a severity-based routing rule. Instead, they motivate evaluation protocols that assess mitigation sufficiency and meaning preservation separately and report both residual harm and over-modification alongside aggregate scores.