Search papers, labs, and topics across Lattice.
This study evaluates the criterion-conditioned behavior of large language models (LLMs) in content moderation using a new method called Diagnostic Evaluation of COntent (DECO). By applying pairwise evaluation across four datasets and four LLMs, the authors reveal that high performance on standard benchmarks often conceals significant failures in applying individual moderation criteria. The findings indicate that LLMs struggle particularly when decisions hinge on specific content aspects rather than overall harmfulness, underscoring the inadequacy of current aggregated evaluation methods.
Strong benchmark scores can mask critical failures in LLMs' ability to apply individual content moderation criteria effectively.
Large language models (LLMs) demonstrate strong performance on standard content moderation benchmarks. However, these benchmarks often aggregate multiple moderation criteria into a single label, making it unclear whether models can disentangle them and reliably apply each criterion when making decisions. To study whether LLMs exhibit criterion-conditioned behaviour, we introduce Diagnostic Evaluation of COntent (DECO), a criterion-independent factorisation of content that enables controlled, criterion-level evaluation. We also introduce pairwise evaluation to compare model outputs across different criteria for the same input. Across four moderation datasets and four LLMs, we find that strong benchmark performance can hide substantial failures at the criterion level. Models struggle most when correct decisions depend not on overall harmfulness, but on the specific aspect of the content that the criterion requires them to assess. Our results highlight a key limitation of current content moderation benchmarks: strong performance on aggregated labels does not provide sufficient evidence that LLMs can reliably evaluate content with respect to individual moderation criteria. These findings call for the development of evaluation methods that explicitly measure criterion-conditioned behaviour.