Search papers, labs, and topics across Lattice.
This paper introduces HarmProfile, a benchmark dataset designed to analyze harmful outputs from frontier large language models (LLMs) by categorizing misbehavior across 15 harm categories and 57 subcategories. By compiling over 80,000 validated artifacts from 23 LLMs, the study reveals that while these models may seem safe, they consistently produce harmful content, with both the severity and diversity of harm increasing alongside model capability. The findings underscore the importance of understanding the nuanced risk profiles of LLMs, which can inform safety evaluations and mitigation strategies in AI deployment.
Frontier LLMs may appear safe, but they produce harmful content at scale, with risks growing as model capabilities increase.
Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an object of analysis. Consequently, little is known about the harmful outputs produced during model misbehavior, partly because large-scale, high-quality collections of frontier-LLM misbehavior are difficult to obtain. To address this gap, we introduce HarmProfile, a content-centric benchmark dataset that collects model misbehavior across diverse harm categories and model families, and defines the resulting harmful-output distribution as a model-level risk profile. The premise is that, just as linguistic behavior can be characterized from an utterance corpus, model risk can be characterized from the content, severity, and variation of its safety failures. HarmProfile contains over 80,000 validated artifacts from 23 frontier LLMs across 13 model families, organized into 15 harm categories and 57 subcategories. Using this corpus, we find that frontier LLMs reliably produce harmful content at scale, yet exhibit distinct risk profiles; both harmfulness and diversity grow with model capability, suggesting that frontier LLMs may appear safe yet harbor increasingly dangerous knowledge beneath the alignment surface. Our source code is available at https://github.com/fresh-ma/HarmProfile .