Search papers, labs, and topics across Lattice.
To unify safety auditing across disparate conditioning interfaces, this work formalizes diffusion model governance through a probabilistic framework that maps concept generation to Bernoulli semantic events via a Concept Risk Operator. This addresses the blind spots of heuristic or task-specific erasure benchmarks that fail to rigorously compare risks across varied conditioning channels and prompt distributions. Evaluations across Stable Diffusion 1.5, 2.1, and SDXL reveal that embedding-level conditioning systematically exposes safety violations missed by standard text prompts, while uncalibrated probability estimates flip policy enforcement actions near regulatory thresholds.
Standard text-prompt evaluations severely underestimate diffusion safety risks, which reliably resurface under embedding-level control and trigger policy enforcement failures unless probabilities are explicitly calibrated.
Diffusion models have become a core paradigm for multimedia generation, offering powerful concept-driven controllability for personalization, semantic editing, and selective unlearning. However, as semantic control extends beyond natural-language prompts to learned embeddings and intervention pipelines, the safety and governance of these systems become increasingly difficult to evaluate in a unified manner, especially for safety-sensitive, identity-linked, and other privacy-relevant concepts. Existing studies mainly rely on heuristic audits, adversarial probing, or task-specific erasure benchmarks, and therefore provide limited support for systematic comparison across models, conditioning channels, and deployment conditions. We present a concept-level probabilistic audit and reporting framework for diffusion models. We formalize governance-relevant concept behaviors as Bernoulli semantic events induced by stochastic generation, and define a Concept Risk Operator that maps model-channel configurations to structured risk profiles, enabling comparison across prompting interfaces, learned embedding channels, models, and recorded conditions. We apply sample-level post-hoc calibration and configuration-level risk aggregation, and show that probability error can change thresholded actions near policy boundaries. Experiments on SD1.5, SD2.1, and SDXL reveal consistent yet non-uniform operational risk patterns across concept families, channels, recorded conditions, and shifted protocols. In particular, embedding-based access and obfuscated prompts expose risks often understated by standard-prompt evaluation. A pooled multi-protocol calibrator improves held-out probability reliability, but we do not claim transfer from a standard-only calibrator. CLRC provides a common audit schema for probabilistic and decision-aware governance of multimedia generation systems.