Search papers, labs, and topics across Lattice.
This paper introduces a six-attribute taxonomy for characterizing threat actors in the context of pre-release risk management for open-weight AI models. By grounding the taxonomy in empirical research from terrorism, biosecurity, and cybersecurity, the authors aim to provide a structured framework that enhances the interpretability and comparability of risk evaluations. The key finding emphasizes that explicit adversary characterization is essential for effective risk assessments, particularly for developers of irreversible AI systems.
Explicitly defining threat actors could transform how we assess risks associated with the release of open-weight AI models.
Pre-release risk management for frontier AI misuse risks routinely leaves threat actor assumptions implicit, inconsistently specified, or ungrounded. This capstone argues that explicit adversary characterization should be regarded as a prerequisite for evaluations that are interpretable, comparable, and faithful to the risks they target. We propose a six-attribute taxonomy (covering technical sophistication, prior domain knowledge, organizational capacity, operational infrastructure, financial capacity, and time horizon) with empirically grounded tiers derived from existing terrorism, biosecurity, and cybersecurity literature. The taxonomy is designed to function as research infrastructure: a common language for pre-specifying adversary assumptions before evaluations are conducted, analogous to pre-analysis plans for randomized controlled trials (RCTs) in medicine and economics. Its application is particularly urgent for open-weight model developers, for whom release decisions are irreversible and must anticipate adversarial reasoning.