Search papers, labs, and topics across Lattice.
This paper introduces a structured framework for identifying behavioral indicators that signal the progression of artificial intelligence systems towards catastrophic risks. By leveraging methodologies from cybersecurity and national security, the authors establish clear metrics and thresholds to facilitate systematic monitoring of AI capabilities and behaviors. The key result is a pragmatic approach that empowers researchers and policymakers to implement evidence-based protocols for risk assessment and management in AI development.
A structured framework for monitoring AI progression could be the key to preventing catastrophic risks before they escalate.
This article presents a structured framework of behavioral indicators that may signal progression toward potentially catastrophic threats from artificial intelligence systems. We adopt a pragmatic approach, inspired by established methodologies in cybersecurity and national security. By establishing clear metrics, indicators, and thresholds across multiple dimensions of AI capability and behavior, this framework enables researchers and policymakers to implement evidence-based monitoring protocols.