Search papers, labs, and topics across Lattice.
This paper critiques the prevailing epistemic practices in AI capability research, arguing that they inadequately address the unique challenges of AI safety and alignment, particularly under conditions of sparse evidence and high-stakes risk. It introduces the Epistemic Code for AI Safety and Alignment (ECAISA), which outlines eight principles and a structured framework aimed at improving the documentation and verification of safety-relevant research claims. The key finding is that by shifting focus from certification to auditability, ECAISA provides a more robust governance mechanism for ensuring AI systems do not lead to catastrophic failures.
ECAISA shifts the paradigm from certifying AI safety to ensuring rigorous auditability, addressing critical gaps in alignment research.
Mainstream AI research emphasises capability growth and tolerates low failure rates when average-case performance is high. AI safety and alignment research has a different mission: to ensure that catastrophic failures never occur, under sparse evidence, adversarial dynamics, and fat-tailed risk. We argue that the two domains differ along two analytically independent axes---{\it capability profile}, demonstrating the absence of hazardous behaviours rather than the presence of positive capabilities, and {\it risk profile}, bounding worst-case outcomes under fat-tailed uncertainty rather than optimising average-case performance---and that mainstream epistemic practices are inadequate on both. Building on a structured synthesis grounded in a preregistered bibliometric baseline, we identify five cross-cutting gap dimensions in current alignment research, including the near-absence of institutionalised independent verification. To address these gaps, we propose {\sc ECAISA}, an Epistemic Code for AI Safety and Alignment comprising eight principles, a three-level scoring rubric, a four-level disclosure ladder that reconciles transparency with information-hazard and commercial-confidentiality constraints, a tiered applicability scheme, an information-hazard adjudication procedure, and seven anti-gaming mechanisms. {\sc ECAISA} does not certify that any AI system is safe; it constrains how safety-relevant research claims are documented, checked, and relied upon, with auditability rather than certification as its governance target.