Bletchley Declaration on AI Safety
UK Government & 28 Nations — 2023
Landmark international agreement recognizing the potential for serious, even catastrophic, harms from frontier AI and committing to collaborative safety research and governance.
KILLSwitch Library
Foundational policy documents, academic research, and governance frameworks on AI safety and human oversight — curated for policymakers, researchers, and decision-makers.
UK Government & 28 Nations — 2023
Landmark international agreement recognizing the potential for serious, even catastrophic, harms from frontier AI and committing to collaborative safety research and governance.
European Parliament — 2024
The world's first comprehensive legal framework for artificial intelligence, establishing risk-based obligations for AI systems deployed in the European Union.
The White House — 2023
U.S. federal directive establishing safety standards, evaluation requirements, and oversight mechanisms for advanced AI systems developed or deployed in the United States.
G7 — 2023
G7 framework establishing guiding principles and a code of conduct for advanced AI developers, emphasizing transparency, accountability, and human oversight.
Amodei, Olah, et al. — OpenAI & Google Brain — 2016
Foundational paper identifying five core technical problems in AI safety: avoiding negative side effects, avoiding reward hacking, scalable oversight, safe exploration, and robustness.
Paul Christiano — ARC — 2022
Overview of the alignment problem and why current machine learning systems may pursue goals misaligned with human values as they become more capable.
Anthropic — 2022
Research introducing a method for training AI systems to be helpful, harmless, and honest using a set of principles — a step toward scalable human oversight.
Hubinger et al. — 2019
Analysis of mesa-optimization and inner alignment — the risk that learned models may develop internal objectives diverging from the objectives they were trained to pursue.
Nathan Benaich & Air Street Capital — 2024
Annual comprehensive review of the most significant developments in AI research, industry, politics, and safety — widely read by policymakers, researchers, and executives.
Stanford HAI — 2024
Stanford's annual data-driven analysis of AI progress, investment, policy, and societal impact — the most cited empirical benchmark for AI development trends.
UK Department for Science, Innovation and Technology — 2023
Policy paper outlining the UK's approach to regulating frontier AI systems, including proposed evaluation frameworks and international coordination mechanisms.
National Institute of Standards and Technology — 2023
Voluntary framework for organizations to manage AI risks across four core functions: Govern, Map, Measure, and Manage — widely adopted across U.S. federal agencies and industry.
DeepMind Safety Team — 2023
Framework for evaluating whether AI models exhibit dangerous capabilities including CBRN uplift, cyberoffense, and deceptive alignment — informing responsible deployment decisions.
Anthropic — 2023
Commitment framework tying AI capability development to safety evaluations — models may only be deployed if they pass defined safety thresholds at each capability level.
Contribute
KILLSwitch maintains this library as a public reference for anyone working on AI governance, safety policy, or human oversight. Submit a recommendation via the contact channel.