Lessons from External Review of DeepMind’s Scheming Inability Safety Case

Safety cases for frontier AI systems should provide a convincing argument, supported by evidence, that the risk of harm is within an acceptable bound. When developers author their own safety cases, confirmation bias and conflicted incentives can affect the quality of argument. External review can help to address this.

In this paper, we apply the Assurance 2.0 framework to perform an external review of Google DeepMind’s public scheming inability safety case. We surface substantive new concerns that materially affect the scope of the safety case and its applicability for decision-making. Based on this experience, we provide concrete recommendations for how external review should be conducted and what information AI developers should provide to support it.

Expert Partner: Henry Papadatos (SaferAI)

Supporting expert advisor: Professor Robin Bloomfield (City St George’s, University of London) - author of ‘Assurance 2.0’, the method underpinning this project.

Alumni

Meet the authors

(Research Team Lead)

Steve Barrett

Steve is a Research Team Leader at Arcadia Impact and has previously worked for SaferAI on AI risk management.

He has worked in both automotive and enterprise cybersecurity as well as safety assurance in the automotive sector. He brings a strong track record in innovation and has spent 20+ years in research team leadership, standardization and systems engineering roles in the ICT sector.

Steve has an MBA and a PhD in communication engineering.

Javier Campos

Javier is Group CTO of Data & AI at CAPE.IO, bringing over two decades of technology leadership spanning Experian, WPP, and Accenture alongside hands-on alignment research at Cambridge AI Safety Hub.

His recent work includes co-authoring research on measuring misaligned behaviour in LLM-based agents, and contributing to EU AI Act and Bank of England AI governance frameworks. He combines deep technical fluency in agentic AI with first-hand regulatory experience to advance evidence-based AI safety governance.

Sean Fillingham

Sean is an AI safety researcher and former astrophysicist applying systems thinking and risk modeling frameworks to frontier AI governance. With a PhD in Physics and postdoctoral experience, he brings an empirical, systems thinking perspective to AI risk assessment.
He is currently a Research Associate at Arcadia Impact, where he is stress testing frontier AI safety cases and developing methodology recommendations grounded in the Assurance 2.0 framework, and a Research Fellow at SPAR characterizing loss of control scenarios.

Umair Siddique

Umair is a safety engineer with a PhD in formal methods, where his research focused on formalizing physics to enable machine-checked reasoning using theorem provers.

He works across autonomous systems including cars, mobile robots, and humanoids, developing safety analysis and verification methods.

His work spans industry and research, with contributions to safety architecture, risk modeling, and the development and interpretation of safety standards.

James Walpole

James has spent the last two years working on AI safety in the UK civil service, project managing alignment evaluations, helping develop testing strategy, and coordinating a joint testing exercise across an international network of AI safety institutes. Before that, he worked on animal welfare policy at Defra, where he designed financial incentive schemes for pig welfare.

He has co-authored research on AI control protocols and safety cases for AI manipulation risk. He has a Master's in Astrophysics from Cambridge.

Alumni

Meet the authors

Steve Barrett

(Research Team Lead)

Steve is a Research Team Leader at Arcadia Impact and has previously worked for SaferAI on AI risk management.

He has worked in both automotive and enterprise cybersecurity as well as safety assurance in the automotive sector. He brings a strong track record in innovation and has spent 20+ years in research team leadership, standardization and systems engineering roles in the ICT sector.

Steve has an MBA and a PhD in communication engineering.

Javier Campos

Javier is Group CTO of Data & AI at CAPE.IO, bringing over two decades of technology leadership spanning Experian, WPP, and Accenture alongside hands-on alignment research at Cambridge AI Safety Hub.

His recent work includes co-authoring research on measuring misaligned behaviour in LLM-based agents, and contributing to EU AI Act and Bank of England AI governance frameworks. He combines deep technical fluency in agentic AI with first-hand regulatory experience to advance evidence-based AI safety governance.

Sean Fillingham

Sean is an AI safety researcher and former astrophysicist applying systems thinking and risk modeling frameworks to frontier AI governance. With a PhD in Physics and postdoctoral experience, he brings an empirical, systems thinking perspective to AI risk assessment.
He is currently a Research Associate at Arcadia Impact, where he is stress testing frontier AI safety cases and developing methodology recommendations grounded in the Assurance 2.0 framework, and a Research Fellow at SPAR characterizing loss of control scenarios.

Umair Siddique

Umair is a safety engineer with a PhD in formal methods, where his research focused on formalizing physics to enable machine-checked reasoning using theorem provers.

He works across autonomous systems including cars, mobile robots, and humanoids, developing safety analysis and verification methods.

His work spans industry and research, with contributions to safety architecture, risk modeling, and the development and interpretation of safety standards.

James Walpole

James has spent the last two years working on AI safety in the UK civil service, project managing alignment evaluations, helping develop testing strategy, and coordinating a joint testing exercise across an international network of AI safety institutes. Before that, he worked on animal welfare policy at Defra, where he designed financial incentive schemes for pig welfare.

He has co-authored research on AI control protocols and safety cases for AI manipulation risk. He has a Master's in Astrophysics from Cambridge.

Programme

AI Governance Taskforce

The AI Governance Taskforce is a career development programme for experienced professionals looking to transition careers into AI governance, focussed on reducing risks from advanced AI.
Participants work around existing commitments during our 12 week, remote, part-time cohorts, producing policy research in teams of 4, led by our Research Team Lead staff in partnership with recognised experts in the field. Teams write an academic-style paper and accompanying blog post to build knowledge, skills and work portfolios.