Testing and Evaluation of Agentic AI Systems in Military Command and Control

AI Governance Taskforce

Summer 2026

Agentic AI systems are being procured for military command and control (C2) under public commitments to rigorous testing and human oversight. Whether such commitments can be discharged depends on their supporting assurance case, which requires three elements: claims specifying the conditions for acceptability, evidence bearing on those claims, and an argument connecting the two. Through a structured review of 240 documented Testing and Evaluation (T&E) practices, spanning eight evaluation dimensions and three lifecycle phases, we identify eight assumptions that established methods make about their test article, grouped into four clusters: system specifiability, stability, composability, and supervisability. Agentic properties weaken all eight assumptions. This erosion affects the argument connecting evidence to claims, not the claims or evidence themselves. As a result, test results may satisfy process requirements, but they do not warrant the inference from tested to fielded behavior. We derive ten assurance claims for the first three assumption clusters and assess whether current and emerging methods can address each, mapping operational consequences through five C2 scenarios. Supervisability is identified but not assessed here, since evidencing it depends on system stability results and human factors T&E methods beyond the present scope. The documented record does not support broad claims about system-level behavior, but narrower claims remain recoverable: bounded mission envelopes, trajectory-grounded correctness, executable runtime constraints, and characterized run-to-run variance. Part of the evidentiary burden shifts into deployment, making the determination to field a continuing act. Where evidence cannot be generated, the residual uncertainty can be governed through defined expiry conditions and assigned ownership.

Alumni

Meet the authors

(Research Team Lead)

Ulysse Richard

Ulysse is a policy practitioner and researcher helping shape international governance frameworks to meet the security challenges posed by frontier AI. He leads high-level dialogues on military AI governance at the UN Office for Disarmament Affairs, supports Track II P5 dialogues on AI-nuclear risk at INHR, and researches AI Safety Institutes collaboration to advance capability red lines with Arcadia Impact.

Previously, he worked across technology policy, cyber threat intelligence, and nuclear cybersecurity. He holds a dual MA in International Security from Sciences Po and Peking University.

Sarah Cao

Sarah Cao is a researcher working at the intersection of AI governance and international security. She is pursuing an MPhil in International Relations at the University of Oxford, is a Fellow at the Oxford China Policy Lab, and is a researcher on Arcadia Impact's AI Governance Taskforce. She brings experience in multilateral security cooperation and defense research, and holds degrees in International Affairs and Chinese.

Di Cooke

Di is an AI governance expert, with a focus on AI assurance and risk mitigation approaches as well as responsible AI operationalisation in the defence and intelligence domains.

She is a UKRA AI and Cyber Associate in the UK Cabinet Office's Resilience Directorate, a Non-Resident Fellow at CSIS, and a doctoral candidate at King's College London.

Sebastian Kwon

Sebastian currently supports the Department of the Air Force Chief Data and AI Officer as AI/Data Readiness Lead and is a nonresident fellow in the Indo-Pacific Security Initiative at the Atlantic Council’s Scowcroft Center for Strategy and Security.

He is a recently retired US Air Force cyberspace effects operations officer with 20 years of experience across IT, DevSecOps, cyber warfare, future concepts, defense innovation, and AI/ML.

Adrianna Tan

Adrianna is an AI researcher working on understanding what it means for AI to go from demo to pilot and deployment and beyond.

Previously, she was AI Red Team lead at Humane Intelligence (founded by Dr Rumman Chowdhury) where she led the world's first participatory defense medicine red teaming project with the Department of Defense. She was also the project lead on IMDA Singapore's world's first multilingual, multicultural red team evaluation and benchmarking project involving 9 countries in Asia Pacific.

She speaks and lectures frequently on AI Safety and Security with international and US federal and state government leaders, and also serves as a AI testing and evaluations consultant with the State of Maryland's AI Innovation Lab.

Alumni

Meet the authors

Ulysse Richard

(Research Team Lead)

Ulysse is a policy practitioner and researcher helping shape international governance frameworks to meet the security challenges posed by frontier AI. He leads high-level dialogues on military AI governance at the UN Office for Disarmament Affairs, supports Track II P5 dialogues on AI-nuclear risk at INHR, and researches AI Safety Institutes collaboration to advance capability red lines with Arcadia Impact.

Previously, he worked across technology policy, cyber threat intelligence, and nuclear cybersecurity. He holds a dual MA in International Security from Sciences Po and Peking University.

Sarah Cao

Sarah Cao is a researcher working at the intersection of AI governance and international security. She is pursuing an MPhil in International Relations at the University of Oxford, is a Fellow at the Oxford China Policy Lab, and is a researcher on Arcadia Impact's AI Governance Taskforce. She brings experience in multilateral security cooperation and defense research, and holds degrees in International Affairs and Chinese.

Di Cooke

Di is an AI governance expert, with a focus on AI assurance and risk mitigation approaches as well as responsible AI operationalisation in the defence and intelligence domains.

She is a UKRA AI and Cyber Associate in the UK Cabinet Office's Resilience Directorate, a Non-Resident Fellow at CSIS, and a doctoral candidate at King's College London.

Sebastian Kwon

Sebastian currently supports the Department of the Air Force Chief Data and AI Officer as AI/Data Readiness Lead and is a nonresident fellow in the Indo-Pacific Security Initiative at the Atlantic Council’s Scowcroft Center for Strategy and Security.

He is a recently retired US Air Force cyberspace effects operations officer with 20 years of experience across IT, DevSecOps, cyber warfare, future concepts, defense innovation, and AI/ML.

Adrianna Tan

Adrianna is an AI researcher working on understanding what it means for AI to go from demo to pilot and deployment and beyond.

Previously, she was AI Red Team lead at Humane Intelligence (founded by Dr Rumman Chowdhury) where she led the world's first participatory defense medicine red teaming project with the Department of Defense. She was also the project lead on IMDA Singapore's world's first multilingual, multicultural red team evaluation and benchmarking project involving 9 countries in Asia Pacific.

She speaks and lectures frequently on AI Safety and Security with international and US federal and state government leaders, and also serves as a AI testing and evaluations consultant with the State of Maryland's AI Innovation Lab.

Programme

AI Governance Taskforce

The AI Governance Taskforce is a career development programme for experienced professionals looking to transition careers into AI governance, focussed on reducing risks from advanced AI.
Participants work around existing commitments during our 12 week, remote, part-time cohorts, producing policy research in teams of 4, led by our Research Team Lead staff in partnership with recognised experts in the field. Teams write an academic-style paper and accompanying blog post to build knowledge, skills and work portfolios.