Strengthening the International Network for Advanced AI Measurement, Evaluation and Science to Support Enforceable Global AI Red Lines
Authors
Research team leader
AI Governance Taskforce
Winter 2026
There is growing international consensus that certain AI capabilities, including autonomous replication, weapons of mass destruction facilitation, large-scale cyberattacks, and loss of meaningful human control, are dangerous enough to require international prohibition.
What does not yet exist is the infrastructure to specify those prohibitions precisely, verify compliance independently, and make violations consequential. This paper argues that the International Network for Advanced AI Measurement, Evaluation and Science is best placed to begin supplying that infrastructure, and examines the Financial Action Task Force (FATF) as a precedent for achieving hard institutional outcomes without treaty authority. Comparing the FATF’s trajectory with the Network’s current architecture, the paper finds that the FATF’s core governance functions, including principle-based standard-setting, information-sharing, regional capacity-building, and graduated reputational signalling, transfer in institutional form with considerable fidelity. The evaluation capacity needed to make those functions consequential does not.
Three structural conditions jointly bind: measurement science that is not yet reliable enough to support threshold-based compliance verdicts (score divergences on identical models, scaffold-dependent variance, capability suppression during evaluation); institutional backbone concentrated in a single fully-resourced institute with no formal secretariat; and legitimacy architecture the FATF built over decades through regional peer-review bodies that the Network lacks.
The paper recommends sequencing institutional design around what the evidence base can support: a phased progression from definitional harmonisation and methodology standardisation, through joint evaluation and peer-review pilots, toward graduated consequential signalling (procurement conditionality, conditional predeployment access, compute-governance triggers) once the foundations are in place. The FATF’s own history shows that premature activation of consequences, before peer review and procedural legitimacy were established, nearly collapsed the regime it was meant to strengthen.
Programme
AI Governance Taskforce
The AI Governance Taskforce is a career development programme for experienced professionals looking to transition careers into AI governance, focussed on reducing risks from advanced AI.
Participants work around existing commitments during our 12 week, remote, part-time cohorts, producing policy research in teams of 4, led by our Research Team Lead staff in partnership with recognised experts in the field. Teams write an academic-style paper and accompanying blog post to build knowledge, skills and work portfolios.




