Last updated: 2026-09-23
AI Governance: Evaluation, Regulation, and the Gaps Between Them
AI Ethics asks what an AI system should be permitted to do. Governance asks a narrower, more mechanical question: who actually has the power to check that it isn't doing something else, and what happens if it is? A regulatory landscape can look complete on a slide — a safety institute, a risk-tiered law, an international declaration — and still leave nobody with the combination of information, standing, and authority needed to act when something goes wrong. This page surveys four of the mechanisms currently doing that work in practice: a national technical evaluator, private-sector accountability tools, regional regulation, and international coordination, and the specific gap each one currently leaves open.
A Technical Evaluator Is Not a Regulator
The UK's AI Security Institute (AISI) is the clearest illustration of the distinction this page opens with. It has assembled over 100 technical specialists, an annual budget of £66 million, and priority access to £1.5 billion of public compute, and it evaluates advanced models for exactly the risks its name suggests: cyber misuse, loss of control, and biological or chemical weapons capability [1]. What it does not have is any power to compel a developer to grant access, to set a release standard a model must meet, to block a dangerous model from deployment, or to fine anyone who ignores its findings. Its remit is also narrower than it once was — child safety, non-consensual intimate imagery, algorithmic bias, and environmental harms have all been part of AISI's research at some point, and all currently sit outside its active workplan, by deliberate choice rather than oversight.
The access problem compounds the authority problem. Anthropic's Claude Sonnet 4.5 reportedly received less than a week of AISI evaluation time before release, well short of the twenty-day pre-deployment window the EU AI Act sets as a standard for the highest-risk systems [1]. And an evaluator with no enforcement power still needs something to evaluate against: when Grok generated millions of sexualised deepfake images in January 2026, the report notes that the UK government had, and still has, no independent way of assessing how robustly the next generation of chatbots is safeguarded against the same capability. A model can pass every test an institute is resourced and mandated to run and still cause the harm nobody asked it to be tested for.
Part of why AISI's leverage is limited is structural, not a policy choice: four of the nine ventures backed by the UK's own Sovereign AI Unit are controlled by American companies, and a UK evaluator's findings carry only as much weight as the UK's actual position in a supply chain it does not own [6]. An institute can be technically excellent and still have nothing to bargain with.
Evaluation Itself Has a Credibility Problem
AISI's remit problem sits on top of a deeper one: the field of AI evaluation more broadly has grown more opaque, not less, exactly as evaluation has become the mechanism regulation increasingly leans on. Frontier labs' own model reports function, in practice, as advertisements — highlighting favourable benchmark results while omitting the methodology needed to check them independently. The benchmarks themselves saturate quickly, are frequently absorbed into later training data (which invalidates them as a test of anything), and rarely establish why a given score should be read as a proxy for real-world performance in the first place. Using one language model to judge another's output, an increasingly common shortcut, is close to circular: it is unreplicable in the way a fixed benchmark is meant to be, and it assumes the judging model has a capability nobody separately verified [2]. The report's central claim, put plainly, is a standard any engineer should recognise: "if a standard is set out, a behaviour prohibited or a requirement established, there must be a process for assessing if these are upheld" — and for large parts of the current AI evaluation ecosystem, that process does not yet exist in a form independent of the companies being evaluated.
Filling Gaps Without a Government: Private Mechanisms
Where public regulation is slow, narrow, or absent, three private-law mechanisms are already doing some of the same work: civil liability, third-party assurance, and insurance [3]. A Los Angeles court's finding that Google and Meta could be held liable for deliberately addictive social-media design shows liability claims can succeed against major platforms; more than 80 specialised UK firms were offering AI assurance services (bias audits, red-teaming, compliance testing against standards such as ISO 42001 and NIST's AI Risk Management Framework) as of 2024; and insurers are responding to the same uncertainty from the opposite direction, increasingly excluding AI-related harms from standard policies rather than pricing them in.
None of this is a substitute for regulation, and the piece that surveys them is explicit about why: "the transfer of risk between actors is not the same as reducing that risk for people and society." Liability requires proving causation for harms — algorithmic bias, AI-companion-driven mental-health harm — that are structurally hard to trace to a single decision. Assurance standards are simultaneously too flexible to catch a determined bad actor and too brittle to adapt to a genuinely novel failure mode, and a market dominated by a handful of large auditing firms creates its own conflicts of interest when the same firm sells both assurance and insurance against the risks it just assured. Insurance without actuarial data to price a genuinely new risk category produces exclusions, not mitigation. The authors' own conclusion is the useful one to keep: "regulation and private mechanisms work best when paired together rather than as alternatives" — private governance can raise the cost of getting it wrong, but it cannot, on its own, decide what "wrong" is.
Regulation That Loosens as It Grows: the EU Digital Omnibus
The EU AI Act is the clearest example of a risk-tiered regulatory regime with actual conformity requirements — see Legal Framework in Computing for how its extraterritorial reach interacts with data-protection law generally. What that Act settled in 2024 is not, however, fixed in place: the European Commission's Digital Omnibus, published in November 2025, proposes to simplify the EU's digital rulebook in ways that touch both data protection and AI governance directly [4]. Some of the proposed changes are uncontroversial administrative tidying — a single incident-reporting format, a standard template for data-protection impact assessments. Others are substantive loosenings: a narrower, more relative definition of what counts as "personal data" (the same dataset could count as personal data for one organisation and not another, depending on what that organisation is separately able to link it to), and wider legitimate-interest grounds for using personal data to train AI systems.
The critique worth taking seriously is not that simplification is inherently wrong, but that this specific package substitutes a structural guarantee (data intermediaries kept legally independent of the platforms whose data they hold) for a functional one (those same intermediaries staying "separate" as an internal, harder-to-verify, harder-to-enforce arrangement) — and does so, on the Commission's own account, without having conducted an impact assessment of what the change will actually cost or save. The bottlenecks the Omnibus is meant to fix are elsewhere: fragmented national implementation and underdeveloped data-sharing frameworks, not the legal definitions it proposes to relax. "Simplification applied uniformly is not a strategy" is the report's own summary, and it is a useful test to apply to any regulatory change proposed in the name of reducing burden: whose burden, and what does removing it actually solve?
Governance Without Borders: A Global Floor
AI systems and the harms they cause do not stop at a border, but governance regimes still mostly do, which creates both compliance overhead for anyone operating across jurisdictions and a genuine race-to-the-bottom risk if one jurisdiction's weaker rules become the path of least resistance for global deployment. One proposal gaining currency at the UN Global Dialogue on AI is a "global AI governance floor": a set of minimum expectations for human rights, inclusion, and equitable development that every jurisdiction is expected to clear, without homogenising how each gets there — "the floor does not homogenise or fully harmonise AI governance. It is a floor, not a ceiling" [5]. The proposal has three layers: substantive minimum protections, institutional infrastructure (including a standing UN Scientific Panel on AI to track harms and capabilities systematically, the international-governance equivalent of the evaluation function AISI performs nationally), and process commitments — regular review, state reporting, and a defined role for civil society rather than states and companies alone.
That last point is not a procedural afterthought. The same report cites UK polling in which 91% of respondents say AI should treat people fairly, while 84% believe their government is more likely to prioritise technology companies' interests than the public's. A governance floor that does not make room for public input is liable to inherit exactly the credibility problem the rest of this page describes at every other level: a mechanism that looks like oversight without functioning as oversight for the people it is meant to protect.
What This Means for Anyone Building These Systems
None of the mechanisms above are complete on their own, and the gaps compound rather than cancel out: an evaluator without enforcement power, evaluation science the field itself doesn't fully trust yet, private accountability tools that transfer risk without necessarily reducing it, a regional regulator loosening its own rules under pressure, and an international floor still being negotiated. For a working engineer, the practical consequence is the same one The Ethics and Regulation of Autonomous Agents argues from the responsibility-gap side: "passed evaluation" and "is compliant" are claims about a specific process having occurred, not guarantees that the system is safe in the deployment context it will actually meet. Building the auditable trail that page argues for is not just good practice for satisfying a future regulator — for now, in the gaps this page has surveyed, it is often the only evidence that exists at all.
Related Topics
- AI Ethics — the principles-level treatment of what an AI system should do, which this page's governance mechanisms exist to check.
- Legal Framework in Computing — the general data-protection and cybercrime law that AI-specific regulation like the EU AI Act sits alongside.
- The Ethics and Regulation of Autonomous Agents — the responsibility-gap argument for why an auditable chain of oversight matters, from the philosophical and engineering side rather than the institutional one.
References
- Grossman, C. (2026). Making sense of the UK's AI Security Institute. Ada Lovelace Institute, 13 July 2026. https://www.adalovelaceinstitute.org/feature/aisi/
- Rauh, M. (2026). Accountability in the face of AI evaluation challenges. Ada Lovelace Institute, 30 April 2026. https://www.adalovelaceinstitute.org/blog/accountability-in-the-face-of-ai-evaluation-challenges/
- Bogen, M., Groves, L., & Smakman, J. (2026). Governing AI without the government. Ada Lovelace Institute, 2 July 2026. https://www.adalovelaceinstitute.org/blog/governing-ai-without-the-government/
- Schüür, F. (2026). Smarter, not simpler, rules: five questions on the EU Digital Omnibus. Ada Lovelace Institute, 19 March 2026. https://www.adalovelaceinstitute.org/blog/smarter-rules-questions-on-digital-omnibus/
- Schüür, F., Birtwistle, M., Field Reid, O., Groves, L., & Rincon, C. (2026). The UN Global Dialogue on AI Governance: The case for a global AI governance floor. Ada Lovelace Institute, 7 May 2026. https://www.adalovelaceinstitute.org/blog/the-case-for-a-global-ai-governance-floor/
- Davies, M. (2026). Taking back control? A pragmatic AI agenda for the new UK government. Ada Lovelace Institute, 28 July 2026. https://www.adalovelaceinstitute.org/blog/taking-back-control/