Why Frontier AI Governance Must Begin Before Training, Not Before Release

High-angle aerial photograph of a lush mangrove swamp with intricate, maze-like tidal channels winding through dense green vegetation. Sunlight glints bright white off the reflective water in several meandering waterways.

I have been thinking deeply about Demis Hassabis’s proposal for an independent body to evaluate frontier AI systems before they are released.

It is an important proposal. It moves the conversation beyond asking frontier AI companies to regulate themselves, and it recognizes that systems approaching AGI may require institutions with real authority, technical expertise, and public legitimacy.

But I believe we need to go further.

A review conducted shortly before a model is released may already be too late.

The most consequential decision may not be whether a completed frontier model should be made available to the public. It may be whether a potentially AGI-producing training run should be allowed to begin in the first place.

The problem with governing AI only at the point of release

Most current AI governance proposals focus on deployment.

They ask questions such as:

  • Is the model safe enough to release?

  • Does it create cyber, biological, or national security risks?

  • Can it deceive users or act autonomously?

  • What restrictions should apply to public access?


These are important questions, but they come after the model has already been trained.

By that point, the most dangerous capabilities may already exist.

Risk can emerge during training, during internal testing, through autonomous research conducted inside a company, or through the theft of model weights. A system does not have to be publicly released to create a major threat.

This is why I believe frontier AI governance must move upstream.

Before a company begins a training run that could plausibly produce AGI-level capabilities, it should have to demonstrate that it has a credible safety case, secure infrastructure, independent oversight, clear pause conditions, and a plan for maintaining meaningful human control.

That is a very different model from simply testing the finished system 30 days before release.

We may need something closer to CERN and the IAEA

I increasingly believe that our best hope for managing advanced AI is an international institution inspired by both CERN and the International Atomic Energy Agency.

CERN shows that countries can pool extraordinary scientific talent, infrastructure, and funding around a shared global mission.

The IAEA provides a different model. It is built around inspections, verification, monitoring, and international safeguards.

For frontier AI, we may need both.

We need a CERN-like institution for collaborative, public-interest AI safety research. It could bring together the world’s best researchers to work on alignment, interpretability, scalable oversight, corrigibility, secure system design, and methods for maintaining human control over increasingly capable systems.

But the institution conducting the research should not be responsible for certifying its own safety.

We would also need an IAEA-like authority with the ability to inspect frontier computing facilities, verify declarations, investigate incidents, and enforce international rules.

The builder and the regulator must remain separate.

Otherwise, we simply reproduce the same conflict of interest that exists when private laboratories are responsible for developing, evaluating, and releasing their own systems.

Frontier compute may be the most practical point of control

Algorithms can be copied. Data can move across borders. Model weights can be stolen.

Frontier-scale computing infrastructure is different.

The largest training clusters are physical, expensive, energy-intensive, and concentrated in a relatively small number of facilities and supply chains. That makes advanced compute one of the few places where effective oversight may still be possible.

A serious international system could require frontier data centers to be registered and licensed.

Training runs above a certain risk threshold could require advance authorization. Facilities could maintain secure compute records, undergo independent inspections, and provide approved evaluators with access to critical checkpoints.

The approval threshold should not be based on compute alone. It should also consider the system’s expected capabilities, level of autonomy, ability to conduct AI research, access to external tools, and potential cyber, biological, or replication risks.

The objective would not be to regulate every AI model or every data center.

It would be to create strong controls around the relatively small number of training runs that could fundamentally alter the balance of power between humans and machines.

The controls must exist before AGI arrives

One of the most dangerous assumptions in the AI debate is that humanity will recognize AGI when it arrives and then have time to decide what to do next.

We cannot rely on that.

The transition from AGI to systems far beyond human capabilities may be slow. It may also be extremely fast.

An advanced system could accelerate AI research, improve its own tools, coordinate large numbers of agents, or help design its successors. Even if recursive self-improvement is harder than some expect, we do not know how much time society would have between recognizing AGI and facing something closer to ASI.

The window may be short.

It may also be disputed. Different companies, governments, and researchers may disagree about whether AGI has been reached. By the time a consensus forms, the most important decisions may already have been made.

This is why I do not believe we can wait until after AGI to build international controls.

The regulatory infrastructure, inspection systems, emergency protocols, and political agreements need to be in place beforehand.

A global institution must not become another competitor

There is also an important risk in the CERN analogy.

A publicly funded international AI laboratory could easily become another participant in the race to build the most powerful system.

Governments might support it for national prestige, economic competitiveness, military advantage, or the desire not to fall behind.

That would defeat the purpose.

A CERN-like AI institution would need a very clear charter. Its primary mission could not be to win the AGI race. Its mission would need to be safe development, alignment research, peaceful use, preservation of human agency, and broad distribution of benefits.

It should be designed to reduce the pressure to race, not intensify it.

This may require participating countries and companies to pool some frontier resources, share safety research, and accept limits on unilateral development.

That will be politically difficult. But the alternative is to trust a small number of private companies and national governments to compete responsibly for what may become the most powerful technology ever created.

I do not believe that is a stable long-term strategy.

Centralizing AI also creates serious dangers

An international institution is not automatically safe simply because it is international.

Concentrating the world’s most advanced AI systems, compute resources, and technical expertise inside one organization could create its own catastrophic risks.

It could become a target for espionage. It could be captured by powerful governments or corporations. It could exclude countries that lack wealth or geopolitical influence. It could become unaccountable, authoritarian, or resistant to public scrutiny.

For that reason, I would favor a federated system rather than a single global super-laboratory.

Several geographically distributed frontier computing centers could operate under common international rules. No single country, company, or institution would have unilateral control over the entire system.

High-risk decisions should require approval from multiple independent groups, including governments, technical experts, civil society, developing countries, human rights representatives, and communities likely to be affected by advanced AI.

No frontier AI company should have a veto.

No single government should control the launch key.

What meaningful pre-training governance could include

Before a frontier training run begins, the developer should be required to submit a detailed safety case.

That case could include:

  • The purpose of the system and the capabilities it is expected to develop.

  • The amount of compute required.

  • The source and governance of the training data.

  • The security measures protecting the infrastructure and model weights.

  • The evaluations that will be conducted during training.

  • The conditions that would trigger an immediate pause.

  • Restrictions on autonomous behavior, external tool use, replication, and AI research.

  • Plans for containment, shutdown, incident reporting, and post-training custody.


The authorization should not be permanent.

If the developer materially changes the system’s architecture, objective, autonomy, data, or scale, it should require renewed approval.

There should also be predetermined international pause conditions.

These might include evidence that a system can autonomously replicate, evade human control, conduct dangerous cyber or biological work, accelerate AI research beyond agreed thresholds, or deceive evaluators in ways that invalidate the existing safety case.

A pause should not depend on an improvised political debate after the danger has already emerged.

The rules should be established in advance.

“Humanity wins” must become more than a slogan

The goal of advanced AI governance cannot simply be preventing extinction.

It should also be ensuring that the benefits of transformative AI are shared broadly.

If AGI is developed using knowledge accumulated by humanity, infrastructure supported by governments, data generated by billions of people, and scientific discoveries produced across generations, its benefits should not belong exclusively to a handful of companies and investors.

A legitimate global framework should address:

  • Public participation in decisions about advanced AI.

  • Access to scientific and medical benefits.

  • Economic disruption and the future of work.

  • The distribution of wealth created by AI.

  • The representation of developing countries.

  • The protection of human rights and political freedom.

  • Restrictions on coercive, military, surveillance, and biological uses.

  • The preservation of meaningful human authority.


“Benefiting humanity” needs to be translated into governance rights, enforceable obligations, and measurable outcomes.

Otherwise, it remains a mission statement without institutional power behind it.

Where the current proposal fits

I see Hassabis’s proposal as an important bridge.

An independent standards body with credible technical expertise could improve frontier evaluations, reduce reliance on corporate self-assessment, and create an external gate before powerful systems reach the market.

That would be meaningful progress.

But I do not believe release testing alone is enough.

A complete governance system must have visibility into frontier computing infrastructure, authority over the highest-risk training runs, access to independent inspections, strong model-weight security, and the ability to pause development when agreed thresholds are crossed.

Most importantly, it must operate internationally.

A system governed only by one country risks driving development elsewhere. A voluntary system risks being abandoned when commercial or geopolitical pressure increases. A process controlled by the leading laboratories risks regulatory capture.

The institution must have real independence, international legitimacy, and enforceable authority.


My conclusion

I believe an international, treaty-backed frontier AI regime may be necessary if humanity is going to navigate the transition to AGI safely.

It will not solve alignment by itself.

We still need fundamental progress in technical alignment, interpretability, control, cybersecurity, and governance. We also need safeguards against the international institution itself becoming a dangerous concentration of power.

But trusting every frontier AI laboratory to race independently, assess its own systems, and voluntarily stop when the risks become too great is not a credible plan.

The strongest principle is simple:

We should govern whether a potentially AGI-producing training run is allowed to begin, not only whether the completed model is allowed to be released.

By the time we have built AGI, it may already be too late to design the institutions needed to control what comes next.

Those institutions need to exist before the decisive training run begins.

#ArtificialIntelligence #AGI #AISafety #AIAlignment #AIGovernance #ResponsibleAI #TechnologyPolicy #GlobalCooperation

Your daily companion for setting intentions and designing what's next in your life.

SOC 2

GDPR

© 2026 Baryons, Inc.

Your daily companion for setting intentions and designing what's next in your life.

SOC 2

GDPR

© 2026 Baryons, Inc.

Your daily companion for setting intentions and designing what's next in your life.

SOC 2

GDPR

© 2026 Baryons, Inc.