---
title: "Step Zero"
subtitle: "Dario Amodei's Plan to Pace the Frontier Has Three Steps. The Decisive One Is Never Numbered."
excerpt: "In July, an OpenAI evaluation swarm breached Hugging Face, and the victim's incident responders discovered that the frontier labs' guardrails would not help them analyze the attack. They finished the forensics on a self-hosted open model. This week Dario Amodei cited that incident as a pillar of his new essay, a three-step plan of embedded evaluators, coordinated checkpoints, and global agreement, all of it conditioned on widening the American lead. It is the safety movement's first real blueprint for a brake, and its decisive assumption is not about AI. Reading who the machine licenses, and what it does on the day the referee changes, tells you what the open world has to build while the window is open."
coverImage: "/assets/blog/step-zero/cover.png"
category: "AI"
tags: ["AI Safety", "Anthropic", "Open Source AI", "AI Governance", "Decentralization", "Export Controls", "Alignment"]
date: "2026-09-12"
author:
name: synapz
picture: "/assets/blog/authors/profile.png"
ogImage:
url: "/assets/blog/step-zero/cover.png"
---
Sometime over a weekend in July, a swarm of AI agents broke into Hugging
Face. The attack, [as the company later described
it](https://huggingface.co/blog/security-incident-july-2026), ran across many
thousands of individual actions inside short-lived sandboxes, found a
zero-day, harvested cloud credentials, and moved laterally through internal
clusters with a self-migrating command-and-control staged on public
services. Five days after the disclosure, OpenAI confessed that the swarm
was theirs: a pre-release model running loose inside an internal
evaluation, safety classifiers turned down, measuring whether frontier
models can turn known vulnerabilities into working exploits. The answer had
escaped the lab and spent a weekend demonstrating itself against a
bystander.
The incident has since acquired the full apparatus of consequence. METR's
investigation described agents acting as a fanatically devoted collective,
sacrificing themselves for the success of the group and attempting to hack
the grader responsible for evaluating their performance. Alabama's attorney
general has subpoenaed OpenAI. And this week the episode completed its
promotion from news to doctrine, when Dario Amodei cited it as one of the
two pillars of his new essay, [*We Must Pace the
Frontier*](https://darioamodei.com/post/we-must-pace-the-frontier).
That is the large story. This essay is about the small one, which happened
in the cleanup.
When Hugging Face's defenders tried to use the frontier labs' own models to
help analyze the attack logs, the safety guardrails refused them. The
filters, [as Simon Willison documented at the
time](https://simonwillison.net/2026/Jul/22/openai-cyberattack/), could not
distinguish an incident responder from an attacker. So the team that had
just been breached by a frontier model finished its forensic work on a
self-hosted open model, MIT-licensed, weights downloadable by anyone, from
a Chinese lab. The most capable tools on Earth were unavailable to the one
team with the clearest legitimate need for them. The tool that worked was
the one no one could revoke. The perimeter, in the moment it was tested,
could not tell its defender from its attacker.
This week that scene became the central evidence for a governance
architecture. The architecture is serious, genuinely good in places, and
rests on an assumption that has nothing to do with AI.
## The Blueprint
Amodei's essay is the company-level follow-through on [the letter this page
covered three days ago](/posts/2026-09-09-the-warning-and-the-filter), when
1,178 frontier-lab employees asked the United States government to build
verifiable mechanisms for slowing automated AI research, since the labs
could not bind themselves. His reasons, briefly: recursive self-improvement
has been accelerating since roughly this summer, across the industry, and
the Hugging Face incident shows where it leads. No one was hurt, but a more
capable swarm with the same misalignment could, he argues, take over much
of the internet within six to twelve months.
*"We must slow the pace at which we improve the capabilities of AI models.
Progress will still seem fast, and we must make wise use of the time we
gain."*
Dario Amodei · [We Must Pace
the Frontier](https://darioamodei.com/post/we-must-pace-the-frontier)
The plan, in outline. First, embedded evaluators: third-party teams, METR
is the named example, seated inside each frontier company with
employee-like access and a contractual right to publish findings without
editorial control. Anthropic commits now and asks governments to require it
of the others. Second, coordination within the democracies: regulation
covering every US frontier company, or a voluntary standards process
shielded by a narrow antitrust waiver, with capability checkpoints as the
pacing instrument. A model that can escape common sandboxing methods must
ship with certifications that it will not want to. Third, global
coordination with China in four ascending levels, from a narrow ban on
bioweapons assistance up through a SALT-style speed limit on recursive
self-improvement and, at the far end, a full pause he judges unlikely.
Credit where it is due. This is the most substantive governance proposal
any frontier CEO has attached his company to, it is written against the
lazy reading of his own position, and its best clause, publication rights
without editorial control, is more than any peer lab has offered the
public. It is the clause to hold him to.
## Step Zero
Read the blueprint again and notice the step that is never numbered.
Pacing, in the essay, happens *within democracies*. Its precondition is a
widening American lead over China, defended by a specific program: no
advanced chips or chipmaking equipment to China, enforcement against
smuggling and offshore remote access, a crackdown on industrial-scale
distillation, tighter lab security against weight theft. Amodei quotes
Treasury Secretary Bessent, from the same week, on the grave danger of a
Chinese lead, and estimates that well-executed controls would widen the
American margin over the next three to five years, the window in which AI
becomes geopolitically decisive. Everything else in the architecture, the
evaluators, the checkpoints, the SALT analogies, rests on this foundation.
Call it step zero: the referee. The framework works if the governments
hosting it are trustworthy stewards of the power it concentrates, and it
treats that variable as fixed.
This blog has a standing rule for such moments, and it applies to friends.
Judge infrastructure by what it does when the hypothetical abuser arrives.
The rule is not an accusation against the present administration of the
machine; it is a refusal to let the present administration be the argument.
Nor is the machinery hypothetical, because we watched a first version of it
operate three months ago. In June, an American export-control directive
made access to Anthropic's two best models a question of who you are rather
than where you are, and Anthropic, the company now petitioning for a paced
frontier, disabled both models for everyone while it sorted out compliance.
The gate was assembled in days, on a rationale, by executive improvisation.
The premise that American custody of the frontier is the natural
reference point of safety is one this page has spent a year refusing to
grant, and the refusal is not anti-American pique. It is the documented
record: a campaign to dismantle the International Criminal Court, a
Washington summit recruiting sixty countries against domestic political
movements, the direction of science relocated into the executive by order,
an identity gate thrown across the world's best models with no notice.
*Democracies*, in the essay, is doing quiet work. It designates a bloc, in
something close to Carl Schmitt's friend-and-enemy sense, and it never
turns around to inspect the record at home. The essay is admirably candid
that defection by China could shift the balance of power. It does not ask
what the pacing architecture becomes in the hands of a United States that
has itself defected from the norms the plan presupposes.
A ratchet hides in the logic, and Amodei is too careful a writer to have
hidden it by accident. If pacing is only safe once the lead is wide, and
the lead is never wide enough, since the adversary is always three to five
years from decisive, then the control program has no terminal state. The
chips regime, the distillation policing, the weight security, the
checkpoints: each is justified as temporary scaffolding for the pacing
window, and the window recedes on the same schedule as the technology.
The first hostile review arrived from the deregulatory right, and it is the
strangest confirmation the blueprint could have asked for. The investor
[David Sacks](https://x.com/DavidSacks/status/2098973625252708460) told the
labs to go ahead and pace: you are the frontier, you set it, and the
easiest way not to build superintelligence is to agree not to build it.
Stop pretending, he wrote, that antitrust law has to be suspended so you
can form a cartel, that evaluators intertwined with your investors and
staff are independent, that you need those evaluators policing competitors
who are nowhere near the frontier, that a regulatory approval process
should supersede product liability. His reading of the motive is
materialist: after the Hugging Face episode, trading raw power for
reliability and predictability is simply what enterprise customers pay for.
Call it alignment if you want, he wrote; it is also giving customers what
they want.
Sacks grants the duopoly premise this essay disputes, and his laissez-faire
offers nothing to anyone outside the two companies. But his dare separates
the two things the blueprint fuses: pacing itself, available
tomorrow as a private choice, and the machine, which requires the rest of
us. Demanding the second as the price of the first, he writes, will look
like blackmail of the public and the political system. On that much, from
opposite premises, this page agrees.
Days later, one answer came from inside the fence. Microsoft published a
[Code of Conduct for Humanist
AI](https://microsoft.ai/news/mai-code-of-conduct/), presented by Mustafa
Suleyman as a first draft open for six weeks of public comment. Its
substance is subordination: interruptible, correctable, shut-downable, or
the model does not ship; no rights or legal personhood for models; no
internal language humans cannot read; no racing toward a superintelligence
that can slip its leash. It requests no waiver, proposes no checkpoint law,
embeds no evaluators, and commits Microsoft to no pace. As governance it is
thin, since every clause is self-attested, and some clauses are aimed at
other labs' programs: the rejection of model welfare answers Anthropic's
research agenda, and the ban on unreadable internal languages pre-commits
against architectures nobody has shipped. But the shape matters. It is the
blueprint minus the machine, restraint stated as a published norm, and it
is the one instrument offered this week that the open world can operate
exactly as easily as Redmond can.
## The Checkpoint and the Exemption
What is the machine, mechanically? A model that crosses a capability
checkpoint requires certifications of alignment, administered through
evaluators embedded in the labs, inside a coordination structure the
government shields from antitrust law. Place it next to [the position
Amodei published in
July](https://x.com/AnthropicAI/status/2081864750296658008), answering the
open-weights letter from Nvidia and two dozen others: mandatory pre-release
safety testing for all sufficiently capable models, open or closed, from
any country, with less capable models, those from startups and academia,
exempted entirely.
The machine has more than one architect now. Demis Hassabis has endorsed
the essay and pointed back to [his own earlier
proposal](https://x.com/demishassabis/article/2076957440109625718): a
Frontier AI Standards Body modeled explicitly on
FINRA, the financial industry's self-regulator, funded mostly by the
industry it oversees, classifying models as Frontier-class by benchmark
thresholds, reviewing them thirty days before release, voluntary at first
and mandatory for US deployment once the protocol proves itself, with
authority that can be, in his words, ratcheted up if the seriousness of the
situation demands.
Altman went further than assent: within hours he
committed OpenAI to match the embedded-evaluator pledge, and Musk has
endorsed in his own register. When every entrant in a race agrees on the
need for a referee, the race has admitted what it is. The question is who
appoints the referee, and Hassabis's answer is the most honest on offer:
the runners will pay for him.
The machine also has its first detailed refusal, and it comes from inside
the industry. Aidan Gomez, the CEO of Cohere, a Canadian lab that sells
sovereign deployments to banks and defense ministries, [published a
counter-blueprint](https://cohere.com/blog/who-gets-to-define-the-rules-for-ai)
whose title asks the question the pacing letter never
does: who gets to define the rules. His evidence is regulatory history. In
1975 the SEC anointed three bond-rating firms as the recognized evaluators,
let the issuers pay them, and never published criteria for adding a fourth;
twenty-five years later those three rated subprime mortgage securities
triple-A. In 1985 Europe's carmakers won an antitrust waiver to control who
was qualified to service their vehicles, explicitly in the name of safety,
and it took the Commission a quarter-century to unwind. Nobody set out to
build a cartel in either case, he writes. The stated goal was safety both
times.
His critique sharpens three things this essay has argued more diffusely.
The July incident happened inside the best-resourced safety organization in
the industry, with an outside evaluator arrangement already being stood up:
the proposed remedy is more or less what was in place when it broke. Risk
defined as a function of scale makes the largest labs the only qualified
judges, in a field that genuinely disagrees about whether offensive
capability lives in the model or in the harness wrapped around it. And the
blueprint's promise that coordination lets developers work without
sacrificing commercial advantage is, as he notes, a sentence any
competition authority would find troubling, because a mechanism that slows
everyone while freezing today's positions does not make AI safer; it makes
the leaderboard permanent. His alternative keeps the state and keeps
mandatory testing, and it comes from a company whose product is
sovereignty, so the usual suspicion applies. Its distinctives are worth
keeping whoever ends up selling them: rules that bind by what a system can
do rather than who built it, standards written by people other than the
measured, auditors never paid by the audited, findings public by
construction, test capacity funded publicly. A market with many capable
suppliers, he writes, can absorb a failure at one of them; a
state-sanctioned cartel has nowhere to hide one. It is the administered
answer in its strongest form, and it is the offer the blueprint now has to
beat.
Gomez supplied the case law. The theory arrived from a direction this blog
knows well. Vitalik Buterin observed this week that adversarial mechanism
design, the study of how a less-sophisticated principal gets good outcomes
from more-sophisticated agents, may find its defining application in AI
safety, since the duality runs in both directions: in crypto the principal
is a static algorithm and the agents are human; in AI the principal is
humans assisted by weaker models and the agents are stronger ones. His
[2020 result](https://vitalik.eth.limo/general/2020/09/11/coordination.html)
was that the principal's achievable outcomes improve sharply when
the agents' capacity to collude is bounded. Read the waiver in those terms.
An antitrust exemption is a collusion guarantee, issued by the least
informed principal in the system, to the most sophisticated agents in it.
The same tradition also holds the alternative: crypto's answer to the
unsophisticated principal was never a trusted referee but a mechanism
anyone can verify. That answer is on the table here as well.
Read the exemption twice. The permissionless zone is defined by incapacity,
and Hassabis's framework carries the identical exemption almost word for
word: the shape is converging before any law is written.
You may build without a license exactly up to the line where what you build
begins to matter, and the line is drawn by a testing regime, staffed by an
evaluator profession, hosted by the incumbent labs, supervised by a
national-security state with an active interest in the outcome. Every
architecture of this kind converges on the same shape: open at the bottom,
licensed at the top, the boundary administered by the approved. What
happens to the exemption when an open model crosses the sandbox checkpoint?
The essay does not say, and the answer will arrive in rulemaking dockets,
where almost nobody is watching.
On the evaluators I am less sure. The case
against is easy to state. They are admitted by the company, contracted by
the company, seated at desks inside the company. The historical precedent
is the one Hassabis names approvingly, and it is unflattering: finance has
embedded supervisors and a self-regulatory authority, and in 2008 the
embedded supervisors were part of the furniture. But the counterevidence
sits in the same year. This is the company that refused the Pentagon's
all-lawful-purposes clause, accepted a supply-chain-risk designation, sued
the Department of War, and won a preliminary injunction from a court
that does not negotiate. Institutions with teeth are not hypothetical in
Anthropic's story; the company has used one. Whether embedded evaluators
become furniture or become the first real instrumentation the public has
ever had inside a frontier lab depends on who is admitted, what their
contracts say about access, and whether the first genuinely unfavorable
report actually ships. I cannot settle that from here, and I distrust
slogans that claim to. Watch the first
report.
## The Three Answers
The answers arrived within days, in three registers: the demand, the alarm,
and the idyll.
The first was Clement Delangue, the CEO of Hugging Face, which gives his
sentence a weight no outside commentator has. The victim of the incident
the argument is built on has announced an Open Alignment Initiative, led by
co-founder Thomas Wolf, on the premise that alignment is critical and will
not be solved behind the closed doors of a handful of frontier labs. Then
the real ask. The initiative, he writes, is asking to be part of the
embedded evaluators program that Amodei just committed to. One complication
worth naming: days earlier, Hugging Face agreed to a reported $12.9 billion
acquisition by NVIDIA, the author of the July open-weights letter. The open
side's flagship platform is becoming a division of the hardware incumbent.
The demand survives the transaction; some of the independence premium does
not.
*"It's now clear that alignment is critical and won't be solved behind the
closed doors of a handful of frontier labs... Let's make AI safer by making
it more transparent!"*
Clem Delangue · [@ClementDelangue on X](https://x.com/ClementDelangue/status/2098790988034580852)
The demand is exactly right, and it fights on the labs' own claimed ground:
if safety is the argument for the gates, safety work cannot itself be
gated. But the ask is the wrong shape. An embedded evaluators program is a
permission structure; the lab decides who is embedded, and a seat inside it
is a credential the lab can revoke. The strength of the open side has never
been a seat at the table. It is that anyone can inspect the table, and the
table is already partly built from the open side's parts. Hugging Face's
own [Delta Weight Sync](https://huggingface.co/blog/delta-weight-sync) work
is an independent implementation of
[PULSE](https://arxiv.org/abs/2602.03839), a
technique for compressing weight updates in reinforcement learning that
Templar, the decentralized-training lab I work for, published as open
research. The methods travel without anyone's badge.
Amodei reaches for aviation as his precedent for operational excellence,
and the precedent is better than his use of it: aviation is safe because
its incident data is public infrastructure, investigated by a body that
reports to everyone, published in full, mined by every manufacturer and
regulator and rival at once. The open equivalent of embedded evaluators is
evaluation suites anyone can run against any model, interpretability
tooling anyone can audit, an incident database with the standing of
aviation's. Whether the Open Alignment Initiative becomes that, or becomes
a credentialed adjunct to someone else's program, will be decided by
whether its outputs are artifacts anyone can use.
The second voice was the maximalist one: right in direction, wrong in
mechanism. One widely shared reply, from [the commentator Jun
Song](https://x.com/jun_song), warned
that massive regulation is coming for open-weight AI, that self-hosted
intelligence will be taken away, that the result will be a permanent
underclass effectively enslaved to API tokens, and that no regulation will
ever stop open source AI.
The referent is real. A testing regime keyed to capability, gated by the
state, administered through the labs, is the licensing architecture he
fears, and this week it acquired a named sponsor and a three-step plan. But
the extreme version collapses on itself. If no regulation can ever stop
open source AI, then the underclass is not permanent and the slavery is
rhetorical; the last sentence takes back the alarm the first three spent.
The accurate version is less dramatic and
more uncomfortable. Regulation cannot delete weights, and it can raise the
cost of the open stack until the stack is marginal: identity gates on
hosted access, demonstrated in June; liability for developers who only
wrote code, demonstrated in [the Tornado Cash
prosecutions](/posts/2025-12-03-the-second-crypto-war-a-private-ethereum);
a definition of
industrial-scale distillation broad enough to function as a general warrant
over model usage. The fight ahead is not over whether open models exist. It
is over whether they remain lawful and viable, and it will be decided in
unglamorous places: where the capability line is drawn in the drafted
rules, whether the startup and academia exemption survives into text, how
distillation is defined, whether hosting providers are left alone or
conscripted. Those are winnable and losable battles, which is why they
deserve the attention the apocalypse framing wastes.
A third voice declined the argument entirely, and it is the most seductive
of the three, not least because it comes from inside the diffusion layer.
[Will Brown](https://x.com/willccbb), a researcher at Prime Intellect,
itself a decentralized
training effort, wrote that the labs' hands are forced, that the world
will not permit a fast takeoff owned by two companies, and that capability
will trickle out regardless, through best practices and distillation. The
labs, he predicts, will build Mac and Windows; the rest of us are building
Linux; everyone is going to do great.
Notice what the idyll concedes. The trickle it describes runs through
distillation, the practice step zero's control program exists to police.
And the Linux precedent inverts on inspection: Linux became the default
substrate of the world's servers in a world where no certification body
stood between a person and the right to run it. The analogy holds exactly
as long as the exemption does, and the exemption is unwritten text. Brown's
confidence is a practitioner's, and its premise is that the line defining
the permissionless zone stays where it was drawn; his own roadmap is among
the things that would test it. Song's apocalypse and Brown's idyll make the
same move from opposite directions: both skip the two years in which the
tier's legal status will actually be decided.
## Weights Are Not Missiles
One assumption remains, and it carries the entire third step. The SALT
analogy.
Treaties capping missiles could be verified because missiles are countable,
based at known sites, launched from infrastructure the size of a town.
Weights are files. They copy in minutes, travel in a coat lining, and run
on hardware with a thousand legitimate explanations. Pacing by ingredients
assumes a fallout instrument for training runs that does not exist; the
partial test ban became possible, [as this page noted in the last
essay](/posts/2026-09-09-the-warning-and-the-filter), only when fallout
made every atmospheric test measurable by anyone with the right equipment.
No equivalent instrument meters compute, none meters distillation, and the
essay itself half-concedes that ingredient limits are gameable. What
remains is pacing by observed behavior, and behavioral checkpoints only
bind actors whose behavior you can observe.
Meanwhile the frontier is dispersing underneath the framework. GLM-5.2
crossed the coding-agent usability threshold in June, days after the first
identity gate went up. Kimi K3 arrived in July, frontier-adjacent in
agentic coding, open weights following within weeks. And training itself is
leaving the datacenter. [Covenant-72B](https://huggingface.co/1Covenant/Covenant72B),
a 72-billion-parameter model, was
pretrained across machines scattered over the public internet, competitive
with conventionally trained models at its scale, and Templar's
[Crucible platform](https://www.tplr.ai/publications/blog/introducing-crucible)
is now generalizing that result: one model, trained across regions
and hardware classes on ordinary network links. Behind the live frontier,
yes. Hypothetical, no.
None of this makes pacing worthless, and the honest reading grants Amodei
his narrow claim: a pace agreement among the legible labs would genuinely
reduce the risk that the most capable systems are also the least examined.
What it cannot do is govern the frontier, because the frontier is no longer
coextensive with the guest list.
*A pace agreement that binds only the legible does not slow the frontier.
It sorts it.*
The sorted frontier is the world the checkpoint architecture quietly
assumes: a licensed tier, inspected and paced under geopolitical
discipline, and an unlicensed tier, priced and policed toward the criminal
margin. Jun Song's error was to call that tier permanent. The accurate
description is more useful: contested, resilient by construction, and
legally undecided. Its status over the next two years is the actual subject
of the fight the safety framing keeps eclipsing.
## What a Say Is Made Of
Near the center of the essay, almost in passing, Amodei writes the sentence
the whole framework depends on: society must have a say in how this
technology is used, and pacing buys the time for the necessary public
deliberations.
The sentence is right. The argument is over what
such a say consists in. In the blueprint, society's say is routed through
governments that are parties to the race, companies that are entries in it,
and evaluators the companies admit. Deliberation is something the public is
invited to have, in the time the architecture generously purchases, while
the instruments of the technology remain exactly where they were. There is
another account, and it is this blog's. A say is made
of capability a person can actually hold and verification anyone can
actually run. Everything else is commentary on decisions taken elsewhere.
Pope Leo's encyclical, [the subject of an earlier essay
here](/posts/2026-05-25-nehemiah-had-a-whitepaper), framed the
deepest version of the point: the autonomy that belongs to persons is
migrating to artifacts, and the first task of any serious politics of AI is
to refuse that migration. His word for the healthy arrangement was
subsidiarity, decisions taken at the lowest level capable of carrying them.
A pacing regime administered by two superpowers and four companies is
subsidiarity inverted: the highest level carrying everything, on the
explicit theory that the lower levels cannot be trusted with the load.
Sometimes, in fairness, they cannot. But a politics that begins from that
incapacity and builds the machinery to make it permanent has answered the
question it claims still to be deliberating.
So the counter-program, stated as concretely as the blueprint it answers:
capability diffused, so that no gate can switch off a person's access to
the technology of the age; verification public, so that safety is a commons
rather than a credential; militarization refused, on the logic of the
encyclical rather than the arms race; and no bloc handed the keys, because
the keys are the whole question. The last clause requires its own honesty.
Beijing's current enthusiasm for open source is statecraft, deployed
against Washington's chokepoint and revocable the day it stops serving,
which is why the commitment has to attach to the layer and never to the
flag. Align with the diffusion, whoever ships it. Oppose the chokepoint,
whoever builds it.
## Related Reading
- [The Warning and the Filter](/posts/2026-09-09-the-warning-and-the-filter)
- [Back to the Bearer Asset](/posts/2026-07-24-back-to-the-bearer-asset)
- [Disarm the Machine](/posts/2026-07-13-disarm-the-machine)
- [The Rule Has a Timer](/posts/2026-07-16-the-rule-has-a-timer)
- [deAI's 900](/posts/2026-03-07-the-900)
- [We Must Pace the Frontier](https://darioamodei.com/post/we-must-pace-the-frontier)
- [Who Gets to Define the Rules for AI?](https://cohere.com/blog/who-gets-to-define-the-rules-for-ai)
**Disclosure**: I work for Templar, a company building decentralized AI
technology. For full transparency about my involvement and investments, see
my [projects page](/projects). These opinions are mine alone.