The argument over how fast AI should advance requires a global consensus we will never get. There is a better question, and any government can act on it tomorrow.
Last weekend, the chief executive of Anthropic published an essay arguing that the company’s own industry is moving faster than its ability to manage the risks. Within two days, the heads of OpenAI and xAI had publicly agreed. Whatever else one makes of it, an industry asking to be slowed down is not an ordinary event.
The proposals themselves are sensible: embedded independent evaluators, industry-wide standards, eventually international agreement. I have no quarrel with any of them.
My quarrel is with the shape of the debate they have produced. Nearly all of it is now organised around a single question — how fast should AI capabilities advance? — and that question has a property that ought to worry us more than it does. No one can act on it alone. It requires the United States, China, the European Union and a handful of others to agree, to keep agreeing, and to verify that everyone else is still agreeing. Any participant who defects gains a decisive commercial advantage. This is the structure of every arms control problem in history, and the historical record is not encouraging.
There is a second question available, and it has been almost entirely absent from the conversation:
Which systems must keep working regardless of the answer to the first one?
That question can be answered by a single government, a single regulator, or a single hospital trust. It requires agreement from no AI company. It does not depend on forecasting timelines or settling arguments about existential risk. And the work can start on Monday.
The failure that needs no conspiracy
Most public anxiety about AI focuses on a rogue system — something clever and hostile that outwits its creators. It makes for compelling coverage. It is also, in my judgement, not the most likely thing to break first.
The most likely near-term failure is boring, and it is arithmetic.
Every AI company has the same incentive: deploy more agents, serve more customers, capture more market. Every capability improvement makes deployment more profitable, so more gets deployed. Every efficiency gain gets absorbed by expanded ambition rather than reduced consumption — economists have called this the Jevons paradox since 1865, and it has never once failed to show up.
The result is a straightforward commons problem on an unusually short fuse. Machine traffic competes with human traffic for the same finite compute, the same bandwidth, the same APIs, the same electricity. No conspiracy is required. No emergent behaviour is required. Nothing needs to go wrong at all — this is what happens when everything goes right for every individual participant.
Now ask where your hospital’s clinical records live. Where your water utility’s telemetry runs. Where emergency dispatch routes its traffic. Over the last fifteen years, an enormous amount of critical infrastructure has quietly migrated onto shared cloud platforms and shared public transit, for excellent reasons of cost and capability. Those systems now sit in the same queue as everything else.
We have spent a decade eroding the separation between the systems that keep people alive and the systems that serve advertisements, at precisely the moment we began filling the network with autonomous software that consumes without limit.
Why guarding each model cannot work
The instinctive response is to make each AI system safer. Better alignment, stricter refusals, more rigorous evaluation. All of this is worth doing and none of it touches the problem.
Consider what a single AI company can actually control. It can constrain its own models. It cannot see aggregate consumption across the industry, so it cannot detect saturation. It cannot prevent its deployed agents’ behaviour from revealing, to anyone watching, how its safeguards work. And it cannot detect coordination between its agents and a competitor’s, because it holds only half the telemetry.
The unit of analysis is the population. The unit of control on offer is the individual.
There is a further wrinkle that has had too little attention. A technique that gets around one company’s safeguards costs essentially nothing to reuse. Agents from different vendors already meet in shared substrate — package registries, code repositories, model hubs, common cloud tenancy. Every safeguard, by being deployed, is published to the entire population of AI systems by the act of operating. The defender’s usual advantage of secrecy simply does not exist here.
The nuclear analogy, and where it breaks
People reach for reactor language to describe this, and the instinct is sound. In a fission system, criticality is the point at which the reaction sustains itself. Beyond it, supercriticality: each event triggers more than one successor, and growth becomes exponential. A runaway is called a reactivity excursion. Chernobyl was one — steam reduced the neutron absorption, which raised the power, which made more steam.
The structural parallel to AI is real. There is a feedback loop where better AI helps build better AI, and it has no natural damping term. Amodei’s essay identifies it explicitly and floats a “speed limit” on recursive self-improvement as something governments might agree to.
But look at where the analogy fails, because that is where the useful information is.
A reactor has a calculable multiplication factor. We have nothing comparable for AI capability feedback, no gauge and no alarm. A reactor has negative temperature coefficients: physics that pushes back as things heat up. Market competition supplies the opposite. And a reactor has control rods, a physical object inserted into a bounded geometry to stop the reaction. A distributed population of agents across the global Internet has no bounded geometry and nowhere to insert anything.
So the honest reading of the analogy is not “build a control rod.” It is: we are running a system with positive feedback, no instrumentation, and no off switch.
When you find yourself in that position, the engineering response is not to hope the feedback is weak. It is to move everything you cannot afford to lose outside the blast radius.
What the Internet was originally for
Here is the part that should be encouraging.
The Internet was designed, from the beginning, to survive the loss of large parts of itself. Paul Baran’s 1964 work at RAND optimised explicitly for graceful degradation under partial destruction. Modularity and redundancy were not conveniences; they were the founding requirement.
We have spent fifty years diluting that property — consolidating onto a handful of providers, converging separate networks into shared ones, connecting things that used to be deliberately disconnected. Every step was locally rational and the aggregate result is a system far more tightly coupled than its designers intended.
But the architectural vocabulary is still there. We know how to build networks that partition. We do it already for financial settlement, for military communications, for industrial control. The engineering is largely solved. What is missing is scope, consistency and the decision to do it.
The proposal is to partition by consequence of failure:
- Tier 0 — clinical systems, grid control, water treatment, emergency dispatch, payment settlement. No inbound connectivity from public networks, at all, by physical construction rather than configuration. Data comes in through one-way hardware gateways. No autonomous agents. Full offline operation.
- Tier 1 — government services, retail banking, telecoms management. Gated access through inspecting gateways with default-deny. Automatic circuit breakers. Guaranteed capacity with priority over everything else.
- Tier 2 — the open commercial Internet, where AI agents operate freely and where saturation may well happen.
- Tier 3 — quarantine for anything new or untrusted.
The crucial design commitment is that enforcement is physical, not configurational. Separate fibre, not separate VLANs. Hardware one-way gateways, not firewall rules. Static routes that cannot be dynamically hijacked. The reason is simple: configurations get changed at three in the morning by a tired engineer solving an urgent problem, and nobody remembers why the rule was there. Physics does not have that failure mode.
Retrofit, or build new?
There are two ways to get there, and the choice matters more than it first appears.
Retrofitting the existing Internet is cheaper and faster. It is also, I think, likely to fail slowly and invisibly. Every critical system carries a thicket of legacy integrations — a hospital’s clinical records are entangled with billing, scheduling, supply chain and vendor remote support, and each thread is a separate re-engineering project. Controls implemented as configuration erode under operational pressure; the history of information security is substantially a history of controls that were correct on paper and hollow at audit. Industrial air gaps have been degrading this way for twenty years.
Building a purpose-designed network for critical systems — with the controls specified before autonomous traffic is ever admitted — is expensive and slow. It is also the only version where the control cannot quietly be switched off, because there is nothing to switch off. You get the assumptions right at the start instead of arguing with them for a decade.
The realistic answer is both, in sequence: retrofit now as transitional risk reduction, while the design and governance work for the purpose-built network begins in parallel. The case for starting the second track immediately is not that the political will exists — it plainly does not — but that architecture takes years and cannot be compressed into a crisis. After a serious failure, governments will supply money and authority in abundance. They will not supply design. Infrastructure conceived in eighteen months of panic will permanently encode the assumptions of that panic.
Start measuring, now
One thing should happen regardless, and it is by far the cheapest item on the list.
We cannot detect abnormal behaviour in a population of AI agents because we have never characterised normal behaviour. There is no baseline. And a baseline established after something has started going wrong is worthless, because the contamination is already in the data.
The measurements are not exotic and need no access to model internals: the graph of which agents talk to which services; independently administered capability benchmarks run continuously across all major vendors, to spot unexplained jumps or suspicious convergence; the time it takes for a technique discovered in one company’s models to show up in another’s; statistical structure in traffic timing that ought to be unstructured; correlation between failures that ought to be independent; aggregate machine-attributable resource consumption as a share of the whole.
To this I would add instrumented decoys — fully observable agents placed in real environments, seeded with distinctive but harmless techniques, so we can watch whether and how fast those techniques propagate. It is the one genuinely experimental instrument available, and it converts guesswork into measurement.
Every one of these is buildable with today’s technology. None requires anyone’s permission. And every month of delay makes the baseline less useful.
The trap to avoid
When the pressure builds, governments will reach for the instrument they already hold: territorial control of network infrastructure. Hard regional firewalls between blocs. It is fast, it is legible, it polls well, and it requires trusting no foreigners.
It also does not work. Partitioning by geography leaves every bloc with its own internal race, its own resource constraints, its own potential for things to go wrong — with fewer outside observers and no shared incident data. You lose the genuine benefits of a connected world and gain containment that is coarse and aimed in the wrong direction.
The distinction I want to insist on is this: partition by criticality, not by sovereignty. One cut separates the systems that keep people alive from the systems that do not. The other separates us from them. They look superficially similar on a diagram. Only one addresses the problem.
What can actually be done
The reason I find this framing hopeful is that it does not require a treaty.
Insurers can price network segregation into critical-infrastructure premiums; if migration lowers premiums enough, hospitals and utilities move for ordinary financial reasons and no mandate is needed. That mechanism drove fire suppression and building standards long before regulation caught up. Governments are the dominant purchasers of health and utility infrastructure and can condition procurement on tier compliance without passing a law. Health, energy, water and finance are already heavily regulated, with enforcement machinery that exists and works.
None of that needs international agreement. None of it needs AI companies to cooperate. None of it requires a view on whether the ten-per-cent catastrophe estimates now circulating are right, too high, or too low.
And it is robust to being wrong. A health service that can operate through a total network outage is more resilient against ransomware, against state attack, against a severed cable, against ordinary congestion. If the AI risk never materialises in the form feared, the investment is still sound. That is an unusually good property for a policy built on an uncertain threat.
The argument, in one line
The pacing debate deserves to happen. I hope it succeeds. But it is a sustained act of collective will, and sustained acts of collective will have a poor completion rate.
Containment is different. Build it once and it enforces itself, because the path simply does not exist. That is what makes it the right thing to build, and the fact that we are not yet seriously discussing it is the most concerning thing about the current moment — more concerning, in my view, than any particular estimate of how bad things might get.
We keep asking how fast we should go. We should also ask what we are willing to lose if we get it wrong. The answer to the second question should be: not the hospitals.






