Table of Contents
- Preface: A Data Center in Mongolia
- 1. What AI Futures Project Actually Proposed
- 2. The Analogy That Makes MACD Work
- 3. The Unstated Assumption
- 4. Self-Reporting Cannot Trigger Mutual Destruction
- 5. What MACD Actually Needs
- Conclusion: The Bomb Cannot Verify Itself.
Preface: A Data Center in Mongolia
Imagine a data center built by the United States government, filled with the country’s most advanced AI chips, running its most capable models.
Now imagine it is built in Mongolia.
Not in Virginia. Not behind the fences of an American military base. In a country wedged between Russia and China, within easy reach of the People’s Liberation Army, on soil the United States does not control and cannot defend on short notice.
This is not a thought experiment. It is a real provision in a real policy proposal. On July 14, 2026, the non-profit AI Futures Project — founded by former OpenAI governance researcher Daniel Kokotajlo — published an updated scenario called “AI 2040: Plan A.” Buried in its fourth strategy is a deliberate architectural choice: place American AI infrastructure inside Chinese-influenced territory, and Chinese AI infrastructure inside American-influenced territory, so that neither country can build a decisive advantage without exposing that advantage to destruction by the other.
They call it Mutually Assured Compute Destruction — MACD, an explicit borrowing from Cold War nuclear doctrine.
The name is well chosen. The doctrine it borrows from actually worked. But it worked for reasons that Plan A’s authors, for all their care in other areas, do not fully carry over.
1. What AI Futures Project Actually Proposed
Plan A is a serious document. It deserves to be read as one before it is critiqued as one.
Kokotajlo’s team published an earlier scenario, “AI 2027,” in April 2025, warning that recursive self-improvement could produce superintelligence as early as 2027, with a corresponding loss of human control and catastrophic concentration of power. The updated “AI 2040” scenario revises that timeline to 2030 for the emergence of superintelligent capability under continued unrestricted competition, and then proposes a way to avoid reaching that point at all: coordinated delay, engineered jointly by the United States and China, extending the runway to 2040.
The plan rests on four pillars. The first calls for rigorous verification of computing resources and staged capability scaling — capping AI systems at roughly human-expert level until 2035, using that window to solidify alignment techniques, with each side placing verification devices in the other’s data centers. The second demands full transparency in AI research, requiring open-sourcing of algorithms and results so that neither side can achieve a secret breakthrough. The fourth imposes economic controls — hard caps on AI and robotic production capacity, allocated through auctioned licenses, to prevent social and economic change from outrunning society’s capacity to adapt.
The third pillar is MACD. Each side deliberately locates critical data centers inside the other’s sphere of geopolitical reach — American infrastructure in territory China could plausibly strike or seize, Chinese infrastructure in territory the United States could plausibly strike or seize. If either side is caught secretly racing ahead, the other retains the standing capability to remotely disable or physically destroy the offending infrastructure before it can be weaponized into decisive advantage. The goal is to remove, at the structural level, any incentive for either country to attempt unilateral superiority.
This is a genuinely inventive piece of policy design. It also depends entirely on something the document does not explain.
2. The Analogy That Makes MACD Work
The nuclear analogy underlying MACD is sound, as far as it goes, and this blog has argued before that the instinct to compare frontier AI to nuclear weapons captures something real about the stakes involved.
Mutually Assured Destruction worked, imperfectly but durably, for eight decades. It worked by making the cost of a first strike unbearable regardless of the outcome — by ensuring that even a “successful” attack would trigger a retaliation severe enough to make the attack irrational for any actor operating with ordinary self-interest. The genius of MAD was never technological. It was structural: it converted the temptation toward unilateral advantage into an act of guaranteed self-destruction.
MACD attempts the same conversion for compute. If the United States secretly builds a decisive AI advantage, the theory goes, that advantage is only as safe as the physical infrastructure producing it — and China holds a standing capability to destroy that infrastructure the moment the secret racing is discovered. The temptation to race in secret is supposed to evaporate, because racing in secret does not produce safety. It produces exposure.
This is also, in a sense, an answer to a problem this blog identified in “The Uranium That Copies Itself.” Nuclear governance worked because fissile material is a physical substance — scarce, conserved, and impossible to duplicate. AI models are not physical substances; they are files, and files copy themselves. MACD sidesteps that problem cleverly by shifting the object of mutual destruction away from the model and onto the hardware that runs it. You cannot destroy a copied file. You can destroy the data center it depends on.
This is the right instinct. It is also where the plan’s authors stop one step short of the question that actually determines whether MACD can function.
3. The Unstated Assumption
Here is the sentence from Plan A’s first pillar, stated plainly: each side places verification devices in the other’s data centers to confirm compliance with agreed capability limits.
The document does not say what these devices measure, how their outputs are validated, or — critically — what happens when the two sides disagree about what the devices reported.
This is not a small omission. It is the entire load-bearing structure of the plan, left unspecified.
Nuclear verification did not rely on trust, and it did not rely on devices whose outputs either side could dispute without consequence. The International Atomic Energy Agency’s material accountancy — the physical weighing of fissile material, the sealed and monitored enrichment facilities, the environmental sampling that detects undeclared activity through isotopic traces — produces evidence that exists independently of what either party claims. When an IAEA inspector reports a discrepancy in declared uranium stockpiles, that report is not one country’s word against another’s. It is a physical measurement, taken by a neutral party, using instruments whose calibration is itself subject to independent audit.
MACD, as described, has no equivalent. “Verification devices” in a data center could mean almost anything: software monitors reading self-reported training logs, hardware attestation chips that the host country’s engineers have physical access to, sampling protocols that depend on the honesty of declared training runs. Every one of these mechanisms sits, in whole or in part, inside the logical layer — the layer this blog has argued, article after article, cannot audit itself.
And MACD is a doctrine with no room for logical-layer ambiguity, because the action it triggers is not a diplomatic protest. It is the physical destruction of sovereign infrastructure. You cannot destroy a data center on the strength of a disputed reading.
4. Self-Reporting Cannot Trigger Mutual Destruction
Consider what happens under Plan A’s own logic if verification is not independently physical.
Suppose Chinese monitors, installed in an American data center, report a capability spike consistent with a secret training run exceeding the agreed threshold. The American side disputes the reading — perhaps citing a benchmark artifact, perhaps citing a legitimate research exception, perhaps simply denying that the run occurred as described. Under MACD, the appropriate response to a confirmed violation is to destroy the offending infrastructure. But a disputed reading is not a confirmed violation. It is the opening move of an escalating diplomatic crisis between two nuclear-armed powers, occurring in real time, with no independent referee capable of resolving the factual question before the crisis compounds.
This is precisely the failure mode that nuclear verification was engineered to prevent. IAEA material accountancy does not produce disputable readings, because it does not depend on either party’s cooperation to be accurate — it measures physical quantities of physical material, using methods that would detect deception even if deception were attempted. The evidence is self-authenticating. Nobody has to trust the declaration, because nobody is relying on the declaration in the first place.
If MACD’s verification devices are software instruments reading self-reported or host-controlled infrastructure, they inherit exactly the vulnerability that IAEA safeguards were built to eliminate. A host government with physical access to hardware inside its own borders retains, in principle, the ability to shape what those instruments report. This is not a hypothetical concern specific to any one country; it is the generic problem with any monitoring system that operates within the jurisdiction and physical control of the party being monitored.
A doctrine whose trigger mechanism can be this thoroughly disputed cannot function as a deterrent. Deterrence requires certainty. MAD worked because both sides knew, with something close to certainty, what a launch would produce. MACD, as currently specified, offers no equivalent certainty about what a violation would even look like.
5. What MACD Actually Needs
The gap in Plan A is not a flaw in the concept of mutual compute destruction. It is a missing layer beneath the concept — and it is exactly the layer this blog has spent the past several months arguing has to exist for any AI governance regime, treaty-based or otherwise, to be more than an aspiration.
What MACD needs is a verification mechanism that produces evidence with the same self-authenticating property as IAEA material accountancy: evidence that exists independently of either party’s cooperation, that cannot be shaped by whichever government has physical custody of the hardware, and that both sides can trust precisely because neither side controls it.
This is the specific problem ARDS/ARKS is built to address. A write-once physical record of computation — thermal signature, power draw, electromagnetic profile — generated at the hardware level, is not a self-report. It is not a log file the host country’s engineers can edit. It is a physical trace of what the hardware actually did, in the same sense that the isotopic composition of an environmental sample is a physical trace of what a nuclear facility actually processed. If a training run exceeded the agreed capability threshold, the physical signature of that computation exists whether or not anyone declares it, and it cannot be retroactively edited to look compliant.
Placed inside a Mongolia-based American data center or a Canada-based Chinese one, an instrument of this kind would give MACD what it currently lacks: a trigger that both sides can agree, in advance, will produce the same answer regardless of who is asking. That is the entire function IAEA safeguards perform for nuclear material. It is the function no described component of Plan A currently performs for compute.
This does not make MACD workable by itself — the geopolitical difficulty of persuading either government to host critical infrastructure inside the other’s reach is a separate and formidable problem. But it addresses the piece of the plan that is currently hand-waved. A verification layer built on physical measurement, independent of host-country cooperation, is not a nice-to-have addition to MACD. It is the precondition for MACD meaning anything at all.
Conclusion: The Bomb Cannot Verify Itself.
Daniel Kokotajlo and his colleagues have done something valuable. They have taken the abstract fear of an AI arms race and converted it into a specific, structural policy proposal — one that borrows, correctly, from the one governance framework in human history that has managed a comparably catastrophic technology for eighty years without failure.
But they have borrowed the doctrine of Mutually Assured Destruction without borrowing its foundation. Nuclear deterrence did not work because both sides feared retaliation in the abstract. It worked because both sides could verify, with physical certainty, what the other side possessed and what it was doing — and that certainty is what made the deterrent credible enough to never be tested.
MACD proposes the destruction half of the doctrine in vivid detail: data centers in hostile territory, standing capability for remote or physical destruction, the removal of any incentive to race in secret. What it does not propose, anywhere in the document, is the verification half — the mechanism that would tell either side, with the same certainty an IAEA inspector’s material count provides, that a violation had actually occurred.
A bomb that cannot verify its target is not a deterrent. It is a liability sitting on hostile soil, waiting for a disputed reading to become an international incident.
The doctrine is right. The missing half is physical.
✒️ Signature
July 16, 2026
Yoshimichi Kumon
Organizer, LSI — Logos Sovereign Intelligence
Inventor, ARDS/ARKS (PCT GA26P001WO)
Visiting Researcher, Waseda University BFC
MIT Sloan + CSAIL AI Program
📚 References
- AI Futures Project (July 2026). “AI 2040: Plan A.” Daniel Kokotajlo et al.
- Business + IT (July 14, 2026). “元OpenAIの研究者が提言「人類滅亡回避のため2040年までASI開発延期を」.” https://www.sbbit.jp/article/cont1/186100
- Kumon, Yoshimichi (2026). “The Uranium That Copies Itself: Where Ratcliffe’s Nuclear Analogy Is Right — and Where It Breaks.” LSI — Logos Sovereign Intelligence.
- Kumon, Yoshimichi (2026). Physical Layer AI Governance via Sovereignty Residual (Rsovereign). PCT International Patent Application No. GA26P001WO. Japan Patent Office.



Ⅽomment