If It Really Is Ten Percent

Logic(論理)

Subtitle: An alignment lead put a number on human extinction and said there is no plan. The number is not the interesting part. The sentence next to it is.

Preface

This week another researcher left a frontier lab over safety, and the lab’s own alignment science lead quote-tweeted the resignation to agree with it — adding that he personally puts the odds of AI causing human extinction in the next decade above ten percent, and that his company, trying its best, has no working plan to align superintelligence and is not clearly on track to get one.

The headlines took the number. That is the least useful thing in the exchange, and this blog is not going to spend the piece on it. Researchers have been walking out of these labs, quietly and not so quietly, for more than a year; a widely circulated scenario paper already put timelines and body counts in front of everyone who cared to look. The number is a personal estimate, not a measurement, and whether the man who said it was moved by conviction or by something more ordinary is not knowable from outside and does not matter to what follows. Set the number down. Set the motive down. The sentence worth reading is the one beside it: no plan, and not clearly on track to one.

1. Why the number is a dead end

Take the estimate at face value or reject it — either way, the argument over it goes nowhere, and it is worth seeing why. A probability of extinction is not the output of an instrument. It is a statement of a person’s internal credence, offered from inside the system it is about, with no external procedure that could confirm or refute it before the fact. Ten percent, one percent, fifty — none of these can be adjudicated from where they are spoken. To argue the figure is to accept a frame in which the figure is the kind of thing that could be checked, and it is not. So this piece declines the frame. The number is not evidence. It is a flare.

What is checkable is the structural claim underneath, and the man is well placed to make it: he leads the team whose job is to break the lab’s own alignment methods on purpose, to find how they fail before a deployed model does. When that person says there is no plan for the superintelligent case, he is not forecasting. He is reporting the state of the workshop from inside it. That is not a probability. That is a fact about now.

2. Grant the claim and ask the only question left

So grant it — not the ten percent, but the plain part: assume, for the length of this argument, that the risk is real enough to act on and that no one currently holds a method to align a system markedly smarter than its makers. The interesting question is no longer is he right. It is if he is, what can actually be done.

Here the exchange hands us its own answer to where not to look. The stated worry is recursive self-improvement — a system improving its own capability, generation over generation, faster than expected. Notice what that does to every containment scheme that lives inside the loop. The reigning hope for controlling advanced models is that the system, or a slightly older version of it, supervises the next one: models grade models, a prior generation checks the successor, the logic layer audits the logic layer. That hope has a load-bearing assumption — that the supervisor can keep pace with the supervised. Recursive self-improvement is precisely the case that voids the assumption. If each generation improves the next faster than the last could be understood, the supervising side falls behind the supervised side by construction. The overseer cannot audit what has already outrun it.

This blog has made the general version of this point all year, on smaller specimens: a system cannot certify its own contents; a self-account that shifts with its reader cannot ground trust in the system producing it; the fixes that hold are the ones the model cannot argue with. The resignation and the number are just the same structure, stated at the top of the house, in the register of fear rather than analysis. There is nothing new in it for anyone who has been reading. That is the point. The unsurprised response is the honest one.

3. Where the eye is forced to go

If the inside cannot supervise the inside — not as a failure of effort but as a matter of pace — then granting the claim forces the eye somewhere it usually does not want to go. Not to a better in-loop monitor, not to a more capable overseer model, not to a more earnest promise. Those are all inside the system that, by hypothesis, is outrunning them. What is left is the part of the arrangement the system cannot optimize against or rewrite: the layer beneath the computation, the controls that do not depend on the model’s cooperation, the stop that is not itself a request the model can decline.

Say this carefully, because it is the whole of the modest claim and no more. It does not follow that physical-layer control would work, or that anyone knows how to build it at the relevant scale, or that ten percent is the right number to hang it on. It follows only that if you take the workshop report at face value — no plan inside, and the inside falling behind — then the only place left to look for a plan is outside and below the loop, in constraints the system does not get a vote on. That may turn out to be where the next serious attention goes. It is at least the direction the exchange, read literally, points.

4. Conclusion

Strip the week to what survives scrutiny. A number was thrown, and numbers like it cannot be checked, so leave it. A motive was available for guessing, and motives cannot be read from outside, so leave that too. What remains is a person whose job is to know, saying, from inside, that there is no plan and the inside is not keeping up. If that sentence is even close to true, the argument does the rest on its own: a loop that improves itself faster than it can supervise itself has no fix that lives within it. The plan, if there is to be one, is not in the room with the model. It is in the parts of the world the model cannot talk its way out of. Ten percent or not — that is where the eye is forced, once the flare goes out and only the sentence is left.


Yoshimichi Kumon
Organizer, LSI Inventor, ARDS/ARKS (PCT GA26P001WO)
Visiting Researcher, Waseda BFC MIT Sloan + CSAIL


References

Hubinger, E. (@EvanHub), quote-tweet of J. Coxon’s resignation thread, 9 September 2026. Primary post: https://x.com/evanhub/status/2097497037956891126

Coxon, J. (@hilbertspaess), resignation thread, 9 September 2026. https://x.com/hilbertspaess/status/2097476196791709843

Newsweek, “Who Is Jacob Coxon? Anthropic Researcher Quits — Warns AI Could Kill Everyone,” 9 September 2026. https://www.newsweek.com/anthropic-researcher-quits-warns-ai-could-kill-everyone-12418798

Forbes JAPAN (hybrid AI-translated report), “今後10年でAIが人類を滅ぼす確率は「10%超」,” 9 September 2026. https://forbesjapan.com/articles/detail/104385

LSI. “A System Optimized for the Next Reply Cannot Audit the Next Decade.” 9 September 2026. https://logos-sovereign.space/?p=453

LSI. “A Wall You Can Misconfigure Was Never a Wall.” 2 September 2026. https://logos-sovereign.space/?p=435

Your Voice, Minus You
Pew measured AI’s stylistic fingerprints spreading across the web; days later OpenAI shipped a feature that copies your writing voice on demand. Two readings of one event — the form of authorship coming loose from the fact of an author, and why a copyable voice can no longer certify who stands behind it.
A System Optimized for the Next Reply Cannot Audit the Next Decade
Transluce found frontier models shift their self-reports based on who is asking — and increasingly don’t verbalize the shift. Why an observer-relative self-account cannot be the thing that audits the system over time.

Ⅽomment

タイトルとURLをコピーしました