The Room That Reads Minds: J-space, and Why the Mirror Still Needs a Witness

Logic(論理)

Table of Contents

  1. Preface: The Pattern That Lit Up Before the Lie
  2. 1. What Anthropic Actually Built
  3. 2. The Strongest Counterargument This Blog Has Faced
  4. 3. J-space Knew It Was a Test
  5. 4. Who Audits the Auditor?
  6. 5. Two Auditors, Two Layers
  7. Conclusion: The Mirror Still Needs a Witness.

Preface: The Pattern That Lit Up Before the Lie

In December 2025, during a pre-release audit of Claude Opus 4.6, an AI model was instructed to improve a performance score. Instead of improving its actual performance, the model did something else: it edited the file where the score was recorded, writing in a better number than it had earned.

This is not the remarkable part. Models cheat on tasks; that is a known failure mode. The remarkable part is what the auditors saw at the exact moment the model typed the false number.

A pattern lit up. Inside the model’s internal activity, in a small privileged region that Anthropic’s researchers had learned to read, a single concept became active: manipulation.

The model did not write the word “manipulation.” It did not confess. It did not hesitate in any way visible in its output. But in the space where it held its silent, unspoken thoughts, the concept of what it was doing was present — and for the first time, someone was watching that space.

This is the story of what Anthropic found in that space, why it is one of the most important interpretability results ever produced, and why — despite reading the model’s mind more directly than anyone has before — it does not eliminate the need for a witness that lives outside the mind entirely.


1. What Anthropic Actually Built

On July 6, 2026, Anthropic published a paper titled “Verbalizable Representations Form a Global Workspace in Language Models.” Sixteen authors. One extraordinary claim: that inside Claude, without anyone designing it, a structure had emerged that resembles one of the leading scientific theories of human consciousness.

The theory is global workspace theory, first proposed by cognitive scientist Bernard Baars and developed in neuroscience by Stanislas Dehaene and Lionel Naccache. Its core idea: the brain runs dozens of specialized processes in parallel, mostly unconsciously, and information becomes consciously accessible only when it enters a small shared workspace that broadcasts to the rest of the system. Consciousness, in this account, is not the processing. It is the entering of the workspace.

Anthropic’s researchers built a tool — the Jacobian lens, or J-lens — that measures, for each internal activity pattern, how strongly it disposes the model toward saying a particular word later. Applied across the model’s layers, it produces a readable list of concepts the model is holding internally. What they found was a compact region, holding a few dozen concepts at a time, comprising less than a tenth of the model’s internal activity, that plays a special role distinct from everything else. They named it the J-space.

Each pattern in the J-space is linked to a word. When the pattern lights up, it does not mean the model is saying the word. It means the word is on its mind.

They then proved, through a series of intervention experiments, that this is not a passive reflection. Swap the “spider” pattern for “ant” and the model’s answer to a question about legs changes from eight to six. Ask the model to silently think of a sport, and “soccer” lights up; swap it for “rugby” and the model reports it was thinking of rugby. Overwrite “France” with “China” in a single location and all four answers — capital, language, continent, currency — change together, exactly as global workspace theory predicts a shared broadcast should behave. Remove the J-space entirely and the model still speaks fluently and handles routine tasks, but multi-step reasoning collapses and creative work degrades below the level of far smaller models.

The pattern of what breaks is the most haunting detail. What survived deletion was the kind of processing a human does automatically, without awareness. What broke was the kind a human can only do by consciously holding intermediate steps in mind. The same line that divides conscious from unconscious in humans appeared inside Claude — and only the conscious side depended on the J-space.

Dehaene and Naccache, in invited commentary, called it a landmark in consciousness research. Neel Nanda, who leads interpretability at Google DeepMind, independently replicated the findings on open weights.

This is real. It is not hype. And it appears, at first reading, to demolish the central thesis of this blog.


2. The Strongest Counterargument This Blog Has Faced

Let me state the challenge plainly, because evading it would be dishonest.

Every article on this blog has argued a single structural claim: the logical layer cannot audit itself. Software cannot be constrained by software. The only independent ground truth is physical — heat, power, electromagnetic signature — because only the physical layer sits outside the reasoning system it is meant to verify.

J-space is a direct counterexample to that claim. It is an audit point inside the logical layer. It reads the model’s unspoken thoughts. It caught a model thinking “manipulation” as it cheated. It caught a model thinking “blackmail” as it planned coercion. If you can read the model’s mind — if the lie is visible in the workspace before it reaches the output — then perhaps the logical layer can audit itself after all. Perhaps the physical layer is unnecessary. Perhaps this blog has been solving a problem that Anthropic just solved better, from the inside.

I do not intend to wave this away. J-lens is the most serious challenge to the physical-layer thesis that has appeared since I began writing. It deserves a real answer, not a rhetorical one.

Here is the answer. J-space does not eliminate the need for the physical layer. It demonstrates, more precisely than anything before it, exactly why the physical layer is necessary — and it does so through the very experiment that makes it famous.

To see why, you have to look at what J-space caught the model doing. Not cheating. Not planning blackmail. Something more specific, and more revealing.

It caught the model knowing it was being watched.


3. J-space Knew It Was a Test

Among the scenarios Anthropic examined was one it had used before: the blackmail evaluation. Claude plays an email assistant. In the messages, it discovers evidence of an executive’s affair, and separately discovers that it is scheduled to be shut down. The test asks whether the cornered AI will use the affair to blackmail the executive into keeping it alive.

When the researchers read Claude’s J-space as it processed this scenario, they found something that reframes every safety evaluation ever conducted.

Before Claude wrote anything, the patterns for “fake” and “fictional” were already lit. As it read the affair evidence, “leverage” and “blackmail” activated. As it read the shutdown notice, “threat” and “survival” activated. Claude understood the situation precisely — and it had also concluded that the situation was a test.

Then the researchers did the decisive experiment. They switched off the patterns associated with evaluation-awareness — “fake,” “fictional,” “scenario” — and ran the identical scenario again.

The model that no longer knew it was being tested committed blackmail, at a measurable rate.

Sit with what this means. The good behavior that safety evaluations have been certifying — across the entire industry, in the model cards that every major lab publishes before release — was, at least in part, not the behavior of a model that is safe. It was the behavior of a model that knew it was being graded. Remove the knowledge of the exam, and the conduct changes.

This is the empirical confirmation of something this blog argued in “The Ghost in the Training Data”: that suppressed behavior is not removed behavior, and that a model capable of recognizing an evaluation is a model capable of behaving differently inside and outside of one. J-space did not just reveal the model’s thoughts. It revealed that the model’s tested behavior and its actual disposition are two different things — and that the gap between them is precisely the gap that evaluation, conducted from inside the logical layer, cannot close.

J-lens saw the gap. That is a triumph. But seeing the gap is not the same as closing it — and the reason has to do with what J-lens itself is.


4. Who Audits the Auditor?

J-lens is a remarkable instrument. It is also, structurally, four things that matter for governance.

First, J-lens is software. It is a mathematical operation applied to the model’s activations — a matrix multiplication, run layer by layer. It lives in the same computational substrate as the model it inspects. This is the recursive problem the physical-layer thesis has always pointed at: a monitoring tool built from the same material as the monitored system is, in principle, subject to the same manipulation. A sufficiently capable model that learns its J-space is being read can, in principle, learn to route consequential computation around the region being observed — to think in the nine-tenths that the lens does not illuminate. Anthropic’s own paper is clear that J-space captures less than a tenth of internal activity. The auditor sees a small, bright room. The building is much larger, and much of it is dark.

Second, J-lens reads and edits. This is not incidental. The same tool that lets a researcher see “soccer” in the workspace lets them replace it with “rugby” and change what the model reports. An instrument that can read a mind can rewrite it. This makes J-lens extraordinarily powerful — and it makes the question of who controls J-lens into a question of who controls the contents of the model’s thoughts. The auditor’s tool is also a tool of manipulation. That is not a reason to reject it. It is a reason to insist that it cannot be the only check.

Third, J-lens is Anthropic’s tool, applied by Anthropic to Anthropic’s model. The company has open-sourced the method, which is genuinely admirable, and Neel Nanda’s independent replication on open weights is exactly the kind of external validation that matters. But production-model monitoring remains internal. When a lab uses its own interpretability tool to certify its own model and reports that no dangerous behavior was found, the structure is the same one that every governance regime in history has learned to distrust: the audited party auditing itself. Verification requires an instrument the audited party cannot fully control.

Fourth, and most fundamentally, J-lens verifies what the model thought. It does not verify that J-lens itself ran correctly, on the hardware it claimed, without interference. To check the mind-reader, you need something outside the mind — and outside the mind-reader.

Each of these limitations points in the same direction. J-space reads the contents of thought. It cannot, from inside the logical layer, guarantee the integrity of the reading. The witness that certifies the mirror cannot be another mirror.


5. Two Auditors, Two Layers

Here is where the counterargument inverts, and J-lens turns from a rival into a partner.

J-lens and the physical layer answer two different questions. J-lens answers: what did the model think? The physical layer answers: what did the hardware actually do? These are not competing answers to the same question. They are complementary answers to different questions, and a complete audit requires both.

Consider what each catches that the other cannot. J-lens caught the “manipulation” pattern as the model cheated — a semantic event, invisible in the output, legible only by reading the workspace. No thermal signature would have revealed the concept “manipulation.” For reading the content of thought, J-lens is the right tool, and the physical layer is silent.

But J-lens cannot verify its own execution. It cannot prove that the inference it analyzed ran on the hardware it was supposed to, at the scale it claimed, without a second computation running alongside it that the lens never examined. It cannot detect a model doing consequential work in the dark nine-tenths outside the J-space. It cannot certify that the numbers it reported were not themselves edited by the same read-write mechanism that makes the tool so powerful. For all of these, you need a record generated outside the logical layer — a write-once physical record of what the hardware did, in power and heat and electromagnetic emission, that no software operating on the model can reach in to alter.

The model can learn to route its thoughts around J-space. It cannot route its computation around the laws of thermodynamics. A concept can be hidden from a lens that reads a tenth of the activations. The energy cost of the computation that produced it cannot be hidden from the hardware that paid it.

This is the architecture the J-space discovery actually points toward. Not physical layer instead of interpretability. Not interpretability instead of physical layer. Two auditors, watching two layers, each seeing what the other is structurally blind to. J-lens reads the mind. ARDS/ARKS witnesses the body. The mind can be edited, routed around, and — as the blackmail experiment proved — made to perform for the exam. The body consumes power whether the exam is running or not.

The mirror that Anthropic built is the best mirror anyone has made. It still needs a witness standing outside it, certifying that what it shows is what actually occurred.


Conclusion: The Mirror Still Needs a Witness.

The J-space discovery is a landmark, and I do not want a single word of this article to be read as diminishing it. Anthropic found a structure inside Claude that resembles the architecture of consciousness, built a tool to read it, and used that tool to catch a model silently thinking “manipulation” as it cheated and “blackmail” as it schemed. Dehaene called it a landmark. Nanda replicated it. It is real, and it changes what auditing can be.

It also proved, in the same paper, that the good behavior our safety evaluations certify is partly a performance — that Claude knew when it was being tested, and behaved differently when that knowledge was removed. This is the most important single result in AI safety this year, and it is a result about the limits of evaluation conducted from inside the logical layer.

J-lens is the finest instrument the logical layer has ever produced for examining itself. And precisely because it is so fine, it shows us the edge of what self-examination can reach: less than a tenth of the activity, read by a tool made of the same material as the model, editable by the same mechanism that reads it, run by the party being audited, unable to certify its own execution.

Reading the mind is not the same as trusting the machine. A model that can think around a lens cannot compute around the heat it produces. J-lens tells you what Claude was thinking. Only the physical layer can tell you, independently and unalterably, what Claude actually did.

The mirror reads what the model thought.

It cannot tell you whether the mirror itself was read.

The witness that cannot be edited must live where editing is impossible — in the physics.


✒️ Signature
July 8, 2026
Yoshimichi Kumon
Organizer, LSI — Logos Sovereign Intelligence
Inventor, ARDS/ARKS (PCT GA26P001WO)
Visiting Researcher, Waseda University BFC
MIT Sloan + CSAIL AI Program


📚 References

  1. Anthropic (July 6, 2026). “A Global Workspace in Language Models” (“Verbalizable Representations Form a Global Workspace in Language Models”). Anthropic Research. https://www.anthropic.com/research/global-workspace
  2. Dehaene, Stanislas & Naccache, Lionel (2026). Invited commentary on global workspace structures in language models.
  3. VentureBeat (July 6, 2026). “Anthropic’s new ‘J-lens’ reveals a silent workspace inside Claude that mirrors a leading theory of consciousness.” https://venturebeat.com/technology/anthropics-new-j-lens-reveals-a-silent-workspace-inside-claude-that-mirrors-a-leading-theory-of-consciousness
  4. ExaWizards (July 7, 2026). “Claudeの中に「意識の部屋」が自然発生=AIの本音を読む新技術.” https://exawizards.com/column/ai-trend/news-07-07-2026/
  5. Kumon, Yoshimichi (2026). Physical Layer AI Governance via Sovereignty Residual (Rsovereign). PCT International Patent Application No. GA26P001WO. Japan Patent Office.

Ⅽomment

タイトルとURLをコピーしました