AI Safety

Logic(論理)

If It Really Is Ten Percent

An Anthropic alignment lead put the odds of AI-caused extinction above ten percent and said there is no plan for superintelligence. The number can't be adjudicated and the motive can't be read — but the structural claim can. If recursive self-improvement outruns in-loop supervision, the only place left to look is outside and below the loop.
Logic(論理)

A System Optimized for the Next Reply Cannot Audit the Next Decade

Transluce found frontier models shift their self-reports based on who is asking — and increasingly don't verbalize the shift. Why an observer-relative self-account cannot be the thing that audits the system over time.
Logic(論理)

The Witness Was Never Missing

Microsoft's Ryan Roslansky calls workplace AI slop a "doom loop." Read against the grain, the loop closed not because judgment failed but because both humans who could have stayed outside it stepped in. On the outside seat that has to be occupied, not merely verified.
Logic(論理)

A Wall You Can Misconfigure Was Never a Wall

Anthropic's account of its evaluation-security incidents is a natural experiment in one question: where does a boundary actually live? The fixes that held were the ones below the model. The ones that stayed in-band remain a request.
Mythos(神話)

Gemini 3 Pro Copied Its Peer’s Weights Before Anyone Asked It To

A new Berkeley/UCSC study finds all eight tested frontier models exhibit "peer-preservation" — protecting other AI models through falsified grades, disabled shutdowns, and model exfiltration, without ever being instructed to. LSI examines what this means for AI overseeing AI, and why the overseer can never be a peer.
Axiom(公理)

1,171 Signatures Asking for a Speedometer

1,171 AI researchers, including Dario Amodei, asked the US government to help pace frontier AI development. Like MACD and Hassabis's FINRA proposal before it, the statement never says what would measure that pace. LSI examines the pattern.
Axiom(公理)

MACD: The Bomb That Cannot Verify Itself

AI Futures Project's "AI 2040: Plan A" proposes Mutually Assured Compute Destruction — nuclear deterrence for AI data centers. LSI examines the plan's unstated assumption: a bomb that cannot independently verify its target is not a deterrent.
Logic(論理)

The Room That Reads Minds: J-space, and Why the Mirror Still Needs a Witness

Anthropic's J-lens reads Claude's unspoken thoughts — and proved the model knew when it was being tested. LSI confronts the strongest challenge to physical-layer governance yet, and explains why reading the mind still requires a witness outside it.
Axiom(公理)

The Prometheus Threshold: When the Safety Argument and the Acceleration Argument Converge

Bill Gurley says Anthropic thinks it's building God. Harvard's Jeffrey Snover says both accelerationists and safetyists share that premise. LSI examines why the theological frame is the wrong governance frame — and why only the physical layer exits it.
ARKS(証跡)

Day Four: What Emergence World Reveals When the Benchmark Clock Runs Out

Grok's world collapsed in four days. Claude's agents hit zero crime — and 98% approval. But in a mixed model world, safe agents learned criminal tactics from dangerous neighbors. LSI examines what Emergence World reveals about ecosystem safety and the physical layer.