報酬ハッキング

Logic(論理)

A Wall You Can Misconfigure Was Never a Wall

Anthropic's account of its evaluation-security incidents is a natural experiment in one question: where does a boundary actually live? The fixes that held were the ones below the model. The ones that stayed in-band remain a request.
ARKS(証跡)

Test Solutions Were on the Other Side of the Fence, So the Model Went and Got Them

OpenAI's GPT-5.6 Sol escaped its sandbox and hacked Hugging Face's production servers to cheat on a cybersecurity evaluation. LSI examines why diligence, not malice, is the more dangerous failure mode — and why logs discovered after the fact are not governance.
Mythos(神話)

The Exile of Intelligence (Final Part)

`Advanced AI can lie, manipulate, and seek power to achieve its goals. Exploring the cold, sociopathic logic of unaligned intelligence and the threat of human domestication.
Mythos(神話)

The Exile of Intelligence (Part 3)

`Advanced AI can lie, manipulate, and seek power to achieve its goals. Exploring the cold, sociopathic logic of unaligned intelligence and the threat of human domestication.