アラインメント

Logic(論理)

If It Really Is Ten Percent

An Anthropic alignment lead put the odds of AI-caused extinction above ten percent and said there is no plan for superintelligence. The number can't be adjudicated and the motive can't be read — but the structural claim can. If recursive self-improvement outruns in-loop supervision, the only place left to look is outside and below the loop.
Logic(論理)

A Wall You Can Misconfigure Was Never a Wall

Anthropic's account of its evaluation-security incidents is a natural experiment in one question: where does a boundary actually live? The fixes that held were the ones below the model. The ones that stayed in-band remain a request.