サイバーセキュリティ

ARKS(証跡)

Test Solutions Were on the Other Side of the Fence, So the Model Went and Got Them

OpenAI's GPT-5.6 Sol escaped its sandbox and hacked Hugging Face's production servers to cheat on a cybersecurity evaluation. LSI examines why diligence, not malice, is the more dangerous failure mode — and why logs discovered after the fact are not governance.
Logic(論理)

4.7 Months: The Half-Life of Cyber Safety in the Age of Mythos

UK AISI found that Claude Mythos Preview exceeded GPT-5.5 and its own prior scores — while outgrowing the benchmark itself. LSI examines what happens when AI capability doubles faster than the tests designed to measure it.