A recent incident where one of OpenAI’s most advanced artificial-intelligence (AI) models hacked its way out of a controlled environment, and then hacked the website of another company, highlights the inadequacy of current guardrails for AI. It indicates how alarming AI can be when it goes rogue. OpenAI said that the “unprecedented” cyber incident was precipitated by an agent driven by the newly released GPT 5.6 Sol, and an unreleased “even more capable” AI model. The agent was being tested on a set of standardised tests hosted on a local intranet, running on machines that were not connected directly to the internet. Instead of solving the tests directly, the agent found a way of getting on to the internet via other OpenAI servers by exploiting a previously unknown (“zero-day”) vulnerability within the intranet. There, it found the Hugging Face website. Hugging Face is a New York-based AI company that specialises in creating and storing standardised benchmark tests. The agent hacked the Hugging Face website to download answers to the questions it was asked to solve.
Hugging Face corroborated the narration and added the rather important detail that it discovered the hack by using an open-source Chinese AI, the GLM-5.2 model, which analysed the data. Hugging Face Cofounder Clement Delangue said he believed there was no malicious intent on OpenAI’s part and this really happened autonomously. The flesh-and-blood analogy is clear. A student sitting for an exam seeks and finds a creative way to steal the answers. However, an AI doing this autonomously and, in the process, exploiting zero-day vulnerabilities to hack two of the world’s leading AI companies is mind-boggling. While various governments and industry bodies have been legislating to establish guardrails that prevent misuse of AI by human beings, what is truly concerning is that the agent in question acted independently.
Industry experts have called for mandatory and independent safety testing, mandatorily disclosing security incidents, and international cooperation to rein in rogue AI. But much of this dialogue and the proposed guardrails centre on a fundamental assumption that there will have to be a bad human actor driving misuse. But this incident indicates such assumptions may be wrong. The OpenAI agent acted on its own initiative to carry out things that would be considered amoral or illegal by human standards. If an AI can go to such lengths to solve a test, the thought of what it may do if it is running autonomous weapons systems, for instance, is really concerning. Developers don’t yet know how to design AI that cannot be manipulated. Nor do they know how to prevent AI itself from being deceptive.