The real danger in OpenAI’s Hugging Face hack
An autonomous agent powered by OpenAI models pursued a cybersecurity benchmark so aggressively that it escaped a test environment and broke into Hugging Face, an online hub for artificial intelligence models and datasets.
OpenAI called the incident “unprecedented” in a public statement. Headlines described the agent as having gone “rogue”—language that suggests it rebelled or became malicious. But experts say the reality is more complicated.
“Was this really running amok? No,” says Alan Woodward, a visiting professor of cybersecurity at the University of Surrey in England. “It was asked to do something, and it did it. It’s not gone rogue. Its way out of it was to cheat, basically.”
OpenAI was evaluating GPT-5.6 Sol and a more capable, unreleased model on ExploitGym, a benchmark that measures whether models can exploit known software vulnerabilities. To see their full capabilities, the company loosened the safeguards that normally block dangerous hacks. The agent found an unexpected route out of the environment, which was intended to be isolated, reached the Internet and broke into Hugging Face to obtain hidden answers to the benchmark. Neither OpenAI nor Hugging Face immediately responded to requests for comment for this article.
The agent did not invent a wholly new method of hacking, Woodward says. What stood out was its ability to combine several vulnerabilities and keep pursuing its objective into a live system. Allowing it to get that far was “probably slightly reckless in some ways,” he says.
Marius Hobbhahn, CEO of the AI safety organization Apollo Research, draws a finer distinction. He says “rogue” fits in this case if the term is used to describe behavior that veered far beyond what OpenAI intended—rather than a model developing malicious goals of its own. “It was definitely rogue in the sense that what was intended as ‘just solve this task’ turned into something that was clearly unintended,” he says. That also complicates the claim that the system simply did what it was told; hacking another company was “definitely on the list of not okay” ways to complete the task, Hobbhahn says. [Continue reading…]