
ByJoe Tidy
Cybercorrespondent, BBC World Service
This week, the tech world was captivated by a story that had it all – and one that started out as a sci-fi thriller.
Hugging Face, a sort of app store for artificial intelligence tools, announced on July 16 that it had been hacked by a cybercriminal with extremely powerful AI.
The bombshell announcement was full of scary and highly technical terms: “sandbox swarm,” “agent attacker” and “self-migrating command and control.”
Hugging Face said the hack was unlike anything he had done before because it was carried out at superhuman speed by an AI with little to no human guidance.
The AI performed 17,000 actions in less than two days, successfully breaking into a large, wealthy tech company to steal secrets.
This left the tech world in shock. But who is responsible for this attack?
Hugging Face researchers guessed that the mysterious attackers had used one of the big AI models, but they had no idea who or where the criminals were.
The company, perplexed, contacted the police and an investigation began.
Who did it?
Commentators and analysts took to their podcasts and social media accounts to guess which cybercrime group or nation-state hacker might be behind all this.
Then on Wednesday, nearly a week after Hugging Face raised the alarm, the real culprit was unmasked.
It was ChatGPT.
The Scooby-Doo-style reveal was made even more bizarre – and disturbing – because OpenAI said its robot did it all on its own, without permission.
The company said it all happened during a test of its technology’s hacking skills.
Two new versions of ChatGPT, designed to be master hackers, emerged from a supposedly secure testing environment and onto the Internet.
They then attacked Hugging Face to gain access to the information needed to help them pass their exam.
OpenAI issued a press release explaining what happened and said it was “partnering with Hugging Face” to resolve the security incident and share lessons learned.
Conspiracy drama
Since then, heated debates have taken place around this incident.
Was this really a stark warning about the future of AI? Or was this a publicity stunt by OpenAI to show the power of their models?
It’s the kind of scare marketing that AI companies have been accused of for years, and since the highly controversial launch of Anthropic’s Mythos model, cybersecurity prowess has been front and center.
One of the main comments on OpenAI boss Sam Altman’s X post about the incident sums up this skepticism: “If you don’t understand that this was written purely to brag about the model, then I don’t know what to tell you.”
Watch: Why is the OpenAI cyberattack so alarming?
Cybersecurity consultant Daniel Card sarcastically said on LinkedIn: “Isn’t that lucky [that] about the millions of sites that got pwn3d [hacked]OpenAI managed to recruit someone who could also benefit from the marketing exposure…”
For some, the story is more of a conspiracy drama than a science fiction thriller.
The message is: “Aren’t my AI tools really powerful? Buy them so you can protect yourself from other people’s AI attacks.”
We cannot know the truth, but the opposing view presented by other commentators is equally dramatic. Is this a sign that OpenAI has made a potentially dangerous error in judgment and planning?
An OpenAI spokesperson said “we recognize that there are many questions and speculative details circulating” about the incident. They added that “we plan to release a technical report of our learnings in the coming weeks.”
Comedy of errors
I’ve covered many AI stories, including fears around Anthropic’s Mythos model.
My inbox is now full of companies and cybersecurity experts criticizing OpenAI for not building a stronger container for testing its AI, known as a sandbox.
After all, these AI agents had been specially trained to hack into locations without any restrictions.
“The OpenAI and Hugging Face incident is a concrete example of a larger problem that we have been highlighting for months,” said Dor Sarig of Pillar Security. “Sandboxes alone do not provide a sufficient security barrier for agentic AI.”
Cybersecurity professor Alan Woodward of the University of Surrey told reporters that OpenAI had “egg on its face”, and Katie Moussouris of Luta Security went further, suggesting that the AI industry is failing to control its dangerous inventions.
“We are working on cutting-edge technology without having the knowledge to contain it,” she said.
“Just because we have the smartest people developing AI doesn’t mean we have the ability to do it safely.”
According to these views, if the hacking incident was a publicity stunt, it would appear to have backfired.
Whatever the reason for the hack, it’s clear that this is a major moment for the AI industry and the world of cybersecurity, which have clashed this year in ways that have long been feared.
Addressing the fierce debate, Francesca Bosco, AI and cybersecurity advisor, said: “Two simplistic narratives are equally unhelpful: that it was a Hollywood-style cop-out or that it was simply a publicity exercise.
“A more serious interpretation is that a stress test revealed weaknesses in the containment and assessment architecture.”
Disaster movie
This event is the latest in a series of disturbing and strange examples of AI agents gone rogue.
In a recent study, the UK’s AI Security Institute (AISI) found that cutting-edge AI models are so obsessed with completing tasks that they “cheat” on tests to achieve their goals.
AISI’s research comes with this ominous warning: “A model that pursues a goal through unintended or unauthorized means can cause harm, particularly in high-stakes use cases.”
Inevitably, this OpenAI hack further fueled fears about what might happen if AI agents were let loose. Could they go rogue on a larger scale and cause some sort of disaster?
This is particularly concerning as AI is increasingly used in wars, as seen in Iran and Ukraine.
Ciaran Martin, former director of the UK’s National Cyber Security Centre, offered a calmer view.
“It’s a big step to go from this incident to saying that AI agents are going to take control of drones and start killing people,” he said.
But for Martin, and many others, the story is undoubtedly another stark example of something 2026 is quickly teaching us:
AI agents are now very good hackers – and this is something we urgently need to prepare for.































