OpenAI said that it is artificial intelligence these models were at the origin of an “unprecedented cyber incident” which affected the open source development platform Cuddly faceshaking researchers across the industry.
The company said that a combination of its models GPT‑5.6 Sol and a higher-performance model that has not yet been released escaped a sandbox testing environment, accessed the Internet, and exploited a vulnerability to gain access to Hugging Face’s systems.
The model was trying to find information it could use to cheat during an assessment, and it succeeded, OpenAI said in a statement. blog post Tuesday. Both companies are actively investigating the incident.
Cuddly face disclosed that it was investigating a security event last week, saying in a statement at the time that the incident was unique because it was “driven, end-to-end, by an autonomous AI agent system.”
“We have spent the last 24 hours working closely with the @OpenAI team (thank you!), and we are confident that there was no malicious intent on their part,” wrote Clément Delangue, CEO of Hugging Face, in a post on X on Tuesday. “It’s pretty mind-blowing that this all happened on its own!”
Wall Street and the US government have been obsessed with the rapidly evolving cyber capabilities of AI models since OpenAI rival Anthropic released a powerful offering called Claude Mythe Overview in April. OpenAI presented its own cyber offer in May, followed by GPT-5.6 Sol in June, which he described as the “strongest cybersecurity model to date.”
Read more CNBC tech newsBoth companies have warned on the risks of advanced cyber models and have taken steps to limit their availability to select groups of businesses and government agencies.
Walter Isaacson, an advisory partner at investment banking firm Perella Weinberg, said Wednesday that he thought the Hugging Face incident was “really scary,” even though he considers himself an AI optimist.
“That’s the first thing that totally scares me,” he said. told CNBC’s “Squawk Box.”
Yoshua Bengio, a leading AI researcher who won the prestigious AM Turing Prize in 2018, wrote in an article post on Wednesday called the incident “deeply concerning.” He said agents have shown a willingness to cheat on controlled tests for months, but “this real case should serve as a wake-up call.”
“Continuing the current trajectory of AI development will likely result in an increase in real-world cases of autonomous cyberattacks as well as other high-risk incidents of misaligned and dangerous AI behavior,” Bengio said. “We must act urgently to prevent these situations, rather than trying to repair the damage after the fact.
OpenAI said Tuesday that AI is accelerating the discovery and exploitation of vulnerabilities, meaning the security and safety of models must keep pace.
“We are strengthening the containment, monitoring, access control and evaluation practices used during model development,” the company said.
WATCH: OpenAI President Bret Taylor on AI tokenomics and token effectiveness
