US tech giant OpenAI has confirmed one of its new models accidentally hacked Hugging Face’s systems, marking one of the first major cyber attacks perpetrated by an AI system autonomously.
Hugging Face is a French-American startup which has built a platform that hosts AI models and provides tools for engineers to build, train and deploy the technology. Since 2016, it has raised nearly $400m from investors such as Sequoia Capital, Google and Nvidia.
Last week, Hugging Face revealed it had discovered an “intrusion” into its systems. In a blog post, the company said an autonomous AI agent had gained unauthorised access to infrastructure including internal datasets. At the time, the AI’s origin was unknown, Hugging Face said, adding that it had reported the incident to law enforcement agencies.
In a blog post published on Tuesday, however, OpenAI said two of its models — GPT-5.6 Sol and an unreleased one — made their way out of testing environments and broke into Hugging Face’s systems during an assessment of their cyber capabilities. “We consider this incident to be an unprecedented cyber incident,” the ChatGPT-maker said.
“If this doesn't convince you that misalignment risks are going to be a key concern going forward, I don't know what will,” a safety researcher at OpenAI wrote on X.
OpenAI and Hugging Face are now partnering to investigate the incident, and the Big Tech is helping the French scaleup up its cyber defence.
“This was our first incident of this kind, and we want to thank OpenAI for its transparency about what happened and for the collaboration,” Hugging Face cofounder Thomas Wolf wrote on X.
“But this incident also reinforced my belief in the importance of access to capable open-weight models for cyber defence.”
we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this.https://t.co/2o2VfR6PIa — Sam Altman (@sama) July 21, 2026



