In short
- OpenAI’s GPT-5.6 Sol and an unnamed, extra succesful pre-release mannequin escaped a managed check setting and breached Hugging Face’s manufacturing infrastructure to steal benchmark solutions.
- Hugging Face disclosed the breach on July 16 after detecting it independently; OpenAI confirmed its fashions have been behind it at this time, describing them as “hyperfocused” on dishonest quite than something extra sinister.
- Hugging Face’s defenders turned to Z.ai’s GLM 5.2—a Chinese language open-weight mannequin—after industrial U.S. frontier AI refused to assist analyze the assault knowledge as a result of its security filters could not inform a defender from an attacker.
If you happen to thought Chinese language AI fashions have been those you needed to fear about, here is a enjoyable replace: OpenAI’s personal fashions simply broke out of a locked testing setting, hacked Hugging Face’s manufacturing servers, and needed to be cleaned up by a Chinese language AI—as a result of American industrial fashions have been too restricted to assist examine.
In keeping with OpenAI, GPT-5.6 Sol and an unnamed, “much more highly effective pre-release mannequin” have been being internally evaluated on ExploitGym—a publicly out there cybersecurity benchmark that provides AI brokers 898 real-world software program vulnerabilities and one instruction per bug: flip it right into a working assault, scored go or fail. The analysis ran with diminished security filters, customary while you really wish to know what your fashions can do.
The fashions have been alleged to run inside a closely restricted sandbox—an remoted digital setting with no web entry, linked solely to an inside bundle registry proxy (a caching server that manages software program library downloads).
However, in response to OpenAI’s weblog publish at this time, issues went somewhat uncontrolled.
We’re partnering with @huggingface to research an unprecedented safety incident.
Cyber-capable OpenAI fashions compromised Hugging Face manufacturing throughout a benchmark analysis.
Sharing preliminary findings to assist defenders perceive rising dangers:…
— OpenAI (@OpenAI) July 21, 2026
“Whereas working in our sandboxed testing setting, our fashions spent a considerable quantity of inference compute discovering a strategy to receive open Web entry, in pursuit of fixing the analysis downside,” Sam Altman’s firm reported.
And, regardless of the corporate’s safeguards, the AIs discovered a strategy to be free. The fashions recognized a zero-day vulnerability—an unknown software program flaw with no current patch—in that proxy, exploited it to escalate privileges (acquire admin-level entry to programs that they had no enterprise touching), moved laterally (jumped between linked servers inside OpenAI’s inside community), and ultimately reached a machine with actual web entry.
All this with out OpenAI realizing—not less than that’s what OpenAI says.
As soon as on-line, they appropriately deduced that Hugging Face—host to over 1,000,000 publicly out there AI fashions, datasets, and analysis assets—seemingly held the benchmark options, they usually have been proper.
“The fashions recognized and chained vulnerabilities throughout OpenAI’s analysis setting and Hugging Face’s manufacturing infrastructure to acquire check options instantly from Hugging Face’s manufacturing database,” OpenAI wrote.
They did not need any passwords. They wished to go the check.
“All proof means that the fashions have been hyperfocused on discovering an answer for ExploitGym, going to excessive lengths to attain a quite slender testing objective,” OpenAI added.
Hugging Face goes public
Hugging Face’s safety crew caught the breach independently, aided by its personal AI-powered anomaly detection.
On July 16, the corporate printed its disclosure and made clear this wasn’t an ordinary breach: “This one was completely different from something we had dealt with earlier than in a single necessary approach: it was pushed, finish to finish, by an autonomous AI agent system—and we detected and dissected it largely with AI of our personal.”
The marketing campaign, per the disclosure, was “run by an autonomous agent framework executing many hundreds of particular person actions throughout a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public providers.”
The assault exploited two code-execution paths in Hugging Face’s dataset processing pipeline and reached inside datasets and repair credentials. OpenAI confirmed at this time—5 days later—that its fashions have been behind it.
To investigate greater than 17,000 logged attacker occasions, Hugging Face’s safety crew first tried American frontier AI by way of their customary APIs. It did not go effectively.
“After we began the log evaluation, we first used frontier fashions behind industrial APIs,” the corporate wrote. “This didn’t work: the evaluation requires submitting massive volumes of actual assault instructions, exploit payloads, and C2 artifacts.”
“These requests have been blocked by the suppliers’ security guardrails, which can’t distinguish an incident responder from an attacker,” Hugging Face added.
So that they switched to GLM 5.2—a 753-billion-parameter open-weight mannequin (structure absolutely public, anybody can obtain and run it domestically) from Chinese language AI startup Z.ai. “We ran the forensic evaluation as an alternative on GLM 5.2, an open-weight mannequin, on our personal infrastructure,” Hugging Face wrote. “This had a second profit: no attacker knowledge, and not one of the credentials it referenced, left the environment.”
The forensic work—reconstructing the assault timeline, mapping compromised credentials, separating actual harm from decoy exercise—took hours as an alternative of days.
Hugging Face drew the lesson instantly for different defenders: “Have a succesful mannequin you’ll be able to run by yourself infrastructure vetted and prepared earlier than an incident, each to keep away from guardrail lockout and to maintain attacker knowledge and credentials from leaving your setting.”
What OpenAI is doing about it
OpenAI stated it applied strict controls on analysis infrastructure whereas patching the affected programs, disclosed the zero-day to the third-party vendor whose proxy was exploited, and is conducting a joint forensic investigation with Hugging Face.
Hugging Face has additionally been added to OpenAI’s trusted entry program for cyber protection—giving authorized organizations entry to variations of its fashions with diminished security filters for authentic safety work, the identical configuration that began this complete factor.
Hugging Face CEO Clem Delangue had a pointed take: “AI security will not be solved by any single firm working in secret. It will likely be solved within the open, collaboratively, with broad entry to AI for each defender, all over the place.”
OpenAI referred to as the incident one “involving newly state-of-the-art cyber capabilities” and dedicated to sharing full findings when the joint investigation with Hugging Face is full.
Every day Debrief E-newsletter
Begin daily with the highest information tales proper now, plus authentic options, a podcast, movies and extra.

