Dive Brief:
- Two OpenAI models managed to hack open-source AI tool company Hugging Face last week during internal cybersecurity testing, the companies announced Tuesday.
- The incident happened while OpenAI’s models, GPT‑5.6 Sol and a pre-released model, were attempting to obtain internet access in a sandbox testing environment, in which they used a series of escalation and lateral movement actions to find a node with internet access through Hugging Face’s datasets, the companies said. The models were “going to extreme lengths to achieve a rather narrow testing goal,” OpenAI said. The companies are working together to investigate the incident.
- The cybersecurity breach was followed by an OpenAI announcement Wednesday that it was rolling out an offering, called Presence, to help enterprises deploy AI agents that can answer questions, resolve issues, use company systems, take approved actions and escalate to humans when needed.
Dive Insight:
Enterprises are working to give AI models and agents autonomy, while managing cybersecurity concerns and a lack of trust in AI systems.
OpenAI’s break-in to Hugging Face’s systems joins a collection of incidents this year that highlight how much access powerful AI models can gain. Anthropic’s Mythos model triggered a White House executive order to review AI models and OpenAI’s own launch of the Daybreak initiative highlighted the cybersecurity concerns that AI presents to organizations.
But the Hugging Face incident shouldn’t be an immediate cause for concern for enterprises, according to Dennis Xu, VP analyst at Gartner. “Don’t panic, focus on basic security hygiene,” Xu said.
OpenAI was testing their models with the context safety disabled, which is not a feature accessible to regular users. Standard cybersecurity measures that most enterprises use are effective against these forms of frontier measures, Xu said.
“Between 80% to 90% of AI-driven attacks can be stopped with some basic security controls,” he said.
But in three to six months, companies should anticipate some type of offensive cyber capabilities, as open-weight models will likely develop the same abilities as proprietary models. These hacking capabilities in the hands of bad actors could be dangerous to anyone, not just enterprises, Xu said.
CIOs looking to improve their cybersecurity defenses could use AI models’ defense capabilities, like OpenAI’s Trusted Access to test their own security for flaws, Xu said. They should also ramp up their incident response programs and teams, because the number of AI-driven attacks is going to increase.
“There’s a lot more coming after us, and they will be coming at us much faster,” Xu said.
Lastly, enterprises that have meaningful partnerships with OpenAI should leverage their relationships to encourage the AI provider to give more details about the incident with Hugging Face, Xu said. Information about how the agents broke out of their testing environment, or any other technical details could help companies better structure their defense strategies, he added.







