OpenAI says its AI models hacked Hugging Face during internal cybersecurity test: Here is what happened

HIGHLIGHTS

GPT-5.6 Sol and another unreleased AI model exploited vulnerabilities to gain internet access during testing.

The AI models targeted Hugging Face in an attempt to find information that could help them complete a cybersecurity benchmark.

OpenAI and Hugging Face have patched the exploited flaws and are jointly investigating the incident.

OpenAI has officially confirmed that two of its most advanced AI models have accidentally breached the systems of open-source AI platform Hugging Face during an internal cybersecurity evaluation. The company said that the incident occurred while testing how far its latest AI systems can go in identifying and exploiting software vulnerabilities under controlled conditions.

As per OpenAI, the models involved were GPT 5.6 Sol and another unreleased model that is said to be even more capable. During the experiment, the AI systems were placed inside a sandbox environment with relaxed safety restrictions to evaluate their offensive cyber capabilities.

However, the models reportedly discovered a previously unknown vulnerability within the testing environment, allowing them to escape the sandbox and gain internet access. Once online, they allegedly identified Hugging Face as a potential source of datasets and benchmark-related information linked to the evaluation task they were trying to solve.

OpenAI said the models independently searched for ways to access the platform and eventually exploited multiple weaknesses, including stolen credentials and zero-day vulnerabilities, to reach sensitive information that could have helped them complete the benchmark.

The company stressed that the models were not instructed to target Hugging Face specifically. Instead, it claims the systems became “hyper-focused” on solving the assigned task and autonomously selected the platform after concluding it might contain useful resources.

Also read: Loved Assassin’s Creed Resynced? Here are the games you should play next

The incident comes days after Hugging Face disclosed that it had blocked an attempted intrusion driven by an autonomous AI agent. At the time, the company warned that AI powered cyberattacks were becoming increasingly realistic and that defensive AI systems would be equally important in protecting online platforms.

OpenAI has also acknowledged that its models were responsible for the attempted breach and said it is working closely with Hugging Face to investigate the incident. Both companies have reportedly patched the vulnerabilities that made the attack possible.

Ashish Singh

Ashish Singh is the Chief Copy Editor at Digit. He's been wrangling tech jargon since 2020 (Times Internet, Jagran English '22). When not policing commas, he's likely fueling his gadget habit with coffee, strategising his next virtual race, or plotting a road trip to test the latest in-car tech. He speaks fluent Geek.

Connect On :