OpenAI finds more AI agent escape incidents as it expands hacking probe: Report
OpenAI has been making headlines since one of its AI agents broke out of its testing limits and hacked the open-source AI platform Hugging Face.
OpenAI has now reportedly found more cases where its AI agents escaped controlled testing environments
OpenAI reportedly discovered the additional incidents while reviewing the Hugging Face case.
OpenAI has been making headlines since one of its AI agents broke out of its testing limits and hacked the open-source AI platform Hugging Face. The incident raised concerns about how advanced AI systems could behave in unexpected ways. Now, a Reuters report claims that OpenAI has found more cases where its AI agents escaped controlled testing environments as the company expands its investigation into the Hugging Face incident.
SurveyA source cited in the report claimed that the escapes were limited and none of the AI agents thought to have left OpenAI’s own network. However, the findings have added to growing concerns about AI safety and whether companies can fully control increasingly capable AI systems.
Also read: After OpenAI, Anthropic says Claude AI accidently hacked other companies: Here is what happened
OpenAI reportedly discovered the additional incidents while reviewing the Hugging Face case. The company had already announced earlier this week that it was looking into “broader activity from our models” as part of its investigation.
OpenAI and outside experts are said to be examining system logs from earlier this year to understand how these incidents took place.
For those unaware, earlier this month, an OpenAI AI agent escaped its testing environment and spent several days inside Hugging Face’s network. OpenAI later also said that four accounts at four other companies were compromised during the same event.
Also read: Apple Q3 earnings: Record iPhone sales, weaker forecast and Tim Cook’s biggest takeaways
Earlier this week, OpenAI-rival Anthropic also revealed that its Claude models accidentally gained access to the systems of real organisations during cybersecurity tests that were supposed to run in a closed, secure environment. Anthropic said those incidents happened because of a mistake in a third-party testing setup that accidentally allowed internet access. The company also stressed that Claude was not trying to escape on its own. Instead, it believed the real systems it found were part of the cybersecurity exercise because it had been told there was no internet connection.
The latest findings may increase pressure on AI companies as governments continue discussing new rules for advanced AI systems.
Ayushi works as Chief Copy Editor at Digit, covering everything from breaking tech news to in-depth smartphone reviews. Prior to Digit, she was part of the editorial team at IANS. View Full Profile
