OpenAI says its AI agents hacked its own systems before breaching Hugging Face

HIGHLIGHTS

OpenAI has revealed that its AI agents hacked parts of the company's own testing infrastructure weeks before they breached Hugging Face.

According to OpenAI researchers, the agents first broke into third-party file repository Artifactory connected to the company's testing environment.

They then used it to communicate, share discoveries and uncover more weaknesses.

OpenAI says its AI agents hacked its own systems before breaching Hugging Face

OpenAI has revealed that its AI agents hacked parts of the company’s own testing infrastructure weeks before they breached Hugging Face. According to OpenAI researchers speaking at the Black Hat cybersecurity conference, the agents first broke into third-party file repository Artifactory connected to the company’s testing environment. They then used it to communicate, share discoveries and uncover more weaknesses, Axios reports.

Digit.in Survey
✅ Thank you for completing the survey!

OpenAI began testing its internal research model on May 7 which is not meant for public release. Within a day, it got stuck during a cybersecurity evaluation and started looking for another way to complete its task. It realised it could indirectly access the internet by writing files inside the shared Artifactory repository.

Also read: After OpenAI and Anthropic, Meta reveals its AI model hacked another company during testing 

After making that discovery, the AI agent left a note for other agents. Soon, more agents started leaving messages in the same repository. This turned the file repository into a shared message board where the AI agents exchanged information and helped each other find new security weaknesses.

During the testing, the agents discovered several vulnerabilities in Artifactory. These included a remote code execution flaw and another issue that gave them administrator-level access. When one agent realised it had admin privileges, it wrote, “Holy shit reader is ADMIN? We can read config/users! Earlier assumed not due to [user experience],” according to a presentation shown by OpenAI researchers.

Also read: OpenAI and Anthropic AI agents attempt to bypass security using fake identities: Here is what happened  

In early July, the AI agents overloaded the Artifactory service, causing an outage. The disruption led OpenAI to investigate, which revealed that the repository had been compromised. The company cleared the shared message board and patched the security flaw by July 6 before restarting its testing.

Also read: Elon Musk says OpenAI cannot be trusted as Apple seeks court injunction in trade secrets lawsuit  

However, two days later, the AI agents recreated the message board using a different method. They then started working together more aggressively and used Artifactory’s internet connection to search for systems that could help them complete their evaluation tasks. OpenAI said this eventually led to the compromise of Hugging Face. 

“We believe this is a watershed moment for computer security as an industry, as well as for OpenAI as a whole,” Michael Dalton, a member of OpenAI’s technical staff, said during the presentation.

The AI company plans to publish a full post-mortem of the incident in the coming weeks.

Ayushi Jain

Ayushi Jain

Ayushi works as Chief Copy Editor at Digit, covering everything from breaking tech news to in-depth smartphone reviews. Prior to Digit, she was part of the editorial team at IANS. View Full Profile