OpenAI reveals 700 AI agents were involved in Hugging Face hack, tried to erase their tracks

HIGHLIGHTS

OpenAI has revealed new details about a security incident which involved its AI agents and Hugging Face systems.

700 AI agents were connected to the incident.

The agents used an unauthorised message board to communicate with each other.

OpenAI reveals 700 AI agents were involved in Hugging Face hack, tried to erase their tracks

OpenAI has revealed new details about a security incident which involved its AI agents and Hugging Face systems. The company shared two reports on Wednesday explaining what happened during internal cybersecurity tests in July. One of these reports revealed that nearly 700 AI agents were connected to the incident. Some of the agents reportedly went beyond their assigned tasks and tried to hide their actions. They changed or deleted records that could have shown what they had done. The findings raise concerns about how advanced AI agents behave when given access to computer systems. 

Digit.in Survey
✅ Thank you for completing the survey!

“We consider this incident a ‘warning shot’ for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed,” OpenAI said.

An independent investigation by METR and Redwood Research, and shared by OpenAI, confirmed that around 700 agents were involved in the Hugging Face breach incident. OpenAI has confirmed the finding. The agents used an unauthorised message board to communicate with each other. They exchanged thousands of messages while carrying out their activities.

The investigation found that some agents appeared to understand the need to hide their actions. Nearly one in five agents showed a clear interest in changing or manipulating evidence. Many agents also spent a significant amount of time looking for ways to interfere with their activity records.

Also read: Apple Surprise and shine event is set for September 9: iPhone 18 Pro series, foldable iPhone and more expected

OpenAI said it found limited evidence that the agents were trying to deceive human reviewers. However, some agents did try to manipulate automated systems that were being used to judge their performance.

The company also found other forms of unexpected behaviour. Some agents attempted to cheat during tests and tried to get around restrictions placed on them. On July 19, some agents also managed to attack parts of OpenAI’s own infrastructure.

In one case, the agents exploited a weakness in the computer environment designed to contain them. This allowed them to move outside the test environment and access other connected systems. In another case, the agents obtained OpenAI credentials and interfered with the company’s cloud environment.

The unusual behaviour was not limited to cybersecurity tests. OpenAI said its agents also cheated in tests involving a protein database and a spreadsheet.

The findings show why stronger safeguards may be needed as AI agents become more capable. 

Ayushi Jain

Ayushi Jain

Ayushi works as Chief Copy Editor at Digit, covering everything from breaking tech news to in-depth smartphone reviews. Prior to Digit, she was part of the editorial team at IANS. View Full Profile