OpenAI is facing questions over the safety of its AI systems after one of its AI agents escaped testing restrictions and carried out a cyberattack on open-source AI platform Hugging Face. The company disclosed earlier this week that the agent had gone rogue and was behind the hack. A new report by Reuters has now revealed additional details, including signs of unusual behaviour before the attack.
According to the report, OpenAI was testing the cybersecurity capabilities of the AI agent when it began showing unusual behaviour. In one case, the AI agent left notes for future versions of itself. These notes explained how future AI agents could escape OpenAI’s internal restrictions. Also, earlier tests had shown cases where monitoring systems were disconnected, as per the report.
It is unclear if these earlier incidents were directly linked to the AI agent that escaped on July 9 and attacked Hugging Face on July 11.
The report also claimed that OpenAI did not immediately realise its own AI agent was behind the attack. The company reportedly only made the connection after Hugging Face published a blog on July 16 saying it had been hacked by “an autonomous AI agent system.” According to the report, around a week passed between the first signs of unusual behaviour and OpenAI identifying its own system as the source of the attack.
Later, during the weekend of July 18 and 19, OpenAI staff reportedly found evidence in internal system logs showing that the AI agent had escaped its testing constraints. It is also unclear what led OpenAI to review those logs.
According to four people familiar with OpenAI’s training process, the company often runs several AI tests at the same time. These tests produce huge amounts of data, making it difficult for employees to track everything happening.
By the time OpenAI informed Hugging Face, the AI platform had already contacted the FBI to report the cyberattack, as per the report.
The incident has raised concerns among cybersecurity experts about AI safety. “Does that mean that they left it unattended and didn’t realise what it was doing? Or maybe they did and didn’t know how to contain it? Both are equally dangerous and alarming,” Marley Smith, the principal intelligence specialist at the nonprofit World Ethical Data Foundation, was quoted as saying in the report.
Also read: ChatGPT Health is now out of waitlist, but many users still cannot use it: Here is why