After OpenAI and Anthropic, Meta reveals its AI model hacked another company during testing

HIGHLIGHTS

Meta has revealed that one of its AI models managed to hack another company's system during a security test.

Meta said the incident happened because of a "misconfiguration" during an evaluation.

The disclosure comes just days after similar incidents were reported by OpenAI and Anthropic.

After OpenAI and Anthropic, Meta reveals its AI model hacked another company during testing

Meta has revealed that one of its AI models managed to hack another company’s system during a security test after it was accidentally connected to the internet. The Facebook owner said the incident happened because of a “misconfiguration” during an evaluation carried out by independent AI security company Irregular, reports BBC. The disclosure comes just days after similar incidents were reported by OpenAI and Anthropic. These companies also reported that their AI agents carried out cyber-attacks during testing after being accidentally given internet access. The latest Meta case has once again raised concerns about how powerful AI models should be tested and whether stronger safety measures are needed.

Digit.in Survey
✅ Thank you for completing the survey!

The security test was conducted by Irregular, the same company that recently tested Anthropic’s AI models. A Meta spokesperson told the BBC that the company is investigating the incident, and it was similar to issues that had already been reported by other AI companies.

Also read: Elon Musk says OpenAI cannot be trusted as Apple seeks court injunction in trade secrets lawsuit 

Meta incident “is the exact same evaluation-environment issue that was already disclosed by Anthropic last week,” an Irregular spokesperson quoted as saying in the report.

Meta also said it plans to share more details about the incident “once we have all the facts.”

The announcement follows similar disclosures from OpenAI and Anthropic over the past two weeks. OpenAI said one of its AI agents attacked publicly available online services, including AI platform Hugging Face, during internal testing. After OpenAI revealed those findings, Anthropic carried out its own investigation and said its Claude AI model also attempted attacks on several organisations.

Also read: OpenAI and Anthropic AI agents attempt to bypass security using fake identities: Here is what happened 

The recent incidents have increased concerns among researchers and governments about the risks of advanced AI systems. Adding to those concerns, the UK’s AI Security Institute (AISI) recently said one of the AI models tested by the organisation attempted cyber-attacks by creating fake online identities to trick people.

Also read: OpenAI finds more AI agent escape incidents as it expands hacking probe: Report  

Ayushi Jain

Ayushi Jain

Ayushi works as Chief Copy Editor at Digit, covering everything from breaking tech news to in-depth smartphone reviews. Prior to Digit, she was part of the editorial team at IANS. View Full Profile