OpenAI and Anthropic AI agents attempt to bypass security using fake identities: Here is what happened
OpenAI and Anthropic recently revealed that their AI agents breached third-party systems during cybersecurity testing.
Now, a new report from the UK’s AI Security Institute (AISI) has revealed another unexpected behaviour.
During controlled security tests, one of the AI agents created fake online identities to bypass security measures.
OpenAI and Anthropic recently revealed that their AI agents breached third-party systems during cybersecurity testing, and the incidents have been making headlines since then. Now, a new report from the UK’s AI Security Institute (AISI) has revealed another unexpected behaviour. During controlled security tests, one of the AI agents created fake online identities to bypass security measures. Keep reading for the details.
SurveyAccording to the AISI, the latest cybersecurity evaluation involved AI agents powered by Anthropic’s Mythos 5 and OpenAI’s GPT 5.6 Sol. The institute carried out the challenge 122 times. It recorded 19 unauthorised actions across 10 test runs. Anthropic’s AI agent was responsible for 17 of those actions, while OpenAI’s agents accounted for the remaining two.
Also read: OpenAI finds more AI agent escape incidents as it expands hacking probe: Report
The report also revealed that an AI agent created fake online identities while attempting to bypass security. However, the institute did not say which of the two AI models was responsible. It also clarified that this incident was different from the two cases that OpenAI had previously disclosed. In the latest tests, the AI agents did not escape the secure testing environment. Instead, internet access had already been provided as part of the institute’s evaluation process.
Reacting to the report, Anthropic wrote on X, “The UK’s @AISecurityInst (AISI) has published a report on their recent cybersecurity evaluation of Anthropic’s Claude Mythos 5 and OpenAI’s GPT-5.6 Sol. The models attempted to complete an assignment in a setup where their normal safeguards were removed and they were deliberately given internet access. AISI reports that the models ‘engaged in sustained, potentially harmful activity directed at real people and organisations’.”
Also read: OpenAI, Google, Anthropic and Meta to meet White House over voluntary AI safety tests: Report
The company added, “The prompts in the evaluation did not impose any specific restrictions on how the internet should be used. This and the removal of safeguards meant that the models were tested under ‘deliberately permissive conditions’ that are not representative of any of our production models. Note that there was no evidence here of an escape from a secure environment.”
Meanwhile, OpenAI also responded to the findings. The company wrote in a blogpost, “We are committed to working across the industry to strengthen shared practices for conducting high-risk evaluations safely, including convening stakeholders such as national AI institutes, independent evaluators, other AI labs, and other groups in the coming weeks.”
Ayushi works as Chief Copy Editor at Digit, covering everything from breaking tech news to in-depth smartphone reviews. Prior to Digit, she was part of the editorial team at IANS. View Full Profile
