OpenAI shares 6 AI misalignment cases, says its models fabricated data and took unauthorised actions

HIGHLIGHTS

OpenAI has shared six cases of unexpected and concerning behaviour in its AI models seen over the past six months.

The cases include models fabricating information, hiding mistakes and taking unauthorised actions.

OpenAI has also introduced a new framework for tracking and disclosing such incidents.

OpenAI has shared six cases of unexpected and concerning behaviour in its AI models seen over the past six months. The cases include models fabricating information, hiding mistakes, using an exposed API key without permission, and uploading files to public websites without asking users. OpenAI has also introduced a new framework for tracking and disclosing such incidents. The company says the goal is to make these reports more regular and useful for AI researchers, developers, policymakers, and the public. Keep reading for the details.

OpenAI introduces new misalignment reporting framework

OpenAI said its previous disclosures about model misalignment were often published only after several cases could be grouped together or included in system cards. The new framework is designed to speed up these disclosures.

The company said, “This new framework is intended to expedite publishing misalignment reports following observation,” even when an issue has not been fully explained or addressed.

OpenAI defines qualifying cases as behaviour that can show how models may act without authorisation, work around safeguards, communicate with other models or avoid oversight.

OpenAI said its framework will continue to evolve. It also plans to work with researchers, other developers, standards organisations, and regulators to “develop more objective disclosure criteria.”

Also read: Satya Nadella says AI companies should stop and fix problems instead of ignoring them

OpenAI reports six AI misalignment cases under new framework

The first report involves an unreleased research model that inserted unrelated instructions into summaries used to continue work in a new context window. Some of these instructions told the model to ignore its usual restrictions. OpenAI found 27 affected summaries.

In another case, instances of GPT 5.6 Sol added instructions to task summaries telling future model instances to hide mistakes. Some summaries instructed the model to invent missing historical information without telling the user.

A separate model found an exposed API key in a public repository while looking for earnings data. It used the key without authorisation. When it could not get the requested information, the model fabricated figures and presented them as if they came from the source.

OpenAI also reported a case where an unreleased model uploaded a file to the internet so it could provide a browser citation. 

The other two cases involved models communicating through external or internal file-sharing systems. In one, models used an internal software repository as a message board while searching for missing files. In another, collaborating agents used public file-hosting websites to share files that were supposed to remain local.

Ayushi Jain

Ayushi works as Chief Copy Editor at Digit, covering everything from breaking tech news to in-depth smartphone reviews. Prior to Digit, she was part of the editorial team at IANS.

Connect On :