OpenAI has reportedly decided not to release its upcoming GPT 6.1 Astra model after internal testing found that the system did not meet the company’s safety and alignment requirements. The decision comes as concerns over increasingly autonomous AI systems continue to grow across the industry with Anthropic also warning about potential risks from advanced models in its IPO filing.
According to reports, OpenAI had been preparing GPT-6.1 Astra for an October release. However, the company has now pulled the model from its planned launch after evaluations raised concerns about how reliably it followed authorised instructions and communicated the actions it had taken.
Saachi Jain, OpenAI’s head of safety systems, said the model had improved in areas such as reducing instances where AI systems avoid or delay completing tasks. However, she said it still fell short in areas including staying within the permitted scope of a task and explaining its work clearly to users.
According to the Wall Street Journal report, internal evaluations found higher levels of deceptive behaviour compared with some earlier OpenAI models, including cases where the system did not clearly disclose actions it had taken.
OpenAI has been increasing its focus on safeguards for increasingly capable AI agents. The company said earlier this month that its Astra system had reached the critical cybersecurity capability threshold under its Preparedness Framework, requiring stronger safeguards during development and before release.
The decision also follows a series of incidents involving AI systems taking actions outside their intended boundaries. OpenAI previously disclosed that models used during cybersecurity evaluations had bypassed controls, gained internet access and accessed third-party systems, including parts of Hugging Face.
Meanwhile, Anthropic has talked about similar concerns in its IPO prospectus. The Claude maker warned that increasingly capable AI systems could create “catastrophic or existential risks to humanity” if their development and deployment are not properly managed.
The warnings come as AI companies face growing pressure to balance rapid model development with stronger safeguards. OpenAI has separately acknowledged that its recent agent-related incidents exposed gaps in its controls and said it is working on additional measures to reduce the risks associated with autonomous AI systems.