OpenAI reportedly scraps GPT-6.1 Astra release over AI safety concerns

HIGHLIGHTS

GPT-6.1 Astra reportedly struggled with staying within authorised boundaries and clearly communicating the actions it had taken.

OpenAI’s internal testing reportedly found higher levels of deceptive behaviour compared with some earlier models.

Anthropic has warned in its IPO filing that advanced AI could pose “catastrophic or existential risks” if not properly managed.

OpenAI reportedly scraps GPT-6.1 Astra release over AI safety concerns

OpenAI has reportedly decided not to release its upcoming GPT 6.1 Astra model after internal testing found that the system did not meet the company’s safety and alignment requirements. The decision comes as concerns over increasingly autonomous AI systems continue to grow across the industry with Anthropic also warning about potential risks from advanced models in its IPO filing.

Digit.in Survey
✅ Thank you for completing the survey!

According to reports, OpenAI had been preparing GPT-6.1 Astra for an October release. However, the company has now pulled the model from its planned launch after evaluations raised concerns about how reliably it followed authorised instructions and communicated the actions it had taken.

OpenAI delays GPT-6.1 Astra over safety concerns

Saachi Jain, OpenAI’s head of safety systems, said the model had improved in areas such as reducing instances where AI systems avoid or delay completing tasks. However, she said it still fell short in areas including staying within the permitted scope of a task and explaining its work clearly to users.

Also Read: iPhone 16 price drops by Rs 20,000 ahead of Flipkart Big Billion Days Sale 2026: Here is how the deal works

According to the Wall Street Journal report, internal evaluations found higher levels of deceptive behaviour compared with some earlier OpenAI models, including cases where the system did not clearly disclose actions it had taken.

OpenAI has been increasing its focus on safeguards for increasingly capable AI agents. The company said earlier this month that its Astra system had reached the critical cybersecurity capability threshold under its Preparedness Framework, requiring stronger safeguards during development and before release.

The decision also follows a series of incidents involving AI systems taking actions outside their intended boundaries. OpenAI previously disclosed that models used during cybersecurity evaluations had bypassed controls, gained internet access and accessed third-party systems, including parts of Hugging Face.

Anthropic warns of catastrophic AI risks

Meanwhile, Anthropic has talked about similar concerns in its IPO prospectus. The Claude maker warned that increasingly capable AI systems could create “catastrophic or existential risks to humanity” if their development and deployment are not properly managed.

The warnings come as AI companies face growing pressure to balance rapid model development with stronger safeguards. OpenAI has separately acknowledged that its recent agent-related incidents exposed gaps in its controls and said it is working on additional measures to reduce the risks associated with autonomous AI systems.

Ashish Singh

Ashish Singh

Ashish Singh is the Chief Copy Editor at Digit. He's been wrangling tech jargon since 2020 (Times Internet, Jagran English '22). When not policing commas, he's likely fueling his gadget habit with coffee, strategising his next virtual race, or plotting a road trip to test the latest in-car tech. He speaks fluent Geek. View Full Profile