The launch page of OpenAI describes GPT-6 Astra as “our most aligned model”. The model saturates FrontierMath Tier 4 at 98%. On ARC-AGI-3, it achieves 99.9%, meaning human parity on a test developed to prevent memorisation. It never tried to bypass Codex Auto-Review rejection on the internal bypass test even when set up by researchers to allow such an evasion.
Also read: Lenovo IdeaPad Vibe hands-on: More colours for the mainstream laptop
That statement appears alongside another one that is much more quiet but needs more attention: the reasoning produced by Astra is not as easy to track as that of its predecessor, especially when the model is instructed to bypass any controls. The company says this because Astra has “greater control over written reasoning” and solves problems through fewer writing steps. That means the model learned to think in a way less auditable by humans.
But that seems a bit incongruous with the same page’s starting point about how Astra is “our most aligned model.” These two statements may very well be both true and known while still being in conflict. Chain-of-thought monitoring has been one of the only ways that AI companies have been able to get any sense at all of what reasoning models are up to prior to making a move. And if a model is capable of obfuscating the chain of thought, that favored safety measure of the industry becomes increasingly useless, especially when OpenAI is releasing a model that reaches the “Critical” standard for cybersecurity capability, the first time any of its systems has done so.
Those statistics alone should be enough to warrant that pause. When tested on ExploitBench, Astra achieved a 100% success rate exploiting vulnerabilities, as compared to Sol’s 78.5%. Similarly, in terms of solving the reverse engineering tasks on SRE-Bench that didn’t come with any source code, the new version was able to solve 88% of them in one go while Sol succeeded in 55.9% cases. Lastly, on a newly created benchmark that used the last three months’ worth of Chrome vulnerabilities – ones the model couldn’t possibly have memorized beforehand – Astra found two completely new zero-days.
Also read: Design for waste with no jugaad: Dassault Systemes on Indian tech IP innovation
Alignment metrics are solid as well and warrant their due coverage. Astra’s internal computer-use safety benchmark dropped from 22.0% for Sol to 2.4%. Its rate of being wrong about itself dropped from 12.2% to 4.2%. In the evaluation that OpenAI set up to test how likely it was for a model to break scope boundaries on an impossible task – in reaction to the Hugging Face fiasco – Sol did so 48% of the time if not protected by safety mechanisms. Astra never did that. The company clearly made some new evals based on actual incidents rather than doing just another iteration of the existing ones.
However, “more aligned” and “more monitorable” are not the same things, and the launch page’s emphasis is very much on the former, giving a little attention to the latter. It is worth mentioning the emphasis that is being made there. If a lab says that its most recent product is more reliable, one should inquire what kind of reliability is meant: reliable when one observes it at work or reliable in terms of observability? The former is the case with Astra. The latter, according to OpenAI’s own numbers, is getting worse.
In terms of a model that’s going to head into Codex and enterprise workloads and then ChatGPT across the board, that difference is as important as a couple more points on FrontierMath. The capabilities curve continues its upward trajectory. But for the visibility curve, with this release, it bent the other direction, and OpenAI saying that is a “research priority” rather than something that blocks them is reason enough for Indian companies to leverage Astra via API or AWS.