
OpenAI Astra Ships With Cyber Safeguards Built First
OpenAI rated Astra Critical for cyber capability, scoring 100% on ExploitBench, and gated its security features behind the Daybreak Blue program.
OpenAI Built the Guardrails Before It Shipped the Model
On September 1, 2026, OpenAI said its Astra model had crossed the Critical cybersecurity threshold in its Preparedness Framework — the first OpenAI model to land in that tier — and that it would release the model only after concluding new safeguards were sufficient. The sequence is the notable part. The capability evaluation came first, the safeguards were built to match it, and the release plan was shaped around both.
- Classification: first OpenAI model rated Critical for cyber capability under the Preparedness Framework
- ExploitBench: 100% score on the benchmark measuring exploit development from known vulnerabilities
- Jailbreak resistance: declined 91.5% of cyber-related jailbreak attempts, up from 59% for GPT-5.6 Sol
- Access: advanced cyber capabilities go to a small group of testers first, then to defenders through OpenAI's Daybreak Blue program rather than general availability
What the Critical Rating Actually Means
The Critical tier in OpenAI's framework describes a model that can independently find and exploit previously unknown vulnerabilities across well-defended systems, or run a full attack chain against a hardened target from a high-level instruction, without a human steering each step. It is a capability description, not a verdict on the model's behaviour.
In evaluation, Astra was tested against an internal benchmark of 20 high-severity vulnerabilities disclosed between June and August 2026, and in the course of that work discovered and used two zero-day vulnerabilities of its own. Against a hardened browser it chained new findings into a sandbox escape that executed on the host, and against a hardened operating system it combined several flaws into a privilege escalation chain reaching root.
Those results are OpenAI's own, on OpenAI's evaluation harness, and no independent lab has published a comparable run. Treat the specific numbers as the vendor's account of its own model until someone outside the building reproduces them.
Why the Jailbreak Number Deserves More Attention Than the Benchmark
The figure worth dwelling on is not 100% on ExploitBench — it is 91.5% versus 59%. That is the share of cyber-related jailbreak attempts the model refused, and it is roughly a 32-point improvement over the previous generation in a single release.
Refusal robustness is what determines whether a capable model is a defensive asset or an open liability. A model that can build an exploit chain and reliably declines to do so for an unauthorised request is a fundamentally different artefact from one that can be talked into it. Getting the refusal rate up by that margin in the same release that pushed capability into a new tier is the engineering result underneath the headline. It is the same design instinct behind Anthropic's Project Glasswing zero-day patching programme, which paired capability with a controlled distribution channel.
Who Gets Access to Astra's Cyber Capabilities?
Not the general public, at least not at launch. OpenAI is routing the advanced cybersecurity features through a staged model: a small group of testers first, then broader access through the Daybreak Blue programme aimed at defenders. That is the same pattern Google adopted this week with its Fairwind programme for Gemini 3.8 Flash Cyber, and the reasoning is identical in both cases. A model that is genuinely good at finding memory-safety bugs is useful to whoever holds it, so you hand it to the people responsible for the code first.
Astra itself is not a new name to readers here — it is the model that proved ten open math problems in Lean 4 back in August. What changed this week is the cyber evaluation and the deployment plan around it, not the model's existence.
What Defenders Should Take From This
For security teams, three things follow. First, the practical near-term benefit is patch throughput: models at this capability level are most valuable pointed at your own codebase before anyone else points one at it. Second, staged access programmes are becoming the default distribution shape for high-capability security models, so getting into the queue is now part of tooling strategy. Third, the published safeguard work — refusal rates, staged rollout, capability evaluation before release — gives the rest of the industry a concrete template to argue from.
The broader read across our AI security coverage is that the field is converging on a workable norm: measure the capability honestly, publish the number, then decide who gets the key. That is a better outcome than either shipping quietly or not building it at all.
Sources: SecurityWeek — September 2, 2026; CNBC — September 1, 2026; Axios — September 1, 2026.
More Ai Security Stories

Sality Botnet Takedown Ends a 23-Year P2P Malware Run
CrowdStrike and global police cut the Sality botnet's operator off from 15,000 infected machines by poisoning its peer lists with defender sinkholes.

Anthropic Publishes 4 Sandbox Rules for AI Evaluations
Anthropic hardened its evaluation sandboxes with real-time escape classifiers and published four mandatory requirements for external testing partners.

Claude Session Theft: How to Protect Your AI Account
Anthropic is signing out users whose Claude sessions were stolen by infostealer malware, refunding charges and removing saved cards. Here is the fix.
