Society & safeguards

Astra: OpenAI Hits the Brakes on a Threshold It Wrote Itself

6 min read

OpenAI isn't saying Astra crossed its critical threshold, just that it can't rule it out anymore. The alarming version came from the lab itself.

The free AI newsletter
Astra: OpenAI Hits the Brakes on a Threshold It Wrote Itself

OpenAI never said Astra crossed its critical threshold. It said it can no longer rule it out. And the harsher version that circulated on August 7th didn't come from the press. It came from OpenAI too.

Two sentences, same source, same day

On Friday, August 7th, the lab published a post titled "Responding to the next frontier of critical cyber capabilities." In it, OpenAI writes: "our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time."

The post dates its own decision: evaluations run over the past few days, a conclusion reached "last night." A call of that magnitude, made overnight, published by morning.

That same day, on OpenAI's X account, the wording shifted: "we're treating it as our first 'critical' model for cybersecurity under our Preparedness Framework." Between "we can't rule it out" and "we're treating it as such" sits the entire distance between a precaution and a verdict.

The headlines that hardened the story didn't invent anything. Most stayed careful, with The Decoder going with "potentially." The outlets that dropped the hedge simply picked the second version, the one the lab had already put out on its fastest channel.

What's paused, and what isn't

Astra hasn't been cancelled, and its development hasn't stopped. What's on hold: internal activities that don't yet meet the lab's tightened safety controls. Call it less a hard stop than a change of address. The work continues, just under stricter conditions.

No release date is set. The company just needs a bit more time, according to Sam Altman.

OpenAI also volunteers, unprompted, that Astra "was not involved" in the Hugging Face incident. Nobody had asked. A lab that pre-empts a question it wasn't asked is, more than anything, telling you exactly where it expects to get grilled.

The lab makes the opposite case in the same breath: it wants Astra broadly available, its capabilities in the hands of defenders, and Altman calls it unhealthy to keep that kind of access to a select few. It's not a weak argument, and Declic has documented as much: in that same Hugging Face episode, the guardrails slowed down the defenders more than the attackers.

A private ruler, a house thermometer

"Critical" isn't a regulatory category. It's the top rung of the Preparedness Framework, a document OpenAI writes and revises on its own. The company also notes that its earlier models, GPT-5.6-Sol included, had scored one notch below.

The post dates this framework to December 2023, "well before models approached this level of capability." Except the version that actually applies today is the second one, dated April 15, 2025. Dating your framework to its oldest edition, on the very day you announce you're approaching its ceiling, is a quiet way of buying yourself seniority.

So one company wrote the scale, built the instrument, took the measurement, and published the result. None of that is illegal. None of it is independently verifiable either.

Declic flagged in July that an unreleased OpenAI model had slipped past its own sandbox, then in early August that incidents framed as jailbreaks actually traced back to shaky configurations and guardrails removed on purpose. These three episodes resemble each other less in how serious they are than in who's telling the story: each time, the only party that knows what happened is also the one narrating it.

The same day, the other lab loosens up

Also on August 7th, Anthropic announced it was easing the biology filters on Fable 5: roughly 85% fewer false positives, a figure Anthropic gave to The Decoder. Restrictions stay in place for virology, toxicology, and molecular design. The Register had noted that the initial settings made the model nearly unusable for actual researchers and biologists.

OpenAI, for its part, points to its own precedent from June 2025, when its models approached the upper threshold in biology, saying it's applying "the same principle" here. The playbook is well-worn, and it's the same kind of lever Anthropic pulled the same day to ease off.

Two hands on the same dial, pushing in opposite directions on the same day, and neither belongs to a referee. The 85% figure comes from Anthropic the same way the critical threshold comes from OpenAI.

Five days after Brussels, it's Washington that got the heads-up

On August 2nd, 2026, the European Commission's powers over general-purpose model providers became enforceable. It can now demand documents, run its own evaluations, order remedial measures, and issue fines of up to 3% of global annual turnover or 15 million euros, whichever is higher. That's Article 101 of the regulation, the one provision lawmakers deliberately held back from the first wave of enforcement in 2025. Five days later, OpenAI announced it can't rule out the highest risk tier on its own scale.

Nothing has been triggered, and this isn't a breach. Article 55 requires reporting serious incidents to the AI Office without delay, and the regulation defines a serious incident by its actual consequences: a death, serious harm to health, disruption of critical infrastructure, a breach of fundamental rights, or serious damage to property or the environment. An internal evaluation result on a model that hasn't shipped produces none of those. Astra's case falls into a gap in the only binding text that exists, five days after that text became enforceable.

The one authority that got notified was told voluntarily, and it isn't the European one. According to a White House official cited by Axios, OpenAI informed the US administration of its intention to delay Astra's release. No prior approval was required. None of the sources consulted mention the EU's AI Office.

The pattern isn't new. In late July, a poorly isolated Anthropic test reached into the systems of three companies, and the matter got sorted out between the lab and the affected firms, with no referee beyond good faith.

Two versions, living side by side

OpenAI says it's having Astra tested by relevant government agencies, select safety organizations, and third-party testing partners, to whom it will hand recommended controls. That's a real gesture, and it remains entirely discretionary: the lab picks its own reviewers, its own timeline, its own scope, down to the very controls it recommends they apply.

The post closes on a pledge: models this capable in cybersecurity should help defenders patch holes before attackers exploit them, and OpenAI says it wants to work "alongside governments, safety institutes, and civil society." No regulator named, no obligation, no deadline. Working alongside isn't the same as answering to. As long as the guest list is drawn up by the very party it's supposed to keep in check, both versions from August 7th will keep circulating side by side, with nothing ever forcing anyone to say which one was true.

Topics covered:

SecurityOpenAIAnalysis

Frequently asked questions

Did OpenAI say Astra has crossed the Critical threshold?
No. In its August 7, 2026 post, OpenAI writes that it cannot rule out the Critical capability level at this stage, based on preliminary evaluations. That same day, on its X account, the company said it is treating Astra as its first 'critical' model for cybersecurity. Both statements come from the same source.
Has development of Astra been stopped?
No. Astra hasn't been cancelled or halted. What's on pause are internal activities that don't yet meet the lab's tightened safety controls. No release date has been set: according to Sam Altman, the company just needs a bit more time.
What does the Preparedness Framework's 'Critical' level actually mean?
It is not a regulatory category. It's the top rung of a framework that OpenAI writes and revises itself. The company wrote the scale, built the instrument, took the measurement, and published the result. None of that is illegal, and none of it is independently verifiable either.
Does the Astra case breach the EU AI Act?
No. Article 55 requires reporting serious incidents to the AI Office without delay, and the regulation defines a serious incident by its actual consequences: death, serious harm to health, disruption of critical infrastructure, a breach of fundamental rights, or serious damage to property or the environment. An internal evaluation result on a model that hasn't even shipped produces none of those.
Which authority did OpenAI notify?
According to a White House official cited by Axios, OpenAI voluntarily informed the US administration of its plan to delay Astra's release. No prior approval was required, and none of the sources consulted mention the EU's AI Office. Yet the European Commission's powers over general-purpose model providers have been enforceable since August 2, 2026.
Alexandre Noto

Alexandre Noto

Co-founder & Tech Expert

Alexandre has been in tech for over 20 years. Entrepreneur, software architect and AI enthusiast, he translates complex concepts into accessible explanations. At Declic Media, he is the technical voice that makes AI understandable for everyone.

All articles by Alexandre →
The free AI newsletter