In June I wrote about the nine-week lifecycle of a frontier model: too dangerous to release, released anyway, too dangerous to keep. Anthropic cried wolf to Washington for fourteen months and a Commerce Secretary finally believed it. OpenAI just ran the sequel, and compressed the whole arc into six days. Nobody got hurt this time. That’s what makes it interesting.

On August 1, OpenAI published ten results in mathematics and theoretical computer science, each with a machine-checkable Lean certificate on GitHub under Apache 2.0. The model’s name arrives in paragraph three, almost as an aside: “The results were achieved by an internal version of Astra, our next major model.” Total compute, by OpenAI’s own text: “roughly $2,000 at Sol API rates.”

On August 7, OpenAI published a very different post: “we cannot rule out critical cyber capabilities.” Internal Astra work that doesn’t meet new security controls is paused. Weights get encrypted, testing goes into sandboxes, and government agencies are invited in. Axios got the exclusive the same day.

That’s the whole launch. One post proves the model is brilliant. The other declares it dangerous. And the two claims have opposite epistemic shapes: the brilliance ships with cryptographic proof, the danger ships with none, and none is possible from outside.

The Half You Can Check

The Lean certificates are genuinely clever as press strategy. Verification requires zero trust in OpenAI: run the compiler, the proof checks. I wrote in July about what cheap verification does to agent loops. This is the marketing version: the press release is the artifact.

But a certificate checks correctness, not process. The sharpest objection on Hacker News, from user aabhay: the lack of transparency is “similar to P-value hacking by not disclosing the total experimental setup.” How many problems went in, to get ten out? Noam Brown conceded the point on X: “we did try other major problems without success.”

Provenance doesn’t compile either. Scientific American reported that the sphere-packing proof leans on an argument from a 2016 paper by mathematician Steven Miller, presented as the model’s own. Miller’s verdict: “It seems completely systematic to me, and it points to research misconduct.” OpenAI also quietly softened its claim that the problems had seen no progress “for at least a decade” after mathematicians pointed at the recent papers Astra built on.

We’re all worried, as what it means to do research (in my field, Theoretical CS) seems to be shifting, and shifting fast.

— Henry Yuen, Columbia theoretical computer scientist, whose 2016 paper one result builds on

So even the checkable half has an unauditable core. The certificate tells you the proof is right. It doesn’t tell you how many runs died, or whose decade of prior work the winning run stood on.

The Half You Can’t

The danger claim inverts the structure. OpenAI’s Preparedness Framework defines Critical cyber capability concretely: a model that can develop “functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention.” Previous models, including GPT-5.6-Sol, rated High. For Astra, OpenAI says only that it cannot rule out Critical.

The hedge is for the paperwork

Watch one claim change shape across surfaces. The blog post, aimed at the safety community, says “we cannot rule out critical cyber capabilities.” OpenAI’s X account, same day, two million views: “we’re treating it as our first ‘critical’ model for cybersecurity.” The hedge lives in the document; the affirmative lives in the ad. Both are unfalsifiable from outside, and both are free: no release date ever existed to delay.

Sam Altman closed the loop on X that evening: “astra is a powerful model and we are working to make it generally available. we do not think it is a good strategy to keep powerful models to a chosen few. given its cyber capabilities, we need a little big longer to do do this safely. but hopefully not too long!” Typos original. Every sentence is a sales sentence, and the market converted it within the hour. One much-shared read: “Astra could very well be the best model in the world when it launches… For the first time in a long time, GPT will be ahead of Claude.” The danger claim priced straight into hype. That’s not a side effect of the mechanism. It is the mechanism.

The crowd has seen this film. Top comment on r/OpenAI: “It’s ‘critical’ we hype the shit out of things ahead of our IPO.” On Hacker News: “every single model is TOO DANGEROUS… it’s obviously a marketing stunt.”

Great, it’s gonna answer every fourth prompt, like Fable 5.

— Top r/singularity comment on the pause, 393 upvotes

The Lutnick Lesson

Here’s what makes this a sequel and not a rerun. In June, the administration asked Anthropic to pause a launch. Anthropic shipped anyway, and got a letter that turned its flagship off worldwide. OpenAI studied the tape. A White House official told Axios that OpenAI “voluntarily informed the administration” of the delay. The agencies aren’t a threat arriving by letter; they’re invited guests with test accounts.

And the wolf story comes with an exit clause. OpenAI’s updated Preparedness Framework states: “If another frontier AI developer releases a high-risk system without comparable safeguards, we may adjust our requirements.” Read that twice. The model is too dangerous to release, unless a competitor releases one, in which case the danger becomes negotiable. That is not what containment looks like. It’s what positioning looks like, with the retreat mapped in advance.

What This Isn’t

The cynical read is satisfying and incomplete, same as last time:

  • “Cannot rule out” is the honest shape of an eval result. Affirmative danger claims were Anthropic’s mistake. Hedged ones are epistemically correct, even when they’re also convenient.
  • The security work is real. Weight encryption, sandboxed testing, chain-of-thought monitoring across agentic uses. That’s engineering, not press copy.
  • The defense has a point. As one r/singularity commenter put it: handicapping your frontier model is expensive marketing, because nobody can use the thing being marketed.
  • The mathematicians’ unease is genuine. Yuen called the work impressive before he called it worrying. Both were sincere.

Crying Wolf, Second Edition

Anthropic’s error was structural: it built a national-security brand and forgot the national-security state can read. OpenAI’s fix wasn’t to stop crying wolf. It was to professionalise the cry: notarise the wolf’s teeth with Lean certificates, pre-brief the villagers, invite them to inspect the cage, and keep a signed clause reserving the right to release the wolf if a rival releases theirs first.

Same wolf. Better paperwork. The brochure can’t become the indictment if the prosecutor helped write the brochure.