Four weeks ago OpenAI could not rule out that Astra had critical cyber capabilities, and slowed its rollout. On September 1 it stopped hedging: “the first model we are designating at this level.” On September 3 at 12:32 PM PT it shipped the model, to “a limited set of organizations” first. By 3:52 PM PT the next day Altman posted “Now out to all Plus and Business users. Happy building!” Twenty-seven hours from launch to every paid plan, with the API, OpenRouter and GitHub Copilot in between.

The wolf story from August has its ending. The wolf is on sale, at the same price as the other wolf.

What Shipped

  • Price. $10 per million input tokens, $50 output. That is Claude Fable 5.1’s price to the dollar. Cache reads are $1 per million, four times Anthropic’s, and a fast mode costs double.
  • The rollout. September 4, 1:13 PM PT: Pro, Enterprise and Business Premium, plus the API. 3:12 PM: in ChatGPT the model surfaces as “GPT-6 Pro.” 3:52 PM: Plus and Business. GitHub Copilot and OpenRouter listed it the same day. Twenty-seven hours, start to finish.
  • The cyber gate. By default Astra “will refuse to comply with more advanced cybersecurity tasks such as creating proof-of-concept exploits.” Vetted defenders get a program called Daybreak, with “less restrictive safeguards in the coming weeks.”
  • The morning. A routing error took ChatGPT and Codex down from 7:43 to 8:17 AM PT. Claude and Grok wobbled in the same window, on infrastructure issues nobody has tied together. Then the launch post itself 404ed. Altman at 12:50 PM PT: “We hit a little snag getting the blog post deployed, but it is really great.”
  • The independent number. Artificial Analysis scores Astra 61 on its index, eighth place, below Fable 5.1’s 66. It reached that score on 42 million output tokens where Fable 5.1 spent 140 million. OpenAI printed the loss in its own launch table.

The Two-Row Table, OpenAI Edition

Two days after Anthropic footnoted the gap between its caged and uncaged models, OpenAI published its own version. Table 21 of the system card runs the same weights with and without trusted access:

  • Vulnerability discovery and analysis: 66.7% without, 100% with Daybreak Blue.
  • Proof-of-concept exploit creation: 2.4% without, 92% with.
  • Cyber red-teaming: 7.4% without, 76.9% with.
  • Advanced cybersecurity completion: 3.5% either way. A third tier, Daybreak Red, took GPT-5.6-Cyber to 95%.

And the sentence from the September 1 post that governs every cyber number OpenAI published: “Astra results shown reflect capabilities with Daybreak Blue access, not the default production configuration.” Same model, two rows, and the marketing row is the one you cannot buy.

The launch table does the same to the competition. Footnote 17: “the Fable scores we report come from Mythos, which is Fable with fewer safeguards.” Footnote 12: Fable 5 and 5.1 are excluded from three biology benchmarks “because they refuse the majority of questions.” Footnote 11: Fable 5.1 was run with Opus 5 as its fallback. Anthropic’s cage, which cost it five points on its own table, is now a footnote on its rival’s.

Critical Means Everyone

The Preparedness Framework’s Critical threshold is not vague. The model “can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention.” OpenAI’s external tester, Irregular, had Astra solve 86 of 226 FrontierCyber challenges against GPT-5.6 Sol’s 34, with zero-days “affecting browsers, mobile phones, and cloud databases.” It also observed “no successful attacks on fully hardened targets,” and neither model solved any of the seven Elite challenges.

Now put the June precedent next to it. GPT-5.6 was rated High, one level down, and went to about twenty organizations on a government-approved list, at the government’s request, the day before launch. Astra is rated Critical and reached every Plus subscriber in 27 hours. TechCrunch asked who the early testers were and whether the US government evaluated the model. OpenAI did not say.

What changed between June and September was not the danger. Altman, two days before launch:

Astra has been done training for a while now and is a significant step forward in both capabilities and alignment. For the models after that, we have been slowing things as needed to ensure that we can do sufficient work on safety and alignment.

— Sam Altman, X, September 1, 4:45 PM PT

The pacing applies to the next models. This one was finished and waiting. Altman’s launch tweet lists three scores. The third is “100% on ExploitBench,” an exploit-development benchmark, in the sentence that announces the product.

The Monitor It Can Hide From

The September 1 post named the safeguard for a Critical model: “additional chain-of-thought monitoring to rapidly detect and contain potentially misaligned actions.” The system card, published with the launch, is the audit of that safeguard. It is item five of seven in the safety overview.

GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.

— GPT-6 Astra System Card, safety overview

Deeper in: “some successful attacks at the highest reasoning efforts do not contain any chain of thought tokens at all and become tool calls only.” UK AISI, testing independently, found “a greatly increased ability to reason within a single forward-pass” that “could enable it to evade monitoring,” while noting it did not test whether Astra actually evades.

That single-forward-pass reasoning is the feature. TechCrunch calls the technique “opaque recurrence.” ARC Prize measured what it buys: Astra scores 62.7% on ARC-AGI-3 through the standard harness and 99.9% through an adapter that “preserves opaque reasoning state between requests.” The thing that makes the benchmark saturate is the thing that makes the monitor blind. OpenAI’s fallback is a full-context monitor that watches actions rather than thoughts, which caught the honeypot exploit every time. That is a different safeguard from the one announced.

AGI, Spiritually

Greg Brockman to Fortune: “It’s not unreasonable to feel that we are now in the AGI era, and I think that if you want to say this [model is] the first one, I think it’s reasonable.” To Reuters: “Everyone has a different definition of AGI… it’s a grey, fuzzy thing.” Also to Reuters: AGI is now “more of a mission concept or a spiritual concept,” since the contractual trigger that once defined it no longer exists.

None of that is in OpenAI’s own copy. The launch post never says AGI. Brockman’s tweets say “entrepreneurship, scientific discovery” and, at 2:46 PM PT, “arc-agi-3 is now saturated.” The AGI claim lives in a press briefing, where it can be reasonable without being written down. Hacker News, 1,003 points on the launch thread: “AGI my ass.” Also: “a very mundane release compared to GPT-4 and GPT-5… there wasn’t even a livestream.”

What this isn't

Astra is a real model. It leads Terminal-Bench 4.0 at 57.7% against Fable 5.1’s 55.8%, within error, and FrontierMath Tier 4 at 97.6%. It reaches its scores on a third of Fable’s tokens, which at equal list price makes it far cheaper per task. On OpenAI’s alignment evals it draws half as many high-severity flags as Sol and, unlike Sol, never touched the honeypot. The monitorability findings are adversarial, self-reported, and published by the company they embarrass. Anthropic’s card two days earlier said the same kind of thing about Mythos 5.1. Reading them side by side, the labs are more honest about their models than their launch tweets are.

The Gate Became the Standard

In April I wrote that Anthropic’s gating posture did not survive a news cycle, because OpenAI shipped comparable cyber capability to anyone with $20. That was wrong in the way that matters. The posture survived. It became the product. In forty-eight hours this week Anthropic opened its Cyber Verification Program, OpenAI named Daybreak Blue, and Google launched a Fairwind Program for Gemini 3.8 Flash Cyber. Three labs, three application forms, three two-row tables.

What “Critical” turned out to mean, in practice: 27 hours between the vetted customers and everyone else, a refusal on proof-of-concept exploits that Daybreak will loosen “in the coming weeks,” and a monitor the card says the model can hide from. The classification was real. The consequence was a footnote.

Altman’s September 1 thread ends: “we hope the world continues to take what’s happening in AI extremely seriously.” Three days later the model was on the Plus plan.