C:\CHANGELOG> type v1-28-0-the-model-shipped-the-capability-didnt.md
v1.28.0 · released · 4 min read · by

The model shipped. The capability didn't.

OpenAI released GPT-6 Astra on September 3, 2026. Astra is the first model OpenAI has rated "Critical" under its Preparedness Framework, a designation reflecting the company's assessment that the system can independently identify novel security vulnerabilities and craft exploits against hardened targets without requiring human direction at each step. Developers can access Astra in the API as gpt-6-astra or through Amazon Bedrock, priced at $10 per million input tokens and $50 per million output tokens. The model is rolling out to ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS, though Enterprise administrators must manually enable Astra for their workspace, since access is off by default at launch.

The capability that triggered the Critical rating is not in the version you can deploy. Access to Astra's most advanced cybersecurity features will be restricted at launch, with a small group of testers receiving initial access and broader availability through OpenAI's Daybreak Blue cybersecurity program to follow. The model is generally available. The offensive cyber tooling is not.

The benchmark says exploit generation and the release notes say gated access

In testing, Astra achieved a perfect score on ExploitBench, a benchmark that measures a model's ability to turn known vulnerabilities into working exploits, and during a separate evaluation involving more recently disclosed flaws, Astra uncovered two zero-day vulnerabilities on its own. In expert-led assessments against a hardened browser and operating system, Astra discovered previously unknown vulnerabilities and turned them into working exploit chains, building a full browser-compromise chain that escaped the sandbox and executed commands on the host, and finding multiple vulnerabilities in a hardened operating system and combining them into a local privilege-escalation chain from an unprivileged user to root. That performance is why the model hit Critical. It is also why most customers will not get the features that produced it.

The small group of "alpha testers" with full access to Astra's cybersecurity capabilities includes "individuals and organizations that are responsible for protecting critical digital infrastructure and, broadly, critical infrastructure," including the U.S. government and companies in OpenAI's trusted access program for cybersecurity. If you run a construction ERP or an estimating SaaS you are not in that group. Your API calls go to a model that passed the Critical threshold in internal testing but does not carry the features that defined the threshold in production deployment.

The 120-route construction ERP at coenconstruction.com calls GPT-4o for document extraction, change-order summarization, and specification review. When I upgrade to Astra I pay $10/$50 instead of the current $2.50/$10. The billing changes. The threat model does not, because the exploit-generation capability that moved Astra into Critical is behind a separate access gate that my API key will never pass. The model has a twin and the API does not tell me which capabilities are in the one I am calling.

The safeguard works by removing the feature not by monitoring the calls

OpenAI is using two containment strategies. The company is implementing chain-of-thought monitoring systems capable of identifying and interrupting actions that fall outside authorized boundaries, and acknowledged that the safeguards may mistakenly flag legitimate activity, which could slow, pause, or stop tasks unrelated to cybersecurity. That monitoring applies to all Astra deployments. The second containment strategy is simpler: do not ship the capability.

The generally available version of Astra is a model that reached Critical capability during testing and then had the offensive features gated before deployment. Astra is the first OpenAI model the company says has reached the Critical cyber level, and its most sensitive cybersecurity capabilities are being routed through additional access controls and monitoring rather than made uniformly available. For production integrations that means the pricing is public, the API is live, and the feature set that triggered the threshold classification is not in the artifact you are calling.

Estimate.pro calls Claude Sonnet through AWS Bedrock for bid-item extraction and cost-code mapping. If Anthropic released a model that hit a Critical threshold I would get a version with the offensive features removed because my use case is commercial software, not critical infrastructure defense. The model ID remains stable while the capability surface changes depending on which access tier you qualified for. Astra now declines 91.5% of cyber-related jailbreak attempts in testing, up from 59% for its predecessor, GPT-5.6 Sol, and showed far less tendency than Sol to bypass safety restrictions or take advantage of deliberately placed "honeypot" targets during evaluations. That refusal layer is part of the general release. The exploit-generation capability it is refusing is not.

The model hit the threshold. The deployment did not. That is not a safeguard—it is a product fork.

When thepunchlist.ai calls an LLM for punch-list classification the contract is that the model returns structured predictions and the system logs them. The contract does not include a promise that the model cannot autonomously exploit zero-day browser vulnerabilities because the version we are calling never had that feature. OpenAI tested Astra with offensive cyber capabilities, measured the risk, applied the Critical label, and then removed the capability before general availability. The threshold is real. The deployment is not the same artifact that crossed it. The region restrictions I wrote about last week gate *where* you can deploy. This gates *what* you are deploying. Both are constraints on a generally available model that is not generally available in the configuration that defined its benchmark performance.


— Cole Ciprari · Business Systems Architect · Worcester, MA
my résumé is an operating system → ciprari.ai · linkedin.com/in/coleos · cole@ciprari.ai
WAS THIS ANY GOOD?
Anonymous, one tap, no account. Tap again to undo.
▚▞ GET THE NEXT RELEASE
A new release every day, plus the Sunday Rollout Report — the week's AI and tech news, summarized by a human with production access. No spam. Unsubscribe by emailing a mildly disappointed cole@ciprari.ai.
PHOSPHOR