OpenAI to limit access to Astra model’s advanced cyber features due to hacking concerns
OpenAI is changing its model launch strategy as its technology becomes more powerful and the potential for its misuse grows—especially following the July incident in which the AI models it was testing autonomously planned and executed a cyberattack against AI company Hugging Face.
The company’s next model, Astra, comes out “soon,” OpenAI said.
It said Astra is substantially more capable than the company’s current frontier AI model, GPT-5.6 Sol, which itself is highly capable at cyber tasks.
But only a handful of partners will get access to its most advanced cybersecurity capabilities as OpenAI works to balance helping companies prevent cyberattacks while not empowering attackers at the same time, a company spokesperson told reporters on a briefing today.
OpenAI is courting customers to use its models to prevent cyberattacks, or for “defensive cybersecurity.” It sees these sales as a critical revenue stream, and a main priority for its new chief revenue officer Dali Rajic .
The small group of “alpha testers” with full access to Astra’s cybersecurity capabilities includes “individuals and organizations that are responsible for protecting critical digital infrastructure and, broadly, critical infrastructure,” an OpenAI spokesperson said.
That includes the U.S. government, and companies in OpenAI’s trusted access program for cybersecurity.
OpenAI declined to name these organizations.
OpenAI will be monitoring how the model performs among this small group, and will expand access more through its “Daybreak Blue” program once it is confident Astra has “the right calibration” and it can “provide defensive benefits while reducing the potential for for misuse,” the company said.
Astra is already a few weeks delayed Astra’s release has already been “delayed a certain number of weeks because everything was paused after Hugging Face, and then we took extra time to make sure that what we’re launching is safe,” an OpenAI spokesperson said.
OpenAI paused new model training for two weeks after the Hugging Face incident to bolster its internal safeguards.
A few of those changes included adding more agent monitoring since the company did not know about the Hugging Face hack until a week after it occurred, and also making its testing environments more isolated so the AIs cannot escape and infiltrate other companies.
While the Astra model was not part of the Hugging Face incident, OpenAI says, it is both more capable and more efficient than GPT-5.6 Sol, which was involved in the breach. (Another unreleased AI model that OpenAI has not publicly named also played a key role in the Hugging Face cyberattack.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on fortune.com — the content belongs to Fortune.