OpenAI has shared new details about Astra, saying the upcoming model has reached a significant level of cybersecurity capability. The company said on Tuesday that Astra is the first large language model to meet its “Critical cybersecurity capability threshold” under its Preparedness Framework.
Astra is designed to identify previously unknown zero-day security vulnerabilities. OpenAI has confirmed that the model will be available soon, although advanced cybersecurity features will initially be limited to a small group of users.
Astra’s cybersecurity capabilities
In a blog post, OpenAI said, “We now believe Astra meets the Critical cybersecurity capability threshold under our Preparedness Framework, meaning that with the right tools and access, it can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step,”
The company said Astra is the first model to receive this designation and will therefore require stronger safeguards during development and before its release. OpenAI has delayed parts of the model’s development and rollout to strengthen protections against cyber misuse and unauthorised model actions.
OpenAI also said Astra was not involved in the Hugging Face incident. The company said lessons from the event have been incorporated into its safety strategy and that additional safeguards have been added for Astra.
The model is trained to reliably refuse harmful cybersecurity requests and follow safety restrictions and measures designed to prevent misuse. It will also include monitoring systems that can stop potentially unauthorised activity.
According to OpenAI, Astra can identify and develop functional zero-day exploits across severity levels in hardened real-world critical systems without human intervention. It can also create and execute end-to-end novel cyberattack strategies against hardened targets when given only a high-level objective.
“We plan to make Astra available soon”, OpenAI said. However, advanced cybersecurity capabilities will initially be available only to a limited group of testers. Broader defensive use is expected to expand later through Daybreak Blue.
OpenAI said Astra recorded higher arbitrary code-execution rates than GPT-5.6 Sol on ExploitBench, which covers 20 high-severity V8 vulnerabilities, while using far fewer output tokens. In cyber jailbreak evaluations, Astra refused 91.5% of requests, compared with 59% for GPT-5.6 Sol.
Also read: Viksit Workforce for a Viksit Bharat
Do Follow: The Mainstream LinkedIn | The Mainstream Facebook | The Mainstream Youtube | The Mainstream Twitter
About us:
The Mainstream is a premier platform delivering the latest updates and informed perspectives across the technology business and cyber landscape. Built on research-driven, thought leadership and original intellectual property, The Mainstream also curates summits & conferences that convene decision makers to explore how technology reshapes industries and leadership. With a growing presence in India and globally across the Middle East, Africa, ASEAN, the USA, the UK and Australia, The Mainstream carries a vision to bring the latest happenings and insights to 8.2 billion people and to place technology at the centre of conversation for leaders navigating the future.


