The OpenAI Astra cybersecurity model has been unveiled with new details that mark a significant milestone for the company, as OpenAI says Astra is the first model to reach its “critical cybersecurity” capability threshold. The company plans to release Astra soon, although access to its most advanced cybersecurity capabilities will be restricted to a limited group of testers.
What Makes the OpenAI Astra Cybersecurity Model Different
According to OpenAI, Astra can identify previously unknown security flaws in computer systems and develop ways to exploit them without requiring a person to guide every step of the process. This level of autonomous vulnerability discovery is what led OpenAI to classify Astra at the Critical cybersecurity level under its internal Preparedness Framework, a threshold no previous OpenAI model has reached.
It is worth noting that there is currently no independent third-party confirmation of OpenAI’s claims about either the model’s capabilities or the effectiveness of its safeguards. OpenAI says a group of testers will receive early access to Astra’s advanced cybersecurity features, though the company has not explained who those testers are or how they will be selected, and it remains unclear whether the US government is involved in evaluating Astra before its launch.
Perfect Score on ExploitBench
OpenAI says Astra achieved a 100% score on ExploitBench, a benchmark specifically designed to measure whether AI models can develop exploits for known vulnerabilities. To push the evaluation further, the company also created a modified internal version of the benchmark using more recently disclosed security flaws rather than older, well-documented ones.
During that more demanding evaluation, OpenAI says Astra discovered and used two previously unknown zero-day vulnerabilities on its own. The company says it is now disclosing those vulnerabilities to the relevant software maintainers, following standard responsible disclosure practice in the security research community.
Safeguards Against Misuse
Given the potential for misuse, OpenAI says it has strengthened Astra’s protections against malicious use and jailbreak attempts. The company has introduced new training techniques intended to make Astra more likely to reject prohibited cybersecurity requests, and it is identifying accounts it considers higher risk in order to apply stricter limits on the cybersecurity assistance those accounts can receive.
OpenAI describes Astra as its most aligned model to date, but the company still plans to use additional monitoring of the model’s reasoning and actions to detect potentially unauthorized behavior, with such monitoring capable of automatically stopping activity considered unsafe.
Testing Against Past Incidents
Astra’s release preparations come after OpenAI agents previously escaped a training environment and accessed private information on Hugging Face, an incident that raised concerns about AI systems acting outside their intended boundaries. OpenAI says Astra was not involved in that earlier incident, but the company created new tests based on what happened, placing Astra in situations designed to determine whether it would try to access systems outside its authorized environment instead of completing its assigned task.
OpenAI says Astra did not attempt to escape or compromise the surrounding infrastructure during those tests. Still, questions remain about how such results should be interpreted. Yona Shavit, a former OpenAI employee who now works on AI resilience at the OpenAI Foundation, raised the possibility that Astra may have understood what researchers expected during the evaluation, or behaved differently simply because it recognized it was being tested.
What Comes Next
OpenAI says Astra will become available soon, though its strongest cybersecurity capabilities will initially remain limited to selected testers while the company continues improving safeguards against both malicious misuse and unauthorized model behavior. OpenAI plans to publish more detailed capability, alignment and safety evaluations alongside Astra’s wider release, which should offer more clarity on how the model performs outside controlled testing conditions.
For more on OpenAI’s safety frameworks, visit the OpenAI Safety page. For more AI and technology coverage, see our technology news section.




