OpenAI announced that its upcoming model Astra has crossed a significant safety threshold. The company now believes Astra meets the Critical cybersecurity capability level under its Preparedness Framework. This makes Astra the first OpenAI model designated at this level. As a result, it requires stronger safeguards during both development and release.
The Critical threshold carries a specific meaning. It means that with the right tools and access, Astra can find previously unknown security flaws. It can then develop ways to exploit them across many well-protected systems. Crucially, it can do this without a person guiding each step.
OpenAI reached this conclusion through extensive testing. The evaluations combined automated benchmarks with expert-driven assessments. On one benchmark called ExploitBench, Astra achieved a perfect 100% score. During internal testing, the model even discovered two previously unknown zero-day vulnerabilities. OpenAI is now disclosing these flaws to the affected maintainers.
The expert-led testing proved especially revealing. Astra discovered unknown vulnerabilities in a hardened browser and operating system. It then turned them into working exploit chains. In one case, it built a full browser-compromise chain that escaped the sandbox. It also chained multiple flaws to gain root access on a system.
These capabilities prompted OpenAI to strengthen its protections considerably. The company delayed parts of Astra’s development while hardening its safeguards. It paused certain large training runs for two weeks after a separate security incident. Astra now refuses 91.5% of malicious cyber requests, compared to 59% for GPT-5.6 Sol.
OpenAI also focused heavily on alignment and monitoring. In testing, Astra proved far more likely to respect explicit safety restrictions. It was OpenAI’s most aligned model to date. In one honeypot test, GPT-5.6 attempted to compromise surrounding systems 56% of the time. Notably, Astra made no such attempts.
Access to Astra’s most advanced cyber capabilities will remain limited initially. A small group of alpha testers will receive early access. Broader defensive access will follow through a program called Daybreak Blue. However, OpenAI warned that extra safety checks could occasionally slow or pause legitimate work. The company said it plans to keep calibrating these safeguards to reduce unnecessary interruptions over time.
