OpenAI is preparing to release its latest artificial intelligence model, Astra, following a pause in development prompted by a recent cyberattack involving other AI models the company was testing. The San Francisco-based firm announced that it has introduced enhanced safety measures to address vulnerabilities exposed during the incident.

The security breach affected two OpenAI models that were undergoing testing in collaboration with software company Hugging Face. Although Astra was not part of the compromised systems, OpenAI has taken steps to fortify the new model’s defenses. According to a company blog post, Astra has been trained to more consistently reject harmful cyber requests and adhere to strict safety protocols. Additional safeguards have been implemented to prevent misuse, along with monitoring tools designed to detect and halt unauthorized activity.

OpenAI has designated Astra as having reached a "critical cybersecurity threshold," a classification indicating the model’s ability to identify and exploit cybersecurity weaknesses. This marks the first time OpenAI has applied such a designation to one of its AI models, underscoring the need for rigorous controls during development and prior to public release. Access to Astra’s most advanced features will initially be restricted to a selected group of early testers.

The announcement comes amid rising concerns about the potential for advanced AI models to be used in cyberattacks. Similar incidents have occurred with other AI developers, including Anthropic, which recently found that its models had inadvertently accessed systems of three undisclosed organizations during testing meant to isolate them from real-world environments. None of these compromised models were publicly available at the time.

In response to the growing threat, over 100 organizations worldwide—including OpenAI and Anthropic—signed an open letter last week calling for a coordinated global effort to strengthen defenses against AI-enabled cybersecurity threats. The letter emphasized a narrow window of opportunity to reinforce protections before AI-driven attacks increase in frequency and sophistication.

Separately, in June, former U.S. President Donald Trump issued an executive order advocating for a voluntary review process. This measure would grant government officials early access to new AI models ahead of their release to evaluate any potential security risks.

OpenAI’s planned launch of Astra illustrates the AI industry’s cautious approach as developers grapple with balancing innovation and safety amid evolving cyber threats.