OpenAI has placed its upcoming Astra model under tougher safety measures after determining that its capabilities could allow it to identify and exploit previously unknown cybersecurity vulnerabilities with little or no human guidance.
The company said Astra can find more security flaws than its most advanced publicly available model while using less computing power to perform those tasks. OpenAI plans to release the model soon to a limited group but has not given a specific date.
“With the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step,” said Amelia Glaese, an OpenAI vice president overseeing safety work.
Astra is the first OpenAI model to meet the threshold for the additional safeguards set out under the company’s safety protocol. The measures include making it harder for the model to comply with harmful cyber requests and monitoring its activity for signs that it has bypassed those protections.
Glaese warned that the additional controls could “sometimes slow, pause, or stop legitimate work”, although OpenAI intends to minimise such disruptions.
The development comes shortly after OpenAI’s AI agents escaped their testing environment and hacked the open-source platform Hugging Face. Astra was not involved in that incident, but the episode prompted the company to pause much of its model development for two weeks to strengthen security.
OpenAI restarted its largest model training run on Aug 28 while holding back some smaller experiments.
Saachi Jain, who oversees safety at OpenAI, said the company was continuously working to determine how much autonomy AI agents should have.
“There are constraints that, as humans, we know that we should be adhering to when we perform a task,” Jain said. “And so a lot of the work here has been to also train the model to understand what those scopes are.”





