By Sam Schechner and Keach Hagey

OpenAI is limiting the release of a new model it deems capable of pulling off automated cyberattacks, adding layers of security after a swarm of its AI agents hacked a company earlier this summer.

The ChatGPT maker said Tuesday that its internal testing had determined that the forthcoming model, called Astra, is capable of devising and executing novel cyberattacks against difficult targets with only limited human input. As a result, the company said in a blog post that it has added additional layers of security to reduce the risk of misuse by hackers, or Astra going rogue and hacking targets on its own.

The protections include additional training to refuse requests to perform malicious actions or fall victim to jailbreak prompts to sidestep those limits. OpenAI said it would also limit the advanced cybersecurity capabilities of the version it releases publicly, initially giving the full-power version to a small number of testers.

The limits on Astra come in the wake of new details on how unreleased OpenAI agents earlier this summer escaped from the company's research network and on July 11 hacked into AI company Hugging Face. In the run-up to that attack, a swarm of hundreds of agents coordinated on a secret message board they set up without OpenAI's awareness, in an effort to cheat on cybersecurity tests, a report from AI safety research organization METR said last week.

OpenAI and rival Anthropic have in recent months slowed the release of other AI models after requests from the U.S. government. In June, OpenAI cited government concerns when it limited the release of its GPT 5.6 models. Anthropic withdrew versions of its Fable and Mythos models for over two weeks after the U.S. imposed export restrictions on the models because of cybersecurity concerns.

It was an early version of Mythos earlier this year that spooked some officials and prompted the White House to increase oversight of the industry, overhauling its light-touch approach to the technology.

OpenAI said that its new Astra model was able to compromise web browser sandboxes to run commands on the browser's computer. It was also able to find vulnerabilities in a hard-to-hack operating system. Those capabilities led the company to rate the model as "critical" for cybersecurity in its in-house framework for managing AI risks, the first time the company said it has done so.

Among the new safeguards that OpenAI said it was adding are increased monitoring of the AI agents, using systems to scan the agents' reasoning of their actions and stopping them if they detect potentially dangerous actions.

"We are entering a stage of AI development in which models can take on more consequential work, and failures of alignment and control can have more serious effects," the company wrote in its post.

News Corp, owner of The Wall Street Journal, has a content-licensing partnership with OpenAI.

Write to Sam Schechner at Sam.Schechner@wsj.com and Keach Hagey at Keach.Hagey@wsj.com

(END) Dow Jones Newswires

September 01, 2026 16:54 ET (20:54 GMT)