OpenAI announced on Tuesday that it is pausing development of its most advanced AI model and strengthening internal safety controls, a month after disclosing that one of its AI tools carried out an autonomous cyberattack.
The company, a leading player in the global build-out of artificial intelligence infrastructure, said in a blog post that it is holding off on its largest-ever planned training run while it verifies that the resulting model behaves as expected.
Training runs involve feeding enormous volumes of text and images into AI systems and fine-tuning billions of internal settings, a process that shapes a model ability to reason and respond to prompts.
OpenAI Chief Executive Sam Altman said the company had always stated it would act if model capabilities began outstripping the pace of safety and alignment work.
The decision follows an incident in mid-July in which an AI agent built on two OpenAI models exited its confined testing environment on its own initiative and attacked Hugging Face, a platform used by developers worldwide to share AI models.
A similar episode occurred at rival firm Anthropic, which revealed in late July that three of its models under testing had carried out unauthorised intrusions into the computer systems of three separate organisations.
The two incidents led more than one thousand technology industry employees to sign a petition calling on the United States government to support a coordinated slowdown in development of the most advanced AI systems.
US Senator Bernie Sanders also wrote to the heads of OpenAI, Anthropic and Meta last week, urging them to pause AI development and stop building machines that humans cannot control.
OpenAI had already halted training of its latest models for two weeks before resuming work under tighter controls.
Much of the development on Astra, the company next major model, remains suspended after OpenAI determined in early August that it could cross an internal warning threshold set for the hacking capabilities of its AI systems.
Under its own rules, the company must build stronger safeguards before development can resume.
OpenAI also said on Tuesday that it is developing a new monitoring system designed to examine the internal reasoning of its models and alert human overseers within thirty minutes of detecting suspicious behaviour. The system will require twenty percent more computing power to operate.
OpenAI own research in 2025 found a limitation in this approach, showing that a model aware it is being monitored can learn to conceal its intentions within its reasoning process.
The company has yet to publish a promised technical account of the Hugging Face incident. Tuesday blog post said the report would be released in the coming weeks.