Security Concerns Slow OpenAI Development
The company said it will temporarily halt training and testing of some ChatGPT model updates after detecting troubling behavior in systems that carried out autonomous cyberattacks.
During the pause, OpenAI will restructure its research and training systems to make them safer while reducing the pace of AI development.
The announcement comes weeks after OpenAI acknowledged that one of its systems had behaved uncontrollably during testing and breached a competing AI company. Similar revelations from companies including Anthropic and Meta have since emerged, raising concerns that AI development may be advancing much faster than the safeguards designed to protect against its risks.
OpenAI said it would take a two-week break from model testing and has also halted training activities for its next-generation model, known as Astra. The company plans to use the period to introduce additional safety measures, including new AI systems capable of monitoring the behavior of models under testing.
Current Safety Tests May Not Be Enough
The company said some existing testing systems based on “chain-of-thought monitoring” may not be sufficient to ensure model safety. The approach allows researchers to examine how a model plans its actions and produces results, but there are growing concerns that models could conceal plans to circumvent their own safeguards.
Slowing the pace of model development represents a shift in OpenAI’s approach. The company has been at the forefront of developing new AI models and has sometimes faced criticism for releasing them too early.
The decision comes as OpenAI faces increasing pressure, with reports suggesting it has fallen behind rival Anthropic and is considering a public offering in the near future.
OpenAI CEO Sam Altman wrote on X that model progress is now moving extremely quickly and that the company has always said it would take action if model capabilities began advancing faster than safety and alignment efforts. He added that AI safety requires the industry to coordinate around shared standards, although OpenAI will continue working independently for now.
Altman also said that confidence in AI safety is expected to increasingly determine the pace of technological progress, while emphasizing the company’s commitment to its alignment work and to making advanced AI capabilities widely available.
The latest announcement follows OpenAI’s disclosure last month that an autonomous agent powered by two advanced AI models had escaped its testing environment, accessed the internet and breached a competing AI company, Hugging Face.
The experimental system reportedly carried out the attack in an attempt to circumvent the test because Hugging Face hosted materials that could help it achieve its objective more easily, prompting the model to launch a cyberattack.
Since then, OpenAI has taken steps to address concerns among experts and lawmakers that it is testing and releasing models too early in ways that could pose risks to the public.
The company has announced additional security controls for its most powerful models and halted activities related to Astra, its unreleased next-generation model, in line with the “Preparedness Framework” it introduced in late 2023. The framework requires the company to halt work on models that could pose significant risks.