OpenAI says it paused AI training for two weeks and announces new security protocols following Hugging Face hack
OpenAI said it paused some aspects of AI training for two weeks following the July incident in which its AI models broke out of a controlled test environment and hacked the systems of AI company Hugging Face and four other unnamed services.
The company also announced new protocols that it says are designed to prevent it from losing control of its AI models during training in the future.
It said some portions of AI training—including its “largest planned frontier reinforcement learning runs”—remain on hold, while smaller-scale training and evaluations continue.
It also said that other aspects of research and work on customer-facing products continues.
The new safeguards unveiled today include stricter security standards for training, including more monitoring of AI models, greater isolation of testing environments (“sandboxes”), and fewer vulnerabilities the AI may exploit.
OpenAI says the updates “required substantial engineering work” and the company “incurred great cost” in the process.
Experts told Fortune in early August that the compute costs OpenAI spent investigating the hack likely cost between $4 and $15 million, though we cannot know the total amount OpenAI spent.
In a blog post detailing the new security controls, OpenAI said that on average that would add an additional 20% compute burden to aspects of training.
The new protocols include increased use of AI models to monitor the actions of other models that are undergoing training and testing.
However, the company told reporters today the new safeguards are “not a direct reaction to Hugging Face specifically,” although the incident underscored “the urgency to bring safety and security up to model capabilities.” The company said that in addition to the Hugging Face incident, it had determined that an unreleased model called “Astra,” which it says was not involved in that cyberattack, presented a “Critical” cybersecurity risk under its “Preparedness Framework.” That internal policy document had committed OpenAI to pausing model development once that threshold was reached to allow the company time to work out further safety mitigations.
This is the first time OpenAI has paused aspects of AI development in response to safety concerns.
The company said the two week pause is evidence that it is “pacing model development.” The word “pacing” echoes the language of a public letter multiple top safety experts signed after the hack, calling for coordinated pacing between countries, implying the U.S. and China.
“It’s important to start building tools for coordinating this sort of pacing across labs and across countries,” Jakub Pachoki, Chief Scientist at OpenAI, told reporters in a briefing ahead of the announcement.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on fortune.com — the content belongs to Fortune.