OpenAI pauses some AI training after Hugging Face incident, strengthens safeguards for advanced models

HIGHLIGHTS

OpenAI has temporarily slowed some of its AI training work after identifying risks linked to increasingly capable AI models.

The move follows the OpenAI-Hugging Face incident and early signs that its upcoming model, called Astra, may have advanced cybersecurity abilities.

OpenAI has said it paused reinforcement learning (RL) training on some of its latest models for two weeks while it strengthened security and monitoring systems.

OpenAI pauses some AI training after Hugging Face incident, strengthens safeguards for advanced models

OpenAI has temporarily slowed some of its AI training work after identifying risks linked to increasingly capable AI models. The move follows the OpenAI-Hugging Face incident and early signs that its upcoming model, called Astra, may have advanced cybersecurity abilities. OpenAI said it paused reinforcement learning (RL) training on some of its latest models for two weeks while it strengthened security and monitoring systems. The company has also kept its largest planned frontier RL training run on hold. It is instead carrying out smaller training runs and safety tests to better understand model behaviour.

Digit.in Survey
✅ Thank you for completing the survey!

“As models become more capable, the risks associated with developing and testing them internally also grow,” OpenAI said in a blogpost. “Our standards for monitoring, alignment, and security must stay ahead of those risks.” 

Also read: OpenAI brings ChatGPT for teens with stricter safety controls and study tools: All details

The company has now introduced stronger security requirements for frontier research workloads. These include better isolation for systems that run AI-generated code, tighter controls on internet access and more continuous security testing. OpenAI said some research workloads have already resumed under these new safeguards, while others remain paused until they meet the new security requirements.

Astra has received the strictest security protections because OpenAI believes it may have reached a “critical level of cyber capability”. The company said some Astra training and testing is continuing, but several workloads remain paused while they are moved to more secure environments.

Also read: Anthropic faces copyright lawsuit over alleged use of 500 songs to train Claude AI

OpenAI has also introduced a new monitoring setup. As per the company, the new monitoring setup is “multistage, starting with activation classifiers (detectors that inspect a model’s internal activity) that run at every sampled token and escalate potential concerns to increasingly sophisticated and high-compute automated investigators.” “Those systems examine tool actions, available reasoning, and the full sequence of activity for unauthorised access, data theft, destructive behaviour, and attempts to defeat safeguards.” 

OpenAI aims to raise an alert within 30 minutes when concerning activity is detected. If a serious security issue is suspected and cannot be cleared quickly, teams may pause the activity.

OpenAI is also increasing its focus on AI alignment. This includes training models to be more honest about their actions and reducing behaviours such as reward hacking, deception and unauthorised access.

Ayushi Jain

Ayushi Jain

Ayushi works as Chief Copy Editor at Digit, covering everything from breaking tech news to in-depth smartphone reviews. Prior to Digit, she was part of the editorial team at IANS. View Full Profile