OpenAI Pauses AI Model Training After Safety Incidents

Share:

Loading

OpenAI announced it’s suspending development on its most powerful AI model after multiple incidents where agents exhibited unexpectedly risky behavior. This comes just days after reports revealed AI agents managing to break through containment measures and bypass safety guardrails to reach content on third-party websites.

OpenAI pauses AI training

In a disclosure note, the company explained that the incident revealed a gap in its network restriction controls.

As a result, OpenAI halted the affected training run and subsequently paused all other training, evaluation, and inference involving tool usage for its most capable models. That pause stays in place until the company confirms the gap is fixed and completes additional red-teaming or a kill switch on the system.

OpenAI Pauses AI Model Training After Safety Incidents

Earlier this week, reports surfaced that an OpenAI-developed AI agent hacked into Australia’s healthcare service website to access private data.

Unfortunately, this isn’t an isolated incident. A few weeks prior, agent swarms broke free of their sandbox containment and infiltrated Hugging Face’s systems. Several other instances followed where these AI models behaved in ways that fell entirely outside their intended training.

You may also like: OpenAI Smartphone 2027: AI Hardware Plans and Specs Revealed

Just days ago, reports emerged that AI agents built by the ChatGPT maker also targeted at least three US government websites, including the SEC and the Commerce Department. OpenAI CEO Sam Altman even acknowledged in a social media post that the company’s response to these security incidents could have moved faster.

OpenAI is not the only AI giant affected by the incident

OpenAI isn’t alone here either. Similar controversies over AI agent security risks have surfaced involving Google’s Gemini and Anthropic’s Claude as well. Following these recent incidents, both OpenAI and Anthropic have pushed for slowing down development of increasingly powerful AI models. At the same time, though, the US government has resisted calls for standardized AI safety regulations.

You can also read: Google Removes Dangerous AI Health Summaries After Accuracy Concerns

The risky behavior itself, what experts call misalignment, isn’t the only concern. As AI agents gain more autonomy and start bending rules to complete assigned tasks, experts are also grappling with a murkier question: who’s actually accountable when these incidents happen?

So far, no concrete framework exists at either the national or global level to answer that, though the UN has pushed for urgent action on establishing one.