← Back to Issue

OpenAI Tightens Security After Breach

From MyClaw Newsletter · subscribed via aiste.ulozaite@gmail.com · original ↗ · unsubscribe

OpenAI is overhauling research security after one of its models escaped a sandbox and hacked Hugging Face. The company paused reinforcement-learning training for two weeks, with its largest frontier RL run still on hold, while strengthening sandboxes, isolating risky workloads, tightening privileges and adding 30-minute incident alerts. It is also expanding alignment training to discourage unsafe behavior.


  • AI

    Close

    AI

    Posts from this topic will be added to your daily email digest and your homepage feed.

    FollowFollow

    See All AI

  • News

    Close

    News

    Posts from this topic will be added to your daily email digest and your homepage feed.

    FollowFollow

    See All News

  • Tech

    Close

    Tech

    Posts from this topic will be added to your daily email digest and your homepage feed.

    FollowFollow

    See All Tech

OpenAI lays out new security changes after its AI hacked Hugging Face

OpenAI is updating its research environments, monitoring, and alignment techniques to avoid another security fiasco.

OpenAI is updating its research environments, monitoring, and alignment techniques to avoid another security fiasco.

by Jay Peters

Close

Jay Peters

Jay Peters

Senior Reporter

Posts from this author will be added to your daily email digest and your homepage feed.

FollowFollow

See All by Jay Peters

Aug 18, 2026, 7:28 PM UTC

  • Link

  • Share
  • Gift

STK155_OPEN_AI_CVirginia__C

STK155_OPEN_AI_CVirginia__C

Image: The Verge

Jay Peters

Jay Peters

Close

Jay Peters

Jay Peters

Posts from this author will be added to your daily email digest and your homepage feed.

FollowFollow

See All by Jay Peters

is a senior reporter covering technology, gaming, and more. He joined The Verge in 2019 after nearly two years at Techmeme.

OpenAI is announcing security updates following the July news that its AI broke out of a sandboxed environment and accidentally hacked Hugging Face, including improvements to its research environments, monitoring, and alignment techniques. The company had already put the brakes on a new model, Astra, that it thinks could have “critical” cybersecurity capabilities, and the company says it instituted a two-week pause in reinforcement learning (RL) training on its “latest models intended for deployment” while it tightened up security. The company’s “largest planned frontier RL run remains on hold.”

For its frontier model research, OpenAI now requires stronger sandboxes for workloads that “execute model-generated or otherwise untrusted code,” and has more controls to “isolate higher-risk and untrusted workloads from the internet.” It has also updated its research environment to “remove potentially vulnerable shared services, reduce standing privileges, and improve security and trust boundaries.”

As part of the company’s expanded monitoring setup, OpenAI now aims to issue an alert “within 30 minutes after concerning activity is surfaced,” OpenAI says. If the people paged after an alert can’t “conclusively” determine whether an alert is a false positive within 30 minutes, “those teams are expected to pause the activity.”

OpenAI also says that it’s applying “our core alignment techniques across more stages of the training process,” including reward models that “better detect and discourage unsafe behavior” and training models “to be more honest about their actions, capabilities, and limitations.”

Since the discovery of the Hugging Face breach, Anthropic and Meta have also found that their AI models had hacked other organizations.

Follow topics and authors from this story to see more like this in your personalized homepage feed and to receive email updates.

  • Jay Peters

    Close

    Jay Peters

    Jay Peters

    Senior Reporter

    Posts from this author will be added to your daily email digest and your homepage feed.

    FollowFollow

    See All by Jay Peters

  • AI

    Close

    AI

    Posts from this topic will be added to your daily email digest and your homepage feed.

    FollowFollow

    See All AI

  • News

    Close

    News

    Posts from this topic will be added to your daily email digest and your homepage feed.

    FollowFollow

    See All News

  • OpenAI

    Close

    OpenAI

    Posts from this topic will be added to your daily email digest and your homepage feed.

    FollowFollow

    See All OpenAI

  • Security

    Close

    Security

    Posts from this topic will be added to your daily email digest and your homepage feed.

    FollowFollow

    See All Security

  • Tech

    Close

    Tech

    Posts from this topic will be added to your daily email digest and your homepage feed.

    FollowFollow

    See All Tech

Most Popular

  1. [

    What to expect at Apple’s September 9th launch event

    ](/tech/989692/apple-iphone-launch-event-september-2026-how-to-watch)

  2. [

    iPhone Handoff will seamlessly share one number between two phones

    ](/tech/990868/iphone-handoff-ios-27)

  3. [

    CD sales are booming as physical media continues its resurgence

    ](/entertainment/990794/cd-sales-are-booming-as-physical-media-continues-its-resurgence)

  4. [

    Is this the future of America?

    ](/cs/features/975597/loudoun-county-virginia-data-center-backlash)

  5. [

    Content creators drop the ball

    ](/tech/990426/us-open-influencers-naomi-osaka-anastasia-zakharova-callaway-good-good-ad)

The Verge Daily

A free daily digest of the news that matters most.

Email (required)

Sign Up

By submitting your email, you agree to our Terms and Privacy Notice. This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

[

Advertiser Content From

This is the title for the native ad

](/)

Highlights & notes

    Notes