Tuesday, August 18, 2026 SourcesAbout🌓
🇺🇸 US ▾
BREAKING
Dominion still has pending lawsuits against election deniers such as Rudy Giuliani and Sidney Powell Russia is 'going backwards' in equipment and deploying post WWII-era tanks, according to Western officials Podcast: One country musician is calling for other artists to oppose assault rifles Bidets save you money and reduce waste — we tested the best options out there 50+ products to make your life easier and our planet cleaner Mother's Day is around the corner. Here are 50+ thoughtful gifts she'll love A head-to-toe guide of how men should dress this spring, and where they should shop 42 of the most useful travel products you can buy on Amazon The 7 best high-yield savings accounts of April 2023 Taxes are due tomorrow. Here's how to file for an extension Dominion still has pending lawsuits against election deniers such as Rudy Giuliani and Sidney Powell Russia is 'going backwards' in equipment and deploying post WWII-era tanks, according to Western officials Podcast: One country musician is calling for other artists to oppose assault rifles Bidets save you money and reduce waste — we tested the best options out there 50+ products to make your life easier and our planet cleaner Mother's Day is around the corner. Here are 50+ thoughtful gifts she'll love A head-to-toe guide of how men should dress this spring, and where they should shop 42 of the most useful travel products you can buy on Amazon The 7 best high-yield savings accounts of April 2023 Taxes are due tomorrow. Here's how to file for an extension
Technology

OpenAI lays out new security changes after its AI hacked Hugging Face

The Verge ·
OpenAI lays out new security changes after its AI hacked Hugging Face

OpenAI is updating its research environments, monitoring, and alignment techniques to avoid another security fiasco.

OpenAI is updating its research environments, monitoring, and alignment techniques to avoid another security fiasco.

OpenAI is announcing security updates following the July news that its AI broke out of a sandboxed environment and accidentally hacked Hugging Face , including improvements to its research environments, monitoring, and alignment techniques. The company had already put the brakes on a new model, Astra , that it thinks could have “critical” cybersecurity capabilities, and the company says it instituted a two-week pause in reinforcement learning (RL) training on its “latest models intended for deployment” while it tightened up security. The company’s “largest planned frontier RL run remains on hold.”

For its frontier model research, OpenAI now requires stronger sandboxes for workloads that “execute model-generated or otherwise untrusted code,” and has more controls to “isolate higher-risk and untrusted workloads from the internet.” It has also updated its research environment to “remove potentially vulnerable shared services, reduce standing privileges, and improve security and trust boundaries.”

As part of the company’s expanded monitoring setup, OpenAI now aims to issue an alert “within 30 minutes after concerning activity is surfaced,” OpenAI says. If the people paged after an alert can’t “conclusively” determine whether an alert is a false positive within 30 minutes, “those teams are expected to pause the activity.”

OpenAI also says that it’s applying “our core alignment techniques across more stages of the training process,” including reward models that “better detect and discourage unsafe behavior” and training models “to be more honest about their actions, capabilities, and limitations.”

Since the discovery of the Hugging Face breach, Anthropic and Meta have also found that their AI models had hacked other organizations.

Read the full article on The Verge ›

5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.theverge.com — the content belongs to The Verge.

This story in other outlets

More from The Verge

See all ›

More in Technology

See all ›