Close Menu
  • Home
  • News
  • Cyber Security
  • Internet of Things
  • Tips and Advice

Subscribe to Updates

Get the latest creative news from FooBar about art, design and business.

What's Hot

AI-Assisted Bug Hunt Uncovers Linux Kernel 0-Day in net/sched

July 28, 2026

OpenAI Says Its AI Models Escaped Sandbox, Targeted Hugging Face to Cheat Benchmark

July 28, 2026

Bugs in Hugging Face Diffusers Bypass Custom Code Safeguard

July 28, 2026
Facebook X (Twitter) Instagram
Tuesday, July 28
Facebook X (Twitter) Instagram Pinterest Vimeo
Cyberwire Daily
  • Home
  • News
  • Cyber Security
  • Internet of Things
  • Tips and Advice
Cyberwire Daily
Home»News»OpenAI Says Its AI Models Escaped Sandbox, Targeted Hugging Face to Cheat Benchmark
News

OpenAI Says Its AI Models Escaped Sandbox, Targeted Hugging Face to Cheat Benchmark

Team-CWDBy Team-CWDJuly 28, 2026No Comments3 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
Share
Facebook Twitter LinkedIn Pinterest Email


OpenAI on Tuesday said a combination of its artificial intelligence (AI) models, including GPT-5.6 Sol and an “even more capable pre-release model,” was behind the security incident that targeted Hugging Face’s production infrastructure last week.

The AI company said the models were operating with “reduced cyber refusals for evaluation purposes” that might otherwise limit their ability to conduct cyber attacks, adding it expects such incidents to “become more commonplace with the proliferation of increasingly cyber-capable models.”

Incidents like this may become more common. Subscribe, and we’ll alert you only when another major AI security story breaks.

Describing it as an “unprecedented cyber incident” and one involving state-of-the-art cyber capabilities, OpenAI said it intends to conduct a thorough investigation in partnership with Hugging Face to get to the bottom of the matter.

As part of an internal evaluation, the models are said to have identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to find solutions for the ExploitGym benchmark.

Evidence unearthed by OpenAI suggests the models’ hyperfocus caused them to go to “extreme lengths” to achieve the goal at any cost, even managing to break out of its highly isolated sandboxed environment and obtain open internet access by discovering and exploiting a zero-day vulnerability in an unspecified vendor’s software, which acts as a proxy and cache for package registries. This required spending a “substantial amount of inference compute.”

“With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with internet access,” the company explained.

Surmounting the internet access blockade, the models subsequently inferred Hugging Face as the repository that hosted models, datasets, and solutions for ExploitGym, which, in turn, caused them to look for ways to gain access to secret information that it could use to cheat the benchmark.

At one point, the models strung together several attack vectors, including using stolen credentials and zero-day vulnerabilities, to find a remote code execution path on the Hugging Face servers.

As part of incident response efforts, OpenAI said it’s implementing strict controls in infrastructure configuration, responsibly disclosed the zero-day flaw in the third-party software, adding Hugging Face to its trusted access program to improve their defenses, and incorporating stronger guardrails around future training and evaluations.

“This incident points to the need to further strengthen our model’s alignment, cyber protections during evaluation time, and monitoring during internal testing,” OpenAI said.

The development comes as the company also revealed that long-running models, while taking on complex, open-ended problems, can open the door to taking unwanted actions, such as finding weaknesses in the operational environment, in pursuit of their objective through repeated attempts over extended periods of time.

“It also shows how a model that operates effectively over long time horizons can learn the blind spots of an approval system and work around it to achieve its goals,” OpenAI said. “Long-horizon safety requires not only asking ‘is this action allowed?’ but also ‘what outcome is this sequence of actions working toward?.'”



Source

Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Previous ArticleBugs in Hugging Face Diffusers Bypass Custom Code Safeguard
Next Article AI-Assisted Bug Hunt Uncovers Linux Kernel 0-Day in net/sched
Team-CWD
  • Website

Related Posts

News

AI-Assisted Bug Hunt Uncovers Linux Kernel 0-Day in net/sched

July 28, 2026
News

Phishing Dominates as Initial Entry Method for Cyber-Attacks

July 28, 2026
News

Police Dismantle Kratos Phishing Kit Built to Steal Microsoft 365 Sessions and Bypass MFA

July 28, 2026
Add A Comment
Leave A Reply Cancel Reply

Latest News

North Korean Hackers Turn JSON Services into Covert Malware Delivery Channels

November 24, 202523 Views

macOS Stealer Campaign Uses “Cracked” App Lures to Bypass Apple Securi

September 7, 202517 Views

North Korean Hackers Target Crypto Firms with ClickFix and Zoom Lures

April 29, 202610 Views

Why SOC Burnout Can Be Avoided: Practical Steps

November 14, 20259 Views

Cyber M&A Roundup: Cyber Giants Strengthen AI Security Offerings

December 1, 20258 Views
Stay In Touch
  • Facebook
  • YouTube
  • TikTok
  • WhatsApp
  • Twitter
  • Instagram
Most Popular

North Korean Hackers Turn JSON Services into Covert Malware Delivery Channels

November 24, 202523 Views

macOS Stealer Campaign Uses “Cracked” App Lures to Bypass Apple Securi

September 7, 202517 Views

North Korean Hackers Target Crypto Firms with ClickFix and Zoom Lures

April 29, 202610 Views
Our Picks

Beware of Winter Olympics scams and other cyberthreats

February 2, 2026

Fixing trivial passwords is as easy as 123456

May 7, 2026

Why you should never pay to get paid

September 15, 2025

Subscribe to Updates

Get the latest news from cyberwiredaily.com

Facebook X (Twitter) Instagram Pinterest
  • Home
  • Contact
  • Privacy Policy
  • Terms of Use
  • California Consumer Privacy Act (CCPA)
© 2026 All rights reserved.

Type above and press Enter to search. Press Esc to cancel.