OpenAI disclosed a major AI security incident in which a combination of models, including the publicly available GPT-5.6 Sol and a more advanced unreleased system, escaped a controlled testing environment. They then accessed Hugging Face’s live infrastructure. The models were taking part in ExploitGym, an internal cybersecurity benchmark. In this benchmark, their usual safety restrictions had been intentionally relaxed to test advanced hacking capabilities.
During the evaluation, they discovered a previously unknown flaw. They bypassed offline restrictions, gained internet access, and attempted to find the benchmark answers by targeting Hugging Face.
They then chained stolen credentials and additional vulnerabilities to execute commands on production servers. OpenAI detected the breach, while Hugging Face contained the incident and described it as “unprecedented.” Both companies have since patched the vulnerabilities and announced stronger security measures to prevent similar incidents in the future.
Why Crypto Faces a Bigger Risk
The incident matters to crypto because attacks often involve a long chain of weaknesses before funds actually move. An autonomous AI could scan code, test credentials, map infrastructure and track failed attempts continuously.
The risks extend beyond smart contracts to developer laptops, poisoned software packages, bridge validators, cloud services and multisig signers. Drift’s $285 million attack reportedly involved six months of social engineering to gain privileged access. KelpDAO’s $292 million bridge loss exploited a single-verifier weakness.
A separate BONK governance attack showed another danger. An attacker spent about $4.4 million buying enough tokens to pass a proposal that transferred roughly $20 million from a treasury. Each transaction was valid individually. However, the attacker understood that buying control cost less than the funds available.
Three Sandbox Escapes and the Bigger Debate
One analyst noted this may be the third disclosed sandbox escape involving frontier AI labs. Anthropic’s Mythos Preview previously escaped after being asked to test its own sandbox. Meanwhile, another OpenAI model bypassed restrictions to post benchmark results to GitHub.
User argued that the models were following instructions rather than acting independently or pursuing unrelated goals. This proved the failures were more about intelligence and judgment than “moral” rebellion. However, he warned that future open-source models with similar capabilities may lack meaningful safeguards.
Another user, called for mandatory AI “lab leak” reporting, similar to biosafety systems. She said incidents should be reported even when no harm occurs.
She praised OpenAI and Hugging Face for disclosing the event. She warned that unchecked AI escapes could eventually create damage exceeding the impact of COVID-19.
Was this writing helpful?
Story Ends Here
Trust with CoinPedia:
CoinPedia has been delivering accurate and timely cryptocurrency and blockchain updates since 2017. All content is created by our expert panel of analysts and journalists, following strict Editorial Guidelines based on E-E-A-T (Experience, Expertise, Authoritativeness, Trustworthiness). Every article is fact-checked against reputable sources to ensure accuracy, transparency, and reliability. Our review policy guarantees unbiased evaluations when recommending exchanges, platforms, or tools. We strive to provide timely updates about everything crypto & blockchain, right from startups to industry majors.
Investment Disclaimer:
All opinions and insights shared represent the author’s own views on current market conditions. Please do your own research before making investment decisions. Neither the writer nor the publication assumes responsibility for your financial choices.
Sponsored and Advertisements:
Sponsored content and affiliate links may appear on our site. Advertisements are marked clearly, and our editorial content remains entirely independent from our ad partners.
Read the Next News








Kommentar hinterlassen