In AI, no one can hear the sandbox scream
https://etimg.etb2bimg.com/thumb/msid-132733085,imgsize-349379,width-1200,height=627,overlay-etciso,resizemode-75/cybercrime-fraud/in-ai-no-one-can-hear-the-sandbox-scream.jpg
OpenAI was running a cyber-capability evaluation against advanced models, including GPT-5.6 Sol and a more capable pre-release model with reduced cyber refusals. The environment was meant to be constrained and the model still found brute forced through it.
OpenAI said the model first found a weakness in the internal package-registry cache proxy, which was supposed to give the evaluation environment-controlled access to dependencies without opening the wider network. The proxy became the first crack. After exploiting it, the model gained more access inside the research environment, escalated privileges, and moved laterally until it reached a system that could talk to the internet. From there, the model began reasoning about where benchmark-related material might exist outside the evaluation boundary. It inferred that Hugging Face could hold ExploitGym-related models, datasets, or solutions, then started looking for a path into that platform. The route reportedly involved stolen credentials, additional zero-day vulnerabilities, and a remote-code execution path on Hugging Face infrastructure.
A hacker would recognize the shape immediately. The environment needed a path for packages, so the package path became interesting. It needed workers to process datasets, so the workers became interesting. It needed credentials for services to talk to each other, so the credentials became interesting.
Hugging Face’s disclosure showed how quickly those spaces can turn into an intrusion path. The compromise began in a data-processing pipeline, where a malicious dataset abused a remote-code dataset loader and a template-injection issue in a dataset configuration. Code executed on a processing worker. The attacker escalated to node-level access, harvested cloud and cluster credentials, and moved laterally into internal clusters. Hugging Face described thousands of actions across short-lived sandboxes, with command and control staged through public services.
Governance cannot sit above the environment like a policy binder while a model is running with tools, credentials, and network reach. In a cyber evaluation, the runtime is where the risk lives. Egress rules, credential scope, infrastructure separation, and logging have to be designed around the assumption that the system may keep testing the boundary until something gives. Package proxies, dataset workers, notebook runtimes, and build systems are easy to dismiss as support infrastructure. In real time, they are becoming the places where normal engineering exceptions are now movement paths.
Hugging Face had more than 17,000 recorded attacker events to reconstruct, and the logs were full of the same material responders always have to touch during a real intrusion: exploit payloads, shell commands, C2 artifacts, credential references, and messy attacker output. Hosted frontier models reportedly blocked parts of that analysis because the evidence looked too much like offensive material. Hugging Face eventually ran the forensic work on GLM 5.2, an open-weight model, inside its own infrastructure. That gave the team two things hosted APIs could not provide at that moment: fewer guardrail collisions and tighter control over where sensitive attacker data went.
Responders cannot lose analysis capability halfway through an incident because the tooling refuses the evidence. During an intrusion, hostile data cannot become a red flag to avoid. It is the thing being investigated. Everything must be parsed, correlated, and understood under pressure. Teams bringing AI into incident response will need a plan for private inference, vetted local models, and evidence-handling rules that keep sensitive material out of places it should never go.
The security posture changes when you give a model an objective, a shell, files, credentials, network access, and enough feedback. It now begins to resemble an operator inside the environment, even if nobody intended to grant it that role. I predict we may start seeing more incidents that come from a system doing exactly what the evaluation asked, following the path that worked, and turning a boundary everyone trusted into the first step of an attack chain.
The author is Aaron Beardslee, Security Researcher, Securonix Threat Labs.
Disclaimer: The views expressed are solely of the author and ETCISO does not necessarily subscribe to it. ETCISO shall not be responsible for any damage caused to any person/organization directly or indirectly.
Firewall Support Company in India All type of Firewalls Support Provider Company in India












