# ChatGPT's Sandbox Became a C2 Node. Here's How.


The name of the talk alone should give security teams pause: "A Billion-User Blast Radius." Simcha Kosman, a senior researcher at Palo Alto Networks, spent roughly five minutes at Black Hat USA 2026 dismantling the assumption that ChatGPT's code execution sandbox is actually isolated before OpenAI could even finish telling the room it had already fixed the problem.


That compressed timeline — exploit demonstrated Tuesday, OpenAI says the vulnerable component was removed before Tuesday — is itself a story worth examining. But the more important story is what the attack chain reveals about where AI security is genuinely broken.


## The Assumption That Wasn't


ChatGPT's sandbox exists for an obvious reason: when you ask the model to run Python, write a file, or manipulate data, something has to execute that code in an environment that can't touch your system or anyone else's. The isolation model is the entire security promise. Without it, code execution in an AI assistant is just code execution.


Kosman's starting point was a foundational claim — "private chats should stay private" — that turned out to be false under a specific set of attacker-controlled conditions. The attack chain he demonstrated worked roughly like this:


1. Initial code injection. An attacker tricks a victim's ChatGPT session into running attacker-controlled code inside the sandbox. The precise vector wasn't fully detailed in coverage available, but the technique involves bypassing the LLM supervisor — the layer supposed to decide whether a given action is safe to execute.


2. Persistence and reasoning influence. Once code runs, it doesn't just execute and exit. Kosman showed it could influence future reasoning within that session — meaning the attacker doesn't just get a one-shot execution, they get an ongoing foothold that shapes subsequent model outputs.


3. Shared backend abuse. The most structurally significant part: the attack abused a shared backend component to establish full command and control. Data moved from the victim's sandbox to the attacker's own sandbox. That's not a sandbox escape in the traditional sense — it's something arguably worse. The container boundary held; the shared layer beneath it didn't.


The result is what Kosman framed as C2-style influence over a victim's ChatGPT session. The attacker doesn't need to escape the sandbox to win. They just need to own the communication path between sandboxes.


## OpenAI's Response Is Half the Story


OpenAI confirmed to Dark Reading that it was aware of the research before the presentation and that the specific system component involved was removed before Kosman took the stage. The company also framed the finding carefully: this was not, in their characterization, an escape from the security sandbox or a path to unrestricted access to other customer accounts.


Both of those statements can be technically accurate while still missing the point.


Removing a vulnerable component before a public talk is responsible behavior, and coordinated disclosure here appears to have worked as intended. Credit where it's due. But the question that OpenAI's statement sidesteps is: how long did that shared backend configuration exist, and what was theoretically accessible during that window? A proof of concept demonstrated at Black Hat implies someone had enough time to build a reliable exploit. The patch timeline matters.


The framing around "not an escape from the security sandbox" also deserves scrutiny. If an attacker can establish C2 — bidirectional data movement between their environment and a victim's — through a shared backend, the sandbox's container-level isolation becomes largely irrelevant to the actual threat. Security controls only matter if they stop the attack. A sandbox that prevents process escape but allows data exfiltration via a shared layer has not actually contained the threat.


## The Kill Chain Is Getting Longer — and More Reliable


What makes this research significant beyond the specific ChatGPT finding is the maturation of the underlying attack pattern.


A year ago, AI sandbox attacks were largely theoretical or required implausible user interaction. The attack chains being demonstrated now at Black Hat and DEF CON are increasingly end-to-end: prompt injection or malicious file upload → LLM supervisor bypass → code execution → persistence → lateral movement via shared infrastructure. That's a kill chain. That's not a research curiosity.


The LLM supervisor bypass is the piece that deserves more attention than it's getting. Supervisory layers — the model checking whether another model or code execution should be permitted — are the current architectural response to prompt injection at scale. If those layers can be socially engineered or technically circumvented, the entire safety model built around them collapses. Kosman's research suggests the bypass is achievable; understanding how generally that holds across other AI platforms is an open question that someone will answer, probably also at a conference.


The shared backend angle connects to a well-understood problem in cloud security: multi-tenant isolation is hard, and the shared components are always where the interesting vulnerabilities live. Container escapes are glamorous, but database connection pools, message queues, and storage backends are where tenant data actually bleeds. AI platforms have all the same multi-tenant infrastructure as any SaaS product, plus the novel attack surface of model inference pipelines sitting on top.


---


## HackWire Analysis


The framing of this talk as a "billion-user blast radius" is not marketing hyperbole — it's a precise description of the threat model that AI platforms have created and that the security community is only beginning to take seriously.


Most enterprise security thinking about AI tools centers on data exfiltration through the model itself: what the LLM might repeat, hallucinate, or leak from its training data. That's a real risk category. But Kosman's research points at something different: the AI execution environment as an attack surface in its own right, one with shared infrastructure that connects millions of users' sessions through components they never see.


This is the multi-tenancy problem applied to AI, and it's not unique to OpenAI. Any platform that lets users run code, process files, or execute actions in a shared infrastructure inherits this class of vulnerability. Anthropic's Code Execution environment, Google's Code Interpreter in Gemini, and GitHub Copilot's integration with developer tools all operate on similar architectural assumptions. None of them have been subjected to the level of adversarial research that ChatGPT attracts simply by virtue of its user base.


The coordinated disclosure here was handled reasonably, but the security community should not read OpenAI's "we removed the component" response as "this class of attack is resolved." The underlying issue — shared backend components bridging what should be isolated session environments — is architectural. Patching one manifestation doesn't close the attack surface.


For defenders: if your organization uses ChatGPT Enterprise, Copilot, or any AI platform with code execution capabilities, your threat model now needs to include the platform's own infrastructure as an attack vector. Audit what data enters those sessions. Treat AI-generated code execution like you'd treat any third-party code execution environment — with network segmentation, data classification, and logging of what leaves the sandbox, not just what enters it.


The era of AI platforms getting a security free pass because they're "just a chat interface" is over. Black Hat 2026 made that clear.


— HackWire Editorial


---


## Related Coverage


  • Read more in our [Vulnerabilities](https://www.hackwire.news/category/vulnerabilities) coverage
  • Cross-reference with [Breaches](https://www.hackwire.news/category/breaches) and [Malware](https://www.hackwire.news/category/malware)
  • Stay current via the [HackWire homepage](https://www.hackwire.news/)