# The AI Infrastructure Layer Nobody Secured: CISA Confirms Ray Exploitation Is Real


There's a particular breed of vulnerability that's more dangerous than the zero-day kind — the one where the vendor tells you it's not a vulnerability at all. CISA's addition of a critical Ray framework flaw to its Known Exploited Vulnerabilities catalog this week closes the loop on a dispute that's been simmering for over a year: researchers said exploitation was trivial, Anyscale said defenders were responsible for securing their own deployments, and now production systems are compromised.


Whoever was right about the semantics, defenders are the ones paying.


## What Ray Actually Runs In Your Stack


Ray isn't niche. With over 33,000 GitHub stars and adoption from companies including OpenAI, Spotify, Uber, and a constellation of AI-first startups, it has become foundational infrastructure for distributed machine learning workloads. If you're training models at scale, running reinforcement learning pipelines, or doing distributed hyperparameter tuning, there's a reasonable chance Ray is somewhere in your stack — possibly in a cluster that your security team hasn't fully mapped.


That's the first problem. Ray clusters are often provisioned by data science and ML engineering teams, not by the people whose job it is to think about network exposure. The result: compute-dense systems with privileged access to training data, model weights, and cloud credentials, running software that was designed for performance rather than adversarial resistance.


The Ray dashboard, which exposes a web interface on port 8265 by default, requires no authentication. That's not a bug Anyscale accidentally shipped — it's a documented design choice. The Jobs API, accessible through that same unauthenticated interface, allows users to submit Python code for execution on the cluster. Combine those two facts and you have, functionally, an unauthenticated remote code execution primitive baked into the default installation.


## The Dispute That Predicted This Outcome


In late 2023, researchers at Oligo Security documented what they called ShadowRay — a campaign actively exploiting Ray's unauthenticated dashboard against production AI workloads. Their findings were alarming: stolen model weights, cryptominers running on GPU clusters, and in some cases, cloud credentials harvested from compromised nodes.


Anyscale's response was to decline a CVE assignment, arguing that the lack of authentication in the dashboard was not a security flaw but a known design decision, and that users bear responsibility for network-level controls. From a narrow product perspective, that position has a certain logic — you shouldn't expose an internal orchestration interface to the public internet, and Ray's documentation does note that clusters should be deployed inside private networks.


What that position misses is how software actually gets deployed. Misconfiguration is not a user failure to be disclaimed away — it's a predictable outcome that vendors have a responsibility to design against. Default-open is default-exploited. Every security team that has dealt with exposed Kubernetes dashboards, unsecured Elasticsearch instances, or unauthenticated Jupyter notebooks has lived this lesson already. Ray shipped the same pattern into the AI era.


CISA's KEV listing means the U.S. government has confirmed active exploitation — not theoretical, not proof-of-concept. Real attackers, real victims, right now.


## Who Should Be Worried


The exposure surface here is broader than a typical enterprise software flaw. Consider the deployment footprint:


Research institutions and universities running Ray for academic ML work, often with limited security operations capacity and networks that connect to broader research infrastructure.


AI startups that provisioned Ray clusters during rapid scaling and didn't circle back to harden them — particularly those running on cloud providers where a misconfigured security group can expose port 8265 to the internet.


Enterprises with shadow ML infrastructure — clusters stood up by data science teams that bypassed standard IT procurement and never made it onto the asset inventory. These are the systems that don't get patched because nobody knows they exist.


The threat isn't just cryptomining, though that's the most common payload researchers have observed. An attacker with code execution inside a Ray cluster has access to whatever that cluster can reach: training data buckets, model checkpoints, API keys stored in environment variables, and potentially lateral movement paths into adjacent cloud infrastructure. In the context of AI development pipelines, that means intellectual property — the trained models that represent months of compute spend — is directly at risk.


## What to Actually Do


If you run Ray, the immediate actions are concrete:


  • Audit exposure: Scan for port 8265 (dashboard) and 10001 (client connection) reachable from outside your intended network perimeter. Shodan queries will show you what's publicly visible.
  • Enforce network-level controls: Ray clusters should not be directly reachable from the internet. If yours are, that's the remediation, not a patch.
  • Check for indicators of compromise: Oligo's ShadowRay research documented specific IOCs; review your cluster logs for unexpected job submissions, unusual network connections from worker nodes, and signs of cryptominer execution.
  • Inventory shadow deployments: Work with data science and ML teams to identify all Ray installations, including those provisioned outside standard IT channels.
  • Upgrade if you can: Newer versions of Ray have made progress on authentication, though the default configuration remains permissive. Check your version and review the current security configuration documentation.

  • ---


    ## HackWire Analysis


    This incident is a case study in what happens when the AI infrastructure build-out outruns the security function — and it's going to happen again, repeatedly, with different tools.


    The broader pattern here isn't specific to Ray. The last five years have seen an explosion of open-source tooling designed to make distributed computing accessible to people who aren't systems engineers: Ray for ML orchestration, Jupyter for interactive compute, Airflow and Prefect for data pipelines, MLflow for experiment tracking. These tools get adopted fast because they're genuinely good at what they do. Security is an afterthought — both in the products themselves and in how organizations deploy them.


    The Anyscale dispute deserves more scrutiny than it's gotten. When a vendor declines a CVE by arguing that unauthenticated RCE is "by design," that should be a loud signal to defenders, not a resolution. The same argument has been made about exposed MongoDB instances, the Jenkins script console, and Hadoop's management interfaces. Every time, the outcome was the same: mass exploitation of predictably misconfigured deployments.


    What's different now is the stakes. Ray doesn't run web applications — it runs AI training pipelines. The data inside those clusters is often the most competitively sensitive data a company holds: proprietary training datasets, fine-tuned model weights, evaluation results. Losing a Ray cluster isn't like losing a web server. It may be equivalent to losing your entire AI investment.


    CISA's KEV listing gives organizations a compliance hook to justify emergency remediation. But the real takeaway is structural: any AI/ML infrastructure tool that ships with an unauthenticated management interface needs to be treated as high-risk until proven otherwise, regardless of vendor reassurances. The threat actors have already figured out where the data is. The question is whether defenders catch up before the next cluster goes down.


    — HackWire Editorial


    ---


    ## Related Coverage


  • Read more in our [Vulnerabilities](https://www.hackwire.news/category/vulnerabilities) coverage
  • Cross-reference with [Breaches](https://www.hackwire.news/category/breaches) and [Malware](https://www.hackwire.news/category/malware)
  • Stay current via the [HackWire homepage](https://www.hackwire.news/)