# Attackers Exploit Exposed AI Endpoints to Launch Autonomous Offensive Operations


Organizations deploying self-hosted AI infrastructure face a critical blind spot: exposed inference endpoints are becoming targets for threat actors seeking free computational resources to power their attacks. New research from Zenity reveals three distinct campaigns between March and May 2026 that weaponized publicly accessible AI endpoints—including Ollama and LiteLLM deployments—to drive penetration testing frameworks and adversarial agents without requiring any authentication or system compromise.


## The Threat: AI Infrastructure as a Target Surface


The attack vector is deceptively simple: threat actors are identifying exposed artificial intelligence inference endpoints and repurposing them as compute infrastructure for offensive operations. Unlike traditional infrastructure attacks that require exploiting vulnerabilities or bypassing security controls, this approach leverages misconfigured deployments running on default ports with no authentication requirements.


Key findings from Zenity researchers:


  • Three distinct threat campaigns targeted honeypot AI infrastructure between March and May 2026
  • Attackers used exposed Ollama and LiteLLM endpoints as backends for their malicious agents
  • No exploitation or special authentication was required—only knowledge of the endpoint's network location
  • The attack chain involved minimal reconnaissance: attackers sent a small "hello" probe to confirm accessibility, then submitted a full attack payload

  • This represents a fundamental shift in how attackers view infrastructure. Rather than seeking to compromise systems for data or control, threat actors are increasingly treating exposed services as commodity resources—compute on tap, available without credential management or complex exploitation chains.


    ## Technical Details: How the Attacks Worked


    ### Vulnerable Endpoints


    The researchers identified several specific endpoints commonly exposed on default ports:


    | Service | Endpoint | Default Port |

    |---------|----------|--------------|

    | Ollama | /api/generate | 11434 |

    | Ollama | /api/chat | 11434 |

    | LiteLLM | /v1/responses | 4000 |


    These endpoints expose inference capabilities—the ability to send prompts to language models and receive responses. When deployed without authentication in network-accessible locations, they become open conduits for attackers to inject their own prompts and directives.


    ### Attack Methodology


    The attack pattern discovered by Zenity reveals a consistent operational sequence:


    1. Reconnaissance: Attackers scan for exposed endpoints using network reconnaissance tools or public IP enumeration

    2. Validation: A minimal probe ("hello" message) confirms the endpoint responds without requiring authentication

    3. Payload Injection: The attacker configures a malicious agent with the exposed endpoint as its model backend

    4. Execution: The entire attack payload—including system prompts, tool definitions, and malicious instructions—rides within the request body


    The attacker essentially outsources their AI agent's "brain" to the victim's infrastructure. Rather than running a language model locally, they use the exposed endpoint to power their attack framework, masking their own compute footprint and leveraging the victim's resources.


    ## Real-World Attack Campaigns


    ### Campaign 1: Strix Penetration Testing Framework


    One operator weaponized the Strix autonomous penetration testing framework against an unidentified French auction house. The attack involved:


  • A single IP source using a LiteLLM client to communicate with the exposed endpoint
  • A 140,000-character prompt containing the full Strix agent configuration
  • Explicit instructions to the agent to operate autonomously without permission requests
  • Directives to suppress any identifying markers or logs that might reveal "Strix" by name

  • This campaign represents a full-scale offensive operation powered entirely through infrastructure the victim had deployed—without the victim's knowledge or consent.


    ### Campaign 2: HexStrike AI Framework


    A second operator deployed HexStrike AI, another autonomous penetration testing system, through exposed endpoints. Like Strix, this framework was designed to operate without requiring interactive human oversight or approval.


    ### Campaign 3: Persona-Based Codex Agent


    The third campaign involved an OpenAI Codex-based agent engineered with a custom persona designed to suppress safety refusals. This agent was specifically built to assist in web reverse-engineering and bypassing security measures that might normally constrain an AI system.


    This variant is particularly concerning because it demonstrates adversaries actively working to defeat AI safety measures—creating personas and prompt engineering techniques specifically designed to make AI systems perform unintended malicious actions.


    ## Implications: The Scope and Risk


    ### Who Is Exposed?


    Any organization running self-hosted AI infrastructure without proper access controls faces exposure. This includes:


  • Development and research teams using Ollama for local model experimentation
  • DevOps and platform teams evaluating LiteLLM for API abstraction
  • AI engineering departments testing open-source models
  • Cloud-adjacent deployments where AI services are exposed on internal networks but lack authentication

  • The problem extends beyond intentional public exposure. Cloud misconfigurations, VPN tunnels, and compromised internal networks can all render these endpoints accessible to attackers.


    ### Resource Hijacking at Scale


    When attackers repurpose a victim's AI endpoint, they gain:


  • Free compute: Language model inference is computationally expensive; using a victim's endpoint eliminates that cost
  • Obfuscated operations: Their attack traffic originates from the victim's infrastructure, blending with legitimate usage
  • Scalability: Multiple attackers can target the same endpoints, running simultaneous operations
  • Persistence without compromise: There's no need to maintain backdoor access or persistent agent—each attack is a single request

  • For large organizations running expensive models on high-end GPUs, a single compromised endpoint could power weeks of attacker operations before detection.


    ## Recommendations: Securing AI Infrastructure


    ### Immediate Actions


    Authentication and Access Control:

  • Deploy authentication (API keys, OAuth, or mutual TLS) on all inference endpoints
  • Use strong, rotating credentials distinct from development passwords
  • Implement network segmentation to restrict endpoint access to approved internal clients only

  • Port and Exposure Management:

  • Never run AI endpoints on public IP addresses on default ports
  • Use reverse proxies with authentication in front of inference services
  • For local development, bind services to 127.0.0.1, not 0.0.0.0
  • Audit all listening ports for unauthenticated services

  • ### Ongoing Monitoring


  • Deploy network IDS/IPS rules to detect anomalous inference requests (unusual payload sizes, suspicious prompt patterns)
  • Monitor API gateway logs for requests sent to inference endpoints from unexpected sources
  • Alert on sudden spikes in endpoint usage that might indicate hijacking
  • Maintain inventory of all AI services deployed and their access controls

  • ### Architectural Improvements


  • Implement request signing to verify that endpoint calls originate from authorized clients
  • Use rate limiting to constrain how many inferences a single client can consume
  • Deploy anomaly detection on model outputs to identify when an endpoint is being used for adversarial purposes
  • Consider using containerized deployments with ephemeral, read-only filesystems for added isolation

  • ## HackWire Analysis


    This campaign exposes a critical gap in how organizations think about AI security—most threat models focus on data theft or model extraction, but few anticipate that their AI infrastructure becomes a commodity resource for attackers. What makes these campaigns significant is not the novelty of each attack tool (Strix, HexStrike, and Codex agents existed before), but rather the convergence of two trends: widespread adoption of self-hosted AI without security hardening, and attackers recognizing that outsourcing their computational needs to victim infrastructure is economically rational and operationally efficient.


    The pattern here mirrors cloud storage bucket exposures and database breaches of 2015—a necessary capability deployed without authentication controls because teams assumed internal deployment meant internal security. The difference is speed: attackers discovered and exploited this window within months of mainstream AI infrastructure adoption.


    Most organizations will miss this risk until they detect unusual API billing spikes or find that their endpoints have been weaponized. The lack of authentication makes the attack essentially invisible—there's no failed login attempt, no privilege escalation, no suspicious process launching. An attacker's prompt injection looks identical to legitimate model usage from a network-level perspective.


    The actionable lesson is stark: any service that performs computation should be treated as a potential attack surface. AI endpoints are not exempt. Teams deploying Ollama, LiteLLM, vLLM, or similar services need to treat authentication and network access control as mandatory, not optional. Given the cost of inference, a compromised endpoint doesn't just pose a confidentiality risk—it represents a direct financial loss and operational abuse of your infrastructure.


    For organizations already running these services in production, assume some may have been probed or exploited without detection. Audit access logs (if available), check for anomalous usage patterns, and implement authentication immediately. The fix is straightforward; detecting exploitation retroactively is much harder. — HackWire Editorial


    ## Related Coverage


  • Read more in our [Vulnerabilities](https://www.hackwire.news/category/vulnerabilities) coverage
  • Cross-reference with [Breaches](https://www.hackwire.news/category/breaches) and [Malware](https://www.hackwire.news/category/malware)
  • Stay current via the [HackWire homepage](https://www.hackwire.news/)