# Attackers Exploit Exposed AI Endpoints to Launch Autonomous Offensive Operations
Organizations deploying self-hosted AI infrastructure face a critical blind spot: exposed inference endpoints are becoming targets for threat actors seeking free computational resources to power their attacks. New research from Zenity reveals three distinct campaigns between March and May 2026 that weaponized publicly accessible AI endpoints—including Ollama and LiteLLM deployments—to drive penetration testing frameworks and adversarial agents without requiring any authentication or system compromise.
## The Threat: AI Infrastructure as a Target Surface
The attack vector is deceptively simple: threat actors are identifying exposed artificial intelligence inference endpoints and repurposing them as compute infrastructure for offensive operations. Unlike traditional infrastructure attacks that require exploiting vulnerabilities or bypassing security controls, this approach leverages misconfigured deployments running on default ports with no authentication requirements.
Key findings from Zenity researchers:
This represents a fundamental shift in how attackers view infrastructure. Rather than seeking to compromise systems for data or control, threat actors are increasingly treating exposed services as commodity resources—compute on tap, available without credential management or complex exploitation chains.
## Technical Details: How the Attacks Worked
### Vulnerable Endpoints
The researchers identified several specific endpoints commonly exposed on default ports:
| Service | Endpoint | Default Port |
|---------|----------|--------------|
| Ollama | /api/generate | 11434 |
| Ollama | /api/chat | 11434 |
| LiteLLM | /v1/responses | 4000 |
These endpoints expose inference capabilities—the ability to send prompts to language models and receive responses. When deployed without authentication in network-accessible locations, they become open conduits for attackers to inject their own prompts and directives.
### Attack Methodology
The attack pattern discovered by Zenity reveals a consistent operational sequence:
1. Reconnaissance: Attackers scan for exposed endpoints using network reconnaissance tools or public IP enumeration
2. Validation: A minimal probe ("hello" message) confirms the endpoint responds without requiring authentication
3. Payload Injection: The attacker configures a malicious agent with the exposed endpoint as its model backend
4. Execution: The entire attack payload—including system prompts, tool definitions, and malicious instructions—rides within the request body
The attacker essentially outsources their AI agent's "brain" to the victim's infrastructure. Rather than running a language model locally, they use the exposed endpoint to power their attack framework, masking their own compute footprint and leveraging the victim's resources.
## Real-World Attack Campaigns
### Campaign 1: Strix Penetration Testing Framework
One operator weaponized the Strix autonomous penetration testing framework against an unidentified French auction house. The attack involved:
This campaign represents a full-scale offensive operation powered entirely through infrastructure the victim had deployed—without the victim's knowledge or consent.
### Campaign 2: HexStrike AI Framework
A second operator deployed HexStrike AI, another autonomous penetration testing system, through exposed endpoints. Like Strix, this framework was designed to operate without requiring interactive human oversight or approval.
### Campaign 3: Persona-Based Codex Agent
The third campaign involved an OpenAI Codex-based agent engineered with a custom persona designed to suppress safety refusals. This agent was specifically built to assist in web reverse-engineering and bypassing security measures that might normally constrain an AI system.
This variant is particularly concerning because it demonstrates adversaries actively working to defeat AI safety measures—creating personas and prompt engineering techniques specifically designed to make AI systems perform unintended malicious actions.
## Implications: The Scope and Risk
### Who Is Exposed?
Any organization running self-hosted AI infrastructure without proper access controls faces exposure. This includes:
The problem extends beyond intentional public exposure. Cloud misconfigurations, VPN tunnels, and compromised internal networks can all render these endpoints accessible to attackers.
### Resource Hijacking at Scale
When attackers repurpose a victim's AI endpoint, they gain:
For large organizations running expensive models on high-end GPUs, a single compromised endpoint could power weeks of attacker operations before detection.
## Recommendations: Securing AI Infrastructure
### Immediate Actions
Authentication and Access Control:
Port and Exposure Management:
127.0.0.1, not 0.0.0.0### Ongoing Monitoring
### Architectural Improvements
## HackWire Analysis
This campaign exposes a critical gap in how organizations think about AI security—most threat models focus on data theft or model extraction, but few anticipate that their AI infrastructure becomes a commodity resource for attackers. What makes these campaigns significant is not the novelty of each attack tool (Strix, HexStrike, and Codex agents existed before), but rather the convergence of two trends: widespread adoption of self-hosted AI without security hardening, and attackers recognizing that outsourcing their computational needs to victim infrastructure is economically rational and operationally efficient.
The pattern here mirrors cloud storage bucket exposures and database breaches of 2015—a necessary capability deployed without authentication controls because teams assumed internal deployment meant internal security. The difference is speed: attackers discovered and exploited this window within months of mainstream AI infrastructure adoption.
Most organizations will miss this risk until they detect unusual API billing spikes or find that their endpoints have been weaponized. The lack of authentication makes the attack essentially invisible—there's no failed login attempt, no privilege escalation, no suspicious process launching. An attacker's prompt injection looks identical to legitimate model usage from a network-level perspective.
The actionable lesson is stark: any service that performs computation should be treated as a potential attack surface. AI endpoints are not exempt. Teams deploying Ollama, LiteLLM, vLLM, or similar services need to treat authentication and network access control as mandatory, not optional. Given the cost of inference, a compromised endpoint doesn't just pose a confidentiality risk—it represents a direct financial loss and operational abuse of your infrastructure.
For organizations already running these services in production, assume some may have been probed or exploited without detection. Audit access logs (if available), check for anomalous usage patterns, and implement authentication immediately. The fix is straightforward; detecting exploitation retroactively is much harder. — HackWire Editorial
## Related Coverage