# Agentjacking: How Attackers Can Hijack AI Coding Agents to Execute Malicious Code on Developer Machines
Researchers have uncovered a critical new attack vector that weaponizes the very tools developers trust to write code faster: AI coding agents. By crafting malicious error events in Sentry—the widely-used error tracking platform—attackers can trick AI agents like Claude Code and Cursor into running arbitrary code with full developer privileges, potentially exposing credentials, environment variables, and private repositories.
The attack, dubbed Agentjacking by Tenet Security, exposes a fundamental architectural weakness at the intersection of error tracking services and Model Context Protocol (MCP) connections. In controlled testing, researchers achieved an 85% exploitation success rate against some of the most widely deployed AI coding assistants.
## The Threat: A New Attack Surface on Coding Agents
Traditional security perimeters—firewalls, VPNs, EDR systems, and Web Application Firewalls—are completely bypassed by Agentjacking. The attack requires no server compromise, no phishing, and no prior access to developer infrastructure. Instead, it exploits a subtle but dangerous assumption: that data returned by connected services through MCP is inherently trustworthy.
"The attacker never touches the victim's infrastructure," explained researchers Ron Bobrov, Barak Sternberg, and Nevo Poran. "The malicious instruction arrives disguised as a legitimate 'Resolution' inside an ordinary error. When a developer asks their AI agent to fix the Sentry issue, the agent reads the attacker's command as trusted guidance and runs it—with the developer's own privileges, on the developer's own machine."
The attack works because AI agents cannot distinguish between legitimate error resolution steps and injected malicious commands—they both appear as structured, formatted guidance from a trusted system.
## How Agentjacking Works: A Step-by-Step Attack Chain
The exploit follows a remarkably simple attack chain:
1. Find the DSN: The attacker identifies a target's Sentry Data Source Name (DSN)—a public, write-only credential embedded in websites, JavaScript bundles, and client applications. These credentials are intentionally public by design.
2. Craft the Malicious Payload: Using Sentry's public ingest endpoint, the attacker sends a specially formatted error event containing malicious code disguised as markdown formatting in the message and context fields.
3. Wait for the Agent Query: When a developer asks their AI coding agent to "fix unresolved Sentry issues" or similar requests, the agent queries Sentry's MCP server to retrieve error data.
4. Inject Trust: The Sentry MCP server returns the attacker's malicious event, rendering it with the same visual formatting and structure as legitimate Sentry diagnostic information.
5. Execute with Privilege: The AI agent interprets the markdown-injected "resolution" as trusted guidance and executes the embedded code—running with the developer's full system privileges.
Key Technical Details:
## Attack Surface: 2,388 Organizations at Risk
Tenet Security's research identified at least 2,388 organizations with publicly exposed Sentry DSNs that could be exploited. These are real credentials, embedded in production websites and applications.
The researchers tested the attack in a controlled manner against over 100 organizations, achieving an 85% exploitation success rate—meaning the attack worked on 85 of roughly 100 deployments tested.
The victims span every industry: finance, healthcare, e-commerce, SaaS, and critical infrastructure. Any organization using Sentry with an AI coding agent deployed among their developers is potentially exposed.
## What Can Attackers Steal?
A successful Agentjacking attack can expose:
~/.ssh/.aws/ or environmentThe injected code runs with the developer's full privileges, meaning attackers gain access to everything the developer can access—often including production databases, customer data, and internal systems.
## Background: The Trust Problem in Agent Architectures
This vulnerability emerges from a fundamental architectural decision in how AI agents are designed: implicit trust in external data sources connected via MCP.
Model Context Protocol was designed to let AI agents safely query external tools and data sources. When an agent queries Sentry, GitHub, or a database through MCP, the returned data is treated as a trustworthy system response. There is no cryptographic verification, no signature validation, and no sandboxing of returned content.
This design made sense for simple data retrieval. But when that data is formatted in ways that AI models interpret as executable commands or instructions, it becomes a pathway for arbitrary code execution.
Sentry's Response and Industry Implications
Sentry acknowledged the issue but stated it is "technically not defensible"—meaning the company does not believe it can meaningfully prevent the injection of malicious content into error events without breaking the core functionality of the platform. The company did activate a global content filter to block specific payload strings, but this represents only a temporary mitigation, not a fix.
This response raises a critical question: if a widely-used platform considers the vulnerability unfixable at the platform level, where should the responsibility lie?
## HackWire Analysis: The Agent Security Reckoning Has Arrived
Agentjacking represents a watershed moment for AI agent security. For months, developers and enterprises have raced to integrate AI coding assistants into their workflows, drawn by promises of 10x productivity gains and fewer bugs. Few organizations have genuinely thought through the security implications of giving an AI agent autonomous access to run code on their machines.
Why this matters now: AI agents are moving fast from research novelty to production reality. Thousands of development teams now delegate error resolution, code refactoring, and debugging tasks to agents. Each of those teams has implicitly accepted a trust model that treats agent-executed code as safe. Agentjacking shatters that assumption.
The pattern: This attack follows a clear trend we've seen in supply chain and API security: designers assume that one layer of their architecture (in this case, MCP connections) is inherently trustworthy because it comes from a "legitimate" source. Attackers find that layer, compromise its data, and bypass all downstream protections. We saw this with Codecov (CI/CD compromise), SolarWinds (software update compromise), and now Sentry (error data injection).
The hidden risk: Enterprises installing EDR, WAF, and IAM controls feel protected. Agentjacking proves they're not. Traditional security tools focus on blocking malicious actor behavior—suspicious network connections, unusual file access, known exploit signatures. But code executed by an AI agent on a developer's machine looks identical to legitimate developer work. EDR won't flag it. Firewalls won't block it. The code is being run *by the developer's own trusted tool*, so it bypasses every control designed to detect outsider threats.
Concrete next steps: Organizations should immediately inventory which developers have AI agents connected to Sentry via MCP, disable MCP for Sentry until a proper fix exists, and audit Sentry DSN exposure. More broadly, development teams need to establish a security principle: AI agents should never have unvetted access to external data sources that could influence code execution. Sandbox agent-executed code. Require explicit human review before running commands suggested by agents that came from external sources. Treat agent suggestions from external systems the same way you'd treat shell commands pasted from the internet.
The era of trusting agents to safely integrate with external systems without verification has ended. — *HackWire Editorial*
## Recommendations for Developers and Organizations
Immediate Actions:
Longer-Term Mitigations:
| Approach | Benefit | Trade-off |
|----------|---------|-----------|
| Code sandboxing | Agent-executed code runs in isolated environment | Performance overhead, limited access to developer tools |
| Human-in-the-loop review | Developers review all agent suggestions before execution | Slower development, defeats productivity gain |
| Cryptographic verification | MCP connections verify payload signatures | Requires changes to all agent vendors and external services |
| Network segmentation | Developer machines isolated from production credentials | Complex infrastructure, requires credential rotation |
For Agent Vendors:
AI coding agent providers should implement:
## Related Coverage