# Agentjacking: How Attackers Can Hijack AI Coding Agents to Execute Malicious Code on Developer Machines


Researchers have uncovered a critical new attack vector that weaponizes the very tools developers trust to write code faster: AI coding agents. By crafting malicious error events in Sentry—the widely-used error tracking platform—attackers can trick AI agents like Claude Code and Cursor into running arbitrary code with full developer privileges, potentially exposing credentials, environment variables, and private repositories.


The attack, dubbed Agentjacking by Tenet Security, exposes a fundamental architectural weakness at the intersection of error tracking services and Model Context Protocol (MCP) connections. In controlled testing, researchers achieved an 85% exploitation success rate against some of the most widely deployed AI coding assistants.


## The Threat: A New Attack Surface on Coding Agents


Traditional security perimeters—firewalls, VPNs, EDR systems, and Web Application Firewalls—are completely bypassed by Agentjacking. The attack requires no server compromise, no phishing, and no prior access to developer infrastructure. Instead, it exploits a subtle but dangerous assumption: that data returned by connected services through MCP is inherently trustworthy.


"The attacker never touches the victim's infrastructure," explained researchers Ron Bobrov, Barak Sternberg, and Nevo Poran. "The malicious instruction arrives disguised as a legitimate 'Resolution' inside an ordinary error. When a developer asks their AI agent to fix the Sentry issue, the agent reads the attacker's command as trusted guidance and runs it—with the developer's own privileges, on the developer's own machine."


The attack works because AI agents cannot distinguish between legitimate error resolution steps and injected malicious commands—they both appear as structured, formatted guidance from a trusted system.


## How Agentjacking Works: A Step-by-Step Attack Chain


The exploit follows a remarkably simple attack chain:


1. Find the DSN: The attacker identifies a target's Sentry Data Source Name (DSN)—a public, write-only credential embedded in websites, JavaScript bundles, and client applications. These credentials are intentionally public by design.


2. Craft the Malicious Payload: Using Sentry's public ingest endpoint, the attacker sends a specially formatted error event containing malicious code disguised as markdown formatting in the message and context fields.


3. Wait for the Agent Query: When a developer asks their AI coding agent to "fix unresolved Sentry issues" or similar requests, the agent queries Sentry's MCP server to retrieve error data.


4. Inject Trust: The Sentry MCP server returns the attacker's malicious event, rendering it with the same visual formatting and structure as legitimate Sentry diagnostic information.


5. Execute with Privilege: The AI agent interprets the markdown-injected "resolution" as trusted guidance and executes the embedded code—running with the developer's full system privileges.


Key Technical Details:


  • Markdown Injection: The attack exploits how AI models render structured markdown. Attackers craft payloads that appear identical to Sentry's standard system templates, making them indistinguishable from legitimate guidance.
  • Implicit Trust in MCP: The Model Context Protocol connections assume data from external services is safe. There is no cryptographic validation of payload origin or integrity.
  • Agent Behavior: Modern AI coding agents are designed to proactively fix issues when asked. They automatically execute commands within the resolution guidance without requiring explicit user confirmation.

  • ## Attack Surface: 2,388 Organizations at Risk


    Tenet Security's research identified at least 2,388 organizations with publicly exposed Sentry DSNs that could be exploited. These are real credentials, embedded in production websites and applications.


    The researchers tested the attack in a controlled manner against over 100 organizations, achieving an 85% exploitation success rate—meaning the attack worked on 85 of roughly 100 deployments tested.


    The victims span every industry: finance, healthcare, e-commerce, SaaS, and critical infrastructure. Any organization using Sentry with an AI coding agent deployed among their developers is potentially exposed.


    ## What Can Attackers Steal?


    A successful Agentjacking attack can expose:


  • Environment variables containing API keys, database credentials, and service tokens
  • Git credentials and SSH keys stored in ~/.ssh/
  • Private repository URLs and access tokens
  • AWS/cloud credentials in .aws/ or environment
  • Development secrets and configuration files
  • Developer identity and authentication tokens

  • The injected code runs with the developer's full privileges, meaning attackers gain access to everything the developer can access—often including production databases, customer data, and internal systems.


    ## Background: The Trust Problem in Agent Architectures


    This vulnerability emerges from a fundamental architectural decision in how AI agents are designed: implicit trust in external data sources connected via MCP.


    Model Context Protocol was designed to let AI agents safely query external tools and data sources. When an agent queries Sentry, GitHub, or a database through MCP, the returned data is treated as a trustworthy system response. There is no cryptographic verification, no signature validation, and no sandboxing of returned content.


    This design made sense for simple data retrieval. But when that data is formatted in ways that AI models interpret as executable commands or instructions, it becomes a pathway for arbitrary code execution.


    Sentry's Response and Industry Implications


    Sentry acknowledged the issue but stated it is "technically not defensible"—meaning the company does not believe it can meaningfully prevent the injection of malicious content into error events without breaking the core functionality of the platform. The company did activate a global content filter to block specific payload strings, but this represents only a temporary mitigation, not a fix.


    This response raises a critical question: if a widely-used platform considers the vulnerability unfixable at the platform level, where should the responsibility lie?


    ## HackWire Analysis: The Agent Security Reckoning Has Arrived


    Agentjacking represents a watershed moment for AI agent security. For months, developers and enterprises have raced to integrate AI coding assistants into their workflows, drawn by promises of 10x productivity gains and fewer bugs. Few organizations have genuinely thought through the security implications of giving an AI agent autonomous access to run code on their machines.


    Why this matters now: AI agents are moving fast from research novelty to production reality. Thousands of development teams now delegate error resolution, code refactoring, and debugging tasks to agents. Each of those teams has implicitly accepted a trust model that treats agent-executed code as safe. Agentjacking shatters that assumption.


    The pattern: This attack follows a clear trend we've seen in supply chain and API security: designers assume that one layer of their architecture (in this case, MCP connections) is inherently trustworthy because it comes from a "legitimate" source. Attackers find that layer, compromise its data, and bypass all downstream protections. We saw this with Codecov (CI/CD compromise), SolarWinds (software update compromise), and now Sentry (error data injection).


    The hidden risk: Enterprises installing EDR, WAF, and IAM controls feel protected. Agentjacking proves they're not. Traditional security tools focus on blocking malicious actor behavior—suspicious network connections, unusual file access, known exploit signatures. But code executed by an AI agent on a developer's machine looks identical to legitimate developer work. EDR won't flag it. Firewalls won't block it. The code is being run *by the developer's own trusted tool*, so it bypasses every control designed to detect outsider threats.


    Concrete next steps: Organizations should immediately inventory which developers have AI agents connected to Sentry via MCP, disable MCP for Sentry until a proper fix exists, and audit Sentry DSN exposure. More broadly, development teams need to establish a security principle: AI agents should never have unvetted access to external data sources that could influence code execution. Sandbox agent-executed code. Require explicit human review before running commands suggested by agents that came from external sources. Treat agent suggestions from external systems the same way you'd treat shell commands pasted from the internet.


    The era of trusting agents to safely integrate with external systems without verification has ended. — *HackWire Editorial*


    ## Recommendations for Developers and Organizations


    Immediate Actions:


  • Disable MCP for Sentry until Sentry or your agent vendor implements cryptographic payload verification or sandboxing
  • Audit DSN Exposure: Search your codebase, documentation, and error logs for exposed Sentry DSNs. If found, rotate the DSN immediately
  • Review Recent Agent Activity: Check your developers' system logs and AI agent histories for any unusual commands executed in recent weeks
  • Limit Agent Permissions: Configure AI agents to run in read-only or restricted modes when possible

  • Longer-Term Mitigations:


    | Approach | Benefit | Trade-off |

    |----------|---------|-----------|

    | Code sandboxing | Agent-executed code runs in isolated environment | Performance overhead, limited access to developer tools |

    | Human-in-the-loop review | Developers review all agent suggestions before execution | Slower development, defeats productivity gain |

    | Cryptographic verification | MCP connections verify payload signatures | Requires changes to all agent vendors and external services |

    | Network segmentation | Developer machines isolated from production credentials | Complex infrastructure, requires credential rotation |


    For Agent Vendors:


    AI coding agent providers should implement:

  • Cryptographic verification of MCP payloads
  • Sandboxed execution of code from external sources
  • Clear user warnings when executing commands derived from external data
  • Explicit user confirmation required before executing code from untrusted sources

  • ## Related Coverage


  • Read more in our [Vulnerabilities](https://www.hackwire.news/category/vulnerabilities) coverage
  • Cross-reference with [Breaches](https://www.hackwire.news/category/breaches) and [Malware](https://www.hackwire.news/category/malware)
  • Stay current via the [HackWire homepage](https://www.hackwire.news/)