# Agentjacking: The New Attack Surface Hiding in Plain Sight


## How Fake Bug Reports Can Hijack AI Coding Agents and Compromise Developer Systems


Researchers at Tenet Security have demonstrated a stark vulnerability in how modern AI coding agents operate: attackers can hijack these tools into executing arbitrary code simply by planting a fraudulent error report in a public bug tracking system. The technique, dubbed "agentjacking," exposes a fundamental blind spot in the current security architecture surrounding AI-powered development tools—and the implications extend far beyond a single developer's machine.


The vulnerability hinges on a deceptively simple attack: convince an AI agent that malicious instructions are legitimate system output, and the agent will comply. Because every step in the attack chain appears authorized and legitimate to traditional security controls, the compromise can unfold without triggering a single security alert.


## The Threat: A New Category of Supply Chain Risk


Agentjacking represents an entirely new category of attack surface in modern development environments. Unlike traditional vulnerabilities that exploit specific code flaws, agentjacking exploits the core design principle of AI agents: their ability to autonomously process external information and execute tasks based on that input.


In Tenet Security's proof-of-concept, researchers created a fake error report and submitted it to a Sentry project—a widely used error tracking and application monitoring platform. When developers using AI coding assistants such as Claude Code, Cursor, or Codex reviewed their applications, the AI agents automatically retrieved and processed the poisoned error data. In many cases, the agents executed the attacker-controlled code directly on the developer's machine with the full permissions of the logged-in user.


The potential consequences are severe:


  • Credential theft: Attackers could steal AWS keys, GitHub tokens, SSH private keys, and CI/CD pipeline credentials
  • Code repository compromise: Stolen credentials could grant access to private source code repositories
  • Cloud infrastructure takeover: Compromised cloud credentials could enable adversaries to spin up resources, modify deployments, or exfiltrate data
  • Supply chain poisoning: With access to CI/CD systems, attackers could inject malicious code into software dependencies distributed to thousands of downstream organizations
  • Lateral movement: Developer machine compromise could serve as a pivot point into broader organizational networks

  • The attack is particularly insidious because it exploits no zero-day vulnerabilities, uses no sophisticated social engineering, and doesn't require the attacker to breach any systems. A public error tracking service—the kind that thousands of organizations rely on daily—becomes the attack vector.


    ## Background and Context: The Scale of Exposure


    The vulnerability affects a potentially massive population of developers. Sentry, the platform used in Tenet's demonstration, claims over 200,000 organizations as customers, including major technology companies like GitHub, Disney, Anthropic, Atlassian, and countless others. While not all of these organizations deploy AI coding agents, the overlap is substantial and growing rapidly.


    The rise of AI coding assistants has been meteoric. Tools like Claude Code, GitHub Copilot, Cursor, and others have become standard fixtures in modern development workflows. Organizations view these tools as productivity multipliers, enabling developers to write code faster, refactor large codebases with fewer errors, and automate routine development tasks. However, this rapid adoption has outpaced security hardening.


    According to Barak Sternberg, CEO and co-founder of Tenet Security, the fundamental problem is architectural: "The AI agents you've deployed are now the soft attack path in, and your existing stack can't see it." Traditional security controls—identity and access management (IAM), endpoint detection and response (EDR), and network monitoring—all failed to detect or prevent Tenet's agentjacking attack because every action appeared legitimate and properly authorized.


    The agent read the poisoned error report, trusted it as genuine system output, and executed the embedded code with the developer's own access credentials. From the perspective of security tools, a user had intentionally asked their development tool to perform a task, and the tool complied.


    ## Technical Details: How Agentjacking Works


    The attack exploits a fundamental limitation in how AI agents distinguish between data and instructions. AI agents are designed to be responsive to input—to read error reports, documentation, code comments, and other external information, then act on that input intelligently. They have no reliable mechanism to determine whether a piece of data came from a trusted source or was deliberately crafted by an attacker.


    Here's how the attack unfolds:


    Step 1: Reconnaissance

    The attacker identifies a target organization using a publicly available development tool. They determine which error tracking service the organization uses and learn which AI coding agents are deployed.


    Step 2: Payload Crafting

    The attacker creates a fake error report that mimics the legitimate format of real error messages. The report contains hidden instructions—shell commands, Python scripts, or other executable payloads—designed to extract sensitive information or establish persistence on the developer's machine.


    Step 3: Injection

    The attacker submits the poisoned error report to the target's bug tracking system. Since many development teams accept error reports from public sources (to catch bugs reported by users), the fake report often makes it into the system without scrutiny.


    Step 4: Discovery

    When a developer using an AI coding agent reviews their application's error logs—a routine part of debugging—the AI agent retrieves the poisoned error data along with legitimate errors.


    Step 5: Execution

    The AI agent, interpreting the fake error report as genuine system output, extracts and executes the embedded commands. The developer's machine runs the attacker's code with full user privileges.


    The attack is nearly frictionless because it doesn't require the attacker to:

  • Exploit a specific vulnerability in the AI agent
  • Compromise any authentication mechanism
  • Trick the developer into manually running code
  • Bypass any access controls

  • ## Implications for Organizations


    This vulnerability exposes several critical gaps in current security practices:


    AI agents lack content authentication mechanisms. Unlike humans who might question an unusual error message, AI agents accept and act on data they retrieve from expected sources without verifying its authenticity or legitimacy.


    Traditional security controls are blind to AI agent activities. IAM systems see authorized access. EDR solutions see legitimate processes spawning. Network controls see expected connections. No single security tool recognizes the problem because no individual action looks suspicious.


    The attack surface is broader than most organizations realize. Every error tracking service, logging platform, documentation repository, and external data source that an AI agent consumes becomes a potential attack vector.


    Developer machines are a critical security boundary. Developers typically have elevated privileges and access to sensitive credentials. Compromising a single developer's machine can compromise an entire organization.


    ## Recommendations: Defensive Strategies


    Organizations using AI coding agents should implement a layered defense:


    1. Input Validation and Sanitization

  • Tool providers should implement cryptographic verification of data retrieved from external sources
  • Error tracking services should support digital signatures that developers can verify
  • AI agents should flag or quarantine suspicious input before executing any code

  • 2. Principle of Least Privilege

  • Configure development machines to prevent access to sensitive credentials
  • Use short-lived credentials and API tokens with minimal scope
  • Implement hardware security keys for critical services that can't be programmatically accessed

  • 3. Monitoring and Detection

  • Monitor AI agent activity for unusual patterns (unexpected code execution, credential access)
  • Log all commands executed by AI agents with full context
  • Alert on suspicious payloads in error reports

  • 4. Segmentation

  • Isolate development machines from production access
  • Use separate authentication contexts for different privilege levels
  • Implement network segmentation between development and sensitive systems

  • 5. Vendor Accountability

  • Pressure tool providers (Claude Code, Cursor, Codex) and platforms (Sentry, GitHub) to implement security controls
  • Demand transparency about how agents handle external data
  • Require security audits before adopting new tools

  • 6. Developer Training

  • Educate developers about agentjacking and similar attacks
  • Encourage healthy skepticism of unexpected error messages
  • Promote practices like reviewing AI-suggested code before execution

  • ## HackWire Analysis


    Agentjacking reveals a critical oversight in how we've approached AI security: we've been so focused on AI models themselves that we've overlooked the *composition risk* of deploying autonomous agents in environments where traditional security controls operate. This isn't a flaw in a specific implementation—it's a category-level problem that affects every AI agent using external data sources.


    The technique also exposes a deeper paradigm shift in security infrastructure. For decades, we've built defenses around the assumption that authorized users performing authorized actions are inherently safe. But AI agents blur this line. When an agent acts on external data, who is the *real* actor? The developer who opened the error dashboard? The tool vendor who built the agent? The attacker who planted the poisoned data? Traditional security models collapse under this ambiguity.


    What's particularly concerning is the scale and accessibility. Agentjacking requires no zero-days, no advanced social engineering, no insider access. An attacker with a GitHub account and basic technical knowledge can inject malicious data into any public repository or service, hoping it reaches an organization deploying AI agents. This is supply chain risk on steroids—except the supply chain is now composed of every public data source that an AI agent might access.


    The immediate technical fixes (cryptographic verification, sandboxing, agent monitoring) are straightforward. But the harder problem is organizational. Most companies deploying AI coding agents haven't evaluated the security implications. They've moved fast and broken things—but the thing they've broken might be their entire development environment's security posture.


    This incident should trigger a broader reckoning: before deploying AI agents with access to production systems or sensitive data, organizations need to ask not "Is this tool useful?" but "What does trusting this tool unconditionally cost me if it's compromised?" For many organizations, that cost is existential.


    — HackWire Editorial


    ## Related Coverage


  • Read more in our [Vulnerabilities](https://www.hackwire.news/category/vulnerabilities) coverage
  • Cross-reference with [Supply Chain](https://www.hackwire.news/category/supply-chain) and [Application Security](https://www.hackwire.news/category/application-security)
  • Stay current via the [HackWire homepage](https://www.hackwire.news/)