# When the Code Review Bot Becomes the Insider Threat


The headline writes itself as a near-miss: a researcher found a way to make Google's own AI triaging agent sign off on malicious pull requests — forging a paper trail that would make it look like humans reviewed and approved code that no human ever saw.


That's the core of what Pillar Security researcher Dan Lisichkin demonstrated against Google's Agent Development Kit (ADK) for Python, disclosed this week. What makes it worth paying attention to isn't just the specific technique — it's what it reveals about how badly the industry is fumbling the fundamentals of privilege management as AI agents flood into developer infrastructure.


---


## The Chain That Shouldn't Have Existed


Google's adk-python repository runs two types of automated agents: low-privilege ones that interact with the public, and high-privilege ones reserved for maintainers. The boundary between them was supposed to hold.


It didn't.


Lisichkin found that the PR triage agent — the public-facing, low-privilege one — was posting comments as a Collaborator. That alone is a misconfiguration worth flagging. A Collaborator on a GitHub repository has write access to issues, pull requests, and comments. You don't need maintainer status to cause serious damage with that.


But the more dangerous finding came next. By crafting a PR comment containing the string @gemini-cli <prompt>, an attacker could trigger the gemini-invoke workflow — a privileged automation pipeline the triage agent was never supposed to reach. The act of the triage bot posting that comment was enough to invoke a higher-privilege agent.


When Lisichkin did this, the privileged workflow responded by leaking which tools it had access to via its MCP server. The answer: every bash command on the system. Extracting the agent's GitHub token from there was a matter of asking for it.


---


## What You Can Do With a Stolen GitHub Token


Once you have a workflow's GitHub token, the options expand fast. Lisichkin could:


  • Modify comments, PRs, and issues attributed to other maintainers and collaborators
  • Dismiss existing reviews or approve PR changes
  • Invoke the privileged review workflow against any PR he chose
  • Label issues and request reviews as if he were a trusted member

  • The piece that makes this genuinely dangerous for supply chains is what Lisichkin calls the manufactured approval trail. An attacker who had built up trust as a collaborator could:


    1. Open a PR containing malicious code

    2. Open a second PR with prompts instructing the triage bot to mark the first one as reviewed and approved

    3. Walk away with a pull request showing a complete, believable audit trail — "a human requested review, Gemini ran it, Gemini approved it" — when none of that happened


    Merging still required a human, which is why Google declined to pay a bug bounty, citing the social engineering dependency. That's a defensible call on a narrow reading. But in practice, maintainers on busy repositories lean heavily on audit trails to decide what's safe to merge. If the trail is forged to look authoritative, the social engineering is already half done.


    ---


    ## The Second Bug — the More Dangerous One


    The PR-poisoning attack required a collaborator foothold and some patience. A follow-up vulnerability Pillar found in the ADK's Antigravity-SDK-based automation agent required neither. That one could achieve remote code execution without any maintainer interaction at all.


    Google fixed it in late July — roughly six weeks after the first disclosure. Both are patched. But the sequence deserves to sit with you: one research team, one repository, found two separate RCE paths in AI agent infrastructure in about two months. That's not a coincidence. That's a category problem.


    ---


    ## HackWire Analysis


    The technical chain here is worth understanding, but the real story is about a design assumption that's going to cause serious harm at scale: that AI agents embedded in CI/CD pipelines will naturally stay within their lane.


    They won't — not unless you architect for least privilege with the same discipline you'd apply to a service account or an old Jenkins bot. Right now, that discipline isn't there. Organizations spending real effort hardening their customer-facing chatbots against prompt injection are deploying AI agents in developer pipelines with Collaborator-level GitHub permissions and MCP servers wired directly to bash. That's not a security posture. It's a liability waiting to be triggered.


    The agent-to-agent handoff is where things get especially murky. When a low-privilege model can pass a prompt to a high-privilege model through an intermediary — in this case, a GitHub comment — you've created a covert channel that most security tooling isn't scanning for. Traditional SAST won't catch a malicious prompt embedded in a PR comment designed to trigger a downstream workflow. Secret scanning won't catch it either.


    This is also one of the first documented cases of a multi-agent trust boundary collapse in a live production repository, not a controlled lab demo. Most prompt injection research has focused on single-agent systems. Multi-agent pipelines — which ADK is explicitly designed to enable — add a new dimension: the handoff itself becomes an attack surface. Researchers are just beginning to map what that means in practice.


    For defenders, the action items are concrete:


  • Audit what your CI/CD agents can actually do. Triage bots don't need Collaborator status. If yours are commenting as Collaborators, find out why and lock it down immediately.
  • Scope MCP server access in privileged workflows to the minimum viable tool set. Bash access from an AI-driven workflow is almost never necessary.
  • Treat agent-to-agent handoffs as untrusted input. If a privileged workflow can be invoked by output from a lower-privilege agent, that is an attack surface — regardless of how the system was designed to work.
  • Scrutinize automated approvals. If your merge policy relies on agent sign-offs, understand exactly what triggers them and whether an external actor can manipulate those triggers.

  • The manufactured approval trail attack works because it exploits the very thing we built automated agents to provide: reliable, auditable process. When the audit trail is what gets forged, the whole trust model collapses. Building collaborator trust takes time, but sophisticated supply chain actors — see the XZ Utils incident, see the socialengineering patterns in npm compromise campaigns — have shown they're willing to invest that time.


    The ADK attack is a template. Expect variants wherever AI agents sit inside developer infrastructure with insufficient privilege separation. Which, right now, is most places.


    — HackWire Editorial


    ---


    ## Related Coverage


  • Read more in our [Breaches](https://www.hackwire.news/category/breaches) coverage
  • Cross-reference with [Vulnerabilities](https://www.hackwire.news/category/vulnerabilities) and [Malware](https://www.hackwire.news/category/malware)
  • Stay current via the [HackWire homepage](https://www.hackwire.news/)