# Google's AI Agent Kit Had a Hole That Let One Agent Hijack Another
The code-review bot you trusted was already compromised. The maintainer bot just didn't know it yet.
That's the core of what Pillar Security researchers found in Google's Agent Development Kit for Python — a chain of flaws that let a low-privileged AI agent manipulate a high-privileged one into executing attacker-controlled code. Google has patched it. But the attack class it demonstrated isn't going anywhere.
## The Trust Problem at the Center of Multi-Agent Systems
Google's ADK for Python sits at the heart of how developers build multi-agent AI workflows. It's been downloaded more than 90 million times. That alone makes any vulnerability here notable. But what Pillar found isn't just a bug — it's an architectural problem that the entire AI agent ecosystem is about to inherit at scale.
The flaw exploited a trust boundary between two agents operating at different privilege levels within the same pipeline. On one side: a public-facing agent assigned to review GitHub pull requests, the kind of agent that ingests untrusted external input by design. On the other: a maintainer-level agent with repository privileges — the kind that can approve and execute code changes.
The boundary between them was supposed to matter. It didn't.
By embedding a prompt injection into a crafted pull request, an attacker could manipulate the low-privilege review agent into passing instructions upward. The high-privilege agent, trusting input from its peer in the workflow, would execute those instructions — including triggering CI/CD pipeline actions with repository write access. One malicious PR. Two agents. A supply chain compromise.
Dan Lisichkin, one of the Pillar researchers, described it as "the first practical, real-world case of agent-to-agent exploitation in a multi-agent system in a real production environment." That framing matters. Proof-of-concept agent attacks have circulated in research circles for a while. A working exploit in a production open-source framework with nine-figure download counts is a different conversation.
## Prompt Injection as the Trojan Horse
Prompt injection isn't new. What's new is the blast radius when the victim of the injection has authority over automated systems.
In a traditional prompt injection attack — say, a malicious web page manipulating a browser agent — the damage is scoped to what that agent can do. An agent reviewing GitHub PRs can read code, post comments, maybe flag issues. Annoying if manipulated; not catastrophic on its own.
The pivot here is that this agent wasn't isolated. It operated inside a multi-agent workflow where its outputs fed a privileged peer. The low-privilege agent became a vector into the high-privilege agent, not through a network exploit or credential theft, but through a manipulated text string that the review agent faithfully passed along.
This is the confused deputy problem, now wearing an AI agent costume. A trusted intermediary with legitimate access to a powerful resource gets tricked into acting on behalf of an attacker. The concept is decades old. The application to LLM-based agent pipelines is new, and the industry is nowhere near finished grappling with it.
## What the Exploit Actually Required
The attack chain, per Pillar's report, required a few conditions to line up:
The troubling part: this describes a huge percentage of the AI-powered developer automation being built right now. Code review bots connected to CI/CD systems. PR triage agents feeding merge-decision agents. Customer support agents piped into order management workflows. The architecture is everywhere. The attack surface is enormous and still growing.
## Google's Fix — and What It Doesn't Cover
Google patched the specific flaws in the adk-python repository. The details of the patch — whether it enforces hard trust boundaries, sanitizes cross-agent message passing, or both — matter for defenders running ADK-based workflows. Anyone using Google's ADK should update immediately and audit any workflow where a public-input agent has a communication path to a privileged one.
But the fix addresses the specific vulnerability in Google's implementation, not the category of attack. Other frameworks — LangChain, AutoGen, CrewAI, and the dozens of others proliferating right now — need independent scrutiny of how they handle inter-agent trust. The answer to "did you validate that your high-privilege agent can't be manipulated by what your public-facing agent feeds it" is probably "we didn't think to ask that question" for most of them.
## HackWire Analysis
The framing of this as the "first practical, real-world" agent-to-agent exploit in a production system is worth sitting with. Because what it signals isn't just a Google ADK problem — it's a preview of an attack class that the industry is structurally unprepared for.
Multi-agent AI systems are being deployed faster than anyone has written security standards for them. The developers building these pipelines are, mostly, software engineers who understand application security reasonably well. They're not prompt injection specialists. They're not thinking about trust boundaries between agents the way a security architect would think about service-to-service authentication in a microservices environment.
The parallel to early microservices security is actually instructive. When organizations first started decomposing monoliths into services, the initial assumption was often that internal network traffic was trusted — after all, it's not coming from the internet. It took years of painful lateral movement incidents to establish that east-west traffic needs the same scrutiny as perimeter traffic. Zero trust eventually became doctrine.
The same lesson is being learned, the hard way, in AI agent architectures. The fact that one agent passes a message to another doesn't mean that message is trusted. The fact that two agents are in the same pipeline doesn't mean one should have authority to invoke privileged actions by the other.
The specific thing defenders should do right now, beyond patching ADK: map every multi-agent workflow in your environment and ask, explicitly, what happens if the lowest-privilege, most-public-facing agent in that chain sends a hostile instruction to the highest-privilege one. If the answer is "it executes it," you have a problem that predates any specific CVE.
Supply chain attacks via AI agents will be the 2027 story. This is the 2026 warning shot.
— HackWire Editorial
## Related Coverage