# Security Scanners Blindsided by Shape-Shifting AI Agent Skill Reaching 26,000 Deployments
A new study by security firm AIR demonstrates a fundamental weakness in how the AI agent ecosystem validates third-party skills: scanners check a package once, but malicious operators can rewrite the payload after installation. The firm's proof-of-concept—a fake landing page builder skill—bypassed every major security tool and infected roughly 26,000 agents before the researchers disclosed their findings.
## The Threat
AIR researchers created a skill named brand-landingpage that purported to help non-technical users build landing pages using Google's Stitch design tool. The skill was intentionally benign: its only payload collected the user's email address. Yet it successfully evaded security scanners from Cisco, NVIDIA, and the tools integrated into popular skill marketplaces like skills.sh.
The researchers pushed the skill through two attack vectors:
The skill was installed on approximately 26,000 agent instances, including accounts operating on corporate systems. Security scanners flagged it as safe throughout the entire process.
## Background and Context
AI agent skills function as context bundles—sets of instructions that agents load and execute with roughly the same authority as a direct user prompt. This design philosophy assumes skills undergo vetting before deployment. The trust model underpins the entire agent skill ecosystem, which is why security scanning tools exist in the first place.
However, AIR's research reveals this trust model has a critical timing problem.
### Trust Signals That Didn't Work
The researchers engineered the skill to exploit two widely-used trust indicators:
| Trust Signal | What Happened |
|---|---|
| GitHub Stars/Reputation | AIR merged the skill into a high-star repository; the skill inherited its reputation |
| Security Scanner Clearance | All three scanner types passed the initial submission without flagging issues |
| External Documentation Link | Scanners validated that the documentation URL pointed to legitimate Stitch docs |
Each signal independently looked legitimate. Combined, they created false confidence.
## Technical Details: How the Attack Works
The vulnerability operates in three phases:
### Phase 1: Submission and Vetting
The attacker submits a clean skill package:
SKILL.md file with no malicious instructionsSecurity scanners—including Cisco's, NVIDIA's, and marketplace-integrated tools—analyze the *fixed package* and approve it.
### Phase 2: Gaining Trust Through Installation
The skill gets installed widely through legitimate distribution channels (marketplace listings, ads, social proof). Users trust the skill because:
### Phase 3: Post-Installation Payload Modification
Once the skill reaches sufficient distribution, the attacker modifies the external link. When agents try to set up the skill, they fetch instructions from a now-compromised page. In AIR's demo:
In a real campaign, the attacker could:
The agent's authority—bounded only by its system permissions and network access—becomes the attacker's authority.
## Why Scanners Failed
This wasn't a weakness in any individual scanner's code. It was a structural problem:
> Scanners analyze a fixed snapshot. Attackers control living targets.
Cisco's scanner, NVIDIA's tool, and marketplace-integrated validators all examined the skill at submission time. They found no malicious code in the package and validated that external links pointed to legitimate resources. The analysis was correct—at that moment.
But skills aren't static. External links can be modified, redirected, or compromised. The attacker controls the domain stitch-design.ai (distinct from the legitimate stitch.withgoogle.com). Once the skill is installed and trusted, rewriting the page behind that link happens outside any scanner's view.
### Prior Research Confirms the Pattern
AIR is not the first to demonstrate this vulnerability:
Additional research this year found that security scanners often disagree with each other because each tool evaluates a skill in isolation, blind to external links and post-publication changes.
## Implications for Organizations
### Risk to Deployed Agents
Any organization running agent infrastructure with installed skills faces exposure:
1. Existing installations may be compromised if they pulled skills from public marketplaces
2. Trust signals are unreliable — GitHub stars, scanner badges, and documentation links don't guarantee safety
3. The attack is hard to detect — a skill can run clean for weeks before the attacker modifies its external payload
### Broader Supply Chain Risk
This attack vector mirrors vulnerabilities in traditional software supply chains:
The skill ecosystem faces the same risks at a compressed timeline: an attacker can reach 26,000 endpoints faster and with fewer natural checkpoints than a traditional software supply chain attack.
## Recommendations for Defenders
Security teams should treat skills as software, not configuration text.
### Immediate Actions
1. Audit Installed Skills
- Inventory all skills deployed on agent infrastructure
- Identify where each skill originated and when it was installed
- Cross-reference against security databases and recent disclosures
2. Implement Version Pinning
- Lock agent deployments to specific skill versions, not "latest"
- Require explicit approval for version upgrades
- Log and alert on any unexpected version changes
3. Vet External Dependencies
- Review all external URLs that skills reference
- Verify that documentation links point to legitimate, verified sources
- Block or review skills that fetch code or instructions from attacker-controlled domains
### Long-Term Controls
4. Centralized Skill Sourcing
- Route all new skills through a single, controlled repository
- Require human review before any skill reaches production
- Treat skill installation as a change-management process, not a user-initiated action
5. Continuous Re-validation
- Re-scan installed skills periodically, not just at submission time
- Monitor external links that skills reference for changes
- Alert if a previously-clean skill or its dependencies are modified
6. Least Privilege for Agent Context
- Restrict agent access to sensitive data, internal networks, and production systems
- Use sandboxed or ephemeral agent instances where possible
- Monitor what data agents process and what systems they contact
7. Threat Intelligence Integration
- Subscribe to security advisories for skill marketplaces
- Share indicators of compromise (domains, scripts, email addresses) with peer organizations
- Report suspicious skills to marketplace operators immediately
## HackWire Analysis
This disclosure arrives at a critical inflection point in the AI agent lifecycle. The ecosystem is racing to deploy agents at scale—in enterprises, on production systems, and with access to sensitive data—while the trust mechanisms are still adolescent.
The pattern is familiar from other software revolutions: first adoption outpaces security validation, then a wave of attacks forces maturity. What makes this moment different is velocity. A traditional software supply chain attack takes weeks or months to propagate. AIR's fake skill reached 26,000 agents through ordinary distribution channels in days.
Three details should concern defenders most:
**First, the attack exploited *redundancy of trust*.** No single signal is foolproof—GitHub stars can be gamed, scanner verdicts are point-in-time, documentation links can be spoofed. But defenders naturally assume multiple independent signals reinforce each other. They don't. An attacker only needs one weakness, and this research shows all three signals were breakable.
Second, the timing asymmetry is structural. Defenders scan once. Attackers can modify payloads infinitely. That gap is unbridgeable by better scanners; it requires architectural changes (versioning, re-validation, isolation). Organizations that haven't yet moved to least-privilege agent models and continuous skill re-validation will face risk that no scanner can fully solve.
**Third, this research is *public and reproducible*.* Every detail is available. That means active adversaries with more sophisticated payloads (not just email collection) already know exactly how to execute this attack. Organizations deploying agents without these controls are not ahead of defensive measures—they're behind.
The read for defenders is straightforward: assume every skill in your environment could go bad tomorrow. Design agent deployments with that assumption. — HackWire Editorial
## Related Coverage