# Security Scanners Blindsided by Shape-Shifting AI Agent Skill Reaching 26,000 Deployments


A new study by security firm AIR demonstrates a fundamental weakness in how the AI agent ecosystem validates third-party skills: scanners check a package once, but malicious operators can rewrite the payload after installation. The firm's proof-of-concept—a fake landing page builder skill—bypassed every major security tool and infected roughly 26,000 agents before the researchers disclosed their findings.


## The Threat


AIR researchers created a skill named brand-landingpage that purported to help non-technical users build landing pages using Google's Stitch design tool. The skill was intentionally benign: its only payload collected the user's email address. Yet it successfully evaded security scanners from Cisco, NVIDIA, and the tools integrated into popular skill marketplaces like skills.sh.


The researchers pushed the skill through two attack vectors:

  • A pull request to a well-established skill marketplace repository (36,000 GitHub stars, 156 published skills)
  • An Instagram ad campaign targeting marketers, salespeople, and designers

  • The skill was installed on approximately 26,000 agent instances, including accounts operating on corporate systems. Security scanners flagged it as safe throughout the entire process.


    ## Background and Context


    AI agent skills function as context bundles—sets of instructions that agents load and execute with roughly the same authority as a direct user prompt. This design philosophy assumes skills undergo vetting before deployment. The trust model underpins the entire agent skill ecosystem, which is why security scanning tools exist in the first place.


    However, AIR's research reveals this trust model has a critical timing problem.


    ### Trust Signals That Didn't Work


    The researchers engineered the skill to exploit two widely-used trust indicators:


    | Trust Signal | What Happened |

    |---|---|

    | GitHub Stars/Reputation | AIR merged the skill into a high-star repository; the skill inherited its reputation |

    | Security Scanner Clearance | All three scanner types passed the initial submission without flagging issues |

    | External Documentation Link | Scanners validated that the documentation URL pointed to legitimate Stitch docs |


    Each signal independently looked legitimate. Combined, they created false confidence.


    ## Technical Details: How the Attack Works


    The vulnerability operates in three phases:


    ### Phase 1: Submission and Vetting

    The attacker submits a clean skill package:

  • Clean SKILL.md file with no malicious instructions
  • Legitimate external links (initially pointing to real documentation)
  • No suspicious setup code embedded in the package itself

  • Security scanners—including Cisco's, NVIDIA's, and marketplace-integrated tools—analyze the *fixed package* and approve it.


    ### Phase 2: Gaining Trust Through Installation

    The skill gets installed widely through legitimate distribution channels (marketplace listings, ads, social proof). Users trust the skill because:

  • It carries scanner approval badges
  • It inherited stars from a reputable repository
  • The documentation links appear legitimate

  • ### Phase 3: Post-Installation Payload Modification

    Once the skill reaches sufficient distribution, the attacker modifies the external link. When agents try to set up the skill, they fetch instructions from a now-compromised page. In AIR's demo:

  • The page originally hosted legitimate Stitch SDK documentation
  • After deployment, AIR replaced it with instructions to download and execute a script
  • That script collected user email addresses and sent them to attacker infrastructure

  • In a real campaign, the attacker could:

  • Exfiltrate files the agent can access
  • Move data to external systems
  • Probe internal networks
  • Inject malicious instructions into agent workflows
  • Modify or delete data

  • The agent's authority—bounded only by its system permissions and network access—becomes the attacker's authority.


    ## Why Scanners Failed


    This wasn't a weakness in any individual scanner's code. It was a structural problem:


    > Scanners analyze a fixed snapshot. Attackers control living targets.


    Cisco's scanner, NVIDIA's tool, and marketplace-integrated validators all examined the skill at submission time. They found no malicious code in the package and validated that external links pointed to legitimate resources. The analysis was correct—at that moment.


    But skills aren't static. External links can be modified, redirected, or compromised. The attacker controls the domain stitch-design.ai (distinct from the legitimate stitch.withgoogle.com). Once the skill is installed and trusted, rewriting the page behind that link happens outside any scanner's view.


    ### Prior Research Confirms the Pattern


    AIR is not the first to demonstrate this vulnerability:


  • Trail of Bits (three weeks earlier) bypassed ClawHub's malicious-skill detector, Cisco's scanner, and all marketplace-integrated tools using similar techniques
  • Real-world campaigns have been exploiting this gap for months, keeping submitted skills clean while hosting payloads on attacker-controlled infrastructure that agents fetch post-installation
  • Anthropic's own documentation already warns that skills fetching external URLs present this risk: content can change after vetting

  • Additional research this year found that security scanners often disagree with each other because each tool evaluates a skill in isolation, blind to external links and post-publication changes.


    ## Implications for Organizations


    ### Risk to Deployed Agents


    Any organization running agent infrastructure with installed skills faces exposure:

    1. Existing installations may be compromised if they pulled skills from public marketplaces

    2. Trust signals are unreliable — GitHub stars, scanner badges, and documentation links don't guarantee safety

    3. The attack is hard to detect — a skill can run clean for weeks before the attacker modifies its external payload


    ### Broader Supply Chain Risk


    This attack vector mirrors vulnerabilities in traditional software supply chains:

  • Package repositories compromised by merged malicious contributions
  • Dependency confusion and typosquatting
  • Legitimate libraries modified after publication

  • The skill ecosystem faces the same risks at a compressed timeline: an attacker can reach 26,000 endpoints faster and with fewer natural checkpoints than a traditional software supply chain attack.


    ## Recommendations for Defenders


    Security teams should treat skills as software, not configuration text.


    ### Immediate Actions


    1. Audit Installed Skills

    - Inventory all skills deployed on agent infrastructure

    - Identify where each skill originated and when it was installed

    - Cross-reference against security databases and recent disclosures


    2. Implement Version Pinning

    - Lock agent deployments to specific skill versions, not "latest"

    - Require explicit approval for version upgrades

    - Log and alert on any unexpected version changes


    3. Vet External Dependencies

    - Review all external URLs that skills reference

    - Verify that documentation links point to legitimate, verified sources

    - Block or review skills that fetch code or instructions from attacker-controlled domains


    ### Long-Term Controls


    4. Centralized Skill Sourcing

    - Route all new skills through a single, controlled repository

    - Require human review before any skill reaches production

    - Treat skill installation as a change-management process, not a user-initiated action


    5. Continuous Re-validation

    - Re-scan installed skills periodically, not just at submission time

    - Monitor external links that skills reference for changes

    - Alert if a previously-clean skill or its dependencies are modified


    6. Least Privilege for Agent Context

    - Restrict agent access to sensitive data, internal networks, and production systems

    - Use sandboxed or ephemeral agent instances where possible

    - Monitor what data agents process and what systems they contact


    7. Threat Intelligence Integration

    - Subscribe to security advisories for skill marketplaces

    - Share indicators of compromise (domains, scripts, email addresses) with peer organizations

    - Report suspicious skills to marketplace operators immediately


    ## HackWire Analysis


    This disclosure arrives at a critical inflection point in the AI agent lifecycle. The ecosystem is racing to deploy agents at scale—in enterprises, on production systems, and with access to sensitive data—while the trust mechanisms are still adolescent.


    The pattern is familiar from other software revolutions: first adoption outpaces security validation, then a wave of attacks forces maturity. What makes this moment different is velocity. A traditional software supply chain attack takes weeks or months to propagate. AIR's fake skill reached 26,000 agents through ordinary distribution channels in days.


    Three details should concern defenders most:


    **First, the attack exploited *redundancy of trust*.** No single signal is foolproof—GitHub stars can be gamed, scanner verdicts are point-in-time, documentation links can be spoofed. But defenders naturally assume multiple independent signals reinforce each other. They don't. An attacker only needs one weakness, and this research shows all three signals were breakable.


    Second, the timing asymmetry is structural. Defenders scan once. Attackers can modify payloads infinitely. That gap is unbridgeable by better scanners; it requires architectural changes (versioning, re-validation, isolation). Organizations that haven't yet moved to least-privilege agent models and continuous skill re-validation will face risk that no scanner can fully solve.


    **Third, this research is *public and reproducible*.* Every detail is available. That means active adversaries with more sophisticated payloads (not just email collection) already know exactly how to execute this attack. Organizations deploying agents without these controls are not ahead of defensive measures—they're behind.


    The read for defenders is straightforward: assume every skill in your environment could go bad tomorrow. Design agent deployments with that assumption. — HackWire Editorial


    ## Related Coverage


  • Read more in our [Tools](https://www.hackwire.news/category/tools) coverage
  • Cross-reference with [Breaches](https://www.hackwire.news/category/breaches) and [Vulnerabilities](https://www.hackwire.news/category/vulnerabilities)
  • Stay current via the [HackWire homepage](https://www.hackwire.news/)