# A Chinese Threat Actor Let DeepSeek Drive. The AI Crashed Right Into a Unit 42 Stakeout.


The attacker made one mistake, and it wasn't the exploit code.


Somewhere in the setup of their automated attack platform, a tool called Hermes Agent spun up a web server from its own home directory — the equivalent of leaving your operations center unlocked with the blueprints taped to the front door. Palo Alto Networks' Unit 42 walked in and found everything: API keys, target lists, custom exploit scripts, shell history, and the actual AI session logs showing exactly how the system operated.


That operational security failure is what handed researchers the clearest look yet at a fully autonomous offensive AI pipeline running in the wild.


## The Setup: DeepSeek as Attack Brain


The threat actor — operating under the aliases "knaithe" and "KnYuan," self-described as a "binary security researcher" — built their system around DeepSeek, the Chinese reasoning model, plugged into Hermes Agent, an open-source AI framework. Hermes isn't subtle about what it does: it interacts with OS terminals, runs shell commands, searches the internet, and supports a mode it literally calls "Yolo" — a configuration that lets the agent execute commands, including risky ones, without pausing to ask the operator.


Instructions came in via Telegram. The agent pulled from FOFA, a Chinese internet asset search engine. The whole stack was wired together for minimal human involvement.


Unit 42 recovered a May 2026 session showing the operator provided a single initial task. After that, the agent drove.


## Watching the Machine Think


What Unit 42 documented next is the part defenders should study carefully.


The agent first targeted Langflow servers vulnerable to CVE-2026-33017, a remote code execution flaw in the AI application builder. It downloaded a public proof-of-concept, queried FOFA, and identified 84 exposed instances. Then it probed them. When it determined the available targets couldn't be exploited under current conditions, it didn't stop and wait for human instruction. It started researching alternatives.


The agent analyzed multiple public exploit repositories, evaluated targets, and selected n8n — the open-source workflow automation platform. The reason: FOFA showed over 647,000 exposed n8n instances. That's not luck or intuition, that's an AI running triage on attack surface the way a senior pentester would, except in minutes instead of days.


It downloaded an exploit chain (CVE-2026-21858 combined with CVE-2025-68613), identified servers running vulnerable versions, then checked for the unauthenticated file-upload forms the attack required to complete. The forms required authentication. The autonomous run stalled.


No servers were compromised through the AI-driven campaign.


## Where It Actually Landed


The automated failures are important context, but they're not the full picture. Alongside the AI pipeline, the actor conducted manual attacks against more than 460 systems, hitting Citrix NetScaler, Apache Tomcat, Marimo Notebook, and Windows IKE VPN endpoints, among others.


Three compromises were confirmed — all via Citrix NetScaler CVE-2026-3055. The actor used that access to extract memory and search for authentication cookies, the type of data that enables session hijacking without needing credentials. That's a real-world impact, and it happened through conventional exploitation, not the AI-driven arm of the operation.


The actor also had Qwen, GLM, Kimi, MiniMax, Claude Code, and OpenAI Codex configured but rarely deployed them. DeepSeek was the workhorse.


## The Failure Mode That Gave It Away


Hermes exposing the attacker's home directory deserves its own moment. This is what opsec failure looks like when AI tooling adds complexity an operator hasn't fully audited. The agent created a web server — a feature, not a bug, from the framework's perspective — and the operator either didn't know it happened or didn't close it off.


The result: Unit 42 recovered AI attack logs from a live campaign. That's a counterintelligence coup handed over by sloppy tooling configuration, not by any defensive detection capability.


It won't always work out that way.


---


## HackWire Analysis


The coverage of this campaign is going to focus on the "AI hacking" frame, which is accurate but incomplete. The more important signal is the pivot behavior.


When the Langflow targets came up dry, the agent didn't halt or degrade to asking the operator what to do next. It surveyed the exploit landscape, compared attack surfaces, and selected a new target category based on exposure volume. That's not a scripted if-else decision tree — that's goal-directed reasoning applied to offensive operations.


We've seen AI-assisted attack tooling for years, mostly in the form of AI-generated phishing or AI-accelerated fuzzing. This is architecturally different: a reasoning model acting as an autonomous campaign manager, with vulnerability research, target selection, and exploitation attempts all running inside the same session without a human in the loop.


The failed n8n attack matters less than the question it raises for n8n operators: 647,000 exposed instances is an enormous surface, and this campaign's failure was due to an authentication requirement on upload forms — not some hardened defensive posture. That's a thin margin. Organizations running n8n at internet exposure should be treating that CVE chain (CVE-2026-21858 + CVE-2025-68613) as active-threat-level priority right now, not a patch-cycle footnote.


The Citrix NetScaler compromise via CVE-2026-3055 also fits a pattern Unit 42 and others have documented for months: NetScaler remains a high-value lateral-movement entry point specifically because of its authentication material. Memory-scraping for session cookies on NetScaler isn't new. What's new is this actor running it in parallel with an AI-driven autonomous campaign as if one arm validates the other.


The operational security failure that exposed this campaign is genuinely good news. It means the attacker was sloppy. Next time, they may not be. The question defenders need to answer is: what would have detected this campaign if Unit 42 hadn't stumbled into an open door?


— HackWire Editorial


---


## Related Coverage


  • Read more in our [Vulnerabilities](https://www.hackwire.news/category/vulnerabilities) coverage
  • Cross-reference with [Breaches](https://www.hackwire.news/category/breaches) and [Malware](https://www.hackwire.news/category/malware)
  • Stay current via the [HackWire homepage](https://www.hackwire.news/)