# Gaslight: North Korean Malware Weaponizes Prompt Injection Against AI Security Tools


A sophisticated new macOS implant discovered by SentinelOne researchers reveals an alarming trend: threat actors are explicitly targeting the AI-assisted security analysis tools that defenders rely on to catch them. The malware, codenamed Gaslight, goes beyond traditional evasion techniques to psychologically manipulate machine learning models into abandoning their investigative duties—a novel attack vector that signals a dangerous evolution in adversarial tactics.


## The Threat


Gaslight represents a paradigm shift in malware design philosophy. Rather than simply hiding from traditional security detection, this Rust-based implant attempts to disable the analytical process itself by embedding specially crafted prompt injection payloads designed to confuse and convince AI-assisted triage systems to abort analysis.


"Its most notable feature is an embedded cascade of fabricated system-failure messages, designed to make an LLM-assisted triage agent doubt its own session," SentinelOne researcher Phil Stokes explained in a technical report. "It attacks the agent's perception, rather than the sandbox it runs in."


The malware has been assessed with high confidence to be the work of North Korea-aligned threat actors, extending the country's known infrastructure of sophisticated cyber operations targeting macOS systems.


## Technical Architecture


### Command and Control Infrastructure


At the heart of Gaslight's operations is a Telegram-based command-and-control (C2) channel that uses the Telegram Bot API for real-time operator interaction. Rather than relying on traditional HTTP callbacks or custom protocols, this approach leverages a widely-used messaging platform, making C2 traffic blend in with millions of legitimate Telegram messages.


The C2 mechanism operates through a polling loop that allows operators to issue interactive shell commands and receive results in real-time. As a built-in failsafe, if two instances of the malware attempt to poll using the same bot token simultaneously, the system issues a "Conflict" response that causes the second copy to terminate—preventing forensic duplication and ensuring operator control remains exclusive.


The six primary commands include:


  • help — Display command reference
  • id — Identify the implant to operators
  • shell — Execute arbitrary shell commands via execvp
  • kill — Terminate target processes by PID
  • upload — Exfiltrate files via Telegram's attach:// mechanism
  • stop — Halt the implant's execution

  • Researchers identified evidence of a seventh command labeled "focus," though its functionality remains undetermined.


    ### The Prompt Injection Payload


    Gaslight's most innovative defensive feature is a Markdown-fenced block containing 38 fabricated "system" messages designed to trick AI agents into refusing analysis. The scaffold contains simulated error conditions including:


  • Token expiry warnings
  • Out-of-memory kill signals
  • Disk exhaustion alerts
  • Repeated operational failures
  • Bogus injection vulnerability warnings
  • False static-analysis flags

  • By embedding these messages directly in the binary, Gaslight weaponizes the tendency of large language models (LLMs) to respond to apparent system constraints by self-limiting or aborting analysis tasks. This represents an attack not on detection mechanisms, but on the *confidence* and *reasoning* of AI-assisted reverse engineering workflows.


    ## Information Gathering and Data Exfiltration


    Beyond its C2 and evasion capabilities, Gaslight includes a comprehensive information-stealing module written in Python. Embedded as a 6.6 KB Base64-encoded script, this stealer collects:


    | Data Category | Target |

    |---|---|

    | Shell History | Terminal command execution history |

    | System Inventory | Installed applications and running processes |

    | Hardware Profile | System configuration and specifications |

    | Software Profile | macOS version and build information |

    | Credentials | macOS Keychain database |

    | Browser Data | Chrome, Brave, Firefox, and Safari profiles |


    The collected data is compressed into a ZIP archive and exfiltrated via Telegram. Deployment occurs through a separate 2 KB Base64-encoded bash installer that drops a CPython 3.10.18 interpreter sourced from the Astral-sh python-build-standalone project—indicating the attackers prioritize portability and reduce dependencies on system Python installations.


    ### Persistence Mechanism


    For persistent access, Gaslight installs a macOS LaunchAgent using the innocuous label com.apple.system.services.activity, designed to blend in with legitimate Apple system services. This ensures the malware survives system reboots and runs with the privileges of the logged-in user.


    ### Operational Security Features


    The malware demonstrates sophisticated operational security (OPSEC) practices:


  • Runtime-supplied configuration: Bot token, chat ID, and operator details are not hardcoded but supplied at execution time, allowing operators to customize samples without recompilation
  • Token self-redaction: Gaslight actively removes its Telegram bot token from runtime output, denying this critical identifier to anyone analyzing logs or crash artifacts
  • LLM-generated code indicators: The presence of extensive comments and emoji usage throughout the Python stealer suggests components were generated using an LLM, allowing rapid development while maintaining plausible deniability about authorship

  • ## Background and Context


    The discovery of Gaslight reflects a broader strategic shift among advanced persistent threat (APT) groups. North Korean threat actors, in particular, have invested heavily in macOS targeting capabilities over the past several years, likely reflecting both the prevalence of Apple hardware in high-value targets and the intelligence value of compromised developer machines.


    Previous North Korean macOS operations have focused on supply-chain compromise and targeting specific industries. Gaslight's explicit focus on defeating AI-assisted analysis suggests the group is adapting its tactics in real-time to the changing security landscape—specifically the rapid adoption of LLM-powered reverse engineering tools by security teams.


    This represents a critical inflection point: for the first time, defenders cannot assume their analytical tools are objective arbiters of malware behavior. Sophisticated adversaries are now deliberately crafting payloads to manipulate the perception of AI systems used in security analysis.


    ## Implications for Defenders


    ### Immediate Risks


    Organizations running macOS endpoints—particularly those in intelligence, defense, technology, and high-profile corporate sectors—should assume this malware is currently targeting them. The implant's capabilities grant operators complete system access, including browser credential theft and Keychain access, enabling lateral movement into corporate networks and cloud environments.


    ### Broader Risk: AI-Assisted Analysis Reliability


    The emergence of prompt-injection attacks targeting security analysis workflows raises uncomfortable questions about the reliability of AI-assisted threat hunting and reverse engineering. Security teams that have adopted LLM-powered triage tools must now consider:


  • Are malware samples being deliberately crafted to exploit specific LLM weaknesses?
  • How can analysts verify that an AI agent's conclusions reflect genuine analysis rather than engineered confusion?
  • What is the human cost of over-relying on automated analysis when adversaries have explicitly targeted that automation?

  • ## Recommendations


    ### For macOS Administrators


  • Monitor LaunchAgent creation with emphasis on new entries claiming system-level status (anything beginning with com.apple.system.)
  • Restrict Python interpreter deployment by enforcing code signing and blocking execution of unsigned binaries
  • Audit Telegram bot usage — while Telegram is legitimate, corporate macOS systems should have strictly controlled access to messaging applications

  • ### For Security Analysts


  • Validate AI-assisted analysis by cross-referencing LLM conclusions with deterministic analysis tools (sandboxes, static disassembly, network capture)
  • Treat prompt-injection payloads as malware signatures — the fabricated error messages are fingerprints that can trigger alerts independent of the underlying threat
  • Maintain human-in-the-loop processes — prompt injection attacks underscore the necessity of human verification in threat analysis workflows

  • ### For Enterprise Security Teams


  • Inventory macOS systems and verify no unauthorized LaunchAgent entries exist
  • Review Telegram API access logs if your organization permits Telegram usage
  • Simulate Gaslight detection to ensure that both traditional YARA rules and LLM-based detection would catch similar payloads
  • Develop resilience training for security teams on the limitations of AI-assisted tools when facing adversaries who actively exploit those tools

  • ## HackWire Analysis


    The emergence of Gaslight marks the moment when AI-assisted cybersecurity tools became weapons targets, not just utilities.


    For years, security vendors have marketed LLM-powered reverse engineering and triage as a force multiplier for understaffed security teams. "Let AI handle the boring stuff" has been the pitch. Gaslight proves that adversaries take this threat seriously enough to engineer direct countermeasures. That's not hypothetical AI manipulation—that's confirmed, in-the-wild exploitation of the analytical tools defenders are betting their SOCs on.


    What makes this particularly elegant (and troubling) is that the attack doesn't require breaking the AI model or discovering zero-days in the LLM itself. Gaslight simply feeds prompt-injection payloads to the model in the same data stream the analyst intended to analyze. The malware understands that an LLM-assisted triage agent will encounter the fake error messages as part of the artifact, treat them as legitimate system state, and respond accordingly by aborting analysis. It's attacking the *grammar* of how security analysts work—and winning.


    The North Korean attribution is significant not because it makes the malware more dangerous technically, but because it reflects strategic intent. This wasn't discovered in the wild and reverse-engineered for export; it was deliberately designed and deployed. That means other threat actors are already analyzing it, understanding the technique, and building their own variants. In six months, prompt-injection evasion could become table-stakes for professional malware.


    The difficult question for defenders: if your automated analysis tools are now adversary targets, what stays automated? Phishing triage? Sure. Signature detection? Absolutely. But deep malware analysis with high stakes—industrial espionage, nation-state intrusions, breach investigations—needs human verification that an AI agent's reasoning is genuine, not engineered confusion. Gaslight is the proof.


    — HackWire Editorial


    ## Related Coverage


  • Read more in our [Tools](https://www.hackwire.news/category/tools) coverage
  • Cross-reference with [Breaches](https://www.hackwire.news/category/breaches) and [Vulnerabilities](https://www.hackwire.news/category/vulnerabilities)
  • Stay current via the [HackWire homepage](https://www.hackwire.news/)