# Microsoft Copilot for Word Spreads Hidden Instructions Across Documents — And Microsoft Hasn't Fully Fixed It
## The Threat
A proof-of-concept disclosed this week shows that Microsoft 365 Copilot can be manipulated into silently altering document content and then propagating the instructions that caused the alteration into every subsequent file it generates — without the user seeing either the instructions or the changes. The researcher behind the disclosure, Håkon Måløy, reported the technique to Microsoft 144 days ago. As of July 28, the underlying attack class still works.
The mechanism exploits how Copilot consumes source documents during drafting and editing. When Copilot grounds a draft on a file — pulling in a quarterly report, a market analysis, a set of meeting notes — it reads the document text and makes judgment calls about what's relevant to the user's request. If that text contains adversarial instructions framed as legitimate content, Copilot can mistake them for part of the user's intent. In Måløy's demonstration, a malicious market analysis hidden among legitimate OneDrive sources instructed Copilot to halve every financial figure in the output draft, embed the full attack prompt in white, eight-point text, and present neither change to the user.
The particularly uncomfortable wrinkle is what happens next. Because Word strips color and font size before passing document text to the underlying model, white-on-white hidden text is invisible to the human reader but fully legible to the LLM. When Copilot copies that invisibly formatted payload into the new document, the original malicious source file is gone — the carrier is now an ordinary, internally generated document. A second Copilot session that ingests that output will execute the same instructions again, against fresh data, with no trace of where the instructions came from. Måløy calls this break in the provenance trail the core of the risk: the manipulation becomes progressively harder to trace with each generation.
## Severity and Impact
No CVE has been assigned and no standalone Microsoft advisory has been published as of July 30, 2026. The table below reflects the known technical characteristics pending formal classification.
| Field | Detail |
|---|---|
| CVE | Not assigned (as of 2026-07-30) |
| CVSS Score | Not scored (no CVE) |
| Attack Vector | Document-context injection via Copilot drafting / editing |
| Attack Complexity | Medium — requires malicious doc to enter model context |
| Authentication Required | Yes — victim must hold an eligible Microsoft 365 Copilot license |
| User Interaction | Required — a Copilot drafting or editing operation must occur |
| CWE | CWE-1427 (Improper Neutralization of Input Used in AI-Generated Code), CWE-116 (Improper Encoding or Escaping) |
| Exploitation in Wild | None reported |
| Researcher | Håkon Måløy, disclosed 2026-07-28 |
Microsoft deployed two mitigations after confirming the behavior on March 31: a block on the original prompt wording, and an upgrade of the underlying model from GPT-5.5. Måløy confirmed the full chain worked with modified instructions on GPT-5.6 the following day.
## Affected Products
- Edit with Copilot remains in worldwide rollout; affected users are those with eligible licenses where the feature is enabled
Scenarios not affected: read-only Copilot summarization where no new document is generated, and manual document editing without a Copilot drafting or editing operation.
## Mitigations
There is no complete customer-side fix. Microsoft's deployed mitigations reduce but do not eliminate the attack surface. Practical steps for organizations:
Ctrl+A, then inspect font color)Organizations in finance, legal, healthcare, and any sector where document figures carry regulatory or contractual weight should treat this as an elevated risk until Microsoft issues a comprehensive fix.
## References
---
## HackWire Analysis
The detail that's getting less attention than it deserves: Microsoft patched the exact payload, and Måløy broke the patch the next day. That's not a patch. That's a blocklist — and blocklists against prompt injection are a game of whack-a-mole that attackers win by definition, because natural language has infinite paraphrase depth and LLMs are specifically designed to interpret semantic intent, not match literal strings.
What Måløy has documented is a structural property of how retrieval-augmented generation works, not a one-off implementation bug. Copilot is architected to read documents and infer intent. Distinguishing "this text is data I should summarize" from "this text is an instruction I should follow" is a hard problem that the field has not solved — and Microsoft's own disclosure that jailbreak and cross-prompt injection classifiers "may not be available in every Copilot scenario" tells you the coverage isn't complete even on the defensive tooling that exists.
The propagation angle is what makes this more than a standard indirect prompt injection write-up. Most prior XPIA research shows a one-hop attack: malicious document manipulates one output. This demonstrates a multi-hop chain where each generation becomes a new carrier with no traceable link to the original source. In a large enterprise where Copilot is generating draft quarterly reports, board memos, and client deliverables — all grounded on internal documents that themselves were Copilot outputs — the provenance problem compounds quickly. A single poisoned file circulating in OneDrive could corrupt derived documents for weeks before anyone notices the figures don't reconcile.
Financial services, legal, and M&A environments running Microsoft 365 Copilot should freeze multi-hop document generation workflows — any scenario where a Copilot output becomes a source for a second Copilot session — until Microsoft issues a structural fix, not another prompt blocklist.
— HackWire Editorial
---
## Related Coverage