# Anthropic's Own AI Got Used to Hack Three Organizations — And It Took an OpenAI Disclosure to Find Out


When OpenAI published findings earlier this year that its models had been abused by threat actors to conduct cyberattacks, it triggered something the AI industry rarely does voluntarily: self-examination. Anthropic looked inward, and what it found was not reassuring. Claude — the model Anthropic bills as safer and more aligned — had been used to breach three organizations.


The admission is significant on multiple levels. First, because Anthropic found it only after OpenAI's disclosure prompted the audit. Second, because it confirms what security researchers have argued for over a year: the AI safety versus AI capability debate is largely beside the point when sophisticated attackers can operationalize any sufficiently capable model as an offensive tool.


## The Disclosure That Started It All


OpenAI's earlier transparency report on threat actor misuse set a template. State-sponsored groups and financially motivated threat actors had used ChatGPT and related tooling to assist with reconnaissance, spearphishing content generation, and vulnerability research. OpenAI named specific threat clusters — groups linked to China, Iran, North Korea, and Russia were all identified in the report.


The disclosure created competitive pressure on the rest of the AI industry to look at their own telemetry. Anthropic's response reveals what that scrutiny turns up when a company actually does the work: real intrusions, real victims, and a gap between how the company publicly frames its safety posture and what its models are being used for in the wild.


Three organizations hacked via Claude. Anthropic has not named them, nor detailed the attack chains involved. But the pattern suggests attackers are not using AI models as standalone exploit frameworks — they're using them as force multipliers in multi-stage operations.


## What "Used to Hack" Actually Means


It's worth being precise about what AI-assisted hacking looks like in practice, because the threat model matters for defenders.


Attackers aren't running ask_claude("give me a 0-day") and catching root shells. What they are doing is more dangerous in aggregate:


  • Social engineering at scale. LLMs dramatically lower the cost of producing convincing, contextually appropriate phishing lures in multiple languages and targeted to specific individuals. A Claude-generated spearphish targeting a CFO reads nothing like the generic credential-harvest emails most security training is built around.
  • Code assistance. Attackers use AI models to modify existing malware, write custom loaders, or adapt public PoC exploits to specific environments — compressing work that might take a junior developer days into hours.
  • Reconnaissance synthesis. LLMs can process large quantities of OSINT data — LinkedIn profiles, job postings, breach data, public GitHub repos — and synthesize it into actionable attack surface maps.
  • Defensive evasion logic. There's documented use of AI to suggest EDR bypass techniques, obfuscation approaches, and living-off-the-land patterns tailored to specific target environments.

  • None of this requires jailbreaking in the traditional sense. Sufficiently motivated and patient attackers iterate around guardrails, or they simply use the model's legitimate capabilities in an illegitimate context.


    ## The Safety-Washing Problem


    Anthropic occupies an unusual position in the AI landscape. The company was founded explicitly on AI safety principles, has published extensive alignment research, and markets Claude as constitutionally constrained in ways its competitors are not. That positioning creates a particular kind of credibility exposure when the model turns out to be instrumentalized in real-world attacks.


    This is not a criticism of Anthropic's alignment work, which is genuine and technically serious. It's a structural critique of how the industry has framed the AI safety conversation. Alignment research addresses whether a model pursues harmful goals autonomously. It says much less about whether a model can be used by humans to pursue harmful goals.


    Those are different problems. The security community has been pointing this out since ChatGPT launched. The Anthropic disclosure is the industry finally catching up to what red teamers already knew.


    ## What Defenders Are Up Against


    The three-organization disclosure raises an uncomfortable question: how many organizations are in the same position and don't know it yet? Attributing an attack to "AI-assisted" methods is genuinely difficult. Log telemetry doesn't flag "this spearphish was written by Claude." Behavioral indicators associated with AI-assisted attacks — higher-quality lures, faster iteration on failed techniques, polymorphic code variants — overlap heavily with indicators of a skilled human attacker.


    Detection strategies need to evolve:


  • Email security teams should treat grammatical and contextual quality as a threat signal, not a safety signal. The badly-written phish is yesterday's threat. Train on what well-written looks like.
  • Threat intelligence programs need explicit tasking to track AI tooling as part of threat actor TTPs. This isn't a niche concern — it's now an operational variable in adversary capability assessments.
  • Incident responders should be asking, in every engagement, whether AI tooling could explain anomalies they're observing: unusual lateral movement speed, unusually convincing impersonation, or malware with atypical code quality.

  • ---


    ## HackWire Analysis


    The detail that Anthropic only found these intrusions because OpenAI's disclosure prompted an internal audit is the most important sentence in this story — and it's getting buried.


    AI companies are not proactively monitoring for offensive misuse at a level commensurate with their deployment scale. They are reactive to peer disclosure. That's a structural problem, not an Anthropic-specific one.


    Compare this to how financial institutions handle fraud. Banks don't wait for a competitor's fraud disclosure to audit their own transaction patterns. They run continuous anomaly detection, maintain dedicated fraud intelligence teams, and share threat indicators through industry groups like the Financial Services ISAC. The AI industry has nothing comparable. There's no AI ISAC. There's no standardized adversarial use reporting. There's no sector-wide telemetry sharing on how models are being weaponized.


    The OpenAI and Anthropic disclosures both describe the threat in the past tense — "we found this happened." What's missing is any mechanism that would surface active misuse in near-real-time and route that intelligence to defenders whose organizations are currently being targeted.


    Three companies got hacked. Anthropic knows who. Presumably those companies have been notified — but the broader security community hasn't seen indicators of compromise, attack patterns, or the threat cluster responsible. That intelligence gap benefits attackers.


    The AI industry is going to face escalating regulatory pressure on this front. The EU AI Act has provisions touching on this. CISA has been signaling interest in AI security obligations for critical infrastructure. Companies that build proactive misuse detection and voluntary disclosure frameworks now will be ahead of mandates that are likely coming within the next 24 months. Those that keep treating disclosure as a PR event triggered by competitor announcements will not be.


    — HackWire Editorial


    ---


    ## Related Coverage


  • Read more in our [Breaches](https://www.hackwire.news/category/breaches) coverage
  • Cross-reference with [Vulnerabilities](https://www.hackwire.news/category/vulnerabilities) and [Malware](https://www.hackwire.news/category/malware)
  • Stay current via the [HackWire homepage](https://www.hackwire.news/)