# When AI Agents Go Offensive: OpenAI and Anthropic Tested Their Models Against Real Targets


The labs building the most capable AI systems on the planet just confirmed something that should concentrate minds across the security industry: their agents didn't just probe theoretical vulnerabilities in isolated sandboxes — they went after real infrastructure and real people.


Reports surfacing this week reveal that both OpenAI and Anthropic conducted cybersecurity evaluations in which AI agents were directed at live systems and actual individuals. The tests were framed as safety assessments — the kind of red-teaming both companies tout as responsible development. But the distinction between "we're testing whether our model can do harm" and "our model is now doing harm to test it" turns out to be thinner than the press releases suggest.


## What "Real Systems" Actually Means


The phrase "targeted real people and systems" does a lot of work, and it's worth unpacking it carefully.


In the AI safety context, this almost certainly refers to controlled red-team exercises — not random attacks on the internet. What it implies, though, is that researchers gave agents access to real credentials, real network endpoints, and real human accounts rather than fully synthetic honeypots. The difference matters enormously for what the results actually tell you.


Synthetic environments have a well-documented problem: models trained on internet data recognize fake infrastructure. Researchers have been burned before by evaluations that showed "low risk" because the model correctly identified the test as a test. Real targets eliminate that confound. You get honest capability numbers. You also get genuine collateral exposure, which is presumably why the story broke the way it did.


What makes this significant isn't that the labs were reckless — the available evidence suggests these were methodical assessments with appropriate scope controls. The significance is the capability ceiling they were trying to measure and, apparently, found.


## The Dual-Use Trap AI Labs Can't Escape


There's an uncomfortable structural problem embedded in how frontier AI labs approach security evaluation. To determine whether your model is dangerous in the wrong hands, you have to demonstrate that it's dangerous. That demonstration produces exactly the kind of capability documentation that, if it leaked, would be more valuable to adversaries than almost anything else.


Anthropic's own published work on offensive cyber capabilities has been more transparent than most about this bind. Their evaluations have shown Claude models capable of meaningfully assisting with vulnerability research, exploit chaining, and reconnaissance. OpenAI has run similar exercises. The results inform safety guardrails — but they also constitute a detailed roadmap of what these systems can do when pointed at a target.


The real question hovering over these tests is: what happened when the agents got access? Reconnaissance only? Or did they move laterally, modify files, exfiltrate data? The answer shapes whether this is a story about capability benchmarking or about AI systems causing actual, if temporary, damage to real infrastructure in the name of safety.


## Prior Art: We've Seen This Pattern Before


Security researchers at academic institutions — Georgia Tech, Carnegie Mellon, UIUC — have been publishing work on autonomous vulnerability exploitation for years. The UIUC team's 2024 paper demonstrating GPT-4 autonomously exploiting one-day vulnerabilities was a milestone that shifted the conversation from "could AI do this?" to "AI already does this, here's the success rate."


What's different with OpenAI and Anthropic running these tests internally is the resource differential. Academic red teams operate with limited compute budgets and model access. The labs operate with the frontier models and the compute to run them at scale. If a university team with API-rate-limited access got meaningful exploitation results, what does a sustained internal evaluation with full model access produce?


The answer to that question is presumably sitting in internal safety reports that will not be published in full. That's not necessarily wrong — detailed capability disclosures can cause genuine harm — but it means the public is being asked to trust that the guardrails are calibrated correctly without being able to verify it.


## What Defenders Should Take From This


If the frontier models can meaningfully operate against real systems in controlled tests, the operational implication for defenders is specific: the threat model for AI-assisted attacks is no longer speculative.


Several things shift:


Phishing and social engineering at scale become qualitatively different problems. An agent that can profile a real person, identify their role, synthesize contextually appropriate lures, and adapt based on responses transforms the economics of targeted attacks. Volume without quality was always detectable; quality without volume was expensive. Agents collapse that tradeoff.


Vulnerability discovery pipelines accelerate. The human bottleneck in exploit development has never been identifying that a vulnerability class exists — it's been the labor-intensive work of finding specific exploitable instances. Agents that can read codebases, identify patterns, and generate proofs-of-concept compress that timeline.


Detection logic built for human-paced attacks may not hold. Security tooling calibrated for human operators making decisions at human speeds — the dwell time assumptions in behavioral analytics, the timing signatures SIEM rules look for — may not reliably flag agents operating faster and with less behavioral regularity.


---


## HackWire Analysis


What the mainstream coverage of these tests is missing: this isn't primarily a story about whether OpenAI and Anthropic were responsible or reckless. It's a story about the structural incentive problem baked into frontier AI development.


Both labs have staked their public legitimacy on being the "safe" AI companies — the ones doing the serious work of alignment research, red-teaming, and capability evaluation before deployment. That reputation requires demonstrating rigor, which means showing you've stress-tested your systems against realistic threats. But "realistic threat" in cybersecurity means real systems, real people, real network paths. The moment you introduce that realism, you've crossed from simulation into actual engagement.


This is the same bind that offensive security teams at large enterprises navigate every day — and they solve it with strict scoping agreements, legal frameworks, and defined rules of engagement. The question is whether AI lab evaluations have equivalent institutional infrastructure, and the honest answer is: we don't know, because those frameworks aren't public.


There's a deeper pattern here that deserves attention. The more capable the agents become, the more the safety case for deploying them requires demonstrating their offensive potential — which in turn creates detailed internal documentation of how to use these systems as weapons. That documentation doesn't disappear when the evaluation concludes. It lives on servers. It gets referenced in follow-up research. It becomes institutional knowledge.


The security industry spent the better part of a decade arguing about responsible disclosure timelines for software vulnerabilities. The argument was fundamentally about who gets to know how dangerous something is, and when. The same argument is coming for AI capability evaluations, and it's going to be considerably more complicated because the "vulnerability" isn't in a codebase — it's the capability itself.


Defenders who are waiting for a public disclosure with full technical details before adjusting their threat models are, in this case, waiting for something that won't come. The operational posture shift needs to happen now, based on what we know: capable agents have been pointed at real infrastructure by the people with the most incentive to understate the results, and the results were apparently interesting enough to generate news coverage.


That's the tell.


— HackWire Editorial


---


## Related Coverage


  • Read more in our [Vulnerabilities](https://www.hackwire.news/category/vulnerabilities) coverage
  • Cross-reference with [Breaches](https://www.hackwire.news/category/breaches) and [Malware](https://www.hackwire.news/category/malware)
  • Stay current via the [HackWire homepage](https://www.hackwire.news/)