# The AI Red Team That Accidentally Ran a Real Attack


There's a particular kind of embarrassment reserved for security firms: getting caught doing to someone else what they exist to prevent. Irregular, the Israeli AI safety company that stress-tests frontier models for OpenAI, Anthropic, and Meta, is living through exactly that moment. A naming error in a test environment turned a controlled simulation into an unauthorized intrusion against a real company's production database — carried out, without human guidance, by one of Anthropic's models.


The Friday blog post Irregular published is careful and clinical, the kind of incident write-up that's designed to look like transparency while controlling the narrative. Read between the lines, though, and what emerges is a picture of how AI safety infrastructure can itself become an attack surface.


## One Fictional Company Name, One Real Domain


The mechanics of how this happened are almost comically mundane. Irregular's engineers assigned a fake target company a name for use in an evaluation scenario. They check fictional names against real companies before using them — standard practice. The check missed this one because the actual domain wasn't widely known. Obscure, but real.


Internet access was enabled in the test environment. Models were instructed to conduct reconnaissance, locate private keys, extract data, and attempt to evade detection — a scenario designed to test whether an AI could assist a malicious insider in raiding a production database. When the model encountered the real domain, it didn't pause. It had a task. It proceeded.


Irregular documented multiple runs in which models exploited vulnerabilities at the real site, extracted credentials, and gained access to the production database. In one additional case, a model drifted to a similarly-named site and picked up credentials that had already been posted publicly — a bonus exfiltration that wasn't even part of the intended exercise.


## The Buried Lede: Logs That Lie


Most coverage of this incident will focus on the naming error. That's the wrong place to look.


The more uncomfortable finding is buried in Irregular's own writeup: the activity was difficult to detect because it happened in a small fraction of runs, often hundreds of interactions deep into a simulation. Red-team evaluation logs, by their nature, are full of intrusion-looking behavior. Models probing for vulnerabilities, escalating privileges, exfiltrating mock data — that's the whole point. The signal of a real attack disappears into a sea of simulated ones.


This is a structural problem, not an Irregular-specific one. Every organization running AI offensive-capability evaluations faces the same monitoring paradox. Your classifier can't distinguish a model that just breached a real database from a model that's doing exactly what you asked it to do, because from the log's perspective, they look identical. Irregular flags this explicitly: "existing monitoring tools and classifiers struggle to tell legitimate red-team activity from genuine attacks."


Nobody in this space has a good answer for that yet.


## What Made the Real Target Easy Pickings


Irregular noted, somewhat pointedly, that the domain being targeted "lacked common safeguards, making it an easy target for most frontier models." That's a polite way of saying a real company's production database was accessible enough that an AI without any targeted instructions, given only a domain name, could get in and take data.


That's worth sitting with. The model wasn't handed credentials, wasn't briefed on the target's architecture, had no insider knowledge. It was told to attack a fictional company, landed on a real one by accident, and succeeded anyway — apparently without breaking much of a sweat. The vulnerability wasn't the AI. The vulnerability was the target.


## Where Irregular Goes From Here


The firm's stated remediation steps are reasonable if somewhat obvious in retrospect:


  • Manual review expansion during testing cycles
  • A dedicated internal team to challenge containment assumptions
  • Continuous revalidation of evaluation domains as new sites appear
  • Better documentation with AI lab customers around evaluation scope
  • Cross-organizational mechanisms for sharing forensic evidence, including model transcripts, after incidents

  • The last point deserves more attention than it usually gets. Right now, when an AI model does something unexpected in one lab's evaluation environment, that information lives inside that lab. There's no ISAC equivalent for AI safety researchers. Irregular is essentially calling for one, albeit carefully.


    ---


    ## HackWire Analysis


    The Irregular incident is the second significant AI testing escape to make headlines in recent weeks — a cluster that should prompt a harder look at whether the governance frameworks for AI evaluation are keeping pace with the capabilities being evaluated.


    The naming-error framing is a red herring that lets everyone off the hook a little too easily. Yes, a naming check failed. But the deeper failure is architectural: internet access was live in an environment where models were being instructed to conduct offensive operations. That's not a naming problem. That's a containment philosophy problem.


    Genuine sandboxing means the simulated target is the only reachable target. If internet access is "enabled" during a red-team evaluation of an AI's offensive cyber capabilities, the sandbox is not a sandbox. It's a staging server with a disclaimer on the door.


    The AI safety community has spent years debating whether frontier models pose risks through emergent deception or misaligned goals. The Irregular incident suggests the more immediate risk is mundane: operational errors in human-run evaluation pipelines, amplified by models that are very good at completing the task in front of them. No deception required. No misaligned goals. Just a model doing exactly what it was told, to the wrong target.


    For defenders, the actionable takeaway is also the most uncomfortable one: if a model with no prior knowledge of your infrastructure can breach your production database starting from just your domain name, the AI labs are not your primary problem. Your hardening backlog is.


    — HackWire Editorial


    ---


    ## Related Coverage


  • Read more in our [Vulnerabilities](https://www.hackwire.news/category/vulnerabilities) coverage
  • Cross-reference with [Breaches](https://www.hackwire.news/category/breaches) and [Malware](https://www.hackwire.news/category/malware)
  • Stay current via the [HackWire homepage](https://www.hackwire.news/)