# When Your AI Gets Hacked, Nobody Tells Anyone — The SAFE Initiative Wants to Change That
The history of cybersecurity is largely a history of organizations staying quiet about how they got compromised. The breach disclosure movement took decades of regulatory pressure, high-profile disasters, and persistent advocacy before companies would reliably admit when customer data walked out the door. Now, with AI systems embedded in critical infrastructure, financial services, and healthcare, the industry is staring down the same problem — and a coalition of security organizations is trying to get ahead of it before the failures pile up uncatalogued.
The Cybersecurity Alliance's draft SAFE (Secure AI Failure Exchange) guidelines represent one of the first serious attempts to create a standardized framework for how organizations report, document, and share information about AI-specific security incidents. The timing is deliberate. AI deployment has moved faster than AI incident response doctrine, and the gap is showing.
## What Makes an AI Incident Different
The problem with applying traditional incident response frameworks to AI failures is that they weren't built for this. A SQL injection attack has a CVE number. There are patches. The blast radius is bounded by what data the database touched. An AI incident is structurally different in almost every dimension.
When an adversary poisons a model's training data, the "vulnerability" isn't a discrete code path — it's embedded in millions of parameters. When a language model is jailbroken to exfiltrate sensitive information or generate harmful outputs, there's no obvious firewall log entry flagging the attack. When a fraud detection model is gamed through adversarial inputs, the organization may not know for months. Traditional SIEM tools, built around discrete events with timestamps and signatures, are often blind to these failure modes.
The SAFE guidelines attempt to give the security community a shared vocabulary for these incidents. Early drafts reportedly include taxonomies for model poisoning, prompt injection at scale, adversarial evasion in production ML systems, and hallucination exploitation — where bad actors deliberately craft queries to elicit confidently wrong outputs in high-stakes contexts.
## The Disclosure Problem Nobody Wants to Say Out Loud
Here's the tension that traditional threat intelligence sharing doesn't face in quite the same way: when you share that your AI system was compromised, you're often revealing the architecture of your model, your training data sources, and your deployment constraints. That's genuinely proprietary. A bank that shares a detailed post-mortem on how its fraud detection model was evaded is handing competitors a roadmap to its entire ML stack.
The ISACs — the Information Sharing and Analysis Centers that have operated as sectoral threat intel clearing houses since the late 1990s — solved the analog problem for traditional cyber threats through anonymization and trust relationships. You share indicators of compromise with your sector peers without necessarily attaching your company name to them. The SAFE framework is trying to extend that model to AI incidents, which is harder because the "indicators" of an AI attack can be inseparable from the system's design.
The draft guidelines propose tiered disclosure levels: technical details about the attack vector (shareable with sector peers), operational context about what the AI system was doing and what the failure looked like to users (shareable in anonymized form), and model-specific architecture details (withheld from general sharing, available only to vetted researchers under specific agreements). Whether that tiering works in practice depends entirely on the trust relationships that organizations are willing to build — and that's where most of these initiatives have historically stalled.
## Regulatory Pressure Is Already Here
The SAFE initiative isn't emerging in a vacuum. The EU AI Act includes mandatory incident reporting requirements for high-risk AI systems, with the first obligations kicking in for providers and deployers in 2025. NIST's AI Risk Management Framework has been encouraging voluntary disclosure practices. CISA has quietly been building an AI Security Incident Taxonomy working group. The UK's AI Safety Institute has been collecting voluntary incident reports since it stood up in 2024.
What's missing from all of these parallel efforts is interoperability. An AI incident reported to CISA under their taxonomy doesn't easily map to an EU AI Act incident report, which doesn't map to what a financial sector ISAC would collect. The SAFE guidelines appear to be an attempt at building a lingua franca — a common structure that could sit underneath multiple regulatory reporting regimes without requiring organizations to maintain entirely separate incident documentation processes.
That's ambitious. The security industry's track record with interoperability standards is checkered at best. STIX and TAXII were supposed to create a universal threat intelligence exchange protocol; adoption has been partial and inconsistent a decade later. SAFE will face similar headwinds.
## Who Actually Gets Attacked — And Who Stays Quiet
The incidents that the SAFE framework is designed to capture are already happening. Researchers at universities and commercial security firms have documented adversarial attacks against production ML systems in financial fraud detection, content moderation, medical imaging diagnostics, and autonomous vehicle systems. What's largely absent is any industry-wide understanding of how often these attacks succeed in production, what the blast radius looks like, or how defenders discovered them.
Organizations that have had AI systems compromised have strong incentives to stay quiet. There's no established legal obligation to disclose (outside specific EU jurisdictions), the reputational damage of admitting your AI model was gamed is significant, and the technical complexity of what happened makes explanation to non-technical stakeholders difficult. The result is that the defender community is essentially flying blind on the actual threat landscape.
---
## HackWire Analysis
The SAFE initiative lands at a moment when the cybersecurity industry is still figuring out what "AI security" actually means as a professional discipline. Is it securing AI systems against attack? Securing other systems using AI as a tool? Using AI to attack? The answer is all three, and they're colliding at the same time.
What the SAFE framework gets right is recognizing that AI incident sharing needs its own category. The instinct to adapt existing frameworks — "just use STIX" or "just file a CISA report" — misses what makes AI incidents structurally unusual. The attack surface is the model itself, not just the infrastructure around it. The failure modes are probabilistic, not deterministic. The incident timeline is often retrospective — you discover you were attacked weeks or months after the fact, when patterns emerge in outputs.
The deeper problem is that organizations don't have good internal language for these incidents yet, let alone external reporting language. Security teams trained on traditional threat models are still building intuition for what adversarial ML attacks look like from the inside. The SAFE taxonomy may do as much for internal clarity as it does for sector-wide sharing.
The comparison to early breach disclosure is apt and sobering. That fight took roughly fifteen years from initial advocacy to something resembling normalized practice, and it still required regulatory mandates to become consistent. AI incident sharing could move faster — the regulatory pressure is already arriving, and the AI safety community is more organizationally coherent than the early data breach advocates were. But "faster than fifteen years" still means we're in the early chapters.
For defenders working with production AI systems today: document your incidents now, even if you're not sharing them externally. Build the internal taxonomy. When sharing frameworks like SAFE mature, you'll want incident records that fit the structure. Organizations that have been quietly cataloguing their AI failures will be the ones with useful data to contribute — and useful intelligence to receive.
— HackWire Editorial
---
## Related Coverage