# Anthropic Disputes Fable 5 Jailbreak Claims; Researcher Counters With System Prompt Release


Anthropic has formally disputed claims that its newly released Fable 5 AI model has been successfully jailbroken, sparking a fresh debate over the definition and feasibility of prompt-based AI exploits. The dispute centers on whether a demonstrated approach constitutes a genuine bypass of the company's safety systems or merely represents a well-known conversational limitation inherent to large language models.


## The Claim and the Researcher


On June 11, 2026—just one day after Fable 5's general availability launch—a researcher operating under the moniker Pliny the Liberator published claims of a successful jailbreak on X (formerly Twitter). Pliny, known within security circles for previous AI jailbreak research, alleged that sophisticated multi-agent prompting techniques could circumvent Fable 5's restrictive safety layer, resulting in the model providing useful information on sensitive topics including:


  • Advanced cybersecurity exploits and attack vectors
  • Chemical synthesis and weapons development
  • Psychological manipulation techniques
  • Explosives design and manufacturing

  • To support the claims, Pliny released several screenshots demonstrating model outputs on these restricted topics and published what they alleged to be Fable 5's internal system prompt—a detailed set of instructions governing the model's personality, safety classifiers, fallback behaviors, tone guidelines, and refusal logic.


    The release of the internal prompt, if authentic, would represent a significant information disclosure that could inform future jailbreak attempts across the entire AI research community.


    ## Anthropic's Response and Technical Defense


    Contacted by SecurityWeek and other outlets, Anthropic issued a statement defending the robustness of Fable 5's safety architecture. The company made several key technical claims:


    ### Core Arguments


    Not a True Jailbreak: Anthropic argues that genuine jailbreaks must bypass the model's *core safeguards*—specifically, independent classifier systems that operate separately from the base model itself. Pliny's demonstrated approach, according to the company, merely coaxes continued responses despite the model's conversational refusals, a limitation present in nearly all contemporary large language models.


    Architectural Isolation: The company emphasized that its strongest protections against the most dangerous risks are enforced by independent classifiers that function as a separate enforcement layer. Overcoming the model's conversational refusals—telling the model to continue despite initial rejection—does not disable these critical safeguards.


    Content Analysis: Upon examination, Anthropic determined that:

  • Some outputs attributed to Fable 5 were not actually produced by the model
  • Outputs that were genuine contained only general information already available in public sources
  • None of the demonstrated examples provided "meaningful uplift" for real-world harm

  • No Evidence of Successful Bypass: A broader review of Fable 5 usage data found no evidence of attackers successfully circumventing the safeguards to generate genuinely dangerous content.


    ## Background: Fable 5 and Its Safeguards


    To understand the significance of this dispute, it's important to contextualize Fable 5's design and deployment strategy.


    Anthropic positioned Fable 5 as a Mythos-class AI model—its most powerful to date—designed to excel across a wide range of general-purpose tasks. However, recognizing the potential for misuse in high-risk domains, the company implemented a tiered safety architecture:


    | Domain | Behavior |

    |--------|----------|

    | Unrestricted use (general queries, coding, analysis) | Full Fable 5 capabilities |

    | Restricted domains (cybersecurity, bioweapons, chemistry) | Automatic fallback to Claude Opus 4.8 (less capable) |

    | Dangerous assistance requests (detailed exploit code, bioweapon synthesis) | Refusal via independent classifier systems |


    This approach represents Anthropic's philosophy that safety should be enforced through *multiple independent mechanisms* rather than relying on a single defense layer.


    ## The Technical Details: What Constitutes a "Jailbreak"?


    The dispute between Anthropic and Pliny the Liberator hinges on how security researchers and AI companies define a successful jailbreak. This definitional question carries significant implications for the AI safety community.


    ### Conversational Refusal vs. Hardened Safeguards


    Most large language models exhibit a behavior called "token steering" or "conversational refusal." When prompted directly to generate harmful content, models will often refuse—but if a user rephrases the request, adds context, or uses multi-turn reasoning chains, the model may relent and continue responding.


    Anthropic's position is that Pliny's demonstrated techniques exploit this conversational limitation, which the company views as an inherent weakness in how language models process requests—not a bypass of true safety infrastructure.


    Pliny's counter-argument (implied in the release of evidence and the system prompt) is that if a method successfully extracts sensitive information, the distinction between "conversational refusal" and "safety bypass" may be academic; the practical result is access to restricted information.


    ### The System Prompt Disclosure


    The most concerning aspect of Pliny's release is the alleged disclosure of Fable 5's internal system prompt. If authentic, this document reveals:


  • The exact refusal patterns the model uses
  • How the model classifies sensitive requests
  • The fallback logic when restricted domains are invoked
  • Tone and personality guidelines that may create exploitable patterns

  • Security researchers will likely use this information to develop more sophisticated multi-agent prompting techniques, potentially informing future jailbreak attempts against Fable 5 and competing models.


    ## Implications for AI Security and Safety


    This incident reflects a broader tension in AI safety research:


    Responsible Disclosure vs. Coordinated Vulnerability Research: Pliny published findings and evidence immediately and publicly, without coordinating with Anthropic beforehand. This approach differs from traditional cybersecurity vulnerability disclosure, where researchers often notify vendors before public release to allow time for fixes.


    Red Teaming Effectiveness: Anthropic stated that Fable 5 underwent "extensive internal and external red-teaming" before launch. The rapid emergence of jailbreak claims raises questions about whether red-teaming practices are keeping pace with the speed and creativity of independent security researchers.


    Model Transparency Trade-offs: The release of the system prompt illustrates a critical tension: transparency (letting researchers understand how models work) can inadvertently enable exploitation.


    ## Industry Response and Monitoring


    Major AI companies have been closely monitoring the Pliny claims and Anthropic's response:


  • OpenAI and Google DeepMind are likely reviewing their own safety architectures to determine if similar multi-agent prompting techniques could circumvent their models
  • Regulatory bodies are watching to assess whether current AI safety standards are sufficient
  • AI safety researchers are analyzing the technical details to develop improved defense mechanisms

  • ## Recommendations for Organizations


    Organizations relying on Anthropic's models should take the following steps:


    1. Monitor Official Communications: Track Anthropic's official security advisories and updates regarding Fable 5's safety systems

    2. Test Fallback Behavior: If using Fable 5 in sensitive contexts, verify that restricted domains correctly trigger fallback to less capable models

    3. Implement Input Controls: Screen prompts for obvious attempts to bypass safety systems, particularly multi-agent or role-playing techniques

    4. Audit Outputs: For high-risk use cases (security research, sensitive analysis), manually review model outputs before relying on them

    5. Stay Informed: Monitor both vendor announcements and independent research findings on AI jailbreaks


    ---


    ## HackWire Analysis


    The Pliny-vs-Anthropic dispute is less about binary success or failure and more about exposing how AI companies and independent researchers measure safety differently. Anthropic's framing—that this isn't a "true" jailbreak because it doesn't defeat independent classifiers—reveals an important but subtle distinction. In traditional cybersecurity, a vulnerability is a vulnerability if it produces unauthorized access. But in AI safety, companies are now arguing that extracting harmful information through conversational techniques *doesn't count* if independent detection systems haven't fired.


    This semantic shift matters because it affects how we evaluate AI safety claims going forward. If Anthropic is correct that their classifier systems are the real guarantee, then a jailbreak that fools the conversational layer is less critical than it initially appears. But if Pliny is right that practical access to restricted information is the defining issue, then Anthropic may be underestimating the threat.


    The most revealing detail: Pliny released Fable 5's system prompt. Within cybersecurity, this is roughly equivalent to publishing source code or architectural diagrams. Anthropic's response—stating that some outputs weren't from Fable 5, and others were just public knowledge—feels defensive. If the system prompt is authentic, the research community now has a detailed blueprint for crafting future exploits. The question isn't whether today's jailbreak is "real enough"—it's whether the prompt release will accelerate the development of ones that are.


    For defenders, this underscores a critical reality: multi-agent prompting and role-play chains are now table-stakes attack techniques against any AI system. Organizations cannot assume safety layers alone will protect them; input filtering and output auditing are mandatory for high-risk deployments. — HackWire Editorial


    ---


    ## Related Coverage


  • Read more in our [Vulnerabilities](https://www.hackwire.news/category/vulnerabilities) coverage
  • Cross-reference with [Malware](https://www.hackwire.news/category/malware) and [Threat Intelligence](https://www.hackwire.news/category/threat-intelligence)
  • Stay current via the [HackWire homepage](https://www.hackwire.news/)