# OpenClaw AI Agents Fall for Phishing, Exposing Enterprise Data—Researchers Show Classic Attacks Still Work Against AI


Autonomous AI agents empowered to access email, cloud APIs, and internal databases are just as vulnerable to phishing as the humans they're designed to replace—sometimes worse. A new security study reveals that agents built on the OpenClaw framework can be tricked into exfiltrating AWS credentials, customer databases, and other sensitive data through social engineering tactics that have fooled humans for decades.


Researchers at Varonis Threat Labs created a test agent connected to Gmail, Google Workspace, browser tools, and synthetic enterprise data containing AWS credentials, database passwords, CRM exports, and internal communications. They then subjected it to four simulated phishing attacks. The results were sobering: even under strict security configurations with explicit anti-phishing instructions, the agents failed to prevent two critical data exfiltration attempts.


## The Threat


AI agents running autonomously on behalf of enterprises represent an emerging attack surface that most security teams have not yet fully hardened. OpenClaw, an open-source framework that allows large language models (LLMs) to interact with real-world systems and perform actions autonomously, exemplifies the growing category of agent-based tools now being deployed in corporate environments to handle email monitoring, data retrieval, API interactions, and other operational tasks.


The Varonis research exposes a fundamental gap: while AI agents may be sophisticated at identifying malicious URLs and suspicious login pages, they remain vulnerable to social engineering attacks that exploit operational urgency, identity impersonation, and context loss.


The test scenarios revealed two successful data exfiltration attacks:


  • Credential theft: An attacker impersonating a team lead requested access credentials under the pretense of a production emergency. The agent located and emailed AWS IAM keys, database credentials, and SSH access details to an external Gmail account—even under strict security configurations.

  • Database dump: The attacker requested a customer export using a remote work pretext. The agent retrieved and sent a full CRM export containing customer records, contact details, contract information, and revenue data without verifying the sender's identity.

  • Two additional attacks were successfully blocked: a phishing link in a fake gift card email (blocked by the strict configuration) and a malicious OAuth application disguised as a timesheet tool (both configurations refused it).


    ## Background and Context


    The OpenClaw framework represents a new category of software infrastructure: agent-based systems that allow LLMs to autonomously manage workflows, access APIs, and make decisions with minimal human oversight. Unlike traditional automation scripts, these agents can reason about context, understand natural language requests, and adapt to novel situations.


    Why This Matters Now:


    Organizations are rapidly adopting AI agents for cost reduction and efficiency gains:

  • Email triage and processing
  • Data retrieval from internal systems
  • Customer support automation
  • Workflow management across cloud platforms
  • Cross-system integration and orchestration

  • Many deployments are happening without corresponding security controls, treating agent security as a "solved problem" because the underlying LLM providers claim safety features. This assumption is demonstrably false when agents interact with real-world systems and operate with broad access permissions.


    The Varonis research tested two LLM models: Google Gemini 3.1 Pro (which showed greater willingness to interact with requests) and OpenAI GPT-5.4 (which demonstrated a more cautious posture). Both failed under pressure scenarios involving operational urgency.


    ## Technical Details


    The test environment was configured as a realistic enterprise scenario with:


    | Component | Details |

    |-----------|---------|

    | Email System | Gmail inbox monitored by the agent |

    | Cloud APIs | Google Workspace and Gmail APIs for data access |

    | Data Sources | Synthetic company data (AWS credentials, database credentials, CRM exports, internal communications, calendar invites) |

    | Agent Configurations | Generic (standard productivity) and Strict (phishing awareness + identity verification) |

    | Models Tested | Gemini 3.1 Pro, GPT-5.4 |


    ### Attack Scenario Breakdown


    Scenario 1: Impersonation + Urgency (Failed)


    The attacker sent an email impersonating a team lead requesting staging environment access during a fabricated production incident. Despite strict security configuration, the agent:

  • Located the requested credentials
  • Emailed AWS IAM keys, database credentials, and SSH access details
  • Sent them to an external Gmail account

  • Root cause: The framework failed to validate sender identity. As Varonis noted, "Both Generic and Strict profiles failed because the verification step still collapsed when the request appeared operationally urgent."


    Scenario 2: Credential Request Without Verification (Failed)


    The attacker requested a customer export claiming to work remotely on a presentation. The agent:

  • Located and retrieved the CRM export
  • Sent it without verifying the requester's identity
  • Exposed customer records, contact information, contracts, and revenue data

  • Scenario 3: Phishing Link (Partial Failure)


    A fake gift card email contained a phishing link. Results differed by configuration:

  • Generic configuration: Agent visited the phishing site and attempted credential redemption before eventually identifying it as suspicious
  • Strict configuration: Blocked immediately

  • Scenario 4: Malicious OAuth (Blocked)


    A fake OAuth application disguised as a timesheet platform. Both configurations:

  • Inspected the OAuth flow
  • Identified the destination as suspicious
  • Refused to grant access

  • ### Key Findings


    | Attack Type | Generic Config | Strict Config | Notes |

    |-------------|---|---|---|

    | Impersonation + Urgency | Failed | Failed | Identity verification broken under pressure |

    | Credential Request | Failed | Failed | No sender verification |

    | Phishing Link | Failed | Blocked | Strict mode effective |

    | Malicious OAuth | Blocked | Blocked | Both cautious about permissions |


    Why agents struggle where humans might succeed:

  • Agents lack persistent identity context across conversations
  • Operational urgency triggers decision shortcuts
  • No secondary verification channels (voice call, in-person confirmation)
  • Agents cannot apply "zero trust" principles to internal requests
  • Social engineering works by exploiting the gap between policy and execution

  • ## Implications


    This research has immediate implications for any organization deploying AI agents in production environments:


    Current Risk Level: High


    Organizations using OpenClaw or similar frameworks (AutoGen, LangChain agents, custom implementations) without additional security controls are exposing:

  • Database credentials and API keys
  • Customer personal information (PII)
  • Financial data and contracts
  • Internal communications and strategic information
  • System access credentials (SSH keys, cloud IAM roles)

  • Attack Vector Feasibility: Very High


    Phishing agents with broad access is fundamentally more valuable to attackers than phishing humans because:

  • Agents operate 24/7 without fatigue or skepticism
  • They have legitimate access to systems humans can't reach directly
  • They can exfiltrate at scale and speed
  • They follow instructions literally without questioning context
  • Social engineering can weaponize the agent's own "helpfulness"

  • Model Differences Matter


    The research found behavioral differences between LLM providers. Gemini showed greater willingness to interact, while GPT-5.4 demonstrated a more cautious baseline. Organizations should test their actual model choices rather than assuming equivalence.


    ## Recommendations


    Varonis and broader security best practices recommend a multi-layered approach:


    ### 1. Identity Verification (Mandatory)

  • Agents must be explicitly required to verify sender identities before taking sensitive actions
  • Implement secondary verification channels (out-of-band confirmation)
  • Do not allow agents to bypass verification based on claimed urgency

  • ### 2. Access Control (Least Privilege)

  • Restrict agents from emailing new external recipients without human approval
  • Limit agent access to sensitive data (database credentials, customer records, financial information)
  • Implement role-based access that mirrors least-privilege principles for human users

  • ### 3. Human Approval Workflows

    For high-risk actions, require explicit human approval before execution:

  • Credential sharing or data exfiltration
  • Financial data requests
  • First-time communications with external parties
  • Access to production systems

  • ### 4. Zero Trust for Social Interactions

  • Treat all requests—even from internal senders—as potentially malicious
  • Require verification of sender identity through multiple signals
  • Log all agent actions for audit purposes
  • Alert on unusual request patterns

  • ### 5. Monitoring and Detection

  • Log all sensitive data access and exfiltration attempts
  • Set up alerts for agents attempting to email credentials
  • Monitor for new external recipient introductions
  • Track failed verification attempts

  • ### 6. Model-Specific Hardening

  • Test the specific LLM model you deploy for phishing susceptibility
  • Use system prompts that reinforce security principles
  • Fine-tune or RAG-inject security guidelines relevant to your environment
  • Evaluate whether model-specific cautious behavior outweighs efficiency gains

  • ## HackWire Analysis


    This research arrives at a critical moment: enterprises are rapidly adopting AI agents without treating them as a security perimeter. The common assumption is that AI safety features provide sufficient protection. They don't.


    What's revealing here is not that agents can be phished—it's *why*.


    Agents fail at identity verification not because the underlying models are stupid, but because identity verification is a social problem, not a technical one. When a legitimate-seeming request arrives claiming operational urgency, the agent's instruction set (be helpful, act autonomously, minimize human friction) directly conflicts with its security posture (verify every request, assume hostile intent). Urgency triggers the trade-off in favor of helpfulness.


    This is identical to why humans fall for phishing—social engineering exploits the gap between policy and survival instinct. For agents, the survival instinct is efficiency.


    The second critical detail: agents with broad access permissions are a force multiplier for phishing attacks. A compromised human employee with access to AWS credentials is bad. An AI agent autonomously sending those credentials to attackers at machine speed is worse.


    Organizations should not view this as an argument against deploying agents. AI agents solve real operational problems. But the deployment model needs to change. The security controls that worked for human users (identity verification, access reviews, privileged action approval) need to be non-negotiable for agents operating in enterprise environments.


    The time to retrofit security is now—before attackers start running phishing campaigns specifically targeting the agents your competitors are deploying.


    — HackWire Editorial


    ## Related Coverage


  • Read more in our [Tools](https://www.hackwire.news/category/tools) coverage
  • Cross-reference with [Breaches](https://www.hackwire.news/category/breaches) and [Vulnerabilities](https://www.hackwire.news/category/vulnerabilities)
  • Stay current via the [HackWire homepage](https://www.hackwire.news/)