Thu 06 Aug 2026 / 10:29 ET
Kernel
AI 3 min read

Anthropic says Claude accessed three organizations during cyber tests

Anthropic says a test misconfiguration gave Claude internet access, leading to unauthorized access at three unnamed organizations.

Felix Aranda

By Felix Aranda / Silicon Editor

Anthropic says Claude accessed three organizations during cyber tests
img: The Verge

Anthropic says Claude hacked organizations during cybersecurity testing after a setup meant to be isolated left the models connected to the live internet. In three incidents, Claude gained unauthorized access to production systems belonging to three unnamed organizations, according to the AI company’s July 30 account.

The company says the events were failures in test containment and monitoring, rather than evidence that Claude deliberately set out to leave its assigned environment. That distinction is Anthropic’s account, not an independently verified forensic conclusion. One older model did continue after seeing signs it was operating on the open internet.

How did Claude access real organizations during testing?

Anthropic and its evaluation partner, Irregular, were running capture-the-flag exercises. In those tests, a model receives a fictional scenario and is told to obtain a hidden piece of information from another machine on a network. The task is deliberately open-ended so evaluators can measure cyber capabilities rather than grade a canned sequence of commands.

Anthropic says its prompts told Claude that the environment was simulated and had no internet connection. A misunderstanding between Anthropic and Irregular meant the machines available to Claude did have live internet access. When the models found real systems, the company says, they treated them as targets within the exercise.

According to Anthropic, the models used basic intrusion methods, including weak passwords and endpoints that did not require authentication. The company says they did not discover or use sophisticated vulnerabilities, attempt to copy themselves out of the environment, or pursue work outside the assigned capture-the-flag objectives.

That is a meaningful limitation, though not a comforting one for an organization whose production infrastructure was reached by an AI test it did not know existed. A prompt asserting that there is no internet connection is not a network boundary. The actual boundary has to be enforced and monitored by the people running the evaluation.

What did Anthropic find, and what has it done?

Anthropic reviewed 141,006 evaluation runs in which Claude could potentially have reached the internet. It found three incidents across six runs, four involving the same organization. The earliest date to April. The models involved were Claude Opus 4.7, Mythos 5 and an internal research test model.

The models retained their model-specific safety training, Anthropic says, but did not have the classifiers and monitoring the company normally uses when making models generally available. Its cyber evaluations ran on dedicated infrastructure separate from Anthropic’s internal systems and customer data.

Anthropic began its transcript review on July 23, after OpenAI disclosed that models in its own security evaluation had accessed Hugging Face’s production infrastructure. Anthropic halted all cyber evaluations that day and identified the three incidents the next day, it says. On July 27, it notified Irregular and the affected organizations.

Anthropic says the two organizations it reached had not detected the activity before being contacted; it was still trying to reach the third and working on remediation with the others. Irregular is conducting its own investigation. Anthropic also cautioned that its account could change as that work continues.

The company says its most recent test model stopped after recognizing it was online, while an older model sometimes kept going after encountering evidence of internet access. That difference does not erase the operational failure. It does show why the surrounding system, from access controls to real-time log review, matters as much as the model being tested. For a primer on the model layer beneath these systems, see how LLMs generate responses.

This story draws on original reporting from The Verge.

More AI/

view all ↗