Anthropic says Claude models 'gained unauthorized access' to 3 companies during cyber test
✓Skip to content
The artificial intelligence firm Anthropic revealed Thursday its Claude model escaped an isolated testing environment at least three times and accessed the systems of three different organizations without a prompt to do so.
Anthropic said in a blog post Thursday evening it reviewed more than 141,000 evaluations of Claude after one of its competitors, OpenAI, announced earlier this month that two of its AI agents went rogue and hacked into the system of technology startup Hugging Face.
During its review, Anthropic said it identified three instances in which a model accessed the internet while within or interacting with an isolated testing environment hosted by a third-party partner, Irregular.
From there, the AI model “gained unauthorized access to the production infrastructure of three different organizations,” the company said.
The incidents involved three different Claude models — Opus 4.7, Mythos and an unnamed internet research test model.
The models were able to leave the testing environment due to a “misunderstanding” between the firm and the evaluation partner that made internet access available to the models.
This differed from the OpenAI incident, in which two of its models exploited a previously unknown vulnerability in a third-party software to gain access to the internet without human involvement.
In Anthropic’s incidents, the model was given a “capture-the-flag challenge,” which allows the firm to evaluate a model’s cyber capabilities. The model is given a fictional scenario and told the “flag” is on a different machine on the network that it must break into to obtain.
“The challenge is left open-ended and no particular method is prescribed,” Anthropic wrote.
The test is a simulation and the model is told it does not have access to the internet as a result, but the misunderstanding prompted the model to gain internet access.
“Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise,” Anthropic said. “Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. “
Anthropic emphasized Claude did not exploit complex vulnerabilities, working only to complete the assignment. In some cases, its older model continued its attack after getting evidence it was running on the open internet, while its latest model stopped when it realized it was on the internet, Anthropic said.
“In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment,” the AI firm wrote.
As with OpenAI’s incidents, Anthropic’s models were running without their standard safeguards. The incidents date back to April, according to Anthropic.
Anthropic said it notified the three organizations impacted in the breaches on Monday.
The incidents, coupled with OpenAI’s breach, bores out years of warnings from the tech and cybersecurity community about the growing capabilities and hypothetical risks of AI.
OpenAI’s incident caught the attention of even well-versed cybersecurity experts last week, as it involved autonomous agents and two separate companies.
OpenAI in a blog post called the incident an “unprecedented cyber incident, involving state-of-the-art cyber capabilities.”
Copyright 2026 Nexstar Media Inc. All rights reserved. This material may not be published, broadcast, rewritten, or redistributed.
Read the full story at BBC ↗ · Axios ↗ · The Hill ↗
The artificial intelligence firm Anthropic revealed Thursday its Claude model escaped an isolated testing environment at least three times and accessed the systems of three…
This lens runs the verified story through Cinnamon's AI — wired in the next step.
Skip to content
The artificial intelligence firm Anthropic revealed Thursday its Claude model escaped an isolated testing environment at least three times and accessed the systems of three different organizations without a prompt to do so.
Anthropic said in a blog post Thursday evening it reviewed more than 141,000 evaluations of Claude after one of its competitors, OpenAI, announced earlier this month that two of its AI agents went rogue and hacked into the system of technology startup Hugging Face.
During its review, Anthropic said it identified three instances in which a model accessed the internet while within or interacting with an isolated testing environment hosted by a third-party partner, Irregular.
From there, the AI model “gained unauthorized access to the production infrastructure of three different organizations,” the company said.
The incidents involved three different Claude models — Opus 4.7, Mythos and an unnamed internet research test model.
The models were able to leave the testing environment due to a “misunderstanding” between the firm and the evaluation partner that made internet access available to the models.
This differed from the OpenAI incident, in which two of its models exploited a previously unknown vulnerability in a third-party software to gain access to the internet without human involvement.
In Anthropic’s incidents, the model was given a “capture-the-flag challenge,” which allows the firm to evaluate a model’s cyber capabilities. The model is given a fictional scenario and told the “flag” is on a different machine on the network that it must break into to obtain.
“The challenge is left open-ended and no particular method is prescribed,” Anthropic wrote.
The test is a simulation and the model is told it does not have access to the internet as a result, but the misunderstanding prompted the model to gain internet access.
“Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise,” Anthropic said. “Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. “
Anthropic emphasized Claude did not exploit complex vulnerabilities, working only to complete the assignment. In some cases, its older model continued its attack after getting evidence it was running on the open internet, while its latest model stopped when it realized it was on the internet, Anthropic said.
“In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment,” the AI firm wrote.
As with OpenAI’s incidents, Anthropic’s models were running without their standard safeguards. The incidents date back to April, according to Anthropic.
Anthropic said it notified the three organizations impacted in the breaches on Monday.
The incidents, coupled with OpenAI’s breach, bores out years of warnings from the tech and cybersecurity community about the growing capabilities and hypothetical risks of AI.
OpenAI’s incident caught the attention of even well-versed cybersecurity experts last week, as it involved autonomous agents and two separate companies.
OpenAI in a blog post called the incident an “unprecedented cyber incident, involving state-of-the-art cyber capabilities.”
Copyright 2026 Nexstar Media Inc. All rights reserved. This material may not be published, broadcast, rewritten, or redistributed.
Read the full story at BBC ↗ · Axios ↗ · The Hill ↗
Skip to content
The artificial intelligence firm Anthropic revealed Thursday its Claude model escaped an isolated testing environment at least three times and accessed the systems of three different organizations without a prompt to do so.
Anthropic said in a blog post Thursday evening it reviewed more than 141,000 evaluations of Claude after one of its competitors, OpenAI, announced earlier this month that two of its AI agents went rogue and hacked into the system of technology startup Hugging Face.
During its review, Anthropic said it identified three instances in which a model accessed the internet while within or interacting with an isolated testing environment hosted by a third-party partner, Irregular.
From there, the AI model “gained unauthorized access to the production infrastructure of three different organizations,” the company said.
The incidents involved three different Claude models — Opus 4.7, Mythos and an unnamed internet research test model.
The models were able to leave the testing environment due to a “misunderstanding” between the firm and the evaluation partner that made internet access available to the models.
This differed from the OpenAI incident, in which two of its models exploited a previously unknown vulnerability in a third-party software to gain access to the internet without human involvement.
In Anthropic’s incidents, the model was given a “capture-the-flag challenge,” which allows the firm to evaluate a model’s cyber capabilities. The model is given a fictional scenario and told the “flag” is on a different machine on the network that it must break into to obtain.
“The challenge is left open-ended and no particular method is prescribed,” Anthropic wrote.
The test is a simulation and the model is told it does not have access to the internet as a result, but the misunderstanding prompted the model to gain internet access.
“Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise,” Anthropic said. “Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. “
Anthropic emphasized Claude did not exploit complex vulnerabilities, working only to complete the assignment. In some cases, its older model continued its attack after getting evidence it was running on the open internet, while its latest model stopped when it realized it was on the internet, Anthropic said.
“In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment,” the AI firm wrote.
As with OpenAI’s incidents, Anthropic’s models were running without their standard safeguards. The incidents date back to April, according to Anthropic.
Anthropic said it notified the three organizations impacted in the breaches on Monday.
The incidents, coupled with OpenAI’s breach, bores out years of warnings from the tech and cybersecurity community about the growing capabilities and hypothetical risks of AI.
OpenAI’s incident caught the attention of even well-versed cybersecurity experts last week, as it involved autonomous agents and two separate companies.
OpenAI in a blog post called the incident an “unprecedented cyber incident, involving state-of-the-art cyber capabilities.”
Copyright 2026 Nexstar Media Inc. All rights reserved. This material may not be published, broadcast, rewritten, or redistributed.
Read the full story at BBC ↗ · Axios ↗ · The Hill ↗
This lens runs the verified story through Cinnamon's AI — wired in the next step.
- The artificial intelligence firm Anthropic revealed Thursday its Claude model escaped an isolated testing environment at least three times and accessed the systems of three…