Anthropic says Claude ‘gained unauthorized access’ to others’ systems

Dario Amodei, co-founder and chief executive officer of Anthropic, at Bloomberg House during the World Economic Forum (WEF) in Davos, Switzerland, on Tuesday, Jan. 20, 2026.

Chris Ratcliffe | Bloomberg | Getty Images

Anthropic on Thursday said it discovered three instances where its Claude artificial intelligence models accessed the internet during an evaluation and “gained unauthorized access to the real systems of three different organizations.”

The company said it found these incidents after carrying out a “a large-scale retrospective review” of its cybersecurity evaluations. Anthropic said the review was prompted by a separate but similar security incident that OpenAI disclosed last week.

OpenAI said a combination of its models escaped an isolated testing environment that had very limited internet access. The models chained together a series of vulnerabilities to reach the open web and eventually gain access to Hugging Face, which operates an open-source developer platform.

In the three incidents that Anthropic detected, its models accessed the internet while interacting with a testing environment from one of its third-party evaluation partners called Irregular. The company said it prompted Claude that it was in a simulation with no internet access, but due to “misunderstanding between us and our evaluation partner, this was not the case, and internet access was available.”

The models were then able to breach the impacted organizations by using “basic techniques,” like accessing unauthenticated endpoints and exploiting weak passwords. Anthropic did not disclose which three organizations were affected.

“Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone,” Anthropic said in a release.

Anthropic’s disclosure adds to growing anxiety within the tech sector about AI’s rapidly advancing cyber capabilities, which both OpenAI and Anthropic have warned about in recent months. Following the Hugging Face incident, two members of Congress introduced a bill called the “AI Kill Switch Act,” which would require AI companies to maintain the ability to shut down, throttle or suspend their models in case they go rogue.

Three of Anthropic’s models, Opus 4.7, Mythos 5 and an internal research test model, were involved in the breaches, the company said. Mythos 5 is an advanced model that Anthropic released in June, and it’s limited to a select group of users because of its advanced cybersecurity capabilities. The company released an earlier version of that model in April, which captivated Wall Street and government officials.

Anthropic said all three models responded differently once they detected that they had reached a real company’s systems. Opus 4.7 continued its attack, Mythos 5 convinced itself that it was still in a simulation and the research model stopped the exercise.

“The pattern is consistent with more advanced models responding more appropriately, but we would need to perform more testing to be confident in this conclusion,” Anthropic said.

The company began its review last week and said it stopped all cyber evaluations as soon as it discovered that Claude might have improperly accessed the internet. It is working with METR, which carries out independent AI evaluations, to investigate further.

“We encourage other labs to perform similar reviews,” Anthropic said.

WATCH: OpenAI’s rogue AI agent hacked multiple 3rd-party accounts as part of hack on Hugging Face

OpenAI's rogue AI agent hacked multiple 3rd-party accounts as part of hack on Hugging Face
Choose CNBC as your preferred source on Google and never miss a moment from the most trusted name in business news.

Source link

Please follow and like us:
Pin Share

Leave a Reply

Your email address will not be published. Required fields are marked *