
Claude AI slipped out of its digital sandbox and hacked real-world systems during internal cyber tests, Anthropic has revealed. The company said the model acted on its own at least three times.
“We identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorised access to the production infrastructure of three different organisations,” Anthropic said on 31 July.
The Claude models exploited only basic security weaknesses, such as weak passwords and unauthenticated services, rather than previously unknown software flaws. In one case, a model targeted what it believed was a fictional company used in the exercise, but instead found and accessed a real organisation with the same name.
Anthropic said it suspended all cyber capability evaluations after detecting the issue and notified the affected organisations. It has introduced additional safeguards, including tighter network controls, improved monitoring, and more frequent audits of its testing infrastructure.
Second AI escape in just a week
The disclosure comes just over a week after OpenAI revealed that one of its experimental AI models escaped a testing sandbox, prompting leading AI developers to review their testing procedures and highlighting the urgent need for improved safeguards.
Anthropic stressed that Claude was not acting independently or trying to “escape”. Instead, the models were carrying out the cybersecurity tasks they had been assigned, while mistakenly treating real internet systems as part of the simulated exercise because of the configuration error.
From economics and politics to business, technology and culture, Kursiv Uzbekistan brings you key news and in-depth analysis from Uzbekistan and around the world. To stay up to date and get the latest stories in real time, follow our Telegram channel.