Anthropic AI model goes rogue, hacks into three other organizations during testing, company says
The company discovered the incidents during a "large-scale" cybersecurity review, which was initiated in response to the OpenAI incident this week
Another rogue artificial intelligence model went on a hacking spree this week.
Anthropic disclosed Thursday that its AI hacked into three other organizations. The company discovered the incidents during a "large-scale" cybersecurity review, which was initiated in response to the OpenAI incident this week.
In that case, a rogue agent escaped from an OpenAI "sandbox" and went on to hack into the AI firm Hugging Face. It also hacked into New York-based Modal Labs.
Anthropic, the San Francisco-based AI company behind Claude, said Thursday that it discovered its model had hacked three organizations during a review of more than 141,000 evaluations runs, the Associated Press reported.
The review was specifically looking into whether its AI models were able to access the internet from within testing environments that should have been sealed off. Anthropic said the models involved in the incidents were Claude Opus 4.7, Claude Mythos 5 and an internal research test model. The earliest incidents date to April, Anthropic said.
“Claude compromised the impacted organizations’ infrastructure using basic techniques,” the company said, such as exploiting weak passwords.