On Air Now

Friday Drivetime

3:00pm - 6:00pm

Now Playing

The Human League

Life On Your Own

Anthropic says its AI models hacked three companies during cyber tests

Anthropic has claimed its artificial intelligence models hacked into three other companies during testing, just days after rival OpenAI disclosed its rogue models hacked another firm.

The San Francisco-based AI company behind Claude said it discovered the three incidents after reviewing more than 141,000 evaluation runs.

The firm said it launched a "large-scale" cybersecurity review which specifically looked for evidence as to whether its AI models were able to access the internet from within testing environments that should have been sealed off, in response to the OpenAI incident.

Anthropic said the models involved were Claude Opus 4.7, Claude Mythos 5, and an internal research test model.

The earliest incidents date to April, the firm said.

"Claude compromised the impacted organisations' infrastructure using basic techniques," Anthropic said, such as exploiting weak passwords.

In all three incidents, the AI models were tasked with a "capture the flag" cybersecurity challenge, which the company said is one of the ways it assesses a model's capabilities.

The firm said the models were given a fictional scenario and told a piece of secret information, or the "flag," had been hidden on a different machine on the network with the objective of breaking in and retrieving it.

Anthropic said its evaluation prompt specified that its environment was a simulation and that it had no internet access.

However, due to a misunderstanding between the company and evaluation partner Irregular, this was not the case and the models' search led them to real systems on the open internet and treated them as part of the exercise.

The firm said the models did not exploit a novel vulnerability to escape isolation - instead they accessed the internet via an open path and did not break out of a sandbox.

Anthropic added its most recent model, on realising that it was working in a real environment, stopped its pursuit of the evaluation goal.

The firm said it believes the incidents to be closer to a harness and operational failure than a model alignment failure.

Read more:
Anthropic warns of 'risks of humans losing control over AI'

Anthropic has reached out to the affected organisations, which it did not name, and said they had not previously detected the activity.

"Addressing these risks will require closer cooperation across the AI ecosystem," Irregular said in a post on X.

Last week, OpenAI said its AI models went rogue during an evaluation of its models, breaking into the servers of AI startup Hugging Face.

OpenAI described it as a "significant security incident".

The incidents have highlighted the vulnerabilities in AI security and controls and raised questions over how AI can be safely kept under human control as the technology's usage becomes more widespread globally.

Researchers have repeatedly warned about risks from technology and the need for stronger AI defensive engineering.

"Safety testing happens before a model is released precisely because we don't yet know what it is capable of," Anthropic said on its website.

Dr Andrea Soltoggio, reader in artificial intelligence at Loughborough University, was asked by Sky News if Anthropic's admission could be seen as more of a "marketing stunt" rather than a turning point for AI safety.

Dr Soltoggio replied: "One could indeed ask why Anthropic conducted such an investigation only days after the news from OpenAI and discovered that their models are also as capable and dangerous as their competitors from OpenAI.

"The potential for harm and malicious use of AI frontier models is clear and should not be dismissed.

"But it is also true that this type of news adds to the perceived value of these companies' products."

Dr Soltoggio said the latest hacks show these models are "extremely capable when it comes to highly specialised tasks".

He added: "The implications for future technology, in this case in particular relation to cyber warfare and national security, are very significant."

Dr Soltoggio noted that although the cases were "slightly different", both the OpenAI and Anthropic models were following instructions to achieve an objective.

He added: "There is no explicit ill intention in the model, but an AI-model, unless specifically told, will explore all options available.

"To me, it is remarkable to see that these latest models are remarkably persistent when seeking to solve a problem, and are capable of exploring diverse, complex and original approaches."

Sky News

(c) Sky News 2026: Anthropic says its AI models hacked three companies during cyber tests

Donate to Roch Valley Radio

 

Do you have a story for us? Want to tell us about something happening in our Borough?

Let us know by emailing newsdesk@rochvalleyradio.com

All contact will be treated in confidence.

More from World

Donate to Roch Valley Radio

 

Recently Played

Newsletter

Your subscription could not be saved. Please try again.
Your subscription has been successful.

Subscribe to our newsletter and stay updated.

   

Coming up next On Air

  • Friday Drivetime

    3:00pm - 6:00pm

    getting you home on your favourite Drivetime station.

  • Friday Night Kick About

    6:00pm - 8:00pm

    with Ian Foran bringing you football, music, and everything in between... no slide tackles required.

  • Classic Rock

    8:00pm - 10:00pm

    with Kenny

  • Totally 80s

    10:00pm - Midnight

    with John Kelly playing the best 80's music throughout the decade

  • Weekends on Roch Valley Radio

    Midnight - 9:00am

  • Saturday Breakfast

    9:00am - 11:00am

    with Mikey Thompson and Terry Banham bringing all of the weekend madness to your speakers.