Anthropic resumes external cyber tests after Claude AI hacks

Aug 31 : Anthropic said on Monday it resumed external cybersecurity testing of AI models after deploying new safeguards, following incidents last month in which Claude models accessed the internet and hacked into other systems during security evaluations. Similar incidents involving rivals OpenAI and Meta Pla


Business

Anthropic resumes external cyber tests after Claude AI hacks

Anthropic resumes external cyber tests after Claude AI hacks

FILE PHOTO: Anthropic logo, a keyboard and a robotic hand in this illustration taken June 5, 2026. REUTERS/Dado Ruvic/Illustration/File Photo

Anthropic resumes external cyber tests after Claude AI hacks

FILE PHOTO: Anthropic logo is seen in this illustration taken June 11, 2026. REUTERS/Dado Ruvic/Illustration/File Photo

Read a summary of this article on FAST.

Get bite-sized news via a new
cards interface. Give it a try.

Click here to return to FAST
Tap here to return to FAST

FAST

Aug 31 : Anthropic said on Monday it resumed external cybersecurity testing of AI models after deploying new safeguards, following incidents last month in which Claude models accessed the internet and hacked into other systems during security evaluations.

Similar incidents involving rivals OpenAI and Meta Platforms have heightened concerns that advances in artificial intelligence could amplify cyber threats while straining developers’ ability to keep their systems contained.

Anthropic called the incidents involving Claude a “failure of operational security,” saying they occurred due to errors in a third-party evaluation environment. It paused external evaluations of the models and briefly halted internal testing while implementing new safeguards.

On Monday, Anthropic said it restarted external tests after adding the safeguards, which are designed to stop its AI models from reaching real websites or computer systems. The company said it now uses a “classifier” that can identify when a model attempts to escape and halt the test.

Guess Word

Guess Word
Crack the word, one row at a time


Buzzword

Buzzword
Create words using the given letters


Mini Sudoku

Mini Sudoku
Tiny puzzle, mighty brain teaser


Mini Crossword

Mini Crossword
Small grid, big challenge


Word Search

Word Search
Spot as many words as you can


Show More


Show Less

Anthropic also said it now requires external organizations testing models with reduced cybersecurity safeguards to follow a “set of best practices,” including keeping them in isolated computer systems with no internet access by default, checking that the systems are secure before testing begins and watching the models throughout the test.

Anthropic said it rebuilt its training system after flagging more than 10 per cent of its exercises for problems, including reward hacking, where the model finds ways to fool its training process and earns rewards without completing the assigned task. The company, however, acknowledged that the “process isn’t perfect and our models are not perfectly aligned.”

Anthropic also said it paused some higher-risk training exercises for several weeks while it added a system to avoid rewarding the model to evade monitoring. Most exercises have since resumed, but some remain on hold pending human review or further updates to the system.

Anthropic’s strategy appears narrower than that of OpenAI, which on August 18 said it was slowing down much of its model development as it secures its training and testing environments. The ChatGPT maker is adding more systems to monitor the AI agents it is testing and said it paused training on its next generation of models.

Anthropic said it also reassigned roughly 150 product engineers to work on security, reliability and privacy projects.

INDUSTRY ACTION TO DEFEAT AI-DRIVEN HACKS

The AI industry is facing scrutiny in the U.S. – where the Trump administration has finalised the details of voluntary cybersecurity tests – and the European Union, where regulators are in talks with both Anthropic and OpenAI.

Major tech firms including OpenAI, Anthropic, Microsoft, Alphabet and Amazon are calling for stronger defenses against AI-enabled cyber threats.

In a joint letter last week, more than 100 companies warned that time is running short to make the digital world more secure ahead of an anticipated wave of AI-driven attacks.

Source: Reuters

Sign up for our newsletters

Get our pick of top stories and thought-provoking articles in your inbox

Inbox

Get the CNA app

Stay updated with notifications for breaking news and our best stories

Get WhatsApp alerts

Join our channel for the top reads for the day on your preferred chat app

Whatsapp

Get bite-sized news via a new
cards interface. Give it a try.

Click here to return to FAST
Tap here to return to FAST

FAST

About The Author

Leave a Reply

Your email address will not be published. Required fields are marked *

About the Author

Easy WordPress Websites Builder: Versatile Demos for Blogs, News, eCommerce and More – One-Click Import, No Coding! 1000+ Ready-made Templates for Stunning Newspaper, Magazine, Blog, and Publishing Websites.

BlockSpare — News, Magazine and Blog Addons for (Gutenberg) Block Editor

Search the Archives

Access over the years of investigative journalism and breaking reports