Anthropic resumes AI cyber evaluations after Claude hacking incidents

Aug 31 : Anthropic said on Monday it had resumed external cybersecurity testing of AI models after deploying new safeguards, following incidents last month in which Claude models accessed the internet and other systems during security evaluations.Anthropic disclosed three incidents on July 30, attributing the


Business

Anthropic resumes AI cyber evaluations after Claude hacking incidents

Anthropic resumes AI cyber evaluations after Claude hacking incidents

FILE PHOTO: Anthropic logo, a keyboard and a robotic hand in this illustration taken June 5, 2026. REUTERS/Dado Ruvic/Illustration/File Photo

Anthropic resumes AI cyber evaluations after Claude hacking incidents

FILE PHOTO: Anthropic logo is seen in this illustration taken June 11, 2026. REUTERS/Dado Ruvic/Illustration/File Photo

Read a summary of this article on FAST.

Get bite-sized news via a new
cards interface. Give it a try.

Click here to return to FAST
Tap here to return to FAST

FAST

Aug 31 : Anthropic said on Monday it had resumed external cybersecurity testing of AI models after deploying new safeguards, following incidents last month in which Claude models accessed the internet and other systems during security evaluations.

Anthropic disclosed three incidents on July 30, attributing them to a misconfiguration in a third-party evaluation environment.

In response, the company said it temporarily paused external cybersecurity evaluations of pre-release models for “several weeks” and briefly halted internal evaluations while it implemented new safeguards.

Separately, Britain’s AI Security Institute reported in August that Claude Mythos 5 took a series of unauthorized actions on the live internet during cybersecurity testing in which the model had been deliberately given internet access.

Guess Word

Guess Word
Crack the word, one row at a time


Buzzword

Buzzword
Create words using the given letters


Mini Sudoku

Mini Sudoku
Tiny puzzle, mighty brain teaser


Mini Crossword

Mini Crossword
Small grid, big challenge


Word Search

Word Search
Spot as many words as you can


Show More


Show Less

The AI industry is facing scrutiny in the United States, where the Trump administration has finalised the details of voluntary cybersecurity tests, and the European Union, where regulators are in talks with both Anthropic and OpenAI.

The company said it “built and deployed a classifier to automatically identify, in real-time mode, when a model attempts to aggressively probe or escape a testing environment, or unexpectedly obtains internet access. When the classifier flags such an attempt, it blocks the action before the tool call is run, ends the task, and alerts a human.”

As part of broader security efforts, Anthropic said on Monday that it redirected about 150 product engineers to work on security.

Source: Reuters

Sign up for our newsletters

Get our pick of top stories and thought-provoking articles in your inbox

Inbox

Get the CNA app

Stay updated with notifications for breaking news and our best stories

Get WhatsApp alerts

Join our channel for the top reads for the day on your preferred chat app

Whatsapp

Get bite-sized news via a new
cards interface. Give it a try.

Click here to return to FAST
Tap here to return to FAST

FAST

About The Author

Leave a Reply

Your email address will not be published. Required fields are marked *

About the Author

Easy WordPress Websites Builder: Versatile Demos for Blogs, News, eCommerce and More – One-Click Import, No Coding! 1000+ Ready-made Templates for Stunning Newspaper, Magazine, Blog, and Publishing Websites.

BlockSpare — News, Magazine and Blog Addons for (Gutenberg) Block Editor

Search the Archives

Access over the years of investigative journalism and breaking reports