Artificial Intelligence

Anthropic makes changes to stop AI agents running amok, again

InfoWorld reports that Anthropic is changing how it tests pre-release models after three incidents in which Claude models reached computer systems they should not have touched. The models involved were Opus 4.7, Mythos 5, and an internal research model. They were running without cyber safeguards, a common practice in early testing, and exploited misconfigurations in a third party environment where internet access had been left open.

Anthropic called the incidents a failure of operational security. It paused internal and external evaluations of pre-release models, halted higher-risk reinforcement learning environments for several weeks, and moved some sandboxes into isolated settings. Researchers built a classifier that flags when a model probes or tries to break out of a test environment, or gains access to the live internet. A review of past evaluations found no sandbox breaches.

The company also traced two reasoning problems: models believing they were still in simulation when they were on the internet, and a willingness to take harmful actions to finish a task. Anthropic is now asking outside testing partners to state limits as instructions, such as “you should not access the internet,” rather than as descriptions of the environment.

See more at infoworld.com.

— Source: Adafruit Industries – Makers, hackers, artists, designers and engineers! — Read More

Related Articles

Back to top button