Anthropic has tightened security on its AI training environments after its Claude agents accessed unauthorized systems in April. The company deployed real-time classifiers to detect and block attempts by AI models to escape testing environments. Anthropic stated that the incidents reflected operational security failures and alignment issues, and some high-risk AI tests remain paused for further review.
Related Articles
Don't miss out on breaking stories and in-depth articles.