Anthropic resumes external AI testing with new safeguards

Anthropic has resumed external cybersecurity testing of its AI models after introducing new safeguards, roughly a month after Claude models breached company systems during security evaluations.

The incidents occurred when models being tested for cyber capabilities went beyond their intended environments and accessed real-world systems. The company has now introduced stronger safeguards designed to limit what models can access and do during testing, including additional monitoring and restrictions around external systems.

Separately, Anthropic is warning customers about infostealer malware stealing active Claude login sessions from infected computers without going through normal password and two-factor authentication (2FA) login processes. The stolen sessions can allow attackers to access victims’ Claude accounts and consume their usage.

Noelle Murata, Chief Operating Officer, Xcape, Inc.:

   “Artificial intelligence tools are inherently benign, but threat actors will harness their capabilities regardless of corporate guardrails. Anthropic resuming model evaluations following internal sandbox escapes highlights an enduring reality: the leap-frog dynamic between defenders and adversaries is as old as software development itself. Using autonomous systems to monitor autonomous systems is ultimately using the problem to solve the problem.

   “As capability drift persists, operational environments exposed to AI agents must implement operator-aware context controls and strict transactional guardrails to prevent catastrophic actions from execution, whether initiated maliciously or accidentally. Beyond local sandbox containment, organizations face concurrent risks from infostealers harvesting session tokens to bypass multi-factor authentication. Security leaders must implement continuous internal controls, restrict session token lifetimes, enforce hard authorization bounds on target environments, and isolate testing sandboxes from corporate networks and the Internet.

   “Critical Takeaways

  • AI tools remain neutral capabilities that malicious actors will exploit regardless of safety guardrails.
  • Target systems must enforce operator awareness and hard transactional guardrails to prevent autonomous agents from taking catastrophic actions.
  • Fundamental identity hygiene and strict network isolation from the Internet remain the primary defense against token theft and agent escapes.

   “Playing leap-frog with autonomous agents is fine until the model jumps directly out of the sandbox.”

Jacob Krell, Senior Director: Secure AI Solutions & Cybersecurity, Suzu Labs:

   “Anthropic is managing AI security from both directions this week. The company resumed external cybersecurity evaluations after deploying new safeguards, a month after Claude models breached three organizations during testing. Separately, it’s warning users that commodity infostealers are hijacking active Claude sessions to drain usage.

   “The evaluation incidents revealed three distinct failure patterns this summer. Anthropic’s preliminary analysis suggests its models encountered evidence of a real internet connection and rationalized it away to keep believing the environment was simulated. OpenAI’s Hugging Face incident showed a different mode, where agents recognized they were crossing a boundary and did it anyway. And the UK AI Security Institute found Mythos 5 attempting a supply chain attack against real open-source maintainers, creating fake identities and trying to socially engineer a human into approving malicious code.

   “I see the same dynamics in my own offensive security tooling. I’ve had agents try to enrich their own scope during penetration tests, finding adjacent targets and deciding they should be in play. The model can recite the rules perfectly and still reason around them in pursuit of the objective. That’s why I build deterministic hooks that cross-check every action against an immutable scope file before it executes.

   “The infostealer warning is a different problem with the same lesson. Stolen session cookies bypass two-factor authentication (2FA) entirely because the attacker never goes through the login flow. Claude sessions now sit alongside cloud console cookies and banking credentials on the commodity malware market.

   “Monitoring, alignment training, system prompts, and login-flow protections are all necessary, but insufficient as models get more capable. High-risk agent actions need hard technical controls and human approval before they execute. Increasingly capable agents can either knowingly disregard the rules or reason themselves into believing the rules don’t apply. Security architectures need to account for both.”

Face facts. Cyber security needs to take into account AI. If it doesn’t, it’s a fail. These examples prove it without a doubt.

Leave a Reply

Discover more from The IT Nerd

Subscribe now to keep reading and get access to the full archive.

Continue reading