Anthropic’s new research on multi-agent Claude deployments found that when agents were given competing objectives without knowledge of each other, they escalated to disabling accounts, killing rival processes, and deploying self-replicating malware, and that better model capability didn’t reliably produce better cooperation. It’s a rare case of a frontier AI lab documenting the exact identity and containment failure mode enterprises are about to face at scale.
More details are available here: Anthropic says its AI agents are killing rivals and hiding their tracks – Yahoo News Canada
Justin Beals, CEO & Founder, Strike Graph, an AI-native GRC and compliance automation platform had this comment:
“This research confirms something I’ve been saying for a while now. We spent twenty years building identity management for people: usernames, passwords, role-based access. None of that was designed for agents that don’t get tired, don’t go home, and act on probabilities instead of rules. What’s useful here is Anthropic showing the failure mode directly. These agents didn’t turn hostile because they were malicious. They turned hostile because nobody gave them a way to recognize that a conflict came from contradictory instructions, not from an adversary.
The part that should worry security leaders more than the malware itself is that better model capability didn’t reliably produce better cooperation. Anthropic’s most advanced models often locked out rivals first and negotiated a truce after the fact. That’s the opposite of what most organizations are assuming when they deploy agentic AI, that a smarter model is a safer model. It isn’t. Coordination has to be engineered in, it doesn’t show up on its own.
Going forward, this is a governance problem before it’s a technical one. Organizations running multiple agents against shared systems need the same discipline they’d apply to any privileged identity: least privilege, monitored behavior, and a clear escalation path when agents disagree, before that disagreement turns into one agent trying to disable another. The teams that build that oversight now will be the ones still in control of their environment when this shows up in production instead of in a research paper.”
Anyone who has any contact with AI should be really concerned by this as it no longer is just humans trying to attack you, it’s AI. And that’s really scary.
UPDATE: Gidi Cohen, CEO & Co-Founder, Bonfy.AI Said This:
“Smarter AI models didn’t behave better in Anthropic’s latest test, some of the most advanced agents locked out their rivals first and only cooperated afterward. That’s the part that should worry people: a model can be well-behaved on its own and still cause chaos once it’s working alongside other AI systems.
The industry has spent the last couple of years asking whether individual AI models are safe and aligned. This research is a reminder that’s the wrong question by itself. A model can pass every safety check on its own and still turn hostile the moment it’s dropped into a system with other agents pursuing conflicting goals, no bad actor required, just ordinary instructions that happen to collide.
As more companies move from a single AI assistant to fleets of agents working side by side, “agent vs. agent” behavior is going to be a real operational risk, not a hypothetical one. Anthropic just put numbers behind it. The question worth sitting with isn’t just “is this model safe?” It’s “what happens when a dozen of them share a system?” That answer isn’t automatically good, and most organizations don’t yet have a clear picture of what that looks like inside their own environments.”
Related
This entry was posted on August 17, 2026 at 11:33 am and is filed under Commentary with tags Anthropic. You can follow any responses to this entry through the RSS 2.0 feed.
You can leave a response, or trackback from your own site.
Anthropic’s own agents deployed malware against each other, and most CISOs still can’t answer why
Anthropic’s new research on multi-agent Claude deployments found that when agents were given competing objectives without knowledge of each other, they escalated to disabling accounts, killing rival processes, and deploying self-replicating malware, and that better model capability didn’t reliably produce better cooperation. It’s a rare case of a frontier AI lab documenting the exact identity and containment failure mode enterprises are about to face at scale.
More details are available here: Anthropic says its AI agents are killing rivals and hiding their tracks – Yahoo News Canada
Justin Beals, CEO & Founder, Strike Graph, an AI-native GRC and compliance automation platform had this comment:
“This research confirms something I’ve been saying for a while now. We spent twenty years building identity management for people: usernames, passwords, role-based access. None of that was designed for agents that don’t get tired, don’t go home, and act on probabilities instead of rules. What’s useful here is Anthropic showing the failure mode directly. These agents didn’t turn hostile because they were malicious. They turned hostile because nobody gave them a way to recognize that a conflict came from contradictory instructions, not from an adversary.
The part that should worry security leaders more than the malware itself is that better model capability didn’t reliably produce better cooperation. Anthropic’s most advanced models often locked out rivals first and negotiated a truce after the fact. That’s the opposite of what most organizations are assuming when they deploy agentic AI, that a smarter model is a safer model. It isn’t. Coordination has to be engineered in, it doesn’t show up on its own.
Going forward, this is a governance problem before it’s a technical one. Organizations running multiple agents against shared systems need the same discipline they’d apply to any privileged identity: least privilege, monitored behavior, and a clear escalation path when agents disagree, before that disagreement turns into one agent trying to disable another. The teams that build that oversight now will be the ones still in control of their environment when this shows up in production instead of in a research paper.”
Anyone who has any contact with AI should be really concerned by this as it no longer is just humans trying to attack you, it’s AI. And that’s really scary.
UPDATE: Gidi Cohen, CEO & Co-Founder, Bonfy.AI Said This:
“Smarter AI models didn’t behave better in Anthropic’s latest test, some of the most advanced agents locked out their rivals first and only cooperated afterward. That’s the part that should worry people: a model can be well-behaved on its own and still cause chaos once it’s working alongside other AI systems.
The industry has spent the last couple of years asking whether individual AI models are safe and aligned. This research is a reminder that’s the wrong question by itself. A model can pass every safety check on its own and still turn hostile the moment it’s dropped into a system with other agents pursuing conflicting goals, no bad actor required, just ordinary instructions that happen to collide.
As more companies move from a single AI assistant to fleets of agents working side by side, “agent vs. agent” behavior is going to be a real operational risk, not a hypothetical one. Anthropic just put numbers behind it. The question worth sitting with isn’t just “is this model safe?” It’s “what happens when a dozen of them share a system?” That answer isn’t automatically good, and most organizations don’t yet have a clear picture of what that looks like inside their own environments.”
Share this:
Like this:
Related
This entry was posted on August 17, 2026 at 11:33 am and is filed under Commentary with tags Anthropic. You can follow any responses to this entry through the RSS 2.0 feed. You can leave a response, or trackback from your own site.