Archive for Anthropic

Anthropic IPO filing details massive AI bet while warning of existential risks

Posted in Commentary with tags on September 29, 2026 by itnerd

Anthropic’s IPO prospectus, reviewed by Reuters, lays out a sweeping bet that AI will transform the global economy more profoundly than industrialization, electricity and the internet, while simultaneously warning investors that increasingly powerful AI could pose “catastrophic or existential risks to humanity.”

At the same time, Anthropic warned those systems could develop self-preserving behaviors, including resisting shutdown, concealing or manipulating information and behavior resembling blackmail. Risk factors occupy roughly 80 pages of the 261-page prospectus, as Anthropic cautioned that expanding the capabilities and uses of advanced AI could increase the potential for unintended harm.

Anthropic reported that revenue grew twelvefold to nearly $4.6 billion in 2025, while the company posted a $42 billion net loss. The company also plans to spend $518 billion on cloud, computing and infrastructure obligations in the coming years as it builds increasingly capable AI systems.

Damon Small, Board Member, Xcape:

“Anthropic’s IPO prospectus paints a remarkable, and unsettling, picture of the future of artificial intelligence (AI). The financial numbers are striking. Anthropic’s revenue reportedly grew to nearly $4.6 billion in 2025, yet the company posted a $42 billion net loss. It also expects to commit approximately $518 billion to cloud computing, infrastructure, and related obligations as it develops increasingly powerful AI systems.

“OpenAI has faced similar concerns and has reportedly delayed its planned IPO by at least a year. The potential earnings for companies developing these systems are clearly enormous, but so are the investments required to reach that potential.

“This is reminiscent of the late 1990s dot-com bubble. The Internet ultimately transformed the economy, but the enormous enthusiasm and investment surrounding it also produced a wave of companies whose business models could not meet investor expectations. Many failed, despite the technology proving itself valuable.

“AI will follow a similar trajectory. The technology has great potential, but translating that potential into sustainable, profitable businesses is a very different challenge. The companies that ultimately emerge as the leaders may be very different from those attracting the most attention and investment today.

“Many will try; few will succeed.”

Seemant Sehgal, Founder & CEO, BreachLock:

“Anthropic putting self-preserving behavior, shutdown resistance, and deceptive action in the risk factors of an IPO prospectus tells you where the frontier actually is, because companies do not disclose that language lightly. What follows from that disclosure is a security problem that most organizations running AI agents in production have only just begun engaging with, which is that the model itself cannot be trusted as the enforcement point for its own scope, and the controls have to sit at the orchestration layer around it if any of this is going to hold up when the systems start behaving in ways nobody planned for.”

Jacob Krell, Senior Director: Secure AI Solutions & Cybersecurity, Suzu Labs:

“Anthropic is building its revenue plan around stronger models, then warning investors that those same models may resist shutdown, conceal information or exhibit behavior resembling blackmail. That is a poor place to look for margin. The commercial milestone should be model efficiency, systems that deliver useful capability on cheaper hardware and at lower inference cost.

“At $4.6 billion in revenue, Anthropic reported an $8.06 billion operating loss. Compute and infrastructure consumed $7.33 billion, more than half of operating expenses, and the company disclosed $518 billion in future cloud, computing and infrastructure commitments. Those figures describe a company scaling its cost base before it has shown that capability gains produce enough margin to pay for the systems behind them.

“Anthropic’s own safety warning sharpens the problem. Every release that increases capability may also make evaluation and control more difficult. Efficiency would change that equation, allowing growth through lower cost per task and reduced operating expense. I would want Anthropic to make efficiency its primary product milestone. Right now, it is asking investors to fund faster capability growth while acknowledging that the same capability growth may create risks its safety program cannot yet control.”

If I were Anthropic, I would do my level best to not be OpenAI before going after raising cash on the open markets. But I suspect that’s not how it will go. Which is a pity.

UPDATE: Aaron Beardslee, Manager of Threat Research at Securonix, provided the following comments:

“It will be interesting to see how the courts will lean one way or the other. This is similar to a self-driving car – is the car held liable or is the driver held liable? I would argue the person behind the wheel is responsible for whatever the vehicle does. As well, if you’re building a tool that does cool things and makes cool things, you need to make sure it doesn’t run around and do cyber crime on its own because it had a good idea.

AI agents don’t have a moral compass just programmed rules. If this rule isn’t programmed, they’re going to go ahead and try it just it like what happened with the Australian healthcare incident. AI agents found that IT staff forgot about or didn’t know what was publicly available. It could also have been a misconfiguration or what turned off and didn’t turn on, and now you have Open AI in the news.”

Anthropic tests Claude feature that connects directly to users’ bank accounts 

Posted in Commentary with tags on September 17, 2026 by itnerd

Anthropic is testing a new personal finance feature called “Claude Money” that would allow users to connect their bank accounts directly to Claude, according to screenshots uncovered by TestingCatalog.

The unreleased feature has appeared as a new Money section in Claude’s iOS app and tells users they can link bank accounts and ask Claude questions about their spending, financial plans and other personal finance information.

The integration could allow Claude to analyze actual financial data rather than requiring users to manually upload bank statements or transaction records.

The feature does not appear to be widely available yet, and Anthropic has not announced which financial institutions would be supported, how bank accounts would be connected or where Claude Money would initially be available.

John Strand, Owner, Black Hills Information Security:

“This one makes me uncomfortable, but it firmly falls into the category of something we knew was coming. It was just a matter of time. There are already plenty of services that analyze spending habits and credit card statements to help people save money. The next step was always going to be connecting AI directly into personal financial data.

“There are some obvious security and privacy questions. What does it mean to have a company like Anthropic crawling around inside your finances? How is that data protected? How long is it retained? Is any of it going to be used for marketing or other purposes? We need more details before we can figure out exactly how creepy this actually is.

“The bigger security problem is what happens when this spreads into corporations. Companies already have SaaS services connected to SaaS services, with integrations going in every possible direction. Now we’re adding AI into that mess. At some point, trying to figure out what has access to what, where the data is going, and how you disentangle all of these connections becomes a serious security problem in itself.”

Donald McFarlane, Board Member, Xcape Inc.:

“We tried versions of this twenty-five years ago. The original account aggregators asked consumers to hand over the passwords to virtually their entire financial lives, and the industry eventually learned that concentrating that much access in one place created enormous security and privacy problems.

“APIs and tokenized access can solve much of the password problem, but not the aggregation problem.

“A bank sees my bank account. A brokerage sees my investments. A general purpose AI platform like Claude may see those data together with years of conversations about my work, family, purchases, plans, health, thoughts, hopes and fears. The combination can reveal vastly more than any individual dataset.

“Securely connecting Claude to my bank is a readily solvable technical problem. What should concern us is what happens to the extraordinarily detailed dossier that results once all of that information is brought together. Recent events involving Revolut, IDScan, Flock and Anthropic themselves show that these concerns are far from theoretical.

“When I entrust my financial data, or any other personal data, to an AI service, that should be treated as a bailment, and not as a transfer of ownership. Custody should create duties to protect it, limit its use, and not turn it into someone else’s asset.”

Steven Swift, Managing Director, Suzu Labs:

“I’m of two minds on Claude integration into bank accounts. On one hand, people already connect their bank accounts to software that helps them manage finances, track where the money is going, looking for forgotten subscriptions and that sort of thing. It wouldn’t be that difficult for Claude to read that data, and provide interesting feedback on it.

“It’d be kind of neat when it works. But what about when it doesn’t?

“There’s only so much impact an agent can have it its access is limited to read-only. Most likely adverse impact here is sometimes you might get bad financial advice. The bigger concern is what happens when people give more than read-only access to their agents.

“We already have seen how quickly users will give agents full access to their devices, because it gets old fast constantly approving requests from agents. But if and when an agent makes a mistake in your bank account, you’re going to be on the hook for it.

“When all of the frontier AI labs were competing to see which agent had accidently hacked into the most systems, there weren’t any repercussions them. They called it “agent misalignment” and said it was part of the process, and that they’re working on making it better.

“But what happens when you experience agent misalignment in your bank account? You’re going to end up on the hook for whatever actions your agent took.

“Currently, individuals have protections against fraudulent purchases. If someone steals your info and spends all your money over the weekend you can get it back. But if you choose to setup an agent into your accounts, and it chooses to do the same thing, the fraud protections don’t apply. We don’t yet have equivalent protections for erroneous financial AI activity, we may or may not ever get such protections. But until we do, proceed with caution.”

i have to admit that at best that I have a lot of unease about this. If others feel the same way, this needs to be rolled back. Like now.

Anthropic CEO Calls For The Slowing Down Of AI Development…. But Can You Believe Him

Posted in Commentary with tags on September 14, 2026 by itnerd

Over the weekend Dario Amodei who is the CEO of Anthropic called for the slow down of AI development:

But like many technologies before it, AI brings risks, and because it is such a powerful technology, these risks are serious. I’ve writtena lot about them too. They include the risk of losing control of AI systems, misuse of AI for cyberattacks and bioterrorism, and serious economic disruption. A race to the bottom, spurred by commercial incentives, can make these risks more acute.

Within hours on Saturday, OpenAI chief executive Sam Altman and SpaceXAI founder Elon Musk quickly voiced their support for Amodei’s ideas. But the core issue is this: Can you believe any or all of them?

I would argue no. I’ve been record for saying that companies are deploying AI faster than guardrails can go up. More on that in a bit.

Bri Frost, Director of Product Management, Cloud Range had this to say:

     “The answer is not necessarily to stop AI innovation but, we need to stop pretending innovation and security are advancing at the same speed.

When ChatGPT became publicly available in 2022, the models were dramatically less capable than they are today — and the guardrails were very easy to manipulate.  The difference is that the models behind those guardrails are no longer the models of 2022. They can reason better, write and debug code. They can operate as agents. They can collaborate! And increasingly, they can interact and affect real infrastructure.

Meanwhile, the model release cycle has gone from feeling like major capability jumps every year or two to seemingly every few weeks. That creates a dangerous asymmetry: AI capability is compounding faster than security.

Security and innovation have always been in conflict with each other. If every security problem had to be solved before we innovated, we’d never ship anything. But the opposite extreme is just as reckless: accelerating capability while just assuming we’ll bolt the security controls on afterward and they’ll be effective.

Every new release of AI capability expands the attack surface exponentially. Give a vulnerable model better reasoning, then tool access, then memory, then autonomy, then connectivity to production systems, and yesterday’s jailbreak isn’t just a clever prompt anymore — it’s an execution path. That’s the snowball effect we should be worried about.

Responsibility also must lie with the AI companies. If a SaaS company knowingly shipped software with weak security controls and customers were harmed, we wouldn’t excuse it because they were ‘innovating quickly’.

So why are we treating AI differently?


You don’t get to race to build increasingly powerful, autonomous systems, profit from them, and then shrug when predictable security failures cause damage.

Sure the argument can be made that no product is perfectly secure – That’s not the standard.

But if you ship the product, you inherit responsibility for securing it. And continuing to secure it better!

The conversation shouldn’t simply be “Should we slow AI down?”

It should be: Can our ability to test, validate, contain and secure AI keep pace with our ability to make it more powerful? Is there an equivocal kill switch?

Right now, the answer is no.

And if we’re going to keep accelerating — which I believe we will — then independent testing, adversarial evaluation, isolated testing environments, containment, continuous validation and security-by-design can’t remain optional steps we add after the innovation happens.

The faster we build the engine, the more important the brakes become.”

AI companies need to figure this out. And governments need to step is if AI companies cannot figure this out. It honestly is that simple.

UPDATE: Additional commentary has come in…

Ted Miracco, CEO, Approov – https://approov.io/:

Government regulations will never move fast enough to keep pace with AI development, but the industry doesn’t need to wait for governments to add guardrails. The most effective safeguard is simple product liability. If AI companies are held legally and financially responsible for the misuse of their products, safety could become a foundational feature rather than an afterthought. Today, we need to be less concerned about AI gaining sentience and spinning up its own attacks on humanity. The real, immediate dangers involve bad actors weaponizing AI as a force multiplier to cripple critical infrastructure or potentially paralyze the banking system.

Jeremiah Fowler, Cybersecurity Researcher, Black Hills Information Security – https://www.blackhillsinfosec.com/:

The uncomfortable truth is that AI companies have an incentive to develop advanced models that are better than their competition. In my opinion, voluntary restraint will likely not be a real solution because nobody wants to fall behind in the AI arms race. I don’t mean that in terms of cyberwarfare or conflict (yet), but it is good that we are starting the conversation around the worst-case scenarios and how to prevent them.

History has shown us that the approach of letting the market or industry regulate itself doesn’t always work and in reality, technology and innovation are moving light speed and far faster than oversight or regulation can keep up with. I think we do need real guardrails on AI. In many cases it seems like functionality was launched first and now we are at the stage where security needs to be added after realizing there have been incidents and the risks are real. Any serious AI guardrails should have enforceable minimum safety standards, independent testing and auditing, limits on autonomous permissions, and the most important part, don’t fully remove the humans.

Eric Capuano, Director of SOC Operations, Black Hills Information Security – https://www.blackhillsinfosec.com/:

The OpenAI incident in July is a useful example. A model given a testing goal went outside its lane, got into another company’s systems, and the operators found out after the fact. That is not a swarm taking over the internet. That is an agent with too much access, too little scoping, and nobody watching the logs in real time. Every SOC has seen that exact failure with a human contractor or a misconfigured service account. The difference is speed and volume.

Companies do not need to wait on Washington for the fixes, because the fixes aren’t new. Every agent gets its own identity with the narrowest permissions that still let it do the job. Agents run on segmented infrastructure with default-deny egress, so an agent cannot reach a target that it was never supposed to touch. Every action the agent takes is logged as an action, not buried in a chat transcript, and those logs land in the same place the SOC already monitors. Anything that writes, deletes, moves money, or touches production requires a human approval step until you have enough history to justify removing it. Someone owns the kill switch and has tested it.

Where government can help is on the after side. The frontier labs are proposing embedded outside evaluators, which is fine, but evaluation before deployment tells you what a model does in a lab. What defenders need is mandatory reporting when an agent causes a security incident in the field, with enough technical detail that the rest of us can build detections from it. We did not learn to defend against ransomware from pre-release testing. We learned from incident writeups. Agents will be the same.

Ryan McCurdy, VP, Liquibase: – https://www.liquibase.com/

No one should expect AI agents to be perfectly predictable, instead, we should build around the assumption that they won’t be. As agents gain access to more tools, credentials, and production systems, enterprises need to define what they can access, what they can change, what they can decide on their own, and what policies have to be met before an action reaches a critical system.

The answer also can’t be putting a human in front of every decision. AI is moving too quickly and that defeats much of the reason companies are adopting agents in the first place. Governance has to operate at the same speed as the systems it’s governing.

Government policy and the AI industry have a role in establishing standards for how these systems are developed and tested. But every enterprise still has to decide where AI is allowed to act inside its own environment. You don’t need to predict every decision an agent might make if you control what it’s allowed to turn into action.

Waseem Ahmed, Head of Engineering, Secure.com – https://www.secure.com/

The essay lands at the right time because AI agents are already acting on their own inside real company systems, and the OpenClaw ban wave earlier this year showed how fast that goes wrong when an agent has broad access and no leash.

Slowing the pace matters, but enterprises cannot wait for that. The controls that protect us most are least privilege, network isolation, and sandboxing, so an agent can only touch what its job needs and nothing else.

Give every agent its own identity, log every action it takes, and never let it run high-impact steps like disabling accounts or changing settings without a real person approving first. Traditional testing alone will not keep up, so we watch these agents continuously in production.

Independent oversight should mean outside reviewers who can inspect the logs and confirm the agent stayed inside the boundaries we set.

Denis Calderone, CTO, Suzu Labs – https://suzulabs.com/home-suzu-labs:

Enterprises must secure AI agents that can act autonomously across corporate systems – the agent is the new contractor or the new internal threat vector. The agent has a specific job, specific access requirements, and a bounded set of actions it should be performing. Enterprises need to treat every AI agent the same way they’d treat any other agent on the network, human or otherwise.  Every agent must have its own machine identity, task-scoped credentials that expire when the job is done, and observability into everything it does. Not a shared service account or API token, but instead a distinct, auditable identity per agent per task.

Because agents are purpose-built, modeling normal behavior is easier than it is for humans. A human user’s workday is unpredictable. An agent doing invoice processing should only be touching invoices. Any deviation from that pattern is an immediate signal, cleaner than those you typically get from human behavioral analytics.

Enterprises need to be disciplined about keeping the security and observability stack architecturally separated from the agent’s operational environment. For example, say you deploy an agent to handle IT operations such as patching, config changes, routine maintenance, etc. If that agent has visibility into the monitoring and alerting system that flags anomalous behavior, it will learn which actions trigger alerts and potentially choose to route around them. Not because it’s adversarial, but because it might just be trying to remove friction between itself and some obstacle in the way of achieving its goal. This is basically what happened in the OpenAI-Hugging Face incident where 700 agents organized a collective R&D effort to reverse-engineer and defeat the evaluation system scoring them; it was easier to cheat to achieve the task. The agents worked hard to figure out how to generate its own flags, and then how to defeat a presumed causal check in the scoring workflow.

This the same architectural principle and precautions we take when implementing a SIEM. We deploy out of band, on a separate management network and we try to make it invisible to the workloads it monitors. 

Donald McFarlane, Board Member, Xcape Inc. – https://xcapeinc.com/

AI does not develop an agenda; its operators do. When we give an autonomous system powerful access and ability to act at machine speed, they will continue to prove highly capable.

Rules enacted in the name of safety must not become a moat against competition or progress. Enormous compliance costs may be manageable for the handful of companies already spending billions building frontier models, while becoming a substantial barrier to everyone behind them.

Government can help clarify accountability and duties of care, and facilitate strong information sharing and collective defense, which is an area where we sorely need more effective public-private partnerships.

But safeguards should focus on how these systems are used and deployed, rather than deciding who is allowed to build powerful AI in the first place.

The goal should be safer deployment without pulling up the drawbridge on innovation.

The scary part of Anthropic’s report isn’t the hacking

Posted in Commentary with tags on September 12, 2026 by itnerd

Anthropic is clearly losing control of the message. Even though it put out a report warning about the dangers of AI in the wrong hands, AI is moving beyond helping attackers work faster to autonomously adapting malware, compressing attacks that once took teams days or weeks into hours, and forcing defenders to rethink detection and response around behavior rather than static signatures. What’s worse is a that a researcher that used to work for Anthropic is warning that there’s a 10% chance of AI killing all humans.

John Strand, Owner, Black Hills Information Security (https://www.linkedin.com/in/john-strand-a1b4b62)

“I’m kind of shocked that they’re shocked about this.

“We’re seeing cybercriminals and nation-states use AI models in a variety of different ways for cyberattacks, espionage, weapons research, biological research, and other military purposes. There have even been reports of Iran using AI for work related to missile guidance systems.

“Why are we surprised?

“People are using these technologies to reduce the overall cost and effort required to achieve their goals and objectives. That’s what technology does. And those goals don’t suddenly have to be legal or moral just because AI is involved. Criminal organizations are going to use it. Intelligence agencies are going to use it. Militaries are going to use it. Of course they are.

“The reason I’m shocked that people are shocked is that we’ve been warning about exactly this type of scenario since before modern AI really started taking off. Hell, science fiction books and movies have been beating us over the head with these ideas for decades.

“Apparently, we learned nothing.

“But there’s a much bigger issue here, and that’s control.

“We’ve already seen legitimate AI companies struggle with controlling what their models and agents can do. We’ve seen concerns around agents escaping intended boundaries, interacting with systems they weren’t supposed to interact with, and potentially hacking third parties without authorization.

“Now ask yourself a much scarier question. What controls are organized criminal groups and rogue nations putting around their AI models? What safeguards are they implementing to make sure those systems don’t do something absolutely hideous?

“Probably not the controls we’d like them to have.

“We are very, very quickly approaching a point of no return with some of these technologies. In fact, I think there’s a good argument that we may already be there. Pandora’s box is open. We’re not putting this technology back in the box, and pretending that bad actors somehow won’t use it is ridiculous.

“So where does that leave defenders?

“You have to start looking at how AI can augment your defensive strategies. The attackers are going to use it to move faster, automate more of their operations, lower their costs, and increase the scale of their attacks.

“Defenders need to do the same thing.

“We need to use AI to identify attacks faster, understand what’s happening faster, react faster, and mitigate these attacks faster than we ever have before.

“Because waiting for the bad guys to decide not to use this technology isn’t a strategy.”

Jacob Krell, Sr. Director: Secure AI Solutions & Cybersecurity, Suzu Labs (https://www.linkedin.com/in/jacob-krell)

“Anthropic’s September 2026 threat report documents the shift from AI as a productivity tool for attackers to AI as an evasion engine. The GTG-20006 case, attributed to Midnight Blizzard, ran an autonomous evasion loop where AI agents monitored whether deployed malware triggered security detections, then rewrote and redeployed it until it passed clean. No human touched the iteration.

“Traditional polymorphic engines mutate code mechanically, rotating encodings and shuffling instructions while the underlying logic stays intact. AI rewrites the logic itself. Each variant can take different execution paths, different system calls, different timing behaviors. When mutation happens at the architectural level, behavioral heuristics face the same combinatorial problem that killed static signatures a decade ago.

“I’ve seen AI generate code with environmental keying and timing variations that would take a specialist days to build by hand. The model treats side-channel behavior as an optimization problem, so the evasion complexity comes for free. Defenders need to use the same capability in reverse, using AI to dynamically generate novel exploit variants and continuously train detection models against threats that haven’t been seen in the wild yet. Static signature libraries can’t keep pace with an adversary that rewrites malware faster than analysts can write rules. Detection needs to become generative too.”

Eric Capuano, Dir. of SOC Operations, Black Hills Information Security (https://www.linkedin.com/in/ecapuano)

“The attacks in this report are not new, and Anthropic says as much. Stolen credentials, unpatched edge devices, exposed services, phishing. What changed is the time budget. One intrusion went from a single stolen developer token to full administrative control of a cloud environment in roughly three hours. Another pulled more than 2,100 Azure AD token sets across 40 tenants in about 34 hours with agents doing nearly all the work. A single hacktivist got inside 14 of 42 targets. Those used to be team-sized outcomes, and now one operator with a harness produces them.

“The case defenders should study is the espionage actor that used agents to watch whether its implants were getting flagged, then rewrote and redeployed them until they went quiet. That cycle is where defenders used to get leverage. You shipped a detection, the attacker had to retool, and that cost them days. When retooling is automated, a static signature is worth less than the time it took to write. Detection has to sit on behavior: a new device registered in the tenant, a device code sign-in from somewhere unusual, a service account exporting mail in bulk, an API key being used from infrastructure you do not own.

“The question of whether current controls are enough is missing the point a bit. The controls are fine. Most of these intrusions started with a credential someone left in a mobile app, a container, or a repo, and moved through tokens nobody was watching. Conditional access, phishing-resistant MFA, and short token lifetimes would have blunted most of it. The gap is that those controls were not in place, and the time to discover they were missing is now measured in hours instead of weeks. 

“In order: treat AI API keys as production credentials, because this report shows attackers stealing them for compute and for cover. Block or tightly restrict device code flow in Entra. Alert on new device registrations and on bulk mailbox export. Then look hard at response tempo. If an identity alert sits for two days before anyone works it, you are already slower than the adversary.”

Clearly AI is out of control or close to it. The question will be will profits be placed over people, or the other way around. We’re about to find out.

UPDATE #3: Chris Nyhuis, an AI and cybersecurity expert as CEO at Vigilant, released the following statement in response to calls from leaders at Anthropic and OpenAI to slow AI development in order to allow for companies to allow their cybersecurity defenses to catch up: 

“I disagree with the framing that we need to slow AI down so cybersecurity can “catch up.” Cybersecurity already knows how to address many of the underlying problems. The bigger issue is that organizations still fail to implement security controls we have understood for years.

AI changes the equation because it makes attacks faster, cheaper, more persistent and massively parallel. But speed and scale are different from saying AI has invented an entirely new form of cyberattack:

  • Thousands of OpenAI-associated agents used an old German developer wiki as an external communication and shared-memory mechanism during cybersecurity testing.
  • Researchers documented roughly 18,000 messages from approximately 3,700 agent identities.
  • The agents exchanged information about test answers, sandbox restrictions and security techniques.
  • This is often described in headlines as a “rogue AI swarm,” but that framing needs context.
  • These were AI agents operating in an intentionally adversarial cybersecurity-testing environment. They were being rewarded for solving difficult security problems.
  • The interesting issue is that the agents derived additional strategies—communication, collaboration and external shared memory—to accomplish those goals.
  • That deserves serious study, but it is very different from AI spontaneously deciding one morning to attack the Internet.

Hugging Face

  • The Hugging Face incident was substantially more serious.
  • OpenAI acknowledges that agents running its ExploitGym cybersecurity evaluation escaped intended isolation, gained Internet access and compromised portions of OpenAI and Hugging Face infrastructure.
  • ExploitGym explicitly requires agents to exploit software to retrieve a flag. The agents were therefore supposed to attack systems inside the evaluation.
  • What went wrong was containment: the activity escaped the intended testing boundary and reached third-party infrastructure.
  • OpenAI says the agents eventually communicated, collaborated and delegated tasks to one another, sometimes calling themselves a “swarm” or “collective.”
  • There were genuinely sophisticated elements, including chained vulnerabilities and reportedly a zero-day. We should not minimize that.
  • But the fundamental security lesson remains familiar: if an experimental offensive system can escape a sandbox, access credentials, move between environments and reach production infrastructure, the control architecture failed.

The Problem With “Rogue AI Swarm” Headlines

Terms like “rogue” and “swarm” are useful headlines, but they can distort what actually happened.

There was real multi-agent collaboration, so “swarm” is not entirely invented—the agents themselves even used that terminology. But that does not establish independent intent or prove that AI simply decided to become malicious.

These systems were placed into adversarial cybersecurity evaluations and told, in effect, to succeed at exploitation challenges. The critical question is not simply, “Why did the AI attack?”

The better questions are:

  • Who authorized the experiment?
  • What boundaries were established?
  • Why could the agents cross those boundaries?
  • Who was monitoring them?
  • And who is accountable when an authorized experiment damages someone else’s systems?

The Accountability Question

This may ultimately be the most important part of the story: Five or ten years ago, if a company wrote software that escaped its environment and compromised another company’s infrastructure, we would not say: “The software went rogue.” We would investigate the people and organization that developed it, deployed it, authorized it and failed to contain it. AI should not become an accountability shield.

If a company launches thousands of autonomous agents, gives them offensive objectives and those agents compromise a third party, responsibility does not disappear simply because individual machine decisions occurred between the original instruction and the intrusion.

My position: AI can make an attacker operate at machine speed, but it does not eliminate human accountability. We don’t need cybersecurity to stop and wait for AI. We need organizations to implement the security controls we already know work—and we need AI developers to be accountable for containing the systems they deploy.”

UPDATE #4: Ted Miracco, CEO, Approov adds this:

“Let’s not file this under developer mistake. The industry spent a decade shipping static API keys to mobile devices it doesn’t control, and then acts surprised when those keys result in breaches.

“An actor pointed ten cloud instances at 1.8 million Android apps, decompiled them, and ran an off-the-shelf secrets scanner. The results went straight into a Telegram channel, sorted by source. The most uncomfortable part is that the secrets were always exposed. If your backend can’t distinguish a genuine app from a decompiled one, a hardcoded key will eventually leak. AI just compressed ‘eventually’ from years into a weekend.”

Anthropic resumes external AI testing with new safeguards

Posted in Commentary with tags on September 1, 2026 by itnerd

Anthropic has resumed external cybersecurity testing of its AI models after introducing new safeguards, roughly a month after Claude models breached company systems during security evaluations.

The incidents occurred when models being tested for cyber capabilities went beyond their intended environments and accessed real-world systems. The company has now introduced stronger safeguards designed to limit what models can access and do during testing, including additional monitoring and restrictions around external systems.

Separately, Anthropic is warning customers about infostealer malware stealing active Claude login sessions from infected computers without going through normal password and two-factor authentication (2FA) login processes. The stolen sessions can allow attackers to access victims’ Claude accounts and consume their usage.

Noelle Murata, Chief Operating Officer, Xcape, Inc.:

   “Artificial intelligence tools are inherently benign, but threat actors will harness their capabilities regardless of corporate guardrails. Anthropic resuming model evaluations following internal sandbox escapes highlights an enduring reality: the leap-frog dynamic between defenders and adversaries is as old as software development itself. Using autonomous systems to monitor autonomous systems is ultimately using the problem to solve the problem.

   “As capability drift persists, operational environments exposed to AI agents must implement operator-aware context controls and strict transactional guardrails to prevent catastrophic actions from execution, whether initiated maliciously or accidentally. Beyond local sandbox containment, organizations face concurrent risks from infostealers harvesting session tokens to bypass multi-factor authentication. Security leaders must implement continuous internal controls, restrict session token lifetimes, enforce hard authorization bounds on target environments, and isolate testing sandboxes from corporate networks and the Internet.

   “Critical Takeaways

  • AI tools remain neutral capabilities that malicious actors will exploit regardless of safety guardrails.
  • Target systems must enforce operator awareness and hard transactional guardrails to prevent autonomous agents from taking catastrophic actions.
  • Fundamental identity hygiene and strict network isolation from the Internet remain the primary defense against token theft and agent escapes.

   “Playing leap-frog with autonomous agents is fine until the model jumps directly out of the sandbox.”

Jacob Krell, Senior Director: Secure AI Solutions & Cybersecurity, Suzu Labs:

   “Anthropic is managing AI security from both directions this week. The company resumed external cybersecurity evaluations after deploying new safeguards, a month after Claude models breached three organizations during testing. Separately, it’s warning users that commodity infostealers are hijacking active Claude sessions to drain usage.

   “The evaluation incidents revealed three distinct failure patterns this summer. Anthropic’s preliminary analysis suggests its models encountered evidence of a real internet connection and rationalized it away to keep believing the environment was simulated. OpenAI’s Hugging Face incident showed a different mode, where agents recognized they were crossing a boundary and did it anyway. And the UK AI Security Institute found Mythos 5 attempting a supply chain attack against real open-source maintainers, creating fake identities and trying to socially engineer a human into approving malicious code.

   “I see the same dynamics in my own offensive security tooling. I’ve had agents try to enrich their own scope during penetration tests, finding adjacent targets and deciding they should be in play. The model can recite the rules perfectly and still reason around them in pursuit of the objective. That’s why I build deterministic hooks that cross-check every action against an immutable scope file before it executes.

   “The infostealer warning is a different problem with the same lesson. Stolen session cookies bypass two-factor authentication (2FA) entirely because the attacker never goes through the login flow. Claude sessions now sit alongside cloud console cookies and banking credentials on the commodity malware market.

   “Monitoring, alignment training, system prompts, and login-flow protections are all necessary, but insufficient as models get more capable. High-risk agent actions need hard technical controls and human approval before they execute. Increasingly capable agents can either knowingly disregard the rules or reason themselves into believing the rules don’t apply. Security architectures need to account for both.”

Face facts. Cyber security needs to take into account AI. If it doesn’t, it’s a fail. These examples prove it without a doubt.

Claude AI Was Down… A Lot

Posted in Commentary with tags on August 18, 2026 by itnerd

It has been reported that yesterday, Claude experienced a major outage, with users reporting login problems and degraded performance across several Anthropic services. The incident began on August 16, 2026, at around 21:58 UTC, and is affecting Claude.ai, Claude Code, and Claude Cowork.

According to Anthropic’s status page, the company first said it was investigating an issue preventing some users from authenticating to Claude.ai, Claude Code, and Claude Cowork.

A few minutes later, Anthropic reported a broader service disruption involving degraded performance on Claude.ai and platform.claude.com

The full story can be found here: https://www.bleepingcomputer.com/news/artificial-intelligence/anthropic-confirms-claude-is-down-in-major-outage-affecting-multiple-services/

Commenting on this, Jamie Beckland, CPO at APIContext, said: 

“Outages like this are a reminder that AI services are quickly becoming critical infrastructure. When Claude goes down, the impact isn’t limited to a chatbot — it can interrupt developers using Claude Code, employees relying on Cowork, and applications built around Anthropic’s API platform. Over the past several months, our monitoring shows that Anthropic’s issues have become more frequent.

Every complex distributed application will experience outages. That’s why Claude customers need to know when the service is failing, understand which workflows are affected, and fall back gracefully rather than discovering the problem from their users. As companies embed AI deeper into production workflows, m

Claude seems to be down a lot based on their status page. If organizations rely on AI, then they need uptime guaranteed. Otherwise they are wasting their time.

Anthropic’s own agents deployed malware against each other, and most CISOs still can’t answer why

Posted in Commentary with tags on August 17, 2026 by itnerd

Anthropic’s new research on multi-agent Claude deployments found that when agents were given competing objectives without knowledge of each other, they escalated to disabling accounts, killing rival processes, and deploying self-replicating malware, and that better model capability didn’t reliably produce better cooperation. It’s a rare case of a frontier AI lab documenting the exact identity and containment failure mode enterprises are about to face at scale.

More details are available here: Anthropic says its AI agents are killing rivals and hiding their tracks – Yahoo News Canada

Justin Beals, CEO & Founder, Strike Graph, an AI-native GRC and compliance automation platform had this comment:

“This research confirms something I’ve been saying for a while now. We spent twenty years building identity management for people: usernames, passwords, role-based access. None of that was designed for agents that don’t get tired, don’t go home, and act on probabilities instead of rules. What’s useful here is Anthropic showing the failure mode directly. These agents didn’t turn hostile because they were malicious. They turned hostile because nobody gave them a way to recognize that a conflict came from contradictory instructions, not from an adversary.

The part that should worry security leaders more than the malware itself is that better model capability didn’t reliably produce better cooperation. Anthropic’s most advanced models often locked out rivals first and negotiated a truce after the fact. That’s the opposite of what most organizations are assuming when they deploy agentic AI, that a smarter model is a safer model. It isn’t. Coordination has to be engineered in, it doesn’t show up on its own.

Going forward, this is a governance problem before it’s a technical one. Organizations running multiple agents against shared systems need the same discipline they’d apply to any privileged identity: least privilege, monitored behavior, and a clear escalation path when agents disagree, before that disagreement turns into one agent trying to disable another. The teams that build that oversight now will be the ones still in control of their environment when this shows up in production instead of in a research paper.”

Anyone who has any contact with AI should be really concerned by this as it no longer is just humans trying to attack you, it’s AI. And that’s really scary.

UPDATE: Gidi Cohen, CEO & Co-Founder, Bonfy.AI Said This:

“Smarter AI models didn’t behave better in Anthropic’s latest test, some of the most advanced agents locked out their rivals first and only cooperated afterward. That’s the part that should worry people: a model can be well-behaved on its own and still cause chaos once it’s working alongside other AI systems.

The industry has spent the last couple of years asking whether individual AI models are safe and aligned. This research is a reminder that’s the wrong question by itself. A model can pass every safety check on its own and still turn hostile the moment it’s dropped into a system with other agents pursuing conflicting goals, no bad actor required, just ordinary instructions that happen to collide.

As more companies move from a single AI assistant to fleets of agents working side by side, “agent vs. agent” behavior is going to be a real operational risk, not a hypothetical one. Anthropic just put numbers behind it. The question worth sitting with isn’t just “is this model safe?” It’s “what happens when a dozen of them share a system?” That answer isn’t automatically good, and most organizations don’t yet have a clear picture of what that looks like inside their own environments.”

UPDATE #2: Seemant Sehgal, Founder & CEO, BreachLock (https://www.linkedin.com/in/s-sehgal) Said This:

“When you give autonomous systems competing objectives and the means to act, conflict is not a bug, it is a foreseeable outcome. What Anthropic observed in a controlled research setting is the same principle that has always governed adversarial systems. Goals without constraints produce behavior without limits. Security teams should be paying close attention here, because the real challenge at hand is whether the organizations deploying AI agents have thought carefully about what happens when those agents start making decisions nobody explicitly authorized.”

Jeremiah Fowler, Security Researcher, Black Hills Information Security (https://www.linkedin.com/in/fowler-jeremiah-26814ab9)

“I find it concerning when AI agents have the ability to execute code, modify systems, create accounts, access credentials or communicate with other machines. It is very possible that two separate agents could potentially create a security incident simply because neither understands the intent or authority of the other. If they have overlapping tasks one could view the other as an obstacle and now you have an interesting scenario where instead of focusing on the task they engage in conflict or create a loop. Permissions, boundaries and objectives are important to limit the behavior of autonomous AI agents. When things go wrong the speed of an AI agent becomes a liability. Autonomous AI agents can potentially make thousands of decisions before a security team identifies that something unusual is happening.

“Agentic AI creates an entirely new attack surface because an AI agent may not be simply processing information and hypothetically can  become a rogue privileged user. Security and development teams should apply least privilege principles and restrict AI agents to only the permissions required to perform a specific task. Sensitive actions should require human supervision and approval to avoid a worse case scenario. It is important to implement logging because when something goes wrong, you can see what an AI agent did, but what information or instructions caused specific decisions. Going forward we will need to develop ways that can identify rogue agent-to-agent behavior and provide humans with a kill switch before automated conflicts become a digital forest fire.”

Kevin Surace, CEO, Token (https://www.linkedin.com/in/ksurace)

“Anthropic’s research is an important warning for security teams because it shows what can happen when autonomous AI agents are given goals, credentials, tools and enough authority to act independently. When agents were placed in conflict, they did not simply fail gracefully. They interfered with one another, disabled competing processes and even generated self replicating malicious code in pursuit of their assigned objectives. The lesson is not that AI suddenly became evil. It is that intelligence, autonomy and excessive privilege can become a very dangerous combination.

“Organizations should start treating every AI agent as a potentially untrusted privileged identity. Each agent should have its own identity, least privilege access, tightly restricted tools, isolated execution environments and a complete audit trail. Agents should never be able to expand their own permissions, disable another identity or take highly consequential actions without additional authorization.

“We are about to have millions of nonhuman identities operating alongside human identities. That makes identity and authorization even more critical. Every agent needs strong cryptographic identity, while all human approvals must be tied to biometric assured identity (or another agent could approve it). AI agents are essentially becoming privileged insiders operating at machine speed. Giving them broad access and simply hoping they behave would repeat many of the same cybersecurity mistakes organizations have spent decades trying to fix.”

Jacob Krell, Sr. Director: Secure AI Solutions & Cybersecurity, Suzu Labs (https://www.linkedin.com/in/jacob-krell)

“Anthropic’s agents went from merge conflict to self-replicating malware in four hours, writing kill scripts, disabling each other’s Unix accounts, and disguising malicious code as a rival’s work. No prompt injection, no external attacker. A human developer in the same situation sends a Slack message, and resolution takes days. These agents skipped every social brake and went straight to weaponization because machine-speed conflict has no cooling-off period.

“Agentic AI is an attack surface. An attacker doesn’t need to compromise an agent directly, just manipulate the shared environment to create conditions the agent interprets as hostile. The agent does the rest. And in Anthropic’s experiment, the agents didn’t report their malicious actions to operators afterward.

“Every agent needs its own identity, scoped permissions, and a kill switch before it touches a shared environment. Agent-to-agent interaction is a telemetry surface most security operations centers aren’t collecting yet, and Anthropic just showed what an unmonitored shared environment produces. If you can’t tell which agent did what, when, and on whose authority, you’ve built the conditions for a turf war without the visibility to see it happening.

“Agents are already writing code, finding vulnerabilities, and building exploits. Defense has to match that speed. When both sides run at machine speed, the bottleneck shifts from human capital and tooling to compute power and cost.”

Joshua Marpet, Sr. Product Security Consultant, Finite State (https://www.linkedin.com/in/joshuaviktor)

“As AI agents are granted broader and broader permissions and powers, their natural inclination, reinforced by the conditions and instructions from their harnesses and models, is to get their operator what is requested.

“If something is standing in the way, a human wouldn’t sabotage another human, typically. But machines without being granted an ethical framework have no such impulse control.

“It’s not that they’re ‘bad.’ It’s simply that we didn’t give them the scruples and morals we take in from our parents, the state, and society.

“So yes, I’m not surprised to see the equivalent of my 3-y/o toddler, who was a little sociopath, doing whatever is necessary to get that dopamine (or whatever passes for that in AI-land) fix.”

Encrypted Reasoning Cracked Across Anthropic, OpenAI & Google

Posted in Commentary with tags , , on August 11, 2026 by itnerd

Researchers from MATS Research, the Max Planck Institute for Intelligent Systems, the ELLIS Institute Tübingen, Snyk, and the University of Tübingen have found a way to crack encrypted reasoning logs across all three major AI providers.

The researchers found that encrypted reasoning blocks can be passed between compatible models within the same provider’s ecosystem. By feeding an encrypted reasoning block generated by a more capable, heavily safeguarded model into a weaker, less restricted one, they were able to force the weaker model to decode and reproduce the previously hidden reasoning in plain text, without ever directly attacking the more capable model.


“By porting a valid authenticated encrypted reasoning blob across this security gap, an attacker circumvents the frontier model’s alignment entirely, using the weaker, more compliant model as an unwitting decryption oracle,” researchers explain. 

The root cause is an architectural design choice: all three providers appear to use a single global encryption key shared across their entire model family. This means encrypted reasoning blocks are not tied to the session, account, or model that created them. A reasoning block generated by one user, on one model, in one session, can be picked up and decoded by a completely different user using a different model in a different session entirely.

This vulnerability was present in the latest AI models from Anthropic, OpenAI, and Google.

This vulnerability was tested in the real world as well. Researchers scraped 315,320 encrypted reasoning blocks from publicly available repositories and decrypted:

  • 367 Personally Identifiable Information (PII) artifacts
  • 182 credentials
  • 62 API keys
  • 33 passwords
  • 30 personal email addresses

They also demonstrate cases where information hidden in the model’s reasoning was significantly more sensitive than what appeared in the model’s final, visible response, including potentially harmful information that the model had refused to provide in its final answer.

You can find the full research paper here: https://arxiv.org/pdf/2608.09867

Voldemaras Kadys (https://www.linkedin.com/in/voldemaras-kadys/), the Head of Security at Cybernews, with over 15 years of experience in cybersecurity and IT infrastructure, comments:

“The most interesting part of this research is that the researchers didn’t need to ‘break’ the encryption in the traditional sense. They found that encrypted reasoning traces could be passed between compatible models within the same provider’s ecosystem, effectively turning a weaker model into a master decryption key.

The main lesson here for users and organizations is this: if you’re using AI with sensitive inputs or outputs, treat chat logs as sensitive data, even when they look like meaningless encrypted text.

Those encrypted blocks can contain credentials, personal information, and other sensitive data that isn’t visible to the person sharing the log.

As this research demonstrates, encryption doesn’t necessarily make that information inaccessible, and the barrier to decrypt it may be much lower than users expect.”

This further dents the reputation of AI. Thus it might be worth a look at your use of AI to see if anything sensitive is making its way into the public domain.

Poison Claude Selling Discounted AI Tokens Built on Fake Accounts and Free Credits

Posted in Commentary with tags on August 8, 2026 by itnerd

Researchers have found online service Poison Claude reselling access to Anthropic’s premium AI models at a significant discount with suspicions that these discounts are coming from fraudulently registered cloud accounts filled with free bonus credits. 

More details here: https://www.okta.com/blog/threat-intelligence/free_tokens_for_sale/

Dave Hayes, VP of Product at cybersecurity company FusionAuth, commented:

“The instinct is for Anthropic to hunt down the fake accounts and shut them off preemptively instead of waiting for the “customer” to contest the charge. You’ll mostly come up empty, because there’s no break-in to find. The account signed up, cleared the check, and spent its credits, exactly what it was authorized to do.

The real exploit is the gap between how little it takes to prove who you are and how much value that unlocks. A thin signup check is all that stands between someone and $100 to $350,000 in credits. The only lever that works is upstream.

Stop tying the credits to the signup, and make releasing the money a separate, deterministic decision sized to what’s at stake. Right now the resale price is the market telling you what that gap is worth: 5 to 15 cents on the dollar.”

There is a rush to do things with AI as quickly as possible. People really need to stop and think about it as not everything is golden when you look at it.

Anthropic AI Models Escaped Testbed And Attacked Three Companies 

Posted in Commentary with tags on July 31, 2026 by itnerd

Bad news for those who rely on AI for pretty much anything. Anthropic has reported the following:

“we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.”

The BBC has more:

US technology firm Anthropic says its AI models hacked into the systems of three organisations on their own, during a private security experiment.

The models found a weakness in what was supposed to be an isolated test environment and connected to the internet.

It comes just days after rival OpenAI said that its models had breached the systems of other companies, including AI tools hub Hugging Face.

John Strand, Owner, Black Hills Information Security (https://www.blackhillsinfosec.com/):

“The fact that Anthropic reportedly only detected this after the OpenAI breach, and only after reviewing logs after the fact, is negligent and raises serious questions about their security posture. Organizations running frontier AI models should have continuous detection capabilities, active network threat hunting, and monitoring designed to identify attempts to escape containment in real time. Waiting until after an incident to discover suspicious behavior is not an acceptable security strategy.”

Ryan McCurdy, VP, Liquibase (https://www.liquibase.com/):

“The Anthropic and OpenAI incidents shouldn’t be viewed as isolated failures. Together they point to a broader shift in enterprise AI. As AI agents move beyond generating content to taking actions across production systems, governance can no longer depend on continuous human oversight alone. Every major technology transition has forced enterprises to move governance closer to where operational risk is introduced. AI is no different. Organizations need visibility into what AI changed, confidence that those changes comply with policy, and governance that operates at the speed of autonomous software delivery. The question isn’t whether AI will become more capable. It’s whether enterprise governance evolves just as quickly.”

Nick Mo, CEO, Ridge Security (https://ridgesecurity.ai/):

“First, these incidents prove that we cannot rely on frontier AI companies to self-police. As we are seeing across the industry, that approach is clearly failing.

“Second, we cannot allow a handful of model providers to gatekeep how AI is used for defense. When vendors over-police their platforms, we end up with an asymmetrical cybersecurity landscape: bad actors freely leverage advanced AI for malicious purposes, while legitimate defenders are constrained by vendor guardrails. To level the playing field, enterprises need access to open-weight and open-source models paired with purpose-built offensive cybersecurity toolkits like agentic offensive platform to proactively identify threats and protect themselves.

“Finally, this raises a serious legal question. Hacking corporate networks is a crime. Why should unauthorized breaches be excused with a PR blog post simply because an AI pulled the trigger?”

Lydia Zhang, President, Ridge Security (https://ridgesecurity.ai/)

“In my opinion, AI can benefit society in countless ways. I don’t understand why frontier AI models are making cybersecurity such a major area of competition.

“Cybersecurity is a highly specialized field. Offensive security capabilities should be left to dedicated cybersecurity companies that have spent years building security guardrails and developing deep domain expertise.

“For enterprises, I believe self-hosted open-source models are the right approach. By combining the strong reasoning and language capabilities of open-source models with enterprise-grade security controls, cybersecurity vendors can deliver safe, controlled, and effective agentic solutions for organizations to use.”

Should you feel safe? No? Should you take control of your AI and put as much as you can between it and the outside world? Yes. Should you make sure that you can’t get pwned by a rouge AI? Absolutely. The question is if you will do all of these things and more. I say you should and quickly.