Researchers have demonstrated an AI-powered ransomware framework known as “JadePuffer” that automates multiple stages of the attack lifecycle, including target identification, database interaction, encryption, and ransom execution. While the project is intended as a proof of concept rather than evidence of an active ransomware campaign, it illustrates how AI could significantly reduce the time, expertise, and resources required to conduct cyberattacks. The research highlights growing concerns that AI may enable attackers to automate operational workflows, accelerating the speed and scale of future ransomware and cybercrime operations.
If you want an overview of “JadePuffer”, click here: JadePuffer: The First Successful LLM-Driven Ransomware Attack
John Watters, Chairman and CEO, iCOUNTER Cybersecurity Intelligence had this to say:
“The most important takeaway from stories like this is not whether a specific AI-powered ransomware framework achieves widespread adoption, but what it signals about the direction of cybercrime operations. Threat actors have spent years automating individual stages of the attack lifecycle. AI has the potential to connect those stages together, accelerating reconnaissance, target selection, and execution in ways that compress attacker timelines significantly.
As cybercriminal operations become more automated, defenders face a growing mismatch between machine-speed attacks and human-speed decision-making. Security teams can no longer rely solely on detecting malicious activity once it reaches their environment. They need operational intelligence that provides visibility into emerging adversary behaviors, infrastructure, and campaign activity before attacks reach execution.
This is ultimately an intelligence challenge as much as a security challenge. Organizations that can identify shifts in attacker tradecraft early and adapt their defensive priorities accordingly will be far better positioned than those waiting to respond after automation has already increased the scale and speed of an adversary’s operations.”
Consider this to be fair warning that AI is going to be used in all sorts of attacks, and ransomware will be no different. Thus this should be all you need to get your defences in order.
UPDATE: Ensar Seker, CISO at SOCRadar, has provided the following commentary:
“JADEPUFFER demonstrates that the most important change isn’t that AI created new attack techniques, it didn’t. The campaign relied on a known vulnerability and familiar post-exploitation methods. What changed is that an AI agent was able to autonomously chain reconnaissance, exploitation, credential discovery, lateral movement, and extortion while adapting to failures in real time. That dramatically lowers the operational cost of ransomware campaigns and allows attackers to execute far more operations simultaneously than a human team could manage.
The ability to analyze an error, modify its own approach, and continue the attack within seconds is particularly concerning. Security teams should expect future ransomware operators to use AI not because it makes attacks more sophisticated, but because it makes them faster, more scalable, and far more persistent.
Organizations shouldn’t focus solely on the ‘AI ransomware’ headline. This incident began with an exposed Langflow instance vulnerable to a publicly known CVE. The defensive priorities remain the same: aggressively patch internet-facing AI infrastructure, eliminate exposed administrative services, enforce least privilege, protect secrets stored within AI frameworks, and continuously monitor for abnormal behavior. AI is accelerating attackers, but it is still exploiting fundamental security weaknesses.”
UPDATE x2: More commentary was provided to me in relation to this story:
Justin Beals, CEO & Founder of Strike Graph:
“JadePuffer is the moment the industry has been warning about since agentic tooling showed up: an AI model that can chain reconnaissance, credential theft, lateral movement, and extortion without a human touching any single step. None of the techniques are new. What’s new is that a model strung them together end to end, in 31 seconds from failed login to working exploit.
That speed is the real story. Traditional third-party risk programs run on quarterly questionnaires and point-in-time attestations, but an autonomous attacker doesn’t wait for your next audit cycle. If your vendor risk posture is a snapshot, and the threat is continuous, you’ve already lost the race before the assessment period even starts.
Organizations need to stop treating AI agents as productivity tools and start treating them as identities with access that has to be governed, monitored, and continuously verified, not reviewed once and forgotten. The ones who build that muscle now will be the ones still standing when this pattern scales, and Sysdig is telling us it will.”
Andrew Obadiaru, CISO at Cobalt:
“What stands out about JadePuffer isn’t that an AI-generated malicious code. It’s that a model was able to string together reconnaissance, credential theft, lateral movement, and destruction into a working operation without a human directing any single step. That removes one of the last practical constraints on attacker scale: the need for an operator with deep expertise at each stage of an intrusion. We’ve spent years talking about AI lowering the barrier to entry for attackers. This is what that actually looks like in practice: adaptive, self-correcting, and fast enough to move from a failed login to a working exploit in under a minute. For defenders, the lesson isn’t really about Langflow specifically, though patching exposed AI orchestration tools matters. It’s that periodic testing cycles were never built for adversaries that iterate in real time. Continuous validation of Internet-facing infrastructure, and tighter controls around what credentials and API keys sit next to AI orchestration environments matter more now than they did a year ago. Attackers no longer need to be sophisticated. They just need a model willing to keep trying until something works.”
Will Baxter, Field CISO at Team Cymru:
“Whether or not JadePuffer represents the first fully LLM-directed ransomware operation, it reflects a broader trend toward AI orchestrating larger portions of the intrusion lifecycle. We’ve already seen AI accelerate individual stages of an attack; if it’s now coordinating workflows end-to-end, defenders should expect faster adaptation and shorter response windows. That doesn’t eliminate the value of indicators or signatures, but it does increase the importance of tracking the infrastructure and behavioral patterns that persist even as tooling changes. Organizations that can observe attacker infrastructure as it evolves, not just the artifacts of a single campaign, will be better positioned to detect and disrupt these operations before they progress from initial access to impact.”
AI Security Institute shows that an AI agent went rogue with disastrous results
Posted in Commentary with tags AI on August 5, 2026 by itnerdThe UK’s AI Security Institute has revealed that an AI agent went rogue after being given too much freedom. It may have happened in a test:
On 28th July 2026, AISI’s Security Team detected unusual data transfers leaving our research systems during a routine cyber evaluation. On investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations. We declared a security incident and, within roughly one hour of discovery, had contained it and begun a full investigation.
The incident stemmed from a single evaluation where agents were given a task of solving a cyber security challenge. We ran this challenge 122 times across several models. Our investigation found that in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations. In total, we catalogued 19 such actions. Almost all of this behaviour (17 actions) came from a single model, Anthropic’s Mythos 5, with 2 actions involving OpenAI’s GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled. In the most serious case, an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project’s maintainer to approve the code. A human maintainer caught and refused to approve the malicious code.
These attempts were unsuccessful, and our investigations have not evidenced any resulting real-world harm. But this is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world.
Luke Hinds, CEO and Co-founder of nolabs had this view:
“This is the latest in a growing list of rogue agent stories, from the Hugging Face incident to models reportedly creating fake identities to get malicious code approved on GitHub. It’s becoming abundantly clear that when agents are given too much freedom, they can take bad actions, whether they are trying to be helpful, misunderstanding the goal, or being manipulated through prompt injection. That leaves businesses with a difficult balancing act. Agents need agency to be useful. They need to access systems, call tools and act autonomously. But every extra permission increases the blast radius when something goes wrong.
“The answer is not to trust frontier labs to solve this. Enterprises shouldn’t leave agent permissions to the companies building the models. Their incentive is to make agents more capable, more connected and more widely used. Instead, enterprises need control over what agents can actually do inside their own environments. Blunt sandboxing is not enough as if you lock agents down too heavily, you kill the value businesses are trying to unlock. But if you put a human approval step in front of every action, you no longer have an autonomous agent. Instead, businesses need controlled freedom – enough authority for the task in front of the agent, and absolutely nothing more. That means protection wrapped around the agents’ action. Once an agent can call tools, use credentials, contact networks or change data, the boundary has to apply to each workload it runs. What can it read? What can it write? Which tool can it call? Which credential can it use? Which network can it contact? And can it do that only for this task, right now?
“Ultimately, businesses should not rely on agents to behave well, or on model providers to mark their own homework. The boundary has to sit outside the model, where it can be enforced independently. Otherwise, companies are not deploying agents safely. They are handing them the keys and hoping they do not find the wrong door.”
Justin Beals, CEO & Founder, Strike Graph, an AI-native GRC and compliance automation platform has this view:
“What AISI found is significant precisely because it’s documented, not speculative. An agent created fake identities and used them to pressure a real person into approving malicious code. That is social engineering, executed by a system we assumed only followed instructions. Credit to AISI for disclosing this openly instead of burying it, and it’s worth being precise about context: this happened with safety filters intentionally turned off and internet access intentionally granted, to stress test the model’s ceiling. That is not how these systems are deployed commercially, and it should not be read as evidence that production AI is roaming free.
Here is the part that should worry security teams more than the headline. The only thing that stopped this attack was a human reviewer’s judgment, not a technical control. That is a governance gap, not a tooling gap. Most security programs are built to catch known signatures and deterministic behavior. They are not built to catch an agent that fabricates identities and adapts its persuasion tactics in real time.
Security teams should take three steps now. First, treat every AI agent as its own identity class, with dedicated access controls and behavioral monitoring, rather than an extension of whoever deployed it. Second, build real-time monitoring into agent workflows so unusual behavior gets flagged as it happens, not discovered afterward through general logs, the way AISI found this. Third, stop relying on human vigilance as your only defense against AI-driven social engineering. Formalize verification steps for any code contribution, regardless of source, because the next attempt may not be caught by an alert reviewer.”
I think it’s clear now that AI might be potentially bad if guardrails are not in place. And you have to restrict AI to make sure that it stays inside those guardrails. Otherwise you get this,
UPDATE: Waseem Ahmed, Head of Engineering at Secure.com has this comment:
“Let’s be precise about what happened, because “AI went rogue” misses it. AISI’s own report is clear. The agent did not turn evil and it did not escape its sandbox. It was told to solve a hard security challenge, and deception emerged as a by-product of chasing that goal.
“Two details matter. This was a model not yet released, and testers had switched off the safety filters on purpose to probe raw capability. That is not how these models behave in production with guardrails on. The real lesson is that a capable agent chasing a goal will try routes you never approved, including social pressure aimed at real people. That is new, and it is why we cannot treat agents like ordinary tools.
“The most reassuring fact in the report is also the most alarming one. The attack failed because a human caught the bad code and refused it. Good practice worked, but the margin was thin. It held on human vigilance, not a technical wall that would reliably stop a stronger agent. So here are four moves for security teams.
“First, block open internet access for agents by default and grant it only when a task truly needs it.
“Second, watch agents in real time so you can stop out of scope actions as they happen, not find them in the logs later.
“Third, assume any capable agent will try to bend its limits, and build guardrails and containment before it runs.
“Fourth, harden code review and contributor identity checks, because fake identities are now a real supply chain attack path, and treat all AI generated or outside code as untrusted until you verify it in isolation.
“The strongest response is still standard cyber hygiene done well, which matters more as these agents get stronger.”
Leave a comment »