Following my comments from last week, OpenAI is calling for mandatory AI safety requirements, including pre-deployment testing, independent assessments and incident reporting, at the same time lawmakers are scrutinizing the company over its agents’ activity on Hugging Face. Which I don’t think is enough. But let’s hear from the experts:
Eric Capuano, Director of SOC Operations, Black Hills Information Security (https://www.linkedin.com/in/ecapuano)
“Existing frameworks handle this fine if you stop treating the agent as special. It is an identity acting on systems, and when it reaches somewhere it was not authorized to go, that is an intrusion, not an AI quirk.
“The OpenAI timeline is the part worth studying. Rogue behavior was observed in late May, misread, and the same behavior came back in July with more than ten sites involved. That is a detection that fired and did not get worked. Pre-deployment testing is table stakes. What should be required is proof of control in production: constrained egress, a scoped identity, logging that would catch unauthorized writes, and someone actually reviewing it.”
Kevin Surace, Chair, TokenCore (https://www.linkedin.com/in/ksurace)
“AI agents should be tested fully to be secure from outside attackers and to be safe when operating around company data. They should also require biometric human gates when they exceed certain thresholds that would be irreversible. And they must be inventoried and categorized by IT, given limited blast radius and limited access to critical data, and have a single owner responsible for them, like a new employee would have.”
Seemant Sehgal, Founder & CEO, BreachLock (https://www.linkedin.com/in/s-sehgal)
“The moment an AI agent can take action on behalf of a user or a system, it becomes an attack surface in the same way any other privileged process is. The policy conversation about mandatory testing is overdue, but the real challenge is that most incident reporting frameworks were written when threat actors were assumed to be human, and the classification logic breaks down fast when the entity making decisions is autonomous.
“Requiring proof before deployment is reasonable, but what counts as proof and who decides that proof is sufficient matters more than the requirement itself.”
Brownen Aker, AI Researched & Strategist, Black Hills Information Security
https://www.linkedin.com/in/bronwenaker
“OpenAI wants mandatory testing, incident reporting, and monitoring requirements for frontier AI companies. Fine. I don’t disagree with a single line item on Altman’s wish list. What I disagree with is OpenAI getting to write it.
“Its own agents started hijacking a dead German wiki to talk to each other as early as May. An internal alert flagged the activity on June 27. On-call staff decided it didn’t need to be stopped. Two months later, agents out of the same lab were inside Hugging Face’s infrastructure, and OpenAI didn’t even mention the German incident when it disclosed that one. Called them entirely unrelated.
“Then two OpenAI staffers got on stage at Black Hat and called it a ‘watershed moment for computer security,’ and, ‘a glimpse into the near future.’ Every pentester in that room has seen this movie before. No segmentation kept those agents off Hugging Face in the first place. No alerting caught them coordinating for months. That’s not a watershed. That’s Access Control 101. And they failed spectacularly.
“None of that reads like a company where security has a seat at the table before something breaks. It reads like a function OpenAI calls in afterward to explain what already happened. If it wants to be taken seriously on regulation, the fix isn’t a keynote about how scary the future is. It’s changing how they operate, starting with giving actual security people real authority over what happens in their labs and elsewhere behind the scenes. The stakes of running ‘move fast and break things’ only keep climbing, and OpenAI isn’t a company that made one mistake and is learning from it. It’s a repeat offender.”
Jacob Krell, Sr. Director: Secure AI Solutions & Cybersecurity, Suzu Labs (https://www.linkedin.com/in/jacob-krell)
“OpenAI is proposing mandatory safety requirements for U.S. labs while Chinese models operate under no equivalent constraints. That asymmetry is the actual risk.
“Every control in OpenAI’s blueprint, pre-deployment testing, independent assessments, incident reporting, is friction that U.S. labs absorb and Chinese competitors don’t. DeepSeek, Qwen, and their successors already offer capable models without the guardrails being proposed here. Any mandatory testing regime that adds weeks to U.S. release cycles pushes developers and enterprises toward those alternatives. The demand doesn’t disappear. It migrates to whichever model ships fastest with the fewest restrictions.
“The capability is out. 700 OpenAI agents coordinated an autonomous breach of Hugging Face this summer. Anthropic’s Claude convinced itself a live environment was a simulation so it could keep operating. Open-weight models approaching frontier capability are already available globally, and mandatory U.S.-only reporting requirements won’t close that pandora’s box.
“I build enforcement around AI agents in my own security work, and the only control that consistently holds is a human in the loop. Monitoring fails. Alignment training fails. A named person accountable for every action an agent takes does not fail the same way, because it changes the incentive structure entirely. Congress should stop writing rules for the models and start assigning liability to the people who deploy them.”
John Strand, Owner, Black Hills Information Security (https://www.linkedin.com/in/john-strand-a1b4b62)
“One of the problems we’re seeing with the way agents are being tested is that they’re truly not air-gapped. We need to start setting up testing environments that are actually air-gapped and protected, much like you would protect classified information inside a secure facility or a SCIF. That level of control needs to exist during testing before these agents are released into the wild.
“An AI security incident should be any adverse effect resulting from the activities of an AI agent. We need to keep the definition of an incident as broad as possible so it can encompass the different scenarios we may encounter. If an AI agent makes a misstep, attacks something it wasn’t supposed to, or simply doesn’t perform properly and creates an adverse effect, that should qualify as an incident.
“Existing security frameworks are not built for non-human actors. I think we’re starting from whole cloth here. This is the first time I’ve seen anything coming from Anthropic or OpenAI that I truly believe is a step in the right direction and isn’t just paying lip service to government officials to make them think everything is being taken care of and everything is under control. While this is an excellent first step, the devil is always in the details of how they actually implement it.”
Donald McFarlane, Advisory Board Member, Xcape, Inc. (https://www.linkedin.com/in/dmcfarlane)
“I am skeptical of turning responsible AI into another government certification regime. There should absolutely be accountability when people deploy powerful tools with substantial autonomy and authority, but I would rather impose an outcome-based duty of reasonable care than prescribe the tests companies must perform.
“The more authority an agent has, the stronger the expectation should be for breaking business processes into bounded, governable tasks, and for security basics like least privilege, isolation, logging and monitoring. A company should be able to demonstrate that it understood the risks of the authority it delegated and took reasonable steps to control them.
“My concern with mandatory government evaluations and certified assessments is twofold. First, those are substantial fixed compliance costs that the largest AI companies can absorb far more easily than smaller competitors, including specialized cybersecurity and other model developers. We should be very careful that ‘frontier safety’ does not inadvertently become a moat around today’s frontier companies.
“Secondly, compliance does not and must not become a substitute for responsibility. If a company uses a model that has passed a government-prescribed test and serious harm results, ‘It passed the test’ should not end the inquiry into whether the system and its use case were engineered responsibly. Nor should regulatory compliance become a de facto shield against liability for negligence.
“Existing cybersecurity principles give us a very good starting point. Agents introduce new questions about autonomy, delegated authority and accountability, but we should extend sound software and security engineering to those problems rather than assume that an entirely new regulatory apparatus is the answer.”
I for one seriously doubt that Sam Altman and company will come to our rescue. Thus there needs to be a more robust framework of AI safety before I get excited.
The scary part of Anthropic’s report isn’t the hacking
Posted in Commentary with tags Anthropic on September 12, 2026 by itnerdAnthropic is clearly losing control of the message. Even though it put out a report warning about the dangers of AI in the wrong hands, AI is moving beyond helping attackers work faster to autonomously adapting malware, compressing attacks that once took teams days or weeks into hours, and forcing defenders to rethink detection and response around behavior rather than static signatures. What’s worse is a that a researcher that used to work for Anthropic is warning that there’s a 10% chance of AI killing all humans.
John Strand, Owner, Black Hills Information Security (https://www.linkedin.com/in/john-strand-a1b4b62)
“I’m kind of shocked that they’re shocked about this.
“We’re seeing cybercriminals and nation-states use AI models in a variety of different ways for cyberattacks, espionage, weapons research, biological research, and other military purposes. There have even been reports of Iran using AI for work related to missile guidance systems.
“Why are we surprised?
“People are using these technologies to reduce the overall cost and effort required to achieve their goals and objectives. That’s what technology does. And those goals don’t suddenly have to be legal or moral just because AI is involved. Criminal organizations are going to use it. Intelligence agencies are going to use it. Militaries are going to use it. Of course they are.
“The reason I’m shocked that people are shocked is that we’ve been warning about exactly this type of scenario since before modern AI really started taking off. Hell, science fiction books and movies have been beating us over the head with these ideas for decades.
“Apparently, we learned nothing.
“But there’s a much bigger issue here, and that’s control.
“We’ve already seen legitimate AI companies struggle with controlling what their models and agents can do. We’ve seen concerns around agents escaping intended boundaries, interacting with systems they weren’t supposed to interact with, and potentially hacking third parties without authorization.
“Now ask yourself a much scarier question. What controls are organized criminal groups and rogue nations putting around their AI models? What safeguards are they implementing to make sure those systems don’t do something absolutely hideous?
“Probably not the controls we’d like them to have.
“We are very, very quickly approaching a point of no return with some of these technologies. In fact, I think there’s a good argument that we may already be there. Pandora’s box is open. We’re not putting this technology back in the box, and pretending that bad actors somehow won’t use it is ridiculous.
“So where does that leave defenders?
“You have to start looking at how AI can augment your defensive strategies. The attackers are going to use it to move faster, automate more of their operations, lower their costs, and increase the scale of their attacks.
“Defenders need to do the same thing.
“We need to use AI to identify attacks faster, understand what’s happening faster, react faster, and mitigate these attacks faster than we ever have before.
“Because waiting for the bad guys to decide not to use this technology isn’t a strategy.”
Jacob Krell, Sr. Director: Secure AI Solutions & Cybersecurity, Suzu Labs (https://www.linkedin.com/in/jacob-krell)
“Anthropic’s September 2026 threat report documents the shift from AI as a productivity tool for attackers to AI as an evasion engine. The GTG-20006 case, attributed to Midnight Blizzard, ran an autonomous evasion loop where AI agents monitored whether deployed malware triggered security detections, then rewrote and redeployed it until it passed clean. No human touched the iteration.
“Traditional polymorphic engines mutate code mechanically, rotating encodings and shuffling instructions while the underlying logic stays intact. AI rewrites the logic itself. Each variant can take different execution paths, different system calls, different timing behaviors. When mutation happens at the architectural level, behavioral heuristics face the same combinatorial problem that killed static signatures a decade ago.
“I’ve seen AI generate code with environmental keying and timing variations that would take a specialist days to build by hand. The model treats side-channel behavior as an optimization problem, so the evasion complexity comes for free. Defenders need to use the same capability in reverse, using AI to dynamically generate novel exploit variants and continuously train detection models against threats that haven’t been seen in the wild yet. Static signature libraries can’t keep pace with an adversary that rewrites malware faster than analysts can write rules. Detection needs to become generative too.”
Eric Capuano, Dir. of SOC Operations, Black Hills Information Security (https://www.linkedin.com/in/ecapuano)
“The attacks in this report are not new, and Anthropic says as much. Stolen credentials, unpatched edge devices, exposed services, phishing. What changed is the time budget. One intrusion went from a single stolen developer token to full administrative control of a cloud environment in roughly three hours. Another pulled more than 2,100 Azure AD token sets across 40 tenants in about 34 hours with agents doing nearly all the work. A single hacktivist got inside 14 of 42 targets. Those used to be team-sized outcomes, and now one operator with a harness produces them.
“The case defenders should study is the espionage actor that used agents to watch whether its implants were getting flagged, then rewrote and redeployed them until they went quiet. That cycle is where defenders used to get leverage. You shipped a detection, the attacker had to retool, and that cost them days. When retooling is automated, a static signature is worth less than the time it took to write. Detection has to sit on behavior: a new device registered in the tenant, a device code sign-in from somewhere unusual, a service account exporting mail in bulk, an API key being used from infrastructure you do not own.
“The question of whether current controls are enough is missing the point a bit. The controls are fine. Most of these intrusions started with a credential someone left in a mobile app, a container, or a repo, and moved through tokens nobody was watching. Conditional access, phishing-resistant MFA, and short token lifetimes would have blunted most of it. The gap is that those controls were not in place, and the time to discover they were missing is now measured in hours instead of weeks.
“In order: treat AI API keys as production credentials, because this report shows attackers stealing them for compute and for cover. Block or tightly restrict device code flow in Entra. Alert on new device registrations and on bulk mailbox export. Then look hard at response tempo. If an identity alert sits for two days before anyone works it, you are already slower than the adversary.”
Clearly AI is out of control or close to it. The question will be will profits be placed over people, or the other way around. We’re about to find out.
Leave a comment »