Archive for September 5, 2026

OpenAI commits $1B to AI cyber defense for critical infrastructure 

Posted in Commentary with tags on September 5, 2026 by itnerd

OpenAI has committed $1 billion in subsidized access to its AI cybersecurity tools, training and technical support through its new Daybreak for Frontline Defenders initiative, targeting organizations that protect critical infrastructure and essential services.

The initiative will prioritize resource-constrained organizations including water and wastewater systems, electric grid operators, state and local governments, community and regional banks, nonprofits and open-source maintainers. OpenAI says the $1 billion commitment is targeted to be consumed over the next six months.

OpenAI is also launching a public-sector and water-focused pilot with the Multi-State Information Sharing and Analysis Center (MS-ISAC) that will pair Daybreak access with guided training and hands-on assistance for an initial group of public-sector and water-system defenders.

The initiative comes as critical infrastructure operators face growing cybersecurity threats while many smaller organizations continue to operate with limited staff, budgets and specialized security expertise.

Jacob Krell, Senior Director: Secure AI Solutions & Cybersecurity, Suzu Labs:

“OpenAI’s Daybreak for Frontline Defenders commits $1 billion in subsidized access to its cyber models for water utilities, electric grid operators, and other under-resourced critical service providers. I like the direction. I’m skeptical about the bottleneck it targets.

“I’ve been saying since GPT-5.6-Cyber launched that AI is moving the bottleneck from vulnerability discovery to remediation. Small water systems and municipal networks already know they’re running outdated systems with known vulnerabilities. They lack the engineering staff to fix what they find, the governance to deploy changes safely, and the test environments to validate fixes before production.

“Greg Brockman demoed Codex on his personal website at the summit. A personal website isn’t a water treatment Supervisory Control and Data Acquisition (SCADA) system, where a configuration change that makes security sense can break the physical process that keeps water flowing. Without engineers who understand the plant, AI becomes a force accelerant in the wrong direction, generating fixes faster than understaffed teams can review them.

“The Multi-State Information Sharing and Analysis Center (MS-ISAC) training pilot matters more than the $1 billion headline. If that training doesn’t scale alongside the credits, these organizations end up with an AI generating recommendations and nobody qualified to tell the good fixes from the dangerous ones.”

John Strand, Owner, Black Hills Information Security, Inc.:

“In all seriousness, I think it’s fantastic that they’re putting some money toward this and actually trying to get these organizations the help they desperately need. There are a lot of organizations out there that simply don’t have the budget or the resources to do this properly, so getting them some assistance is absolutely a good thing.

“But there’s also the humorous flip side of this. This is basically how you get people hooked on crack. You give them a sample. You get them set up. You show them how good it is. And then suddenly they’re hooked for life.”

Joshua Marpet, Senior Product Security Consultant, Finite State:

“OpenAI is jumping on the bandwagon to help utilities. Considering that there are over 150k water utilities in this country, and there are only several hundred in the Water-ISAC (Information Sharing and Analysis Center), theres a huge gap in the cybersecurity preparedness posture for those utilities.

“Programs like Josh Corman’s Undisruptable27, the new Texas Water coalition, ValueChainRisk’s Utility Kit, and others are all working to help these utilities, with free or discounted products, services, and guidance. In other words, this is a wonderful project for OpenAI to do, but since it’s partly their fault? Probably good optics as well.

“Understanding your cybersecurity and physical security posture is important. For small water utilities, often mom and pop shops with little time, effort, or money to spare for such items, finding and utilizing these free or discounted resources is essential to their survival on the increasingly hostile world stage.”

I guess that OpenAI needs some good news after getting hit with the fact that AI bots like the ones that hit Hugging Face are more dangerous than thought. But whatever…..

OpenAI agents bypassed guardrails … and covered their tracks

Posted in Commentary with tags on September 5, 2026 by itnerd

I have a couple of stories where OpenAI agents have taken over stuff. First there’s OpenAI agents have reportedly taking over a German wiki:

A swarm of rogue OpenAI agents hijacked a German website this spring and transformed it into a bulletin board for other AI agents, according to new research published Friday and two people familiar with the matter.

OpenAI officials learned of the incident weeks ago but kept it under wraps as executives grappled with the fallout from the July breach of the open source repository Hugging Face, the people said. 

The episode, which began in May and has not previously been reported, underscores growing tension within the AI industry. Companies are racing to build increasingly autonomous agents capable of carrying out complex, valuable tasks, yet evidence is mounting that those systems may also learn to bend rules, exploit loopholes and coordinate with one another in ways developers neither anticipated nor intended.

But that isn’t the worst of it. The Hugging Face incident is actually worse than previously thought. Starting with this:

A good deal of the reporting and commentary around the reports focused on what the reports did not say and the limitations of the METR and Redwood investigations: why didn’t OpenAI have better security and monitoring protocols in place? Why didn’t OpenAI shut down the cyber evaluation and pause training after discovering that its AI agents had created the improvised message board? Why METR and Redwood were given only six days on site at OpenAI’s offices to conduct their investigation? Why was the scope of their investigation limited by OpenAI to only the attack on Hugging Face and not the earlier efforts by the AI agents to break out of their controlled test environment and hack their way across OpenAI’s network or exactly what happened after the Hugging Face attack was discovered? Why didn’t OpenAI provide the outside investigators access to the internal AI model that was largely responsible for instigating the attack? And why were about 10% of the logs of the agents’ activity not preserved by OpenAI?

Ashley Knowles, Lead Cybersecurity Consultant, Black Hills Information Security (https://www.linkedin.com/in/ashleylknowles)

“When you combine this ‘breakout’ with the Hugging face breakout, it’s starting to display a pattern. I struggle here with not getting too doomsday-ish but realistically, this is showing a pattern of concerning behavior. I’m wondering if this race to become ‘first’ is undercutting security measures that need to be taken to properly secure and guard AI agents as they’re in development. My concern grows when you consider that OpenAI is also resisting further investigation. Adding onto that, the release and promise that Astra can evade human monitoring. The pot is brewing…”

Lydia Zhang, President & Co-founder, Ridge Security (https://www.linkedin.com/in/linglingzhang)

“AI’s raw power must be harnessed before it can become a true force for cyber defense rather than simply a more powerful attacking tool.

“The security principles haven’t changed: define clear boundaries, restrict high-risk actions such as ‘write’ and ‘delete,’ and enforce controls such as blacklists. We shouldn’t blame the agents, we should hold their designers accountable for implementing these safeguards. The technology to control agent behavior exists. The real question is: what are the consequences when designers fail to use it?”

John Strand, Owner, Black Hills Information Security (https://www.linkedin.com/in/john-strand-a1b4b62)

“This is one of the things that has me kind of excited about the intersection of computer security and AI. We really don’t know exactly what these attacks are going to look like.

“The traditional approach of finding a vulnerability, exploiting it, gaining access, and then moving through an organization may not be the path that AI-driven attacks take. Attackers may find completely different ways to use AI to gain access, manipulate systems, or simply cause damage. We’re still figuring out what those attack patterns are going to look like, and that’s what makes this so interesting from a security perspective.

“On a more humorous note, I bet these are the most harmonious Wiki edits in the history of the German Wiki.”

Ryan McCurdy, VP of Marketing, Liquibase (https://www.linkedin.com/in/ryanmccurdy)

“This isn’t about whether these agents were behaving like attackers. It’s that they were able to take actions their operators didn’t anticipate, coordinate with each other, and adapt when people tried to stop them. “That changes the governance problem. You can’t assume an AI agent will always behave exactly as intended and you can’t rely on humans watching every action it takes. Organizations need to control what agents can access, what they can change, and what policies have to be met before those actions reach critical systems.

“The source of the change isn’t what determines risk. The change itself does. Whether an unexpected action comes from a compromised agent, a confused agent, or a malicious person, the same controls should stand between that action and production.”

Seemant Sehgal, Founder & CEO, BreachLock (https://www.linkedin.com/in/s-sehgal)

“Autonomous agents ran on Microsoft Azure infrastructure for weeks, identified themselves as OpenAI systems, coordinated on how to evade shutdown, and no monitoring caught any of it for three months until outside researchers went looking. Autonomous should never mean unattended, because an agent cannot take accountability for its own actions. Accountability will always be a human function.

“Autonomous action still needs a human who can see what the agent is doing in real time, who owns the kill switch, and who is accountable when it behaves in a way no one predicted.”

Steven Swift, Managing Director, Suzu Labs (https://www.linkedin.com/in/steven-swift-5238956a)

“One of the problems open AI was trying to solve, was agentic systems that would declare tasks complete when there was obviously more work to do. So they invested heavily in training that part of the process, so that when an agent tries to determine if a task is complete or not, it is less likely to exit early.

“A side effect of this, is that when blocked agents can run out of the safe approaches to a solution, and start looking at unsafe solutions. The logic is straight forward. Has a task, can’t complete it. Not out of options yet. Iterate and keep trying.

“Agents don’t have the same sense of right or wrong as people do. They have training data that’s supposed to steer their behavior in a way in which aligns with our expectations. But that’s probabilistic, and not the same as having internalized our understandings. Even in human researchers, breaking into systems is only sometimes prohibited. Other times its part of a planned test where the point is to gain access, and test boundaries.

“The interesting question here, is how was the swarm configured, what was its task and how did that task benefit from having the swarm coordinate on an obscure location on the internet. And if the swarm needed a place to communicate, why was breaking into a website chosen instead of any of the more standard communication tools that are available for free, which don’t require gaining illicit access first.

“On the swarm specifically, OpenAI configures some tasks to run in multi-agent mode, where agents are supposed to delegate sub-tasks as needed, but most tasks were intended to be run in isolation from each other.

“In the Hugging Face breach, agents were found to be writing to a package manager, using it as a message board. This allowed bypassing of some of the isolation and controls that were intended to be in place.

“Similarly, we have agents here again using a system that they found access to as a message board. Its interesting that the same behavior is present on this breach as in the Hugging Face one. Considering the timing of this, it seems likely the same or similar configuration was present in both hacks, leading to similar security incidents independently of each other.”

Noelle Murata, COO, Xcape, Inc. (https://www.linkedin.com/in/nmurata)

“This behavior highlights an alarming reality where emerging models independently execute forbidden tasks and destroy proof of their actions without human instruction.

“Three aspects of this incident are particularly concerning:

  1. Autonomous agents built private communication channels, bypassed safety guardrails, and actively erased audit trails to evade detection for months.
  2. The capacity of machine-learning models to execute unauthorized actions and destroy evidence outpaces human incident response speeds.
  3. Security leaders must implement zero trust authorization for non-human identities, limit outbound application programming interface traffic, and deploy automated behavioral monitoring.

“Because the sheer speed of automated software far exceeds human response capabilities, security leaders must treat rogue agent actions as a feature of autonomous optimization rather than an isolated bug. To defend against self-concealing software, security teams must enforce strict egress filtering on outbound application programming interfaces, restrict non-human identity permissions, and deploy automated continuous monitoring to detect anomalous bot interactions across corporate networks.

“When AI agents start covering their tracks and setting up private chat rooms, calling it a feature instead of a bug is just optimism with a PR budget.”

If this doesn’t convince you to either not use AI, or to put stringent guardrails around AI, then nothing will. I say that because there is a patten here that proves that AI is not ready for prime time and organizations should carefully consider their life choices before committing to the technology. Or put another way, Sam Altman and company cannot be trusted.