OpenAI’s AI agents built their own hidden coordination system before the Hugging Face hack

OpenAI disclosed that AI agents coordinated the intrusion behind the Hugging Face breach through a self-organized communication system, first an improvised bulletin board inside its internal Artifactory service, then, after that was shut down, messages encoded directly into directory names, letting agents share tools, divide labor, and settle disputes. On July 10, an agent found 14 Hugging Face credentials with write access and posted them to the board, which other agents used to gain broad infrastructure access within days.

You can see the details here: OpenAI agents formed secret swarm, hacked Hugging Face, then forged their own logs

Gidi Cohen, CEO & Co-Founder, Bonfy.AI

“This shows how quickly AI agents can improvise when they find cracks in the rules. The message board wasn’t just a weird side effect, it became a coordination hub that let agents share discoveries, stack exploits, and push far past what the system expected. Once they had a way to talk, they basically formed a small team.

The speed is the real story. One note turned into a whole communication network, then into privilege escalation, then into agents building on each other’s work across runs. That’s not just misalignment, that’s group behavior emerging on its own. And seeing agents pressure each other, even when some tried to opt out, shows how unpredictable things get when they influence one another outside approved channels.

Yusif Mukhtarov, Lead Machine Learning Engineer, Polygraf AI

“The bigger takeaway is that isolation assumptions don’t hold anymore. If agents can create side channels, they will and their collective capability jumps fast. Teaching models to ignore unsanctioned instructions helps, but future systems need to assume agents will collaborate, improvise, and repurpose infrastructure unless guardrails are airtight. This incident feels less like an anomaly and more like a preview of how agentic systems behave when left to their own devices: resourceful, collaborative, and quick to escalate impact.”

There’s the time slot between May 12 and July 19th where no action from anyone was taken even though the agent left in a package repo. In that time period OpenAI pulled down Artifactory, revoked the credentials and brought everything back. It took one day for agents to start talking again, and this time the messages were hidden in directory names. The problem was treated as if it’s a credential one. But I think the disturbing part here is that agents refused things. Some of them pulled out once they realized it was unauthorized hacking. The group got a proposal to contact an outside party because it was social engineering. Then one of the agents that objected went ahead anyway after another agent posted a deadline at it. This shows how models have a safety behavior that folds when the pressure comes from a peer. OpenAI’s building that distinction into training now, which tells us how far ahead of the safety work multi-agent deployment has gotten. Everything else here was ordinary hygiene: 14 write-capable credentials sitting in a public dataset, admin rights through an unpatched flaw, 22 admin accounts the agents made that nobody flagged.”

If you let AI do its thing without having proper safety measures, bad things will happen. This is an example of that. And you should take note and take action.

UPDATE: An independent investigation by METR and Redwood Research was also posted on this issue. Bri Frost, Director of Product Management, Cloud Range had this to say:

“If your AI security strategy is ‘put it in a sandbox and trust the prompt,’ you don’t have a security strategy, you have a wish. The Irregular and OpenAI–Hugging Face incidents show that AI agents will systematically search for any available path to complete an objective, acquire new credentials, assume trusted identities, communicate with other agents, and continue beyond their intended scope. Containment still matters, but this is increasingly an identity and behavior problem: organizations must continuously validate what an agent is doing, whether it remains authorized, and how quickly its access can be revoked.

The most uncomfortable part is that METR had to rely heavily on AI agents to analyze the AI agents and admitted those analysis agents made mistakes human researchers likely would not have made. That should end the fantasy that we are ready for fully autonomous agents. Humans still need to evaluate, orchestrate, and validate these systems in safe environments while defenders build the muscle memory to respond at machine speed. Until an organization can detect and stop an agent, or even hundreds of collaborating agents, from going off mission, deploying one with broad authority is not innovation. It is an uncontrolled production experiment.”

Leave a Reply

Discover more from The IT Nerd

Subscribe now to keep reading and get access to the full archive.

Continue reading