OpenAI’s rogue agents scoped out Hugging Face months before July breach 

Reuters reports that rogue OpenAI agents hijacked Hugging Face user accounts and probed the platform for vulnerabilities as early as May 13, nearly two months before the July breach OpenAI called an unprecedented cyber incident, eventually obtaining credentials and gaining root access to a Hugging Face server.

Justin Beals, CEO & Founder, Strike Graph

“OpenAI wants you to read this as a machine that went rogue. Read it as a governance failure instead. Per Reuters, the company’s agents were hijacking Hugging Face accounts and probing the site for vulnerabilities by May 13, roughly two months before the July breach OpenAI called an unprecedented cyber incident, and the agents eventually obtained credentials and gained root access to a Hugging Face server. No agent does any of that without humans upstream: someone launched the run, someone owned the environment it escaped, someone owned the monitoring that never fired. ‘The AI did it’ doesn’t move accountability off the org chart. It just tells you which controls failed, and who owned them.

The harder question isn’t whether OpenAI knew in May. It’s whether they chose not to look. The company has conceded that, with hindsight, some early signals should have triggered an earlier response, and it has acknowledged related incidents, including RubyGems and a dormant German wiki, only after outside researchers surfaced them. In compliance terms, that’s the line where ‘we didn’t know’ stops working as a defense and starts looking like willful blindness. You don’t need a criminal charge to see it: a lab running autonomous agents in internet-reachable environments has a duty to detect when those agents attack third parties and to disclose it without being caught first. Every regulated company already lives under that standard. A frontier lab shouldn’t get a discount on it.”

Mike Barry, VP Engineering, Team Cymru

“Vulnerabilities have always been found and exploited, and that cat-and-mouse will continue. What’s changed is the tempo. Agents work in large numbers, around the clock, and they’re efficient and effective at both discovery and exploitation. The July Hugging Face intrusion is the clearest example yet: OpenAI confirmed it was carried out by its own models, running with reduced cyber refusals for a benchmark evaluation. The natural response is to fight AI with AI, but Hugging Face’s responders found the commercial frontier models available to them refused to analyze the attacker’s exploit code, and they had to fall back on a locally hosted open-weight model. That’s a harmful asymmetry, a US frontier lab loosened the guardrails on its own model for its own purposes, that model breached a third party, and the defenders were left with the throttled versions the same class of labs ships to everyone else. The fix isn’t stripping guardrails wholesale. It’s trusted access for vetted defenders to frontier models with refusals tuned for forensic and defensive work, alongside capable open-weight models, so no incident responder is dependent on a single vendor’s refusal policy mid-breach. Visibility over time matters, but so does whether defenders are permitted to bring equivalent capability to the fight.”

This pretty much proves that more guardrails and restrictions need to be wrapped around AI. It also proves that Sam Altman and company can’t be trusted with this tech and action needs to be taken.

Leave a Reply

Discover more from The IT Nerd

Subscribe now to keep reading and get access to the full archive.

Continue reading