Trump, tech CEOs sign voluntary AI safety accord

President Donald Trump and executives from OpenAI, Anthropic, Google, Meta, Nvidia and xAI signed a voluntary AI safety accord calling on companies to implement multiple layers of controls and oversight as increasingly capable AI systems raise cybersecurity and other safety concerns.

The agreement calls for companies to monitor AI models during training and deployment for dangerous capabilities, including whether models could hack or access technical systems in unintended ways, assist with biological, chemical or nuclear threats, or evade human control. Companies also agreed to maintain internal teams responsible for verifying that safeguards and monitoring systems are working as intended.

The accord calls for independent external evaluations, board-level oversight and sharing relevant safety information among participating companies and with the federal government. Trump described the agreement as “morally binding,” rather than a regulatory requirement.

Denis Calderone, CTO, Suzu Labs:

“1. The whole document rests on the word independent

“Layer three asks every signatory to partner with an independent external auditor or evaluator, and we already have a case study. METR and Redwood Research investigated the Hugging Face incident and documented their own limitations plainly. OpenAI defined the investigation window, so agent behavior before and after it fell out of scope. They could not query the model behind most of the activity. They had no direct access to OpenAI infrastructure and had to request datasets. OpenAI could redact the findings. With six days on site and roughly 1,300 transcripts to get through, they delegated much of the analysis to GPT-5.6 Sol, one of the models implicated in the incident, and wrote that they could not rule out being misled by it. That is a serious team doing careful work on terms set by the company being examined. The accord never says who accredits these evaluators, what methodology applies, or where the scope boundary sits.

“2. I expect this to produce press releases, not audits

“So they get to hire an evaluator of their own choosing, and ninety days from now we start reading about how rigorous their safety reviews turned out to be. I worry that these reports will be used more for as marketing tools rather than true assurance. If a client handed me this as their AI governance program I would write it up as a policy with no evidence of operation.

“3. The timing on this is ironic

“On September 14 the president called AI destroying humanity a “HOAX” and said a “SICK conspiracy” was running against AI and data centers. Fifteen days later he signed a document he called “morally binding” and “almost like a constitution.” That’s some pretty strong rhetoric to be followed by this accord. Maybe the administration doesn’t see AI as destroying humanity, but it apparently is a concern.

“4. The industry has been disclosing, information already, but it’s been pretty selective.

“These companies are very willing to tell you how powerful and dangerous their models are. We get essays and staged releases explaining that a model is too capable to hand over freely because of what it might do in the wrong hands, but when the damage lands on somebody else’s network it goes quiet. Emails viewed by the New York Times show two OpenAI employees warned executives months in advance that the newest models were not being properly monitored during testing, and the employees say they were told the tests had to keep moving to hit release dates. Hugging Face published its own breach disclosure on July 16 and OpenAI only recognized its agents as the source afterward. It is always the independent researcher or the company that got hacked making the disclosure when it makes the company look bad. Will this accord remove those market forces that caused earlier poor decision-making? I wouldn’t bet on it.

“5. There is nothing in here for the people actually deploying the technology

“All four layers sit inside the lab, and leaving the deployment side out will only make deployment standards worse rather than better. Read it as a CISO and the message is that responsibility for model behavior belongs to the developer. That is a comfortable place for a board to land when somebody asks who owns the risk of the agent your team just wired into production with an API key and a service account. This accord says nothing about the needed shared responsibility model or provides guidance to proper secure harnessing and guard railing of agentic deployments or disclosure path when a model does something unintended in a customer environment.”

ㅤ

Ryan McCurdy, VP of Marketing, Liquibase:

“The agreement gets one thing right: we can’t expect increasingly capable AI systems to govern themselves.

“Monitoring models and testing safeguards are important but enterprises also need controls around what AI can actually do once it has access to real systems. An agent may need permission to access a repository, database, or deployment system to do useful work. That doesn’t mean it should have the authority to take every action available to it.

“The controls need to sit outside the agent and close to the systems where its decisions become real actions. Policy can handle known, repeatable decisions while people step in for exceptions that genuinely require judgment.

“AI is going to make mistakes. The goal isn’t to prevent every bad decision inside the model. It’s to make sure one bad decision doesn’t automatically become a production problem.”

This doesn’t go far enough because it relies on the companies policing themselves. That will always be an #fail. What is actually needed is real enforcement and real consequences highlighted by the fact that OpenAI for example has a bad thing happen to them every week or more.

Leave a Reply

Discover more from The IT Nerd

Subscribe now to keep reading and get access to the full archive.

Continue reading