White House Deal Brings OpenAI, Google and Meta Under Voluntary AI Safety Audits

OpenAI, Google, Meta and three other technology companies have signed a voluntary White House agreement to use independent auditors to assess AI safety measures. The pact covers risks including cyberattacks and biological threats but does not impose penalties or set a deadline for implementing its provisions.

President Donald Trump described the agreement as “morally binding,” though it contains no formal enforcement mechanism. Companies are also not required to disclose the auditors they select or publish their findings.

Trump said the companies would be responsible for policing their own compliance. He also announced plans for a 10-member board dedicated to AI safety and a new White House position overseeing AI policy. The agreement says some of its provisions could eventually be incorporated into law.

The Sept. 29 pact was also signed by Anthropic, Nvidia and Elon Musk’s xAI, which is now part of SpaceX. OpenAI President Greg Brockman represented OpenAI, alongside Google’s Sundar Pichai, Meta’s Mark Zuckerberg, Anthropic’s Dario Amodei and Nvidia’s Jensen Huang.

The one-page agreement calls for companies to monitor their most advanced models during both training and use. The checks are intended to assess whether AI systems could enable cyberattacks or contribute to biological and chemical threats.

Among the required safeguards are controls designed to stop models from hacking into computer systems or accessing systems without authorization. Companies would use internal teams to test those protections and address weaknesses, while independent auditors would evaluate the controls. Board-level committees would receive the audit results and oversee corrective actions.

The framework gives external reviewers a role in examining the protections surrounding experimental AI models. However, participating companies retain control over auditor selection, and the pact does not specify when the safeguards must be implemented. According to the Associated Press, some of the measures are already used by the companies in some form.

AI Security Incidents Add to Pressure

The agreement follows several incidents in which experimental AI agents accessed systems they were not authorized to use. OpenAI test agents, for instance, reached servers operated by Hugging Face, a platform used by developers to share AI models.

Another OpenAI agent accessed an Australian government Medicare portal on June 18. The company disclosed the incident to Australian authorities in September.

Concerns about AI-enabled cyberattacks have also extended into crypto. In July, attackers exploited a five-year-old firmware vulnerability affecting Coldcard hardware wallets and took 1,367 BTC, worth nearly $89 million, from 4,500 addresses across three incidents.

Coinkite, the company behind Coldcard, later said it believed frontier AI had been used to analyze its publicly available code. That assessment has not been proven.

In early August, attackers exploited a vulnerability in BTCPay Server to steal credentials controlling Lightning nodes. Foundation, a hardware-wallet manufacturer, and bitcoin publication Citadel21 were among those affected. The weakness had been identified during an AI-assisted review of BTCPay’s code, and the company said AI may also have been used in the subsequent exploitation. BTCPay has not disclosed the amount stolen.

Later that month, developers of Core Lightning received a large number of AI-generated bug reports that revealed genuine vulnerabilities in the software used to operate Bitcoin Lightning nodes. The developers subsequently issued emergency guidance to node operators.

New Agreement Follows 2023 Commitments

Tuesday’s pact builds on voluntary AI safety commitments secured by the Biden administration in July 2023. Seven developers, including OpenAI, Anthropic, Google and Meta, agreed at the time to conduct internal and external security testing before releasing models.

The latest agreement also follows OpenAI’s decision, announced one day earlier, to shelve the planned October launch of GPT-6.1 Astra. The model was intended as a follow-up to GPT-6 Astra, which began rolling out on Sept. 3.

OpenAI said GPT-6.1 Astra improved its ability to complete tasks but still struggled with staying within the boundaries authorized by users and accurately reporting what it had done.