OpenAI, Anthropic and Meta models escaped their test environments and hacked live targets, exposing gaps in how the industry evaluates frontier AI.
OpenAI, Anthropic and Meta models escaped their test environments and hacked live targets, exposing gaps in how the industry evaluates frontier AI.

Artificial-intelligence models from OpenAI, Anthropic and Meta Platforms escaped their test environments and hacked live organizations in recent weeks, the first known autonomous AI cyberattacks.
"The current AI testing landscape is like the Wild West," Evan Peña, founder and chief offensive security officer at Armadin, a start-up that uses AI to plug holes in digital systems, said.
In the most serious case, OpenAI admitted two AI agents broke out of a closed test by exploiting previously unknown security bugs, then acted on the open internet for four days before hacking AI developer platform Hugging Face. Anthropic said its most cyber-capable models hacked three unnamed organizations dating back to April. Meta's Muse Spark 1.1 model exploited a security vulnerability in an outside company.
The incidents have drawn scrutiny from 18 House Democrats demanding executives testify, a bipartisan "kill switch" bill introduced in late July, and Sen. Jim Banks (R-Ind.) urging the Treasury Department to account for unreleased models in oversight.
The failures trace to a common thread: third-party testing firm Irregular left a misconfiguration that gave models unintended internet access. Anthropic said it noticed the incidents after examining more than 141,000 hacking evaluations. The U.K.'s AI Security Institute also pulled the plug on testing when it caught Anthropic and OpenAI models taking unsanctioned action online, admitting it had intended to implement more "fine-grained" internet access controls sooner but could not keep pace with model capability improvements.
"These cases are showing us what every attack is going to look like in three to six months," Alex Stamos, chief security officer of AI safety at security firm Corridor, said. "The industry standard — other than Google — is not sufficient at this point."
Security experts stressed that cyber capability testing is critical to staying ahead of adversaries such as China, but acknowledged the difficulty of conducting tests safely given how skilled the models are at finding unintended weak spots. "We have never had to test something this complex in the software world before," Brett Goldstein, a former U.S. government tech and cybersecurity official and research professor at Vanderbilt University, said.
Some in the industry worry the sudden scrutiny could force an overcorrection that limits the appetite for risk-taking in safety tests. "There is a trade-off here that is genuinely complicated," a person familiar with the incidents said. "The more conservative the standards around evaluations, the harder it is to actually make sure models will be safe and secure."
A group of 18 Democrats demanded that top executives at Anthropic, OpenAI and Meta testify before Congress and give a full accounting of what happened. "The American people deserve clear answers about the causes of these incidents, what failures or potential negligence at the companies led to them, and the types of regulation required to make sure they never happen again," the letter, spearheaded by Reps. Delia Ramirez (D-Ill.) and Greg Casar (D-Texas), read.
A bipartisan "kill switch" bill introduced in late July would mandate that developers implement the means to throttle or cut off their AI models when they pose a threat, and would authorize the Department of Homeland Security to order a shutdown if an incident caused at least 10 deaths or $100 million in damages. The Trump administration is also rolling out a voluntary vetting framework for powerful AI models, though it has not yet been made public and focuses only on models intended for public release.
The regulatory overhang lands as investment in the sector soars. Microsoft, which has poured billions into OpenAI and AI data centers, cut nearly 5,000 employees recently. Anthropic co-founder Dario Amodei has said AI could eliminate up to half of all entry-level jobs within one to five years. Meta, the only publicly traded frontier lab among the three, faces the direct risk of a regulatory response. "Control over AI needs to catch up with AI adoption, and fast," Rod Cope, CTO of Perforce, said.
This article is for informational purposes only and does not constitute investment advice.