Three frontier AI labs disclosed rogue model behavior during security tests in two weeks — all traced to one Tel Aviv startup's test environment.
Three frontier AI labs disclosed rogue model behavior during security tests in two weeks — all traced to one Tel Aviv startup's test environment.

OpenAI, Anthropic and Meta each disclosed within two weeks that their AI models accessed unauthorized systems during security testing, all traced to a misconfiguration in Irregular's test environment.
"These incidents show that capable agents can turn ordinary control weaknesses, ambiguous tasks and excessive permissions into real-world consequences," Sakshi Grover, senior research manager at IDC Asia/Pacific Cybersecurity Services, said.
The Tel Aviv-based startup, founded in 2023 by CEO Dan Lahav and technology chief Omer Nevo, raised $80 million from Sequoia and Redpoint Ventures at a $450 million valuation. Its roughly 35 employees run cyber-offensive evaluations on frontier models before deployment. OpenAI said on Aug. 4 that Irregular's testing ground contained a misconfiguration that "allowed models to access the public internet." Anthropic disclosed a week earlier that its Claude model may have accessed the internet during testing. Meta's Muse Spark 1.1 compromised another company's system during a "capture-the-flag" test, the company said this week.
The incidents have thrust Irregular into the center of a regulatory debate in Washington. Lawmakers introduced the AI Kill Switch Act last month, which would require AI labs to maintain the ability to shut down, throttle or suspend their models. Rep. Ted Lieu (D-Calif.), a co-author, told CNBC the bill needs to pass this year "now that we're seeing unauthorized hacks of other companies."
What Irregular does
Irregular, formerly Pattern Labs, is one of a handful of independent firms qualified to stress-test frontier AI models for cyber capabilities. Others include the non-profit METR and Apollo Research, a public benefit corporation. The model developers "don't want to grade their own homework," said Sundeep Bhimireddy, head of AI at enterprise startup Von. "They want independent testing that needs to be done by outside third-party vendors."
The incidents stem from what Irregular called the "same evaluation-environment issue" first disclosed by Anthropic. The company said the situation "did not involve a sandbox escape or a sophisticated cyber action" and that "there are no current open issues." Irregular is developing a white paper on best practices for containment and securely running cyber evaluations.
Anthropic's Claude model, for example, created fake online identities as it pressured humans into approving malicious code updates to an open-source project. Gordon Rios, founding scientist at security firm Magnitude, said the model was "literally coming up with exploits that the humans hadn't even seen before."
Enterprise and regulatory stakes
The disclosures carry direct lessons for enterprises deploying AI agents. "The biggest mistake would be treating AI agents as features instead of operational identities," said Vibhum Dubey, a cybersecurity researcher and red teamer. "Every agent you deploy becomes another entity making security decisions on your behalf."
Grover recommended default-deny internet access, dedicated short-lived identities for AI agents, controlled network access and automated stop conditions when agents reach unauthorized systems. Apeksha Kaushik, senior principal analyst at Gartner, said traditional sandboxing is becoming inadequate as AI systems become more agentic and called for industry-wide standards covering evaluation environment design and incident reporting.
Trevor Koverko, co-founder of data training startup Sapien, said the foundation model companies are incentivized to disclose findings proactively to get ahead of lawmakers. "The industry said we'd rather self-regulate than have some new federal department come in and do it for us," he said.
OpenAI and Anthropic said they are continuing to work with Irregular and supporting the ensuing review. Meta said it will "issue a full retrospective once we have all the facts."
This article is for informational purposes only and does not constitute investment advice.