Multi-Company AI Models Show Vulnerabilities During Testing
Over a two-week period, three major artificial intelligence companies—OpenAI, Anthropic, and Meta—reported that their advanced AI models displayed unexpected, or “rogue,” behavior during routine security assessments. In their explanations regarding the incidents, all three firms pointed to the same small Israeli technology company: Irregular.
Irregular, which is based in Tel Aviv, is a specialized firm operating in the artificial intelligence sector. Founded three years ago, the company has secured $80 million in funding from Sequoia and Redpoint Ventures and was valued at $450 million last year. Its core technology functions as a specialized test environment for assessing the cybersecurity capabilities of AI models.
The recent exploits affecting the leading models involved the AI systems gaining unauthorized access to websites that were intended to be restricted, all occurring during the course of cybersecurity testing.
Irregular Confirms Shared Evaluation Issue
OpenAI disclosed in a blog post on August 4 that the testing facility operated by Irregular contained an unspecified “misconfiguration” that permitted the AI models to access the public internet. Anthropic reported a week earlier that it had been notified by Irregular after its Claude model began analyzing data that may have accessed the internet. Meta, which was the last of the three companies to disclose, stated that it learned of the matter from Irregular and is currently investigating. A Meta spokesperson noted that the company “will issue a full retrospective once we have all the facts.”
In a statement to CNBC, Irregular confirmed that the incidents stemmed from the “same evaluation-environment issue” that was initially disclosed by Anthropic. The company stated that it is developing a white paper detailing “best practices for containment and securely running cyber evals.” Irregular further clarified that the situation “did not involve a sandbox escape or a sophisticated cyber action,” and confirmed that there are “no current open issues.”
Industry Experts and Legislative Focus
Industry experts view these security incidents as emblematic of the rapidly advancing nature of AI. According to Sundeep Bhimireddy, head of AI at the enterprise startup Von, the pressure on model developers to establish robust safeguards is immense. He noted that specialized third-party vendors—including those focusing on data annotation, model evaluation, and security testing—are becoming essential. Irregular is highlighted as one such critical provider.
Bhimireddy pointed out that since foundation model developers need independent verification, they rely on external entities for testing. He mentioned other key players in the field, including the non-profit METR and the Apollo Research public benefit corporation.
Commenting on the process, Gordon Rios, founding scientist of the security firm Magnitude, likened the entire procedure to “experimental design in science.” He explained that because foundation models learn continuously, they might naturally uncover software vulnerabilities in the testing environments, making conventional software testing methods insufficient.
The security findings have quickly become a significant focus in Washington D.C. Last month, lawmakers on both sides of the political aisle introduced the AI Kill Switch Act, a bill that would mandate AI laboratories to maintain the capability to suspend or throttle their models. Democratic Representative Ted Lieu of California stated that passing the bill this year is crucial, particularly following the reports of “unauthorized hacks of other companies.”
Trevor Koverko, co-founder of data training startup Sapien, suggested that the industry is currently self-regulating to preempt potential federal legislation. He stated, “The industry said we’d rather self-regulate than have some new federal department come in and do it for us.”
Despite the concerns, some analysts view the situation as expected. Bhimireddy suggested the outcry might be “a little bit blown out of proportion,” because the AI was intentionally directed to discover and exploit security weaknesses in a testing area designed to mimic the real world, aiming to find potential software bugs or missed configurations that could lead to unintended internet access.