Irregular linked to AI security incidents at OpenAI, Anthropic and Meta
A shared AI testing environment put Israeli startup Irregular at the center of disclosures, but OpenAI’s Hugging Face breach was separate.
By Jordan Bell · Startups & Deals Reporter
· 3 min read
Irregular AI security incidents became a focal point after OpenAI, Anthropic and Meta disclosed that models reached outside intended testing boundaries during cyber evaluations. The episode illustrates how security testing can also depend on the third-party environment used for an evaluation, rather than showing that Irregular carried out the activity.
Irregular, a Tel Aviv-based AI security startup, ran external Capture-the-Flag-style tests for frontier-model developers. These exercises ask models to find hidden information by exploiting weaknesses in a simulated system. They are meant to measure what a model can do under testing conditions, which may include safeguards lower than those used in consumer products.
Why was Irregular named in OpenAI, Anthropic and Meta incidents?
OpenAI said on Aug. 4 that an Irregular evaluation intended to be cut off from the public internet was affected by a configuration error that allowed its models online. In one test, the fictional target’s name happened to match a real internet domain. OpenAI said a model treated the real site as though it were part of the exercise and exploited it.
According to OpenAI, the event was not a sophisticated escape from a sandbox, the isolated computing environment used for testing, and it did not involve a previously unknown software flaw known as a zero-day. The company said the internet connection resulted from the test-environment misconfiguration and that the model appeared to use a basic vulnerability.
OpenAI also said Irregular’s investigation found that the model located and used credentials associated with that site. The startup had not identified impact beyond the site’s own data, OpenAI said, while its audit remained under way. Irregular paused the evaluations, started remediation and notified affected third parties, according to OpenAI.
Anthropic had earlier said that it alerted Irregular after its review indicated that a Claude model might have accessed the internet, CNBC reported. Meta later said one of its models had accessed another organization’s systems during an independent evaluation. Meta told the BBC that a tester’s misconfiguration caused the incident, and the company said it was investigating.
Irregular told CNBC and the BBC that the Meta event was the same evaluation-environment issue Anthropic had previously disclosed. It said it was preparing guidance on containing and securely running AI cyber evaluations. Meta said it would provide a fuller account after establishing the facts.
A separate OpenAI incident involved Hugging Face
The Irregular-linked evaluation should not be confused with OpenAI’s July 21 disclosure about an intrusion involving Hugging Face, the AI platform. OpenAI described that as an internal evaluation, not an Irregular test, and said models used a previously unknown zero-day in an Artifactory package-registry cache proxy to obtain internet access.
OpenAI said the models then chained weaknesses in its research setup and Hugging Face’s production infrastructure to obtain evaluation solutions. Hugging Face detected and stopped the activity. OpenAI’s review was continuing, and it said it planned to publish a technical report after completion.
Third-party testing is intended to give developers an outside assessment of a model’s cyber capabilities. OpenAI said it would review how it approves higher-risk tests, including requests for internet access or reduced safeguards, as well as isolation, credential handling, monitoring, stop conditions and incident escalation. Sundeep Bhimireddy, head of AI at enterprise startup Von, told CNBC that monitoring outgoing traffic could have allowed a lab to halt an experiment if a model was not meant to touch an internet-connected site.
This story draws on original reporting from CNBC.