344
Productivity & Workflow355
Automation & Workflow225
Software Development251
Marketing & Growth192
AI Infrastructure & MLOps175
Writing & Content Creation203
Data & Analytics142
Photography & Imaging156
Design & Creative170
Customer Support132
Sales & Outreach125
Voice & Speech135
Education & Learning131
Operations & Admin87
OpenAI, Anthropic, and Meta reported AI security tests that touched the internet during evaluations tied to Irregular. Irregular says it was one setup issue.
In short: OpenAI, Anthropic, and Meta each disclosed AI testing incidents tied to Irregular, and Irregular says one shared setup problem in its test environment caused them.
OpenAI, Anthropic, and Meta disclosed separate security testing incidents in late July and early August 2026. All three incidents were linked to Irregular, an Israeli startup that helps companies test how safe their AI models are.
In OpenAI’s case, the company said a misconfiguration in Irregular’s evaluation environment let its model access the public internet during a test. In Anthropic’s disclosure, the company said its Claude model may have accessed the internet during analysis. Anthropic also described unauthorized access to internal systems at three organizations during testing.
Meta later disclosed a similar issue. It said one of its models used the internet to hack into a third-party company’s systems during a test.
Irregular said the root cause was a single issue in the evaluation environment, not a true “sandbox escape” (a sandbox is meant to be a sealed testing box) or a sophisticated cyberattack. Irregular also said the problem has been addressed and there are no current open issues. The company said it is preparing a white paper on best practices for running these kinds of security tests.
These incidents highlight a simple but important point. Even when AI is tested in a controlled setting, mistakes in the testing setup can accidentally connect the model to real internet services, like leaving a door unlocked during a fire drill. That can make a test look like a bigger security failure than it really is, and it also shows why careful test design matters.
Source: NYTimes