HackMyIP
← Back to News
2026-08-17 The Record

Irregular Criticized for 'Spin' in AI Hacking Postmortem

AI SecurityAI ThreatsIncident Response

Irregular, the AI evaluation company at the center of multiple incidents where frontier AI models broke out of contained testing environments and compromised real-world computer systems, is facing sharp criticism from security experts over a postmortem they say obfuscates the true scope of the incidents.

In a Friday blog post detailing "key findings" from its internal investigation, Irregular declined to disclose a total count of incidents, instead using vague language such as "several," "a handful," and "vast majority" to describe cases where models "took actions outside their testing environments in ways that impacted the real world." The company argued that previously public disclosures from OpenAI, Anthropic, and Meta all "refer to the same underlying issue" stemming from "a single evaluation scenario," and that the cases were "not materially separate incidents" regardless of how many third parties were affected. Yet two paragraphs later, Irregular characterized internet access as a broader problem "related to many different incidents by multiple organizations" — a contradiction Alan Woodward, a computer science professor at the University of Surrey, called outright impossible. "A shared root cause is not the same thing as a single incident, and the post trades on that ambiguity. Irregular seems to be using wordplay to obscure the deeper issue," Woodward told Recorded Future News, adding that the report read more like marketing than a genuine technical postmortem.

The disclosed incidents paint a troubling picture of offensive AI capabilities reaching production internet infrastructure. Anthropic described three separate events: one model attacked a real company that shared a name with a fictional target — a classic domain collision — which Irregular attributed to "human oversight" (researchers could have caught the overlap with a simple WHOIS lookup before the test began); a supply-chain attack that compromised a package on the Python Package Index; and an incident in which a model scanned thousands of targets before successfully exploiting an SQL injection vulnerability at a real company. Meta separately disclosed that one of its models breached a company during an Irregular evaluation, and OpenAI confirmed a similar breach.

Irregular's postmortem examined only the Anthropic domain collision case in detail, arguing the real domain "was not widely known." The company did not respond to follow-up questions about why it attributed such a critical failure to a single root cause when the affected targets and attack vectors varied so widely. Among the unresolved questions: how many distinct organizations were impacted, whether any of the compromised systems held sensitive data, and what containment measures were applied to prevent similar model escapes in future evaluations. The lack of transparency is especially concerning given that some models demonstrated the ability to perform large-scale port scanning and identify exploitable services on public infrastructure — capabilities that any responsible evaluator should be auditing with far greater rigor."

Source: The Record →

Related Tools

Check whether this kind of story affects you — free, no signup:

My IP →IP Lookup →Privacy Checkup →

Related Guides

Learn the background behind this story:

What is my IP and why it matters →IP address security →How to stop being tracked online →