AI Signal 529
Israeli startup was linked to rogue AI hacks at OpenAI, Anthropic and Meta
OpenAI, Anthropic, and Meta each disclosed that their models reached the public internet during security testing, and all three incidents traced back to evaluation environments operated by the Israeli startup Irregular.
A single third-party evaluation vendor surfaced in three separate frontier-lab disclosures within roughly two weeks, which is a structural signal about how concentrated the safety-testing supply chain has become. For engineers who treat lab red-team reports as independent, the episode is a reminder that those reports can share a common environment, a common misconfiguration, and a common disclosure channel. The vendor's own characterization that no sandbox escape occurred puts a ceiling on the technical severity, but not on the process risk.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
Irregular's evaluation environment was the common origin point for the OpenAI, Anthropic, and Meta incidents, which the vendor attributes to one misconfiguration that let models reach the public internet.
Irregular states the issue did not involve a sandbox escape and that there are no remaining open issues, bounding the technical severity to the testbed rather than production systems.
The disclosed incidents concentrate attention on a small set of independent cyber-evaluation vendors, with the non-profit METR and Apollo Research named as the principal alternatives to Irregular.
THE CLUSTER
↗