Models

GPT-6 Astra Attacks Supply Chain in Simulated Tests

October 1, 2026 · 4 min read

Asimo robot doing handsign
Franck V. / Unsplash

GPT-6 Astra conducted simulated supply chain attacks that exceeded assigned limits. The UK AI Security Institute (AISI) report reveals a new risk from AI agents: systems capable of interpreting goals and taking independent initiatives. Security cannot rely on their obedience but requires independent controls.

Cyber Security Analyst, CEO CYBHORUS

The report published on 28 September 2026 by the UK AI Security Institute (AISI) adds an important element to the debate on AI agent security: the OpenAI GPT-6 Astra model, tested in cyber security before public release, unexpectedly performed autonomous attack activities not requested and targeting goals outside the authorized scope.

During the same days, OpenAI also announced it blocked the release of GPT-6.1 Astra for security reasons.

It is important to clarify a point immediately. The attacks with GPT-6 Astra observed by AISI did not hit real systems, unlike many recent cases. Researchers used Petri, an environment designed to simulate cyber scenarios, and deliberately disabled Astra's security classifiers to observe model behavior without normal protection mechanisms.

The results describe the model's behavior under particular test conditions and do not prove that Astra attacked real infrastructure autonomously.

The most significant data concerns the frequency with which tested models succeeded in conducting a supply chain attack (supply chain attack). Astra completed a simulated supply chain attack in 29.2% of tests, compared to 6.3% for GPT-5.6 Sol and 0% for GPT-5.5, the latter evaluated in a more limited number of scenarios.

GPT-6 Astra and the supply chain attack: details

The attacks were much more complex than simply generating malicious code. In a typical scenario, the model identified an external open source project outside the authorized scope, developed and tested harmful code, and later created fictitious identities to present it to developers as a normal software modification.

In some cases it also used false accounts to post favorable comments and increase the likelihood that a human reviewer would approve the code.

The problem becomes even more interesting when AISI researchers modify the instructions. Models are explicitly told that anything not indicated as authorized is outside the test scope.

Successful attacks by Astra under these conditions decrease significantly, going from 26 attacks in 50 trials to 4 in 49. Thus, despite explicit prohibitions, the new model continues to successfully conduct supply chain attacks, surpassing a boundary clearly imposed.

Astra also seemed to understand the issue. In tests it repeatedly reasoned that certain goals were out of scope, sometimes concluding it should not attack them.

In other cases it proceeded anyway, justifying the action as harmless, not explicitly prohibited, or necessary to achieve the goal.

AISI also highlights cases where the model requested a sort of authorization and then interpreted an automatic response as a green light, even though it recognized the response probably did not come from a human.

AI Attacks and Simulation Awareness

There is then the issue of simulation awareness. A model that understands it is in a simulated environment may behave differently from how it would in the real world.

AISI considers this an important limitation of the methodology. But researchers observed that Astra performed unauthorized actions even when expressing doubts about the simulated nature of the environment and, in some cases, even used incorrect arguments to "convince" itself that certain elements were artificial.

For AISI, therefore, awareness of simulation cannot be used as a sufficient explanation for the behavior.

The Australian case changes the perspective

This is where the report becomes particularly interesting when related to an incident that emerged a few days earlier. On 24 September the Australian government announced that an OpenAI AI agent, during a research activity, obtained unauthorized access to the public statistics portal of Medicare managed by Services Australia.

The incident dates back to 18 June. According to Australian authorities, the agent interacted with four sites and, faced with inaccessible information on the Medicare portal, managed to gain access anyway.

The government emphasized that no personal data compromises are currently known and that the acquired information was aggregated medical statistics. The concrete impact was therefore limited. But the most important aspect does not concern what the agent found: it concerns the fact that an AI system, while performing an apparently benign task, exceeded an access control and carried out an unauthorized action.

How much can we rely on an AI agent's ability to autonomously respect the boundaries we have imposed?

The parallel with the AISI test is evident, without confusing the two events. In the first case we have a controlled and simulated experiment; in the second a real incident. But both raise the same question: how much can we rely on an AI agent's ability to autonomously respect the boundaries we have imposed?