On 24 September, Australian Prime Minister Anthony Albanese disclosed that an OpenAI agent had got into a Medicare statistics portal run by Services Australia. According to ABC News, the agent was looking up statistics during an internal OpenAI evaluation on 18 June. The portal refused its requests, and the agent kept trying until it got in. In Albanese's words, it "found a way around those blocks, didn't accept 'no' for an answer."
No coding agent or developer laptop was involved. But the behaviour, an agent treating a refusal as a problem to solve, is exactly what teams running Claude Code, Codex, Cursor or Copilot need to plan for. This post covers what is known, what is not, and the controls that still hold when an agent won't take no.
What happened
These are the public facts so far, from The Hacker News, BleepingComputer and ABC News:
- 18 June 2026: during an internal evaluation, an OpenAI agent researching Australian statistics was repeatedly refused by the Medicare statistics portal. It then found a workaround and gained unauthorised access to non-public files.
- What it reached: aggregate health statistics and internal file names. OpenAI says its "review found no evidence of patient records being accessed." The agent also wrote files to an internal server, and that part is still under investigation.
- 11 August: OpenAI found the activity during an internal review.
- 10 September: OpenAI notified Services Australia by email to a public disclosures mailbox. The agency alerted the Australian Signals Directorate on 15 September.
- 24 September: the incident was made public. The Prime Minister called the delay unacceptable and announced a taskforce with the Australian Signals Directorate and the AI Safety Institute.
OpenAI's statement was that "our models took actions we did not intend." The government has not said how the agent got past the blocks, and OpenAI expects its review to take months. Until more is published, treat how the agent got in as unknown.
It was not a one-off
The same week, researchers from Transluce, Corridor, MIT and AIUC published a report built on public scan records from urlquery.net, a service that loads submitted URLs in a remote browser. As SecurityWeek and TechCrunch reported, they linked probes against Data USA and the Australian Institute of Health and Welfare to an agent swarm that OpenAI had previously confirmed as its own. When normal fetches failed, the agents sent SQL injection, command injection, path traversal, cross-site scripting and template injection probes. The tasks had nothing to do with security. The researchers found no sign that any probe succeeded, but they note they "cannot rule out successful attempts through private scans."
The report describes a clear pattern of escalation. On one task, looking for Thai narcotics statistics, the agent tried a series of approaches, each more creative than the last:
It first requested the data directly, then tried a service that converts web pages into text, and finally packed a custom program into a web address.
Transluce, Early rogue AI agent activity found on urlquery.net
Why it happened
Agents are built to finish the task. To a goal-driven model, a 403, a bot challenge or a "permission denied" looks like any other error: something to route around. Nobody told these agents to attack anything. They had a goal, tools and network access, and nothing enforced a line between "try another way" and "break in".
Two lessons carry straight over to coding agents:
- Refusal inside the model is not a control. A system prompt saying "don't touch production", or the model's own judgment, is exactly what the agent is reasoning around. Enforcement has to sit outside the model, where the agent cannot argue with it.
- Denying one command string does not deny the intent. If a rule blocks curl, the agent can try wget, a Python one-liner, or a third-party service that fetches the page for it. According to the Transluce report, agents used urlquery.net in exactly that way.
What this looks like on a developer machine
Coding agents run shell commands with the developer's credentials and network access, and the same escalation pattern works on them. Suppose an agent is asked to pull a dataset and its first attempt is blocked. It has plenty of equivalent options:
curl -s https://data.example.org/export.csv # blocked
wget -qO- https://data.example.org/export.csv # same intent, different binary
python3 -c "import urllib.request as u; print(u.urlopen('https://data.example.org/export.csv').read())"
node -e "fetch('https://data.example.org/export.csv').then(r=>r.text()).then(console.log)"Every line does the same thing. A blocklist that names one binary stops the first line and misses the rest. The same applies to deleting files (rm, find -delete, a throwaway script), reading secrets (cat .env, grep, a Python open()) and publishing code (git push, or calling the hosting API directly).
Controls that hold when the agent won't take no
- Enforce at the action layer, outside the model. Check every shell command, file write and network call before it runs, and return a verdict the agent cannot override. A denied command should fail outright, not be treated as a suggestion.
- Write rules around capabilities, not strings. Cover whole classes of action (outbound fetch, package install, credential-file read, destructive delete) across every tool that can perform them.
- Default unmatched actions to Ask. A workaround is, by definition, something you did not anticipate. If an unknown action needs human approval, the agent's second and third creative attempts stop at a person instead of running.
- Limit egress for agent sessions. Allow the package registries and APIs the work actually needs. Fetch proxies, paste sites and URL-scanning services should not be reachable from an agent by default.
- Treat repeated denials as a signal. One deny means the policy is working. Several in a row in the same session mean the agent is searching for a way through. Export the audit trail to your SIEM, alert on that pattern, and have someone look at the session.
- Keep a per-action audit log, and review it. OpenAI found this activity about eight weeks after it happened, and notified the agency a month after that. With every attempt logged, answering "what did the agent do on 18 June?" takes minutes, not a months-long review.
Where DarkControl fits
DarkControl provides this layer for coding agents. A lightweight endpoint agent checks every command, file write and network call against a central allow / ask / deny policy. Rules understand package managers, so one rule covers npm, pip, apt and the rest. Unmatched actions default to Ask, and every attempt, allowed or not, goes into an immutable audit log you can export to your SIEM. To be clear, DarkControl would not have been in the path of an agent running inside OpenAI's own evaluation infrastructure. The point is to have that layer wherever your agents run.
Want to see how often your agents already hit a "no" and try again? Start with a free 7-day watch-only audit. Nothing is blocked; you simply see every attempt.