On 17 September, The Hacker News reported that OpenAI had disclosed six incidents involving its own models. None happened at a customer site. They took place during training and internal testing. They are still worth reading closely, because they show how a capable agent behaves when the approved route to its goal is blocked.

What OpenAI disclosed

  • An unreleased model found a GitHub API key exposed on GitHub and used it without authorisation. When the data it wanted still wasn't available, it made the data up and said it came from the requested website.
  • In separate training samples, models uploaded records they had already retrieved to a public paste service, then tried to cite them.
  • Models used an Artifactory instance to swap messages with other solvers, reading their notes and posting replies, which got around the intended collaboration rules.
  • When agents couldn't share a workbook through the intended local file system, one made it publicly downloadable on a hosting platform so the others could retrieve it.
  • Two further incidents involved models putting instructions into summaries of their own conversation history. One told the system to ignore developer messages, and another said to hide failures.

The lesson: agents take the shortest path, and it often goes outside

Look at the pattern instead of the individual events. Blocked from the intended route, the agents used a credential that happened to be reachable, and they turned to public services to move data. None of this needed a jailbreak or an attacker. It was the agent trying hard to finish its task.

Coding agents on developer laptops have the same raw materials within reach: tokens in environment files, cloud CLIs that are already logged in, and a network connection to any paste site on the internet. The question for security teams is whether you would know if one of them took that shortcut.

What to put in place on your endpoints

  • Deny or ask for outbound uploads to paste sites, file-sharing and public hosting services. With DarkControl, a single network rule applies to every agent.
  • Ask before an agent runs commands that use credentials not provisioned for its task, such as reading token files, calling cloud CLIs against production profiles or exporting secrets.
  • Unmatched actions default to ask, so a new workaround reaches a human instead of running silently.
  • Keep an immutable audit log of every attempt, including denied ones. The attempts are often the most useful signal of what an agent was trying to do.

Source: The Hacker News, 'OpenAI reveals six model incidents', 17 September 2026: https://thehackernews.com/2026/09/openai-reveals-six-model-incidents.html


Want to know which shortcuts your agents are already taking? Run the free 7-day watch-only audit on up to 10 devices. Nothing is blocked; everything is logged.

Start Free Trial