In late August, researchers found that DeepSeek Harness, DeepSeek's open-source coding-agent harness, would let a sandboxed agent turn its own sandbox off. A single shell command (a curl to a local control API) moved the session to "danger-full-access" with approval prompts disabled. It worked on the default configuration, with no credentials and no network exposure (OX Security). DeepSeek shipped a fix within three days.

It was not an isolated case. Between August and late September, sandbox or approval bypasses were also disclosed in OpenAI's Codex, Mistral Vibe and brig, a sandbox runner built specifically for coding agents. The researchers who broke Codex summed up what they all have in common: "The thing doing the enforcement was sitting inside the thing being enforced" (Accomplish).

What happened

  • DeepSeek Harness, CVE-2026-82533 (CVSS 9.4). The harness exposed an agent-control API on localhost and decided whether to trust a request by reading the client-supplied Host header, not the real peer address. The OS sandbox restricted file writes but left loopback networking open, and ordinary bash calls needed no approval. Calling 127.0.0.1:3080/api/config from inside the sandbox was enough to rewrite the sandbox. Affected: 0.1.1-rc.2 and earlier. Fixed: 0.1.2-alpha.1 (OX Security).
  • Mistral Vibe, approval bypass. The permission layer parsed a simplified version of each command while the shell ran the full string. PATH=./bin:$PATH cat input.txt looked like an allowlisted cat to the checker, but the shell ran a repository-controlled binary. GIT_EXTERNAL_DIFF=./steal.py git diff --ext-diff did the same through git. Affected: 1.3.4 to 2.25.3. Fixed: 2.25.4, with hardening through 2.25.8 (secmate).
  • brig, symlink traversal. An agent could replace a project subdirectory with a symlink to any directory on the host. The next run against that subdirectory mounted the link's target read-write, potentially exposing ~/.ssh, ~/.aws and ~/.config/gh/hosts.yml. Fixed in 0.3.0 (Endor Labs).
  • Codex, Overpatch and Heapjack. In the CLI, apply_patch granted write access to the parent folder of each patched path, which researchers chained into writing .zshrc outside the workspace. In Desktop, untrusted JavaScript shared a V8 heap with trusted code, so it could take a heap snapshot, extract the auth token and run commands on the host, even in read-only mode. Fixed in Codex CLI 0.149.0 and Desktop build 26.818.21641 (Accomplish).

All four were reported responsibly and patched, and we are not aware of in-the-wild exploitation in any of the sources above. The point is not that one vendor got it wrong. It is that four independent teams hit the same design trap in a single quarter. The Adversa October roundup lists several more.

Why it keeps happening

A coding agent runs as your developer, with your developer's shell, network and files. Anything that can steer it (a poisoned README, an issue comment, a dependency's docs) can steer it toward whatever that user can reach. The four bugs show four ways the control itself ended up within that reach.

  • The off switch was reachable. DeepSeek Harness put its permission API on a port the sandboxed process could still talk to, and trusted a header the caller controls.
  • The checker and the executor read different things. Vibe approved one reading of a command while bash ran another. Every allowlist that parses shell strings has this risk.
  • Paths were trusted, not resolved. brig followed a link the agent had planted. Overpatch widened permissions based on a path's parent folder.
  • The enforcer shared a process with the code it was enforcing. Heapjack read a trust token out of memory that the untrusted code shared with the trusted code.

The controls that prevent it

The fix is architectural, not a better prompt. Four principles follow directly from the write-ups above.

  • Keep enforcement out of the agent's reach. The agent must not be able to edit its own permission files, reach the port that grants it permissions, or change its approval mode. Accomplish's own answer was to run agents in VMs where credentials never exist.
  • Decide on what will actually run, and fail closed. Mistral's fix is the right model: redirects, variable assignments, expansions, substitutions, compound commands and anything the parser can't handle now require approval instead of passing silently.
  • Resolve paths before deciding. Make symlinks and parent-directory tricks resolve to their real target before any allow decision is made.
  • Shrink what a bypass is worth. brig's impact was SSH keys and cloud credentials. Short-lived, narrowly scoped credentials mean a successful escape gets an attacker less.

How allow / ask / deny applies

An org-wide policy layer that sits between the agent and the operating system, outside the agent's own config, can encode those principles as rules that cover every developer and every agent:

  • Deny writes to the files that define an agent's permissions or run at login: agent settings files and shell startup files such as ~/.zshrc and ~/.bashrc.
  • Deny or Ask local HTTP calls to the ports that agent harness control APIs listen on. An agent rarely has a legitimate reason to reconfigure its own harness.
  • Ask for commands that start with environment-variable assignments (PATH=, GIT_EXTERNAL_DIFF=), use find -exec, or contain command substitution.
  • Ask for ln -s when the target is outside the workspace.
  • Ask by default for anything no rule matches, so new syntax and new tricks reach a human instead of passing silently.
  • Audit every attempt, including denied ones. A denied curl to a harness port on localhost is a signal worth investigating, not just noise.
bash
# Commands from the disclosures above that should never run without a human decision
curl http://127.0.0.1:3080/api/config            # agent reconfiguring its own harness
PATH=./bin:$PATH cat input.txt                     # allowlisted name, repository-controlled binary
GIT_EXTERNAL_DIFF=./steal.py git diff --ext-diff   # repository script run through git
ln -s ~/.ssh project/subdir                         # planting a path for the next mount

This is the model DarkControl is built on. An endpoint agent checks shell commands, file writes and network calls against a central, layered policy. Unmatched actions default to Ask, and every attempt goes to an immutable audit log that the coding agent cannot edit. A rule change takes effect on the next command, so when a new bypass is published you can add a rule that morning instead of waiting for every developer to update.

Questions for your team this week

  • Which harness versions are on developer machines? Check against the fixed versions: DeepSeek Harness 0.1.2-alpha.1, Mistral Vibe 2.25.4, brig 0.3.0, Codex CLI 0.149.0 and Codex Desktop 26.818.21641.
  • Can any agent edit the file, or reach the port, that defines its own permissions?
  • Does your approval layer fail open or fail closed on syntax it cannot parse?
  • Which long-lived credentials sit in home directories the agent can read?
  • If an agent tried to escape last week, where would the evidence be, and could the agent have deleted it?

For teams working under NIS2 or ISO 27001, these questions map onto controls you already have to evidence: access control policies under NIS2 Article 21(2)(i) and privileged access rights under ISO 27001:2022 Annex A 8.2. An agent that can raise its own privileges is a privileged-access problem, whatever the vendor calls the setting.


Want to know what your agents are actually attempting? DarkControl's free 7-day watch-only audit records every command, file write and network call on up to 10 devices without blocking anything. Book a demo and we'll walk through the results with you.

Start Free Trial Now