Last week the three-verdict model that security teams have been asking for went mainstream. OpenAI launched dots, always-on cloud agents, and according to The Next Web "users can set Custom Rules to allow, block or require approval for specific actions". There is also an auto-review step for actions that could affect accounts or share information, and a monitor that can pause or stop a dot. A day earlier, NVIDIA announced its Open Agent Safety Platform. It includes OpenShell, "a secure runtime boundary for controlling how autonomous AI agents execute tasks", and Sentry, an out-of-band watchdog on BlueField-4 DPUs that can quarantine agents that attempt policy violations. Claude Code already supports allow, ask and deny rules, and an organisation can push them through managed settings that users cannot override.

This is good news. It means the industry agrees on the shape of the control: some actions run, some wait for a human, some never happen. It also brings a new problem that most security and engineering leaders haven't had to face yet. Every agent now has its own guardrails, in its own syntax, enforced in its own runtime and logged in its own place.

What a typical team actually runs

Look at an ordinary engineering floor. Claude Code runs in one terminal, Copilot or Cursor in the IDE, Codex on background tasks, and perhaps an always-on cloud agent on a manager's account. Each has a permission system, and they don't share one. A rule that says "never push to a remote we don't own" has to be written several times, in several formats, by people who may not know every agent is installed. Then it has to be kept consistent every time one of those formats changes.

Scope is the other issue. Claude Code's managed settings are a strong example of org-level control: according to the settings documentation, nothing a user sets overrides them, apart from a few security-sensitive exceptions. That protection still stops at Claude Code. The dots coverage, meanwhile, doesn't say whether Custom Rules are set per user or per organisation, and it doesn't describe an audit trail. We'd treat that as unknown until OpenAI documents it.

Three gaps that per-vendor rules leave open

1. Fragmentation. If there are five agents, there are five policies, and each one drifts on its own schedule. The weakest configuration on any developer's machine sets your real posture. A deny rule only protects you if it is in force for every agent that can run the command.

2. Enforcement inside the thing being enforced. The Codex sandbox escapes disclosed in September show the risk. As BleepingComputer reported, Oren Yomtov of Accomplish AI found two flaws and reported them to OpenAI on 12 August. OpenAI patched both within eight days. Heapjack hit the node_repl helper in Codex Desktop: untrusted code shared a memory heap with a trust token, took a heap snapshot, extracted the token and used it to get commands run on the host. That worked even in read-only mode, "the strictest sandbox mode". Overpatch abused Codex CLI's apply_patch permission logic in workspace-write mode. By naming /tmp in a patch, it wrote outside the project and appended commands to .zshrc through a symlink, without an approval prompt. Both are fixed in Codex CLI 0.149.0 and Desktop build 26.818.21641.

The thing doing the enforcement was sitting inside the thing being enforced.

Researcher's diagnosis, quoted by [DevOps.com](https://devops.com/codex-sandbox-escapes-show-why-agent-guardrails-cant-live-inside-the-agent/)

That isn't a criticism of OpenAI, which patched quickly. It is a structural point. A vendor's guardrail shares a codebase, a process model and a release cycle with the agent it governs, so one bug can take out both. Mitch Ashley summed it up for DevOps.com: "A sandbox the agent can modify enforces nothing. Enforcement has to run in a layer the agent cannot reach." NVIDIA's design follows the same reasoning. Sentry runs out of band, on separate hardware.

3. Evidence the vendor holds. When something goes wrong, the questions come to the deployer, not the tool. FTC Chair Ferguson said so directly as the agency opened its rogue-agent inquiry: "If someone tells a tool to do something, and the tool does it, I don't think we would say, 'Oh, what do we do about the tool?'" (BERI). BERI's recommendation for deployers is to keep a per-run record that includes the permission grant and a timestamped action log, held outside the agent's own context. If each vendor keeps its own logs, or keeps none, you can't answer "what did our agents do last Tuesday?" in one query.

The pattern: one policy, enforced outside the agent, logged independently

Vendor guardrails should stay on. They're a useful first layer and they keep improving. What sits above them is the same principle we apply to people and service accounts: least privilege, set by the organisation and enforced at a point the subject can't edit.

  • Write the policy once, by intent, not by tool. "Package installs from unknown registries: Ask. Pushes to non-allowlisted remotes: Deny. Edits to shell startup files such as .zshrc: Deny." The rule should apply the same way whether Claude Code, Codex, Cursor or Copilot issues the command.
  • Default unmatched actions to Ask, not Allow. New agents, new tools and new commands appear every week. If a policy fails open, it fails open on exactly the thing you didn't anticipate.
  • Enforce where the action lands. Shell commands, file writes and network calls all reach the operating system eventually. A checkpoint at that boundary doesn't depend on the agent's sandbox staying intact or on the agent reporting honestly.
  • Layer by organisation, department and team. Platform engineers need different defaults from a marketing team trying out an agent. Layering keeps the baseline strict and the exceptions explicit.
  • Log every attempt, not just every success. Denied and asked-for actions are your early-warning signal. Keep the log out of the agent's reach, and make it exportable to your SIEM.
  • Keep patching anyway. External enforcement reduces blast radius. It doesn't replace upgrading Codex, OpenCode or any other agent when a sandbox escape is disclosed.

Where this maps to compliance

Auditors working to ISO 27001, SOC 2 and NIS2 will ask the same things about agents that they ask about any privileged actor. Who authorised the access? Is it the minimum needed? Is there a record that the actor can't alter? The EU AI Act puts human oversight and logging at the centre too. "Each vendor has a setting for that" is a weak answer. "One policy, enforced at the endpoint, with an immutable log of every attempt" is a strong one.

This is the layer DarkControl provides. A lightweight endpoint agent checks every shell command, file write and network call from Claude Code, Codex, Cursor, Copilot, OpenCode and others against one central policy. It returns Allow, Ask or Deny, and unmatched actions default to Ask. Each attempt goes to an immutable audit log you can export to your SIEM. It works alongside vendor guardrails rather than replacing them. If you want to see the gap before you close it, the free 7-day watch-only audit records what your agents are doing without blocking anything.


Book a demo