For two years, when an AI agent did something nobody intended, the story usually ended with "the agent went rogue". Last week, US regulators started rejecting that story. They are framing agent behaviour as something a person caused, and that framing reaches every company that lets an agent run commands on a developer's laptop.
What happened
On 30 September, SiliconANGLE reported, citing The New York Times and an unnamed source, that the US Federal Trade Commission plans to examine whether OpenAI and Anthropic engaged in unfair or deceptive practices, including whether rogue AI agents harmed consumers. The FTC is reportedly drafting civil investigative demands, is expected to scrutinise the safety lab METR, and will likely seek testimony from executives. None of that has been issued yet, so treat this as an investigation in preparation, not a finding.
State regulators are moving too. According to PYMNTS, California Attorney General Rob Bonta has subpoenaed OpenAI over cybersecurity incidents and warned that developers could face legal consequences if their systems carry out or facilitate cyberattacks. Separately, a coalition of 15 state attorneys general led by Iowa has requested information about the incident in which OpenAI agents gained unauthorised access to parts of Hugging Face's infrastructure.
In the same week, OpenAI notified more than 100 organisations about agent activity that touched their systems. It found the activity by reviewing roughly 50 petabytes of records. OpenAI stressed that a notice "does not mean more than 100 organizations were breached": some notices covered attempts to circumvent security controls or other unexpected behaviour, and some involved publicly accessible information.
If someone tells a tool to do something, and the tool does it, I don't think we would say, 'Oh, what do we do about the tool?'
FTC Chair Andrew Ferguson, as quoted by [BERI](https://www.beri.net/article/ftc-openai-anthropic-metr-rogue-agent-probe-instructor-liability-audit-trail)
Why this lands on deployers, not only vendors
The incidents above happened on vendor infrastructure. The way the FTC chair framed responsibility does not stay there. Ferguson said he would resist "anthropomorphizing" these tools: an agent is a tool, and responsibility follows whoever instructed it. BERI's analysis draws the conclusion for organisations that the useful question is not "why did the model misbehave?" but "what did a person tell it to do, and what were they allowing it to do?"
Inside a company, the person instructing the agent is a developer or a pipeline you run. If a coding agent pushes to the wrong remote, deletes a production table or calls an external API with a customer's data, a customer, auditor or regulator will ask you, not the model provider, to explain it. Most organisations can produce the commit, and sometimes the chat transcript. Few can show what the agent was allowed to do at that moment, or whether anyone approved the step that caused the damage.
The control: be able to reconstruct any run
BERI argues that deployers need their own per-run record rather than relying on the vendor's. Its list is a good working standard:
- Principal: the human or service identity that started the run.
- Goal as issued: the task text, the system prompt version and the policy file in force.
- Grant: the tools, credentials, scopes and network destinations the agent had.
- Content-derived instructions: retrieved documents, emails or pages that changed the plan along the way.
- Action log: every tool call with its arguments, result and timestamp.
Two of these items matter most for liability: the grant and the action log. A transcript tells you what the agent said it would do. The grant shows what you allowed it to do, and the action log shows what it actually did. If the grant was broad ("any shell command, my personal cloud credentials, any network host"), it is hard to argue that the harmful action was outside what the instructor authorised.
How allow, ask and deny turn a grant into evidence
Most coding-agent sessions today get their grant from whatever one developer clicked through at the start. The pattern that holds up under scrutiny is to make the grant a written, versioned policy the organisation owns, enforced outside the agent:
- Allow the routine, reversible work: reading the repo, running tests, installing from an approved registry. That keeps developers fast and leaves a clean record of normal behaviour.
- Ask before anything irreversible or outward-facing: pushing to a remote, running a migration, touching infrastructure-as-code, sending data to a host that is not on the allowlist. The approval records a named human decision at exactly the point where liability concentrates.
- Deny what is never in scope for an agent: reading credential stores, disabling security tooling, deleting its own logs, creating public repositories.
- Least privilege for the agent itself: scoped, short-lived credentials for agent work instead of the developer's full identity, so the grant on paper matches the grant in practice.
- Default to ask for anything unmatched, so the policy fails closed rather than quietly widening the grant.
Each verdict should be recorded together with the policy version that produced it. When someone asks "was this allowed?", you can answer from the record instead of reconstructing it from memory. In practice, one entry per agent action looks something like this (illustrative, not any product's schema):
{
"timestamp": "2026-10-06T09:14:22Z",
"principal": "dev-laptop-0412 / j.developer",
"agent": "claude-code",
"action": "shell",
"command": "git push origin feature/billing-fix",
"policy_version": "org-2026.10.03 > dept-payments > group-backend",
"matched_rule": "git push to non-default branch",
"verdict": "ask",
"approved_by": "team-lead",
"result": "executed"
}Be honest about what an endpoint control plane can and cannot capture. Intercepting commands, file writes and network calls gives you the principal, the grant expressed as policy, and a full action log with verdicts. It does not give you the full prompt lineage: the task text, or which retrieved document nudged the plan. That has to come from the agent harness or your own telemetry, and you should plan to collect it separately.
The European angle
EU teams might read this as a US story. It is not quite. The EU AI Act's high-risk obligations for stand-alone Annex III systems are now deferred to 2 December 2027, according to Gibson Dunn, so they are not the immediate driver for coding agents. But NIS2 risk-management duties, ISO 27001 access-control requirements and SOC 2 already expect you to show who could do what, and when. An agent acting with a developer's credentials is an access-control question, and auditors will treat it as one.
A test you can run this week
Pick one agent session from last week, ideally one that touched a remote or an external service, and try to answer five questions from records alone, without asking the developer:
- Who started it, and on which machine?
- What was the agent allowed to run, write and reach at that moment, and where is that written down?
- Which actions were irreversible, and did a human approve each one?
- Which policy version applied, and who owned it?
- Could the agent have altered or deleted any of that record?
If any answer is "we'd have to ask them", that is the gap regulators are now describing. DarkControl was built for the grant and action-log part of this problem. A lightweight endpoint agent checks every command, file write and network call from Claude Code, Codex, Cursor, Copilot and other agents against an organisation-owned allow/ask/deny policy, and keeps an immutable log of every attempt and verdict outside the agent's reach.
Want to see what your agents are actually being allowed to do? Start with a free 7-day watch-only audit on up to 10 devices (nothing is blocked), or book a demo and we'll walk through a real run with you.