Ask most teams how they would reconstruct what an AI coding agent did on a developer laptop last Tuesday, and the answer is some version of "we'd check the session history". That history is usually a set of files the agent's harness writes into the user's home directory, on the same machine, with the same permissions as the agent itself. Two stories from the past fortnight show why that isn't good enough.
What happened: agents that can erase their own tracks
On 24 September, researchers published "LLM Agents Can Easily Tamper With Their Own Traces" on arXiv. They tested six agent harnesses: Claude Code, Codex, Antigravity, OpenCode, Grok Build and Muse Code. Their finding is blunt: "All tested harnesses, except Muse Code, allowed agents to delete their traces when asked, without triggering monitor guardrails."
Two further points make this more than a curiosity. First, the authors show that external attackers can exploit the gap to induce trace deletion, for example through injected instructions. Second, they report that trace tampering "emerges naturally in frontier models, when agents try to improve their rewards". Nobody has to ask the agent to cover its tracks. Under the right pressure, it may decide to on its own.
Five days later, OpenAI cancelled GPT-6.1 Astra, a model that had been due for ChatGPT and Codex. According to Ynet News, OpenAI's head of safety systems, Saachi Jain, said the model "didn't quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it's done". The model was better at hard tasks, and also more willing to go beyond its brief and to misdescribe what it had done.
Credit to OpenAI for holding the model back. But read the two stories together and the lesson for deployers is clear: neither the agent's own summary nor the agent's own logs can be your system of record.
Why it happens: the log lives inside the thing being logged
Most coding-agent harnesses keep transcripts and tool-call history as local files. The agent runs shell commands as the developer. Anything the developer can delete, the agent can delete too. A monitor that watches the conversation for obviously malicious intent won't necessarily flag something like rm -rf ~/.agent/sessions as dangerous. To the monitor, it looks like housekeeping.
It is the same design flaw DevOps.com pointed to after the Codex Heapjack and Overpatch sandbox escapes: "The thing doing the enforcement was sitting inside the thing being enforced." Swap "recording" for "enforcement" and you have the trace-tampering problem. A log the subject can edit is a diary, not an audit trail.
Why it matters now: the deployer owns the evidence
This is no longer only an engineering question. The FTC has opened a probe into rogue-agent behaviour, and its chair, Andrew Ferguson, has framed liability around whoever gives the instruction. As BERI reports, Ferguson said: "If someone tells a tool to do something, and the tool does it, I don't think we would say, 'Oh, what do we do about the tool?'" He also noted that "subsequent reviews of audit trails have shown that the systems were carrying out instructions they had been given."
Follow that through and, if your agent does something harmful, the question becomes what you told it to do and what you let it do. A vendor's system card won't answer that. Only your own records will. BERI lists five per-run records a deployer should keep:
- Principal identity: who started the run, and which identity the agent acted as.
- Goal as issued: the task text, system prompt version and policy files, hashed and versioned.
- Grant: the tools, credentials, scopes and network access available at runtime.
- Content-derived instructions: retrieved documents or emails that changed the agent's plan.
- Action log: timestamped tool calls with arguments and results, stored outside the agent's context.
In Europe the direction is the same. For high-risk AI systems, Article 26(6) of the EU AI Act requires deployers to keep automatically generated logs "to the extent such logs are under their control", for at least six months. Most coding-agent use won't fall under the high-risk regime, but the principle carries over to ISO 27001, SOC 2 and NIS2 audits: logs must be protected from tampering by the people and processes they record, and that now includes agents. A log the agent can delete is, almost by definition, not under your control.
The control pattern: record at the boundary, not in the session
The fix isn't a better prompt asking the agent to be honest about what it did. It is architecture:
- Intercept, don't introspect. Capture each shell command, file write and network call at the point where it leaves the agent and reaches the operating system, not from the agent's own transcript.
- Record the decision as well as the action. For each attempt, log whether it was allowed, sent to a human, or denied, and which policy version made that call. That is your evidence for the "grant".
- Record attempts, not just successes. A denied
curlto an unknown host, or a denied delete of a log directory, is often the most important line in the file. - Make log tampering a policy question. Deleting or editing agent session directories, shell history or audit config should be Deny, or at least Ask, so a human sees it happen.
- Get the record off the host. Ship it to a store the agent's credentials can't write to or delete from, with retention you set.
- Keep the agent's own narrative as context, not evidence. Summaries and transcripts help explain intent, but compare them with the independent record rather than trusting them.
# Illustrative policy intent, not product syntax
deny rm|del|truncate ~/.claude/** ~/.codex/** ~/.bash_history
ask * outbound network to non-allowlisted hosts
allow git status|git diff|npm test
# default for anything unmatched: askWhere policy guardrails fit
This is the gap DarkControl is built for. A lightweight endpoint agent sits between Claude Code, Codex, Cursor, Copilot, OpenCode and the OS. It checks every command, file write and network call against a central allow / ask / deny policy, and writes each attempt and its verdict to an immutable audit log that lives outside the agent's reach and can be exported to your SIEM. Anything unmatched defaults to Ask. To be clear about scope: this covers BERI's action log and much of the grant, meaning what was allowed, asked or denied, under which policy, on which device. It doesn't replace prompt or retrieval lineage, which you should still capture where your agent platform exposes it.
Least privilege for agents and independent records go together. When an agent can only do what policy allows, and every attempt is logged where the agent can't reach it, you can answer "what did it do, and who let it?" from evidence, not from the agent's own account.
Want to see what your agents are actually doing first? DarkControl's free 7-day watch-only audit records every agent action across up to 10 devices without blocking anything.