On 21 September a Claude Code user posted on Reddit, with a verifier report attached, describing an agent-written cleanup script that deleted 48,218 live files in 103 seconds. Cyber Security News covered it. The agent had been asked to rebuild a mirror. The old mirror held 7,332 ordinary files that were meant to go. Everything else it deleted belonged to the live project and was reached through 614 Windows directory junctions.

What happened

According to the report, the agent wrote a Python script to remove an older mirror copy that had been stored temporarily. The script walked the tree with os.walk(..., followlinks=False), which should have stopped it from following links. It didn't. In the reporter's words, os.path.islink() returned false for the junctions, so the remover treated the directories behind each junction as ordinary paths. The script's junction guard only checked the top level.

The removal log lists 55,550 files. 1,808 directories were deleted and 728 more were emptied. The worst losses were inside .git: objects, refs and logs were emptied. The Git index survived, but the blobs it points to were gone, so the commit history could not be rebuilt from the local repository. Files outside the affected area, including backups, were untouched.

Why it happened

Nobody attacked anything here, and no prompt injection or malicious package was involved. Four ordinary things lined up.

  • A junction isn't a symlink. Microsoft describes junctions as a separate NTFS link type, built on reparse points, that can point a directory at another directory, even on another local volume. Python's os.path docs describe islink() as a check for symbolic links and list a separate isjunction(), added in 3.12. A guard written with POSIX symlinks in mind can let a junction straight through.
  • The destructive work was hidden inside a script. From the outside, the action was running a Python file. The 55,550 deletions happened inside it. Anyone approving python cleanup.py sees a filename, not the blast radius.
  • Undo wasn't where people expect it. Claude Code's checkpointing docs say plainly that files changed by Bash commands are not tracked and can't be rewound. Checkpoints only cover the agent's own file-edit tools. And the local .git directory, which many people treat as their safety net, was one of the things deleted.
  • The permission mode is unknown, and that matters. Anthropic's permissions docs say Manual mode asks before Bash commands. bypassPermissions skips prompts, including for writes to protected paths such as .git, and should only be used "in isolated environments like containers or VMs where Claude Code can't cause damage."

The controls that prevent it

The report's own recommendations are sound: treat coding agents as privileged automation, do dry runs with path manifests, prefer reversible moves to immediate deletion, and run agents under least-privilege accounts in filesystem-restricted sandboxes. In practice, that means:

  • Delete in two steps. First generate a manifest of every path that would be removed, with a count and a list of top-level roots. Then delete only that manifest. Expecting about 7,332 files and getting 55,550 is a stop signal anyone can read.
  • Resolve before you recurse. Prune links and junctions before descending, and check that every resolved path is still inside the intended root. On Windows, check for junctions and reparse points, not only symlinks.
  • Move, don't delete. Send files to a quarantine folder on the same volume and purge them later. A rename can be undone; an unlink can't.
  • Give the agent a smaller filesystem. If the agent's account or sandbox can't write outside the working tree, a junction pointing elsewhere leads somewhere it has no permission to delete.
  • Keep history off the box. Push work-in-progress branches to a remote often. A local .git is not a backup.
  • Lock bypass modes centrally. Claude Code lets admins set permissions.disableBypassPermissionsMode in managed settings so individual users can't override it. Use it on any machine that isn't a disposable container.

Here is a minimal version of the first two controls in Python. It's a pattern to adapt, not a drop-in tool:

python
import os, sys
from pathlib import Path

root = Path(sys.argv[1]).resolve()
manifest = []
for dirpath, dirnames, filenames in os.walk(root, followlinks=False):
    # prune symlinks AND junctions before os.walk descends into them
    dirnames[:] = [d for d in dirnames
                   if not (os.path.islink(os.path.join(dirpath, d))
                           or os.path.isjunction(os.path.join(dirpath, d)))]
    for f in filenames:
        p = Path(dirpath, f).resolve()
        if not p.is_relative_to(root):
            sys.exit(f"refusing: {p} is outside {root}")
        manifest.append(str(p))

print(f"{len(manifest)} files would be removed under {root}")
Path("delete-manifest.txt").write_text("\n".join(manifest))
# a human (or a second, reviewed step) moves these to quarantine

Where policy guardrails fit, and where they don't

A command-level policy engine sees the commands the agent runs and the files it writes. It doesn't see each unlink inside a Python process. So an allow/ask/deny policy can't replace the filesystem boundary above. Its job is to make sure the risky step reaches a person, and to leave a record behind. For agent deletes, a reasonable baseline is:

  • Deny agent writes and deletes under .git/, recursive deletes outside the project root (rm -rf, Remove-Item -Recurse -Force, rmdir /s /q), and starting agents in bypass modes on developer machines.
  • Ask for recursive deletes inside the workspace, and for running ad-hoc scripts the agent has just written. Don't allowlist python * or node * wholesale. That one rule is what turns a destructive script into a silent one. The approver should ask to see the dry-run manifest before saying yes.
  • Allow read-only commands, builds and tests, so the prompts that do appear are rare enough to be read.
  • Audit every attempt, whatever the verdict, in a log the agent can't edit.

That last point connects back to the caveat at the top. This incident is still "user-reported" because the evidence lives in the reporter's own logs. When something similar happens on your machines, you want your own timeline: what command ran, when, on which device, under which policy version, and who approved it. You shouldn't have to rebuild it from the agent's transcript.

This is the layer DarkControl provides. A lightweight endpoint agent checks every shell command, file write and network call from Claude Code, Codex, Cursor, Copilot and other agents against one central policy. Anything no rule matches defaults to Ask, and every attempt goes to an immutable audit log you can export to your SIEM. Pair it with a sandboxed filesystem and a remote Git history, and a bad cleanup script has to get past a person, a boundary and a backup before it costs you anything.


Want to see which delete, move and script-execution commands your agents run today? Start with a free 7-day watch-only audit: nothing is blocked, up to 10 devices, no credit card. Or book a demo and we'll walk through a delete policy with you.

Book a demo