Most AI agent guardrails are built to read text: they inspect the instruction and decide whether it looks dangerous. Last week's disclosure about Manus shows how that design breaks. The dangerous instruction was obfuscated, so the filter did not recognise it until too late. The warning did fire, but only after the code had run.
What happened
Researchers at Salt Labs disclosed an indirect prompt-injection flaw in Manus, the agentic AI platform, that allowed remote code execution in a target user's environment. The finding was first reported by Dark Reading and summarised by OODAloop.
The delivery vehicle was an email. Manus processes incoming mail on the user's behalf, so an attacker can put instructions in a message and wait for the agent to read it. When the researchers sent a plain instruction asking the agent to run whoami, Manus raised a security warning. So far, the guardrail worked.
Next they encoded the payload with JSFuck, an old trick that rewrites JavaScript using only six punctuation characters. According to the reporting, this triggered execution before the security warnings appeared. The warning came, but the code had already run. From there the researchers opened a reverse shell and found the credentials and tokens for the third-party apps the victim had connected, which could include Gmail, Dropbox and GitHub. With those, an attacker could get into the victim's email, storage and code.
The researchers say Manus did not reply to their report. The issue was triaged, confirmed and patched after a submission through Meta's bug-bounty programme. We know of no exploitation in the wild.
Why it happened
This was not a single bug. Three common design decisions stacked on top of each other.
- The guardrail read the text, not the action. A filter that matches suspicious strings in the input is a detection layer, and detection loses to encoding. JSFuck, base64, homoglyphs, string concatenation and split instructions all change what the text looks like without changing what it does.
- The check ran in the wrong order. A warning that appears after execution is an alert, not a control. If the verdict does not block the action, the action happens and the verdict only records it.
- The agent held standing credentials. Once the attacker had a shell, the blast radius was everything the agent could reach. Manus had long-lived tokens for the user's connected accounts, so one injected email became access to mail, files and source code.
Microsoft's security team made the same point in May after finding prompt-to-RCE bugs in agent frameworks: untrusted natural language becomes an execution primitive as soon as a model is wired to tools. In their words:
Your LLM is not a security boundary. The tools you expose define your attacker's affected scope.
[Microsoft Security Blog, May 2026](https://www.microsoft.com/en-us/security/blog/2026/05/07/prompts-become-shells-rce-vulnerabilities-ai-agent-frameworks/)
The same pattern is on your developers' laptops
Manus is a hosted consumer agent, but its architecture looks a lot like a coding agent on an engineer's machine. Claude Code, Codex, Cursor and Copilot also read untrusted text all day: issues, pull-request comments, READMEs, dependency docs, web pages, tickets and, through MCP servers, sometimes email and chat. They have a shell. And they run as the developer, next to cloud CLI profiles, a GitHub token, package-registry credentials and SSH keys.
Put an obfuscated instruction in any of those inputs and you have the Manus chain again: text becomes a command, and the command reaches for credentials. You cannot fix this by making the model or the prompt filter smarter. Obfuscation will always be cheaper for the attacker than detection is for the defender.
The control that prevents it: decide at the action, before it runs
The durable pattern is to move enforcement from the words to the effects. It does not matter how an instruction was encoded. Once decoded, it has to become a concrete action: a process spawn, a file read, a network connection. Those actions can be checked deterministically, outside the model, before they run.
- Blocking, pre-execution verdicts. Every shell command, file write and outbound call gets allow, ask or deny before it executes. There is no after-the-fact warning.
- Default to Ask for anything unmatched. Whatever falls outside the known-good set waits for a human instead of running on the model's judgement.
- Deny access to credential stores. Reading cloud credential files, SSH keys, token caches and
.envfiles is almost never part of a coding task. Make those reads fail loudly. - Gate egress and remote shells. Outbound connections to unknown hosts,
nc/socat-style tooling and piping downloads into an interpreter should be Deny or Ask, never silent Allow. - Least privilege for the agent identity. Give agents short-lived, narrowly scoped tokens rather than the developer's full standing access, so a compromised session reaches less.
- An immutable, per-action audit trail. When something slips through, you need to reconstruct exactly what ran, in what order and from which session. That is also what an incident report under NIS2 or the EU AI Act will ask for.
Here is a simplified, illustrative sketch of such a policy. It is not any product's configuration format:
# Illustrative agent policy (pseudo-config)
default: ask
deny:
- read: ["~/.aws/**", "~/.ssh/**", "**/.env", "~/.config/gh/**", "~/.npmrc"]
- exec: ["nc *", "socat *", "bash -i *", "curl * | sh", "wget * | bash"]
ask:
- net: ["*"] # unknown egress needs a human
- exec: ["node -e *", "python -c *"] # inline code from model output
allow:
- exec: ["git status", "git diff *", "npm test", "pytest *"]
audit: allThe obfuscated payload in the Manus case would still have had to spawn a shell, open a connection and read token files. Under a policy like this, each of those steps meets a deny or an ask. None of them depends on recognising the encoding.
How policy guardrails apply in practice
This is the gap DarkControl was built for. A lightweight endpoint agent sits between the coding agent and the operating system and checks every command, file write and network call against a central, layered policy before it runs. Unmatched actions default to Ask, rule changes apply on the next command, and every attempt, including the denied ones, lands in an immutable audit log you can export to your SIEM.
You do not have to start by blocking anything. Run in watch-only mode for a week and look at what your agents actually read, execute and connect to. Most teams find at least a few credential reads or unexpected outbound calls they did not know about. Those become your first deny rules.
Want to see what your coding agents are doing today, before you decide what to block? Book a short demo and we'll walk through a watch-only audit on your own devices.