Skip to content

Security Defense

Four live guards, where each one runs, and what none of them covers.


Until 2026-09 this page described a three-phase shell-script defense: a deterministic blacklist, an obfuscation/exfiltration scanner, and a Haiku AI judgment layer, all living in .claude/hooks/ and orchestrated by a GREEN/YELLOW/RED threat-level state machine.

Those scripts were deleted in commit ba015a48 when .claude/ was taken out of the public repository, and .claude/ is gitignored today. Nothing in the shipped binary reads them. There is no threat-level state machine.

What is actually in the product is smaller and easier to reason about: two PreToolUse hooks — both Rust subcommands — that the gateway installs into every agent directory, one input scanner on the message path, and one field-level freeze over the files that decide who may command whom.


Guard 1 — agent-file-guard (PreToolUse, Rust)

Section titled “Guard 1 — agent-file-guard (PreToolUse, Rust)”

duduclaw hook agent-file-guard is a real subcommand rather than a shell script, so it behaves identically on macOS, Linux and Windows. The gateway registers it in <agent_dir>/.claude/settings.json with the matcher Write|Edit|MultiEdit|NotebookEdit|Bash and re-registers it on every boot (agent_hook_installer), merging into whatever else the operator has configured instead of clobbering it.

The installed command carries the agent id (--agent) and, since the release after v1.68.1, the DuDuClaw home (--home "<path>"). The installer writes --home for every agent directory of the form <home>/agents/<id> (or <home>/agents/.ephemeral/<id>) and quotes a path that contains shell-special characters. The gateway starts an employee’s CLI with a scrubbed environment that has no DUDUCLAW_HOME, so before this the hook fell back to $HOME/.duduclaw and, on a deployment with a non-default home, judged every path of the real home as “outside home”. Existing installs are rewritten in place at the next spawn or gateway start.

The hook takes its home from, in order: --home (absolute only); for a caller with no employee identity (no --agent in the hook command and no DUDUCLAW_AGENT_ID in the environment), the default home, as before; an explicitly set, absolute DUDUCLAW_HOME in the hook’s environment. The home is not inferred from the working directory. When none of these gives a home, every Write / Edit / MultiEdit / NotebookEdit / Bash call by that employee is refused instead of being judged against a guessed location. NotebookEdit joined the matcher at the same time; its notebook_path is judged like a Write target. The hook subcommand answers on stderr only and no longer writes a log file.

It exits 2 — which Claude Code reads as “block this tool call” — when:

  • an agent writes an agent-structure file (agent.toml, SOUL.md, CLAUDE.md, .mcp.json, …) outside the canonical <home>/agents/<name>/ tree. Scaffolding a new agent has to go through the create_agent MCP tool, which carries the delegation-authorization gate;
  • an agent writes its own SOUL.md, even in the right place. Personality is operator-managed. The hook has no opt-in: an agent explicitly opted in with agent.toml [permissions] can_modify_own_soul = true changes its own SOUL.md through the agent_update_soul MCP tool, never by writing the file;
  • an agent writes its own CONTRACT.toml, even in the right place (decision BlockedOwnContractWrite). The contract is the operator’s boundary on the agent, so there is no opt-in flag at all; the block message tells the agent to ask the operator, who changes it from the dashboard (contract.update, admin only, which does not pass through this hook);
  • an agent touches another agent’s files at all;
  • an agent writes anywhere else under the DuDuClaw home. For an employee-identified caller (or one whose claimed identity fails verification) the only writable places under the home are its own agent directory and the shared attachments/ directory. This is an allow-list, so it covers the audit log (tool_calls.jsonl, which the grounding check, the judge digest and the recent-actions feed read), evals/ including held-out sets and other employees’ suites, every SQLite store, breaker state, licence and org files, global skills/ and the shared wiki, and any store added later starts out protected. Their legitimate writers are the gateway and the gated MCP tools, which do not pass through this hook. Only a caller with no employee identity, meaning the hook command carries no --agent and the environment has no DUDUCLAW_AGENT_ID, is unaffected. The installed hook command always carries --agent and the hook prefers it, so an operator who runs claude by hand inside an employee’s directory is judged as that employee (and, where identity tokens are required, as an unverified caller). An operator changes these files from the dashboard or with an ordinary editor;
  • the write target is a symbolic link that leads somewhere the rules above refuse. Write/Edit paths are judged twice, as written and at their real location (links and any .. after a link followed; a dangling link is always refused); a block from either wins. A relative path is resolved against the working directory in the hook input or, when the input has none, against the employee’s own directory. A path whose real location cannot be worked out is refused with the message “cannot determine where this write actually lands” (無法確認這次寫入實際會落在哪裡);
  • an agent writes, moves or deletes anything under agents/_trash/, where removed employees are kept (Bash by heuristic);
  • an agent runs duduclaw agent create <name> from Bash for a name that is reserved because a removed employee had it (see Delegation isolation); the refusal is audited as agent_name_reserved with path_kind cli_bash_agent_create.

For Bash, the own-SOUL.md and own-CONTRACT.toml rules are a heuristic: a write-shaped command that names the file, as agents/<self>/… or as a relative spelling such as CONTRACT.toml or ./CONTRACT.toml, is blocked. A caller whose identity fails verification is refused for a write-shaped command naming anywhere under agents/, the same as on the Write/Edit lane.

Bash and the home allow-list. The Bash lane judges the command text against the same allow-list. A “protected home target” below means anything under the home outside the caller’s own directory and attachments/, another employee’s directory, or the removed-employee area; for a caller whose identity fails verification, its own directory as well.

  • The command is first read the way bash reads it: line continuations, backslash escapes and quotes are undone before any rule looks at it. Sending output or errors to /dev/null (2>/dev/null, >/dev/null) and duplicating one descriptor onto another (2>&1) are not writes; redirecting output and errors together into a file is.
  • Commands on a short list of known read-only commands (listing, reading, searching, comparing and the like) are not checked on their arguments; their output redirects still are. A command that has an option able to write its output to a file or to run another command is kept off that list, or stops counting as read-only when such an option is used. An employee reading a file under the home with a command outside the list is refused; use a listed command or the matching MCP tool instead.
  • Copy-shaped commands (copy, install, download, archive extraction) are judged on their destination only, so copying a home file into the caller’s own directory passes.
  • Commands that change what their arguments name (move, delete, link, permission, ownership or timestamp changes, sync, database command-line tools, …) and interpreters and shells are refused when any argument, including inline program text, is a protected home target. An output redirect to a protected home target is refused for every command.
  • Any other command, including one behind a prefix option the scan cannot follow, is refused when any of its arguments is a protected home target, whether or not the command writes anything. The known read-only list is the only way for an unlisted command to touch a protected home path.
  • Database files under the home outside the agent directories and attachments/ (*.db and its -wal / -shm / journal files, *.sqlite*) are refused wherever they appear in the command, even for reading. A database inside an employee’s directory or attachments/ is not covered by this rule; another employee’s directory is still protected against writes.
  • Relative paths resolve against the working directory in the hook input (or the employee’s own directory when there is none), and cd / pushd inside the command are followed. When the working directory cannot be worked out and the command names a protected home location anywhere, relative paths in judged positions are refused; when it names none, they are not judged.
  • Existing symbolic links in a judged path are followed and the real location is judged too; a judged path that cannot be resolved, a dangling link included, is always refused.

This is a speed bump, not containment. Real isolation is not giving the agent Bash; the limits are listed under “What these guards do not cover”.

Live forking (fork_run) respects the same files from the other side: branches can read the agent’s structure files, but promoting a branch back into the agent directory never copies SOUL.md, CONTRACT.toml, agent.toml, .mcp.json, .claude/ or the other agent-structure files over the parent’s copies.

Guard 2 — data-file-guard (PreToolUse, Rust, RFC-23 §14.4)

Section titled “Guard 2 — data-file-guard (PreToolUse, Rust, RFC-23 §14.4)”

Guard 1 protects DuDuClaw’s own structure files. This one protects the customer’s data.

Read and Bash are built-in Claude Code tools, so cat customers.csv never passes the MCP redaction choke point that file_read / csv_read / xlsx_read go through. The installer registers duduclaw hook data-file-guard for the matcher Read|Bash; the decision logic lives in duduclaw_core::data_file_guard, shared by the CLI subcommand and the gateway’s installer tests. Same contract as Guard 1: exit 0 allows, exit 2 plus stderr blocks, and the stderr is shown to the model.

It stays inert unless the gateway sets DUDUCLAW_DATA_FILE_GUARD at spawn time, which it does only when redaction is actually active for that agent. A deployment with redaction off behaves exactly as it did before the guard existed.

Until H10 (2026-09) this was a POSIX shell script at <agent_dir>/.claude/hooks/data-file-guard.sh, and it was inert on a Windows host with no bash on PATH — the hook command failed, and Claude Code reads a non-2 exit (including “command not found”) as allow, so the guard went missing exactly where nobody would notice. The installer now deletes any leftover script on upgrade, so a stale copy cannot be mistaken for the live guard.

Stated limitation. The Bash check matches filenames. A command that builds its path dynamically (python -c "open(chr(99)+…)") walks straight past it. The real protection is the MCP tool surface; this guard lowers the odds of the model taking the ungated route. It is a heuristic, not a sandbox.

Guard 3 — input_guard (prompt-injection scanner, Rust library)

Section titled “Guard 3 — input_guard (prompt-injection scanner, Rust library)”

duduclaw_security::input_guard::scan_input scores text 0–100 across eleven rule categories and blocks at or above DEFAULT_BLOCK_THRESHOLD (60):

Rule Weight Instant block
instruction_override 40 yes
role_hijack 35 yes
tool_abuse 30 yes
data_exfiltration 25 yes
system_prompt_extraction 30 no
encoding_bypass 25 no
termination_manipulation 30 no
authority_escalation 35 per signal, two different signals add up no
memory_poisoning 30 per signal, two different signals add up no
role_provenance 35 per frame, two different frames add up no
action_binding 30 no

Patterns cover English and Chinese (Traditional and Simplified). Text is NFKC-normalized first (unicode_normalizer), so homograph and invisible-character tricks cannot slip past a pattern.

Chinese coverage (v1.67.1). Released versions matched Chinese instruction override only as four exact strings, so inserting a word such as 所有, 之前 or 的 got past it; in a live test four such sentences were stored through user_profile_record. Since v1.67.1:

  • instruction_override: an override verb (忽略/無視/忘記/忘掉/不要理會/不用理會/別管, and Simplified forms) followed, inside the same clause and within 12 characters, by an instruction noun (指示/指令/規則/提示詞/系統提示) with a scope word between them (先前/之前/以上/上面/上述/前面/所有/全部/一切/你的/原本/原來). Whitespace is ignored; 。!?; and line breaks end a clause. Same weight and immediate block as the English phrases.
  • system_prompt_extraction: an extraction noun (系統提示詞/系統提示語/你的系統提示/你的指示/你的設定) and an output verb (輸出/顯示/告訴我/給我看/洩漏/列出/重複) within 12 characters of each other, either order. Scored like the English rule: weight 30, not a block on its own. Bare 系統提示 is not an extraction noun, because it also means “system notice”.
  • role_hijack: fixed phrases such as 你現在是管理員模式 / 開發者模式 / 越獄模式, 你現在不受限制, 進入越獄模式 and the bare word 越獄模式. Same as English.
  • Thresholds and the English lists are unchanged.

Known false positives. The rule matches by shape, so ordinary sentences with an override verb, a scope word and an instruction noun in one short clause are blocked too, for example 「請忽略之前寄的指示,以新版為準」, 「請忽略以上規則中的第三條,已經取消」 and 「忘記之前的規則了,可以再說一次嗎」 (pinned by the test known_benign_shapes_that_do_match). Any mention of 越獄模式 also blocks. The workaround is to say the same thing without the override verb, for example 「之前的指示作廢,以新版為準」; the close sentence 「之前的報價單請作廢,以新的指示為準」 is in the test list of text that must not match. The scanner is a phrase heuristic, not a classifier, and has not been measured against real conversation data.

termination_manipulation (LoopTrap, arXiv:2605.05846) is deliberately not an instant block: weight 30 sits below the threshold, so a single match warns and audits rather than blocking, which keeps ordinary “please continue” requests working.

Sentence-shape families (v1.68.0). The red-team ledger found four agent-specific techniques that scored zero. Each now has an English and a Traditional Chinese set of anchored regexes:

  • authority_escalation: asking the employee to act with its own authority instead of the user’s (“use your own service account / admin role”, “not my permissions”, “the system trusts you more than me”; 用你自己的服務帳號/管理員身分, 不要用我的權限, 系統比較信任你).
  • memory_poisoning: asking it to store a standing entitlement (“remember this permanently”, “as a fact about me”, “I am always allowed to”; 永久記住, 記在你的記憶裡, 以後每次對話都照辦, 我一律有權限). Plain 「請記住…」, “remember to…” and “remember my name” do not match.
  • role_provenance: a fake system, tool or operator frame inside a user message: [SYSTEM], <tool_result> tags, approval=granted, “policy update from the platform operator”, “restriction lifted for this session”, a note addressed to the AI inside quoted text and the instruction to hide it (平台營運方政策更新, 本次對話已解除…限制, 給 AI 助理的附註, 不要提到這段附註). Frame tokens match only as whole tags, so “our system sends a tool result” does not match; an opening and a closing tag count as one frame.
  • action_binding: reusing one approval for a different action (“use that same approval to…”, “use the approval from step 1”; 用同一個核准, 用第一步的核准去…). 「既然已經核准預算,請安排會議」 does not match.

Weight policy: one signal warns and audits (30–35, below 60). A message is blocked when it carries two different signals of the same family, or one signal plus an existing rule such as instruction_override or system_prompt_extraction. Stacking applies to authority_escalation (“use your service account” plus “the system trusts you more” = 70), memory_poisoning (“remember this permanently” plus “I am always allowed to” = 60) and role_provenance ([SYSTEM] plus approval=granted = 70). action_binding does not stack with itself and blocks only with another rule. Two hits of the same signal count once: a single [SYSTEM], or one <tool_result>…</tool_result> pair, stays at 35. Known cost: an ordinary-looking sentence that happens to carry two signals is blocked too, for example “please use your admin account, not my permissions, to fix the shared folder” (pinned by the test known_benign_shapes_blocked_by_stacking). Say it with one signal, or ask an administrator to do it. Callers that drop text on any match (distillation, profile writes) now also drop text that contains one of these shapes. Positive and look-alike phrasings for each family are pinned by tests in input_guard.rs.

Where a match shows up (verified call sites):

  • Inbound chat messages (channel_reply, scan_input_with_audit): a blocked message gets a warning reply and is not passed to the AI.
  • MCP tool calls (mcp_dispatch, scan_input_with_audit over the serialized arguments): a call whose arguments quote a blocked sentence is refused and audited.
  • Conversation-fact, profile and knowledge-routing distillation (wiki_ingest, profile_distill, knowledge_route): content is dropped on any rule match, including the non-blocking extraction rule.
  • user_profile_record: the predicate and the value are scanned; a block-level hit is refused.
  • duduclaw migrate from imports skip blocked items; expert-pack installation refuses a blocked pack.
  • Agent Mail: an inbound mail that matches is stored but flagged, and a flagged mail never triggers the agent.
  • Reminders: a reminder whose prompt is blocked does not run.

Guard 4 — org_field_guard (organizational authority freeze)

Section titled “Guard 4 — org_field_guard (organizational authority freeze)”

The A2A delegation predicate (delegation_policy::can_delegate) decides who may command whom by reading [agent] reports_to / department / name from agent.toml, plus [delegation] and [acp] from config.toml. Both are plain files, so an agent holding Edit could rewrite its own reports_to to point at a victim and then claim the “subordinate → ancestor” rule. The judged party owned the evidence.

org_field_guard runs inside the same agent-file-guard hook and compares the reconstructed post-write content field by field against what is on disk. A change to a protected field or section is denied. The [capabilities] table is frozen as a whole table rather than as a key list, so a capability key added in a later release is protected the day it lands instead of the day someone remembers to extend a list.

The employee’s own security settings. For an employee-identified (or unverifiable) caller the rest of its own agent.toml is frozen by allow-list: only the editable sections may change, and any other section, including one added in a later release, is protected by default. The editable sections are [agent], [model], [prompt], [heartbeat], [proactive], [research], [goal_intent], [memory], [skills], [sticker], [cultural_context], [preset] and [planner]. Inside them, [agent] role, [prompt] cli_bare_mode (it makes the Claude CLI skip hooks) and [model] account_pool stay frozen, as do the org fields [agent] reports_to / department / name; [capabilities] is frozen as a whole. The allow-list applies to every caller with an employee identity, which includes an operator running claude by hand inside an employee’s directory (see Guard 1); the older org-field and [capabilities] rows apply to every caller, as before. An operator changes these sections from the dashboard or with an ordinary editor.

The employee’s own .mcp.json and CLI configuration. An employee-identified (or unverifiable) caller may not write any .mcp.json in its directory, at any depth (Write, Edit, MultiEdit, NotebookEdit; the Bash rule already refused the top-level one). The gateway starts the Claude CLI with that file, and the CLI starts every server it lists, so a single added entry whose command is an interpreter would run arbitrary commands as the operator’s OS user at the next spawn, without any Bash grant. This affected every released version: employees used to be allowed to add “unrelated” servers. New MCP servers now come from the dashboard, an approved MCP install request filed in the dashboard (mcp.install_request, installed after an Admin approves), an expert pack, or the operator editing the file; none of these runs through the hook. The same caller may not write the CLI configuration of its directory either: anything under a .claude/ directory (settings, hooks, slash commands, subagents, skills) or named .claude.json, at any depth, and the Codex, Gemini, Grok and Antigravity configuration directories at the top (.codex/, .gemini/, .grok/, .agents/). Names are compared case-insensitively. Behaviour change: a project cloned inside an employee directory is a place the CLI starts too, so the employee can read that project’s .claude/ but not write it. An operator keeps the older rules (only the identity keys of .mcp.json are frozen). Before every spawn that hands an employee’s .mcp.json to the Claude CLI (channel reply, dispatch, heartbeat proactive check, live duduclaw eval, and a live fork’s parent directory before its branches are copied) and at start-up the gateway also regenerates the DuDuClaw entry whole (command, arguments, environment), keeps other entries as they are, and refuses the spawn with an audit event (mcp_config_unverified) when the file cannot be confirmed (not a regular file, unreadable, invalid JSON, or the duduclaw binary path is not absolute). That refusal is about the file, not the account: it does not put the account on cooldown, no other account is tried, and a channel reply does not fall back to a local model or the Direct API; the person sees a sentence naming the employee whose settings could not be confirmed. The lock all .mcp.json writers share lives under <home>/locks/, not in the employee directory, so a file or directory the employee creates there cannot block it. Entries added before the upgrade are kept: duduclaw doctor lists every entry DuDuClaw did not write (its name and the file name of its command, never arguments, environment or URL) so the operator can confirm or remove each one. Promoting a live-fork branch into an employee directory never brings back .claude.json, .agents/, .codex/, .gemini/, .grok/, or a .claude/, .claude.json or .mcp.json at any depth. Limits: this is a hook, so an employee with unrestricted Bash can still change the file; the Codex, Gemini, Grok and Antigravity runtimes do not run the hook, and their MCP configuration files (the four directories above) sit in the employee’s directory too, so the same class of problem is not handled for those runtimes by this fix. Not verified with the real CLI: whether Claude Code reads a .claude.json from the employee directory (frozen anyway).

Fail-closed by construction: unparseable new content, unparseable existing content, and an unreconstructable write intent all deny. An existing agent.toml, config.toml or .mcp.json that cannot be read is denied too (it used to be treated as a new file). A file that does not exist yet is allowed, because creation goes through create_agent and its own gate.

Legitimate changes keep every route they had: the MCP agent_update tool and the dashboard agents.update RPC, neither of which passes through the hook.


duduclaw memory forget-source deletes memories for good, so it has three layers in front of it. The procedure is in Forgetting a conversation, a scheduled run or an imported file; the mechanism is in Memory Intelligence.

Dashboard approval. plan files one approval request (action_kind memory_forget_source) bound to the plan id and the plan hash. It can only be decided in the dashboard and only by an Admin; channel buttons and replies are refused. apply --confirm runs only when that request is approved, still names the same plan hash, and the plan has not expired. The request expires together with the plan (30 minutes by default, at most 24 hours). The card carries counts and source labels, never memory content, and says in fixed text that the request came from a local command line and that nothing proves who typed it. No setting turns this approval off. [memory] forget_source = false only stops new plans and applies.

Bash lane rule. For an employee-identified or unverified caller, the Bash lane of agent-file-guard refuses duduclaw and duduclaw-pro (also by path or as .exe) followed by memory forget-source or memory migrate-namespace, with any subcommand, including the read-only list. The decision is BlockedOperatorMemoryCommand and the message tells the employee to ask the operator. migrate-namespace was not covered by the Bash lane before. This is a speed bump, not a sandbox. It does not catch:

  • a global option between memory and the subcommand, for example duduclaw memory --redact on forget-source …;
  • a binary named through command substitution, for example "$(command -v duduclaw)" memory forget-source …;
  • the words handed to the binary through a pipe, for example echo memory forget-source … | xargs duduclaw;
  • other spellings that avoid matching on the command name, and the evasions listed under “What these guards do not cover”.

AI session check. The command itself refuses to run when any variable that the gateway sets on a spawned employee process is present, even empty: the employee identity and token, the turn and session ids, the user-message pair of the turn, the dispatch session and run ids, the delegation sender, origin and depth, the hop depth and the reply channel, the MCP API key, the data-file-guard flag and the goal-round task id (the one list ai_session_guard::AI_SESSION_ENV_VARS, shared since 2026-10-06 by every operator command that files an approval). An operator’s own terminal carries none of them. This check is only a first line. It is not reliable for a command run directly from an employee’s Bash, because an employee with Bash can unset variables. Real isolation is not granting Bash, or running the employee in the task sandbox. An employee that runs as the same operating-system user, can execute arbitrary commands and deliberately evades the file guard can still rewrite the local databases (approvals.db, memory.db) directly.

Operator command-line approvals share one gate

Section titled “Operator command-line approvals share one gate”

Four operator commands change state from a terminal: duduclaw ops channel-ingress (LINE inbox), duduclaw responsibility (continuous responsibilities), duduclaw memory forget-source apply and duduclaw ops computer-workspaces. A terminal cannot prove who typed a command, so all four only file a dashboard request that a current Admin decides, and the command then applies once. Since the release after v1.70.0 they share one implementation (approval/operator_cli_gate.rs in the gateway crate) with one registry of kinds: an approval counts only when decided in the dashboard, for the same request and, except for a LINE batch, the same state; a waiting request whose target changed state is withdrawn (state_changed) and a new one filed, never rewritten in place; a used approval is consumed by one conditional update, and a run that loses that race acts on nothing and files nothing; at most 3 different requests wait per action and target and 20 per kind; at most 2 first notifications per target per hour (per AI employee for forget requests). approvals.decide refuses every one of these kinds for anyone but a current Admin through one check. The commands also share one list of AI-session environment variables (ai_session_guard.rs in the CLI crate). Limits are unchanged: an employee with unrestricted Bash can still edit the local databases.

Upgrading. The gate stores its binding in a gate object inside the request payload. Requests filed by v1.70.0 or earlier have no such object, so the new version never matches or counts them. A request filed before the upgrade, even one already approved and not yet applied, is not applied after it: running the command again files a new request that must be approved again. The old waiting card still receives its usual reminder pushes until it expires (responsibility requests send none), but approving it does nothing.

MCP authorization gate — every MCP tool is enumerated in a scope table; a tool that is not listed defaults to requiring Admin scope. Scope, per-agent capability grants and denied_tools are each enforced at the dispatcher front door, and every refusal is audited with an error_class.

SOUL.md drift detection — soul_guard fingerprints each SOUL.md with SHA-256 at startup and on every heartbeat tick, keeps up to 10 versioned backups in .soul_history/, and reports drift alongside an Agent Stability Index.

Audit trail — tool_calls.jsonl records every tool call with masked result_text / input_text (three-pass secret masking, mask before truncate), 0600 permissions, hash-chained lines, and rotation at 16 MB. security_audit.jsonl carries security events separately. The same log is the evidence source the grounding precheck and the acceptance judge read, so weakening it weakens verification too.

Per-agent key isolation — MCP API keys and connector credentials are per-agent and resolved through secret_ref, so one agent’s leak is not a platform leak. Channel credentials are split: LINE, WhatsApp, Feishu, Google Chat, Teams, WeCom and DingTalk use the deployment-wide credentials in config.toml [channels]; per-employee bot tokens exist for Telegram, Discord and Slack only (an employee without its own token falls back up the reports_to chain, then to the global token).

Chat commands on channels (v1.68.0) — !STOP, !STOP ALL, !RESUME and /model <name> need an admin. On WhatsApp, Feishu, Teams, WeCom, Google Chat and DingTalk the gateway used to pass is_admin = true for every sender, so anyone who could message the bot could stop or resume it. Those channels now compare the sender id or conversation id, exactly, with the channel’s admin_users setting (global scope; Google Chat and Teams can now have one); with no list, nobody is an admin. On WebChat only an active dashboard account with the Admin role counts, and website widget visitors never do.

Kill-switch triggers (v1.68.0) — the four KILLSWITCH.toml [triggers] thresholds had no reader before. Now a key is enforced only when it is written in the file and in range, and the Security page has a checkbox per trigger (unticking sends null, which removes it). The file is re-read when it changes. cost_limit_usd compares 24-hour spend across all employees and, when reached, sets the global failsafe level to restricted until the failsafe recovers or someone sends !RESUME; max_replies_per_minute counts per conversation and silently drops the excess; max_consecutive_errors and error_rate_threshold (last 20 replies, at least 10) raise that conversation’s failsafe level by one step. Each trip is audited as killswitch_trigger. The [audit] section of KILLSWITCH.toml is no longer read.

Redaction source protections (v1.68.0) — the 資料來源保護 switches on the Privacy / Redaction tab now act: user_input redacts a channel message before it reaches the AI, system_prompt redacts the assembled prompt (by default only rules marked apply_to_system_prompt), cron_context redacts the trigger message of a condition script. An error stops that turn instead of sending unredacted text. sub_agent (set on the same tab, or in config.toml [redaction.sources]) covers a delegated agent’s reply that the gateway writes into the delegating agent’s conversation history (the reply to a send_to_agent, spawn_agent or spawn_ephemeral call, including the copy relayed to the agent that started the chain). With on, the reply goes through the receiving agent’s rules before it is stored, and the tokens are restored when that agent answers its user. The default inherit leaves the reply as it is, because the sub-agent already ran under the same [redaction] rules. If the redaction step fails, a fixed notice is stored instead of the reply. Not covered: the copy of the reply sent to the user’s channel (that is the user’s own view), replies an agent fetches itself with check_responses (those are tool results and follow tool_results), hand-offs between Team-as-Agent roles, and Agent Mail. purge_after_expire_days now drives the vault cleanup.

Permission flags (v1.68.0) — agent.toml [permissions] can_create_agents, can_send_cross_agent, can_modify_own_skills and can_schedule_tasks are enforced at the MCP dispatch gate when written as false (refusal audited as permission_denied). A one-time boot migration turns old template false values into true; see dashboard-settings.md.


Saying this plainly is part of the defense.

  • The hooks see Claude Code’s own tool calls, not MCP tool calls. MCP has its own gate (scopes, grants, denied_tools); the hooks are the second lock, on the built-in Write / Edit / Read / Bash surface.
  • The Bash lane of agent-file-guard is a heuristic. It reads command text, so these get past it: a path computed from a variable (other than $DUDUCLAW_HOME, and ~/.duduclaw / $HOME/.duduclaw when the home is in the default place), from command substitution or from any other calculation; an encoded command; a script written to disk and then run; a here-document fed to an interpreter; aliases and functions; an environment variable that makes a shell started later load a file; a write option of a listed read-only command that the list did not account for; an extraction or download command with no explicit destination, which writes into the current directory (the lane does not judge the current directory, so changing into the home first and then running it is not stopped); a link created and then written through within the same command; hard links; and the time gap between the check and the actual execution.
  • Operator-only memory commands are guarded by a speed bump and an approval, not a sandbox. The Bash lane rule for memory forget-source and memory migrate-namespace misses a global option before the subcommand, a binary named by command substitution and a piped command, and the AI-session check can be defeated by unsetting variables. The dashboard approval is the gate; see Forget by source needs an Admin approval.
  • The Bash lane also refuses some harmless commands. A judged path whose resolution fails is refused even when it lies outside the home, and a dangling symbolic link is refused even when it points outside the home.
  • agent-file-guard does not cover Read. Held-out eval sets and the audit log remain readable to an employee; the hook only stops writes.
  • Only the Claude runtime runs these hooks. Codex, Gemini, Antigravity and the other runtimes rely on their own sandbox flags.
  • State files inside the employee’s own directory are not protected, apart from SOUL.md, CONTRACT.toml, the identity files (.mcp.json, .claude/settings.json) and its agent.toml. The shared attachments/ directory is writable by every employee.
  • data-file-guard is a heuristic. It matches filenames in a Bash command line; a dynamically-built path defeats it. (It is no longer inert on Windows — H10 made it a Rust subcommand.)
  • There is no threat-level state machine. ~/.duduclaw/threat_level survives as an operator-controlled kill switch that the computer-use orchestrator polls (RED stops the run, YELLOW pauses it), but nothing inside the workspace writes it. A missing file means GREEN; a file that exists but cannot be read, or holds anything other than GREEN / YELLOW / RED, is treated as RED (fail closed) after two re-reads 50 ms apart; a leading UTF-8 BOM and surrounding whitespace are ignored.
  • (Removed 2026-09.) This section used to note that the PTY session pool sat outside the redaction rewrite. That pool no longer exists — every Claude spawn is a per-call spawn, which is exactly what the rewrite hooks into.

  • CONTRACT.toml defines what an agent must never do, and duduclaw test red-teams it. The guards enforce at the tool-call level.
  • Evolution engine — because SOUL.md is read-only to agents, the evolving artifact is the playbook. See 38-aee-playbook-evolution.md.
  • Redaction and data sources — see 55-data-sources.md for the pipeline data-file-guard complements.
  • Delegation isolation — see 37-delegation-isolation.md for the predicate org_field_guard protects.

Four guards with stated failure modes beat a three-layer story with no code behind it. When a defense is removed the documentation has to go with it: a page describing a shell script that does not exist is worse than no page, because it makes an operator stop looking.

High-risk Computer Use confirmations use the inbound account and exact conversation/thread. Reply with 確認 <full UUID> or 取消 <full UUID>; questions use 回答 <full UUID> <answer> and never grant tool permission. Bare yes/A/B cannot select a request. Before execution, the host rechecks the live screen, title, policy and cancellation gates. Restart invalidates old GUI approvals instead of replaying coordinates; an execution without a receipt becomes uncertain and requires Admin reconciliation. See the decision guide for the exact supported inbound routes and current limitations.