Skip to content

MCP tools: what an agent is shown, and why

DuDuClaw’s MCP server advertises its tools through the standard tools/list call. This page explains two things that are easy to get wrong:

  1. why an agent sees fewer tools than the server actually implements, and
  2. where the long-form details went, now that every tool description is capped.

If you are adding a tool, read custom-mcp-tool.md first; this page is about the declaration surface, not the implementation. To run the server on its own, without the gateway, see mcp-standalone.md.

tools/list only lists the tools the calling agent can actually call right now. Every filter below mirrors a gate the dispatcher already enforces, so the list an agent reads and the calls the server accepts describe the same set.

This matters for cost as much as correctness. The tool schemas are a fixed prompt cost paid on every spawn: the CLI reads tools/list once per session and the whole payload lands in the model’s context before the first user token. A schema for a tool the gate would reject is paid for twice — the tokens, and the model planning around something it cannot use.

Hiding a tool is not an authorization decision. Calling an unlisted tool still reaches the real gate and is still refused, with the gate’s own message.

Filter Source Effect when off / empty
External client whitelist principal.is_external Only the 7 legacy whitelist tools plus tools in explicitly granted scopes; then the scope row below applies
Scopes of a non-employee caller principal.scopes, for a key that holds no admin and is not an AI employee (not the gateway-internal key, not a per-agent key, not a process the gateway spawned for an employee) Only tools whose minimum scope the key holds; tools that act for the process’s own agent (working_state_*, canvas_*, shared_wiki_delete, memory_search_by_layer and three consolidation reads, …) are hidden from it, and from every external key, and the dispatch gate refuses them by the same rule (process_agent_tool). See mcp-standalone.md
Google Workspace config.toml [integrations] google_workspace 19 gmail_* / calendar_* / sheets_* / forms_* / gtasks_* / drive_* / docs_* / slides_* tools hidden
GitHub config.toml [integrations] github 5 github_* tools hidden
denied_tools / allowed_tools agent.toml [capabilities] Denied tools hidden; a non-empty allowlist hides everything else. An entry matches a tool exactly, or by a trailing *: *, mcp__duduclaw__* (every DuDuClaw tool), mcp__duduclaw__odoo_* / memory_* (anchored prefix). mcp__<other server>__… never names a DuDuClaw tool; a * elsewhere is literal. The approval lists and scoped_tools use the same rule
os_native agent.toml [capabilities] 6 os_* automation tools hidden
recording agent.toml [capabilities] 5 browser/desktop recording tools hidden
system_operator agent.toml [capabilities] 19 appliance operation tools hidden
codrive agent.toml [capabilities] codrive_run / codrive_status hidden
computer_use agent.toml [capabilities] 8 computer_* tools hidden (see computer_* below)
computer_use_config.workspace agent.toml [capabilities.computer_use_config] (with computer_use) 3 computer_workspace_* tools hidden (see computer_workspace_* below)
db_sources agent.toml [capabilities] 4 db_* tools hidden
[fork] enabled agent.toml 6 forking tools hidden
[responsibilities] enabled config.toml 3 responsibility_* tools hidden (see responsibility tools below)
scoped_tools agent.toml [capabilities] + live grant Hidden until a task-scoped grant is active

A freshly created agent has none of these opted in, which is the point: the default deployment pays for the tools it can use and nothing else.

How a mid-session change reaches the agent

Section titled “How a mid-session change reaches the agent”

Hiding a tool used to make it permanently unreachable, because an MCP client reads tools/list once when the session opens. The server now declares tools.listChanged in its initialize response and emits notifications/tools/list_changed whenever the caller’s visible set actually changes — compared as a set, so touching a config file without changing it sends nothing.

So this sequence works without a restart:

  1. An agent hits a scoped_tools denial, or an operator grants a capability.
  2. capability_request is approved (or the dashboard saves agent_update, or someone edits agent.toml).
  3. Within a few seconds the server notices the change and notifies the client.
  4. The client re-reads tools/list and the tool is there.

A task-scoped grant is revoked at every task terminal state, and the same mechanism takes the tool back out of view.

Caveat: an agent cannot self-discover a scoped_tools name

Section titled “Caveat: an agent cannot self-discover a scoped_tools name”

A tool listed in scoped_tools is hidden until a grant is active, so the agent cannot read the name off tools/list in order to ask for it. Tell it the name another way — in its SOUL.md, in a playbook rule, or by minting the grant at goal kickoff with a grant:<tool> tag. This is the one place the pruning costs something, and it is deliberate: advertising a tool that is denied right now is the failure mode the rule above exists to remove.

Exceptions since v1.68.0: listed but refused

Section titled “Exceptions since v1.68.0: listed but refused”

Two new gates refuse a call without removing the tool from tools/list:

  • agent.toml [permissions]: a flag written as false refuses create_agent (can_create_agents); send_to_agent, spawn_agent (can_send_cross_agent); create_reminder and tasks_create with a schedule (can_schedule_tasks); skill_hub_install, shared_skill_adopt, skill_graduate, skill_pin, skill_from_recording (can_modify_own_skills). The refusal is JSON-RPC error -32003 and a permission_denied audit event. An agent.toml that exists but cannot be read or parsed refuses these tools; a missing one allows them. tasks_create without a schedule stays allowed, which is why these flags are checked per call. Ephemeral role members are scaffolded with can_create_agents, can_modify_own_skills and can_schedule_tasks set to false.
  • config.toml [odoo] features_*: Odoo tools stay listed and are refused per call for models of a switched-off module (project and hr are off by default).

Record relationship checks: listed but refused

Section titled “Record relationship checks: listed but refused”

Some tools change or trigger a record that belongs to an AI employee. They stay listed for every caller and check the relationship per call:

  • Tasks: tasks_update and activity_post with a task_id pass for the task’s assignee, claimer or creator; tasks_complete and tasks_block pass for its assignee or claimer; tasks_claim passes for an unassigned task or one already assigned to the caller. Anything else needs a delegation relationship with the assignee (same department, reports_to above or below, or a whitelist pair, per [delegation] policy). An unassigned, unclaimed task must be claimed first. Taking another employee’s task for yourself through tasks_update always needs the relationship.
  • Task fields: an AI employee cannot change title or description of a goal-mode task (acceptance_criteria is refused for every MCP caller), and cannot add, remove or reorder tags starting with outcome: or grant: or the auto-research tag, in tasks_update or in tasks_create’s tags.
  • Parent of a new task: the gateway tells the MCP server which task the current round works on (DUDUCLAW_TASK_ID, the same value approval cards carry). When that task is the caller’s own and belongs to a continuous responsibility run (the run’s task or any task under it, including a sub-task woken by the task board), tasks_create (including kind="goal") places the new task under it when no parent_task_id is given, and a parent_task_id the employee gives must be that task or one of its sub-tasks, or the call is refused. A value that is present but empty or malformed is refused, never read as “no round”. In every other case (an ordinary goal round, a task-board wake-up of a task outside any run, no round information) there is no default parent. New in this release for every caller: a given parent_task_id needs a relationship to the parent (assignee, claimer, creator, or the delegation policy), where v1.69 wrote it unchecked; kind="goal" accepts parent_task_id, which v1.69 ignored; and a task can hold at most 200 unfinished sub-tasks created by AI employees (a schedule, routine or reminder, is not a sub-task). The round is not passed to a server started from Bash, or under the Grok and Gemini CLI runtimes; a task created there is not placed under a run unless the employee names the run, and is not counted toward its spending. Correct placement also relies on the employee not being able to rewrite the duduclaw entry of its own .mcp.json, a separate platform fix that must ship first.
  • Routines: update_cron_task, delete_cron_task, pause_cron_task and run_cron_task need the caller to be the employee the routine runs as, or to have the relationship with it. Addressed by name, they act on exactly one routine; a shared name is refused with the candidate ids.
  • Reminders: create_reminder with an agent_id other than the caller needs the relationship with that employee.
  • agent_update on yourself: an AI employee cannot send reports_to, db_sources, db_sources_add, db_sources_remove, budget_cents or role about itself (audit agent_authority_refused). Editing a subordinate is unchanged.

An operator (an MCP key that maps to no AI employee, in a process that is not running for one) is not restricted. The shared internal key in a process with no employee identity owns no record, so it is refused on everyone’s. A process whose identity is a system sender name (dashboard, cron, …) is treated as untrusted. Every refusal is audited in tool_calls.jsonl. Details: task board, delegation isolation.

Each tool’s description is capped at 200 bytes and each parameter’s description at 200 bytes (one documented exception, below). The cap is enforced by a test, not by convention.

Two consequences for anyone adding or editing a tool:

  • Say what the tool does and what will refuse it. Safety-critical clauses — “this does not send”, “this is publicly visible”, “rejected whole, never truncated” — stay in the description. Rationale, worked examples and design labels do not.
  • Put the long version here. Link a section of this page, or the tool’s own spec page, from the description.

team_handoff’s packet parameter carries the full TaskPacket shape (capped at 1,024 bytes). build_tool_schema types every parameter as a bare JSON-Schema string, so a parameter description is the only place its real shape can be stated — and a TaskPacket is rejected whole rather than truncated, so an agent that cannot see the shape cannot produce a valid one. The full reference is ../spec/task-packet.md.

Exceptions live in one list (PARAM_CAP_EXEMPTIONS) with a test that fails when an exemption stops being necessary.

Long-form details moved out of descriptions

Section titled “Long-form details moved out of descriptions”

codrive_run — the three-rung execution ladder

Section titled “codrive_run — the three-rung execution ladder”

Each step is dispatched through three rungs, tried in this order; prefer the highest one that applies.

  • C-L2 — api_action. Calls a registered third-party app’s native API/CLI/D-Bus action for target_app before touching the GUI at all. Prefer this whenever the app and action you need are registered. action is a short registered identifier (chromium’s open_url, networkmanager’s state, …); params is that action’s payload, validated against the action’s own schema at dispatch time. A registry miss or an exec failure falls through to the step’s action field — which is why action is required even here.
  • C-L3 — locate. Resolves a move/click step’s coordinates by (role, name) lookup in target_app’s AT-SPI2 accessibility tree instead of hand-guessed pixels. Far more robust to layout, resolution and theme changes. Ignored for text / key_name / wait / take_over. A miss falls back to the literal x/y.
  • C-L1 — literal x/y. The fallback when the other two are absent or fail.

Other script-level rules:

  • target_app is a single script-level field, not per-step. One script drives one app; issue a second codrive_run call to drive another.
  • Every consequential step (send / submit / delete / purchase / other) pauses for human approval before any rung dispatches it. A refuse-list hit (banking pages, CAPTCHA bypass, …) is rejected outright, before a connection is attempted.
  • A login/password/payment step (take_over, or a credential class) hands the shared desktop to a human. You never send the credential text yourself through any rung; the script resumes when they hand control back.
  • Any human input on the shared desktop immediately freezes the agent’s seat. The dropped step is retried once a person hands control back.
  • watch_mode: true arms idle-based supervision for the rest of the run.
  • Maximum 50 steps.

Two modes:

  • Plain note. Omit status; legacy behaviour, silently truncated at about 1,200 characters.

  • Structured (Ralph-loop style). Pass status and it is validated together with next_steps / evidence / blocker:

    • continue — requires non-empty next_steps, forbids blocker.
    • complete — requires non-empty evidence, forbids both blocker and next_steps. A self-declared “done” with no evidence, or with a leftover next step, is rejected: “I finished” is not evidence.
    • blocked — requires a non-empty, specific blocker.

    The combined payload is capped by config.toml [memory] working_state_handoff_max_bytes (default 16,384, CJK-safe byte count). Going over rejects the whole call; it is never silently truncated, because truncating could delete exactly the evidence that made the handoff authoritative.

Leave source alone unless you already know where the skill lives.

  • all (default) — configured hubs plus this agent’s learned skill bank, de-duplicated by name and labelled with its source. Hub results are ranked by relevance × trust × install count × freshness, with official first-party skills floored into the top results.
  • github — a skill you expect in a public GitHub repo.
  • hub — the curated registries (anthropic-skills, github, clawhub, lobehub, skills-sh); narrow to one with hub.
  • bank — only what this deployment learned on its own. hub is not accepted with this source.

evolution_toggle — the stagnation sub-fields

Section titled “evolution_toggle — the stagnation sub-fields”

Beyond the standard flags, field accepts stagnation_enabled (bool), stagnation_window_seconds (60–604800), stagnation_trigger_threshold (1–1000) and stagnation_action (log_only | suppress). See evolution-switches.md.

The script runs in the script sandbox: a container from the image in config.toml [container.sandbox] image (the same image as the task sandbox, never pulled automatically). The container runs as the host user (1000:1000 when the host process is root, and always on WSL2), with all capabilities dropped, no-new-privileges, a read-only root filesystem, no network, 2 GiB of memory with no swap, 256 processes, one CPU and a small /tmp tmpfs. Only a private directory holding the script is mounted, read-only, at /workspace. timeout_seconds (default 30, at most 300) applies under a hard cap of 600 seconds, stdout and stderr come back as one output (the read is capped at 2 MiB, the reply at 1 MiB), and the container is force-removed if the call is cancelled. Docker is used on macOS and Linux; on Windows WSL2 is tried first, then Docker.

When the sandbox cannot run (no Docker, image missing, invalid [container.sandbox], …) the script is not run: the tool returns Script sandbox unavailable (<code>): … with the docker pull <image> command and writes the audit event script_sandbox_unavailable. Older versions silently ran the script on the host instead. To get that back, set [container.sandbox] script_when_unavailable = "run_unsandboxed" (a separate key from the task sandbox’s when_unavailable); every host run is then audited as script_sandbox_bypassed.

A script cannot call platform tools back: there is no RPC socket inside the container.

computer_* — sessions run by the gateway

Section titled “computer_* — sessions run by the gateway”

The eight tools drive one computer-use session per employee: an isolated container with a virtual display and a kiosk browser. The MCP server only forwards each call to the gateway over loopback (POST /api/internal/computer-use, signed per request); the gateway owns the container and runs every check, so the gateway must be running. Hidden unless agent.toml [capabilities] computer_use = true.

Tool Parameters Notes
computer_session_start task string, optional; width integer 320–1920; height integer 240–1200; workspace string, optional: "new" or a ws-… id from computer_workspace_list One session per employee. The result lists the limits, whether high-risk actions can be confirmed in a chat, and the sites computer_navigate can open. With workspace it also returns workspace_id, mount_path (/workspace/files, read-only, root only), the revision, usage and quota
computer_screenshot none MCP image block (PNG, masked) followed by a text block with actions used and time left. A fully masked picture is reported as such in the text, with the reason (several windows, sensitive or unreadable front window, detection failure) and the next step
computer_click x, y integers (required); button string left/right; double boolean double is left button only
computer_type text string (required), 1–2,000 characters Audited as a character count only
computer_key key string (required): letters, digits, +, -, _ e.g. Return, ctrl+s
computer_scroll x, y integers (required); direction string up/down (default down); amount integer 1–20 (default 3)
computer_navigate url string (required) https:// only, host exactly on the employee’s [capabilities.computer_use_config] allowed_domains and resolved at session start, port absent or 443, no user name or password, at most 2,000 bytes. With no allowlist the session has no network and the call is refused
computer_session_stop session_id string, optional Removes the container

Integer and boolean parameters also accept numeric and "true"/"false" strings. Click, type, key, scroll and navigate each count as one action against max_actions (default 50). Limits, the approval and confirmation rules, the network allowlist and its residual risks are in Browser automation.

computer_workspace_* — durable workspaces

Section titled “computer_workspace_* — durable workspaces”

A folder the gateway keeps for one employee across computer-use sessions. Listed only when [capabilities] computer_use and [capabilities.computer_use_config] workspace are both true; every call also needs config.toml [computer_use.workspaces] enabled = true (except that the owner may still list and read when the feature is switched off). A workspace of another employee, a malformed id and a missing one all get the same “not found” answer.

Tool Parameters Notes
computer_workspace_list none The caller’s workspaces: state, data_revision, bytes_used, files_used, quota, expires_at, leased, and for readable ones files (path, size, sha256, hash_unknown), files_truncated (more than 200 files) and unprocessable_items (a count; names are never shown)
computer_workspace_read workspace_id string (required); path string (required) One UTF-8 text file of at most 48 KiB. The content comes back inside a <computer_workspace_file> data fence with injection-scan flags
computer_workspace_write workspace_id string (required); path string (required); content string (required); expected_revision integer, optional Only into the workspace the caller’s live session attached. At most 48 KiB of UTF-8, and the whole request must fit 64 KiB after JSON encoding, so content full of quotes, backslashes, newlines or control characters fits less. Atomic; refused over quota, on a revision mismatch, and while the session is stopped or paused or the threat level is not GREEN. Content is never restored from redaction tokens

Paths are relative, at most 4 levels, letters, digits, inner spaces and -_.()() only. The audit rows carry a path hash and a character count, never the path or the content. Settings, the lease, operator commands and known limitations: Computer-use workspaces.

belief_stats / belief_settle — verified and self-reported settlements

Section titled “belief_stats / belief_settle — verified and self-reported settlements”

Only a settlement cross-checked against a platform price counts toward calibration. Nothing in production supplies that cross-check today: belief_settle records every settlement as the employee’s own report (settle_source = "agent_unverified"), so calibration reads “no verified settlements” on every current deployment.

belief_settle returns the settled row plus counts_toward_calibration (boolean) and, when it is false, a note saying the settlement was recorded but does not count.

belief_stats returns (the same object as the dashboard’s belief.summary stats, plus a note):

Field Meaning
n_submitted every belief submitted, settled or not
n_settled_all every settled belief (verified.n + self_reported.n)
calibration_status no_verified_settlements, insufficient_samples (1–29 verified) or calibrated (30 or more)
verified n, hits, and hit_rate, hit_rate_wilson_low, mean_brier, overconfidence (all null unless calibrated)
self_reported n and a descriptive hit_rate (null when n is 0); not calibration
per_subject[] subject, verified (n, hits, mean_brier), self_reported (n)

The flat fields of earlier versions (n_total, n_settled, insufficient_samples, top-level hit_rate …) are gone. See Belief loop.

create_agent / agent_remove — removed names are reserved

Section titled “create_agent / agent_remove — removed names are reserved”

agent_remove moves the employee to ~/.duduclaw/agents/_trash/ and answers that the employee was removed, that the administrator can restore it, and that the name is reserved. It returns no path. create_agent then refuses that name for every MCP caller while a trash entry exists, while org.toml still records the id without a directory, or while the trash cannot be listed; a different name works. Operators can reuse the name from the dashboard or a terminal. Over HTTP with a non-internal key, both tools act as that key’s own client id. Details: Delegation isolation.

responsibility_get / responsibility_followup / responsibility_ask — responsibility tools

Section titled “responsibility_get / responsibility_followup / responsibility_ask — responsibility tools”

The three tools for continuous responsibilities. They are listed only while config.toml [responsibilities] enabled is true and the caller is an employee identity (external clients never see them), and every call is refused while the feature is off. Like the tasks_* tools they sit behind the Admin scope, which the gateway’s own key carries.

Tool Who may call Rules
responsibility_get Any employee; an operator key may read any responsibility No responsibility_id: lists the caller’s own. With an id: the summary (schedule, subscriptions, window spend and runs, the open run, cost_not_counted). Another employee’s responsibility reads as not found unless the delegation policy relates the caller to its owner
responsibility_followup The owning employee, only while its own run of that responsibility is open and the responsibility is active One one-shot wake-up at due_at (RFC 3339), at least min_wake_interval_secs from now and no later than stop_at. It counts toward the window’s run limit and is refused when that limit is used up; it is never silently moved to another time
responsibility_ask Same as responsibility_followup One question (at most 1,000 characters) with up to 5 options (each at most 1,000 characters); ttl_secs defaults to one day and is held between 60 seconds and stop_at. The question is scanned for prompt injection and shown to the operator as quoted text cut to 200 characters, with a warning when the scan matched. Any push goes through the responsibility’s notification gates and per-window cap. The answer wakes the next run as data and grants nothing. Stopping the run withdraws the question

A responsibility holds at most one live wake-up armed by its employee, so a pending responsibility_followup and a pending responsibility_ask exclude each other. Operator keys cannot use the two wake-up tools, since that would be acting as the employee. Closed error codes: not_found, not_active, no_open_occurrence, invalid_due_at, period_occurrence_limit, agent_followup_limit, epoch_changed, invalid_question.

There is no tool for creating, changing, resuming or re-enabling a responsibility, and none for steering a task.

memory_store / user_profile_record / decision_resolve / wiki_write — sources and refusals

Section titled “memory_store / user_profile_record / decision_resolve / wiki_write — sources and refusals”

Every memory write records where it came from, and wiki_write stamps the page it writes. The gateway supplies the source through the environment of the employee process (turn id, conversation id, and the number and time of the user message of that turn, or the run key of a scheduled or dispatched run); the model cannot name or change it. A call from an external MCP client is recorded as one call of that client.

A memory write is refused with a tool error in two cases:

  • The source was forgotten by the operator (duduclaw memory forget-source): “Not stored: the conversation this write comes from was forgotten by the operator (forget by source), so nothing derived from it may be stored again”. decision_resolve answers “Not resolved: the conversation this choice comes from was forgotten by the operator”. The refusal is audited as memory_write_fenced.
  • The employee process carries a conversation or run identity from the gateway that is malformed: “Not stored: this employee process carries a malformed conversation or run identity from the gateway, so the write could not be tied to its source. Restart the gateway; if it persists, report it.” The write is not recorded as an anonymous external call.

A write that is derived from other memories and names a parent memory that does not exist is also refused. The memory, key fact and its source rows are written in one transaction, so a write either has its source recorded or does not happen.

wiki_write (both scopes) adds a host_sources key to the page’s frontmatter with the sources of that write, keeping the newest 20 across rewrites. A host_sources value the employee writes itself is dropped. The key is only read by forget by source, which lists such pages for human review and never deletes them. No MCP tool performs a forget: it is an operator command, refused in an employee’s session.

A deprecated tool name keeps appearing in tools/list with a [deprecated → …] prefix on its description, because hiding it would make it uncallable — the opposite of what a deprecation window is for. The full old → new table is deprecations.md. No MCP tool is deprecated at the moment: the aliases of the v1.66.0 window were removed in v1.69.0.