Goal loop
Drop a goal into a conversation and the AI employee plans, executes, and self-checks its own work, then comes back to you when it’s done or stuck. This page covers the /goal command in channels, the AutonomyLevel tiers, the relevant config keys, and what the buttons mean when a task escalates to a human.
The dispatch engine has been on by default since v1.59 (with no goal tasks queued it only does periodic SQLite polling; an LLM call only happens once a task actually reaches review — see “Two-stage acceptance judging” below), so plain question-and-answer conversations are completely unaffected. To turn it off, set [dispatch] enabled = false in config.toml, or flip the “Dispatch engine (派工引擎)” switch under Settings → Automation (設定 → 自動化) on the dashboard (hot-reloads, no restart needed).
The /goal command
Section titled “The /goal command”Type this to an AI employee in any connected channel (Telegram / Discord / Slack / LINE / …):
| Command | Behavior |
|---|---|
/goal <goal description> |
Creates an autonomous goal task assigned to the AI employee in the current conversation. With no separate acceptance criteria, the goal description itself becomes the acceptance basis. |
/goal <goal> || <acceptance criteria> |
Split with ||: the first half is the goal, the second half is the acceptance criteria (what the judge checks against). |
/goal <goal> || <acceptance criteria> || outcome:<spec> |
Adds a layer of structured outcome verification (see “Structured outcome verification” below). A zero-cost deterministic check runs before delivery; anything that fails goes straight back for revision without spending a judge call. |
/goal status |
Lists the AI employee’s in-flight goal tasks (short code / status / round number). |
/goal |
Shows usage help. |
Example
/goal Compile this batch of customer data into a monthly report and send it || The report needs a monthly revenue chart, send to boss@example.com/goal Produce the Q3 monthly report || Include a monthly revenue chart || outcome:files:report.docxCreating a goal returns a confirmation message with the task’s short code, its round cap, and a note that “I’ll notify you here when it’s done or stuck.” Progress updates and human-needed notifications are pushed back to the conversation where you created the goal (the source channel), not just to the AI employee’s [proactive] notification channel.
If the dispatch engine is off (
[dispatch] enabled = false), the task is still created, but the confirmation message tells you it won’t start running on its own.
Goal contract: frozen at creation, never quietly edited later
Section titled “Goal contract: frozen at creation, never quietly edited later”The moment a goal is created, its acceptance criteria are frozen into an immutable baseline (acceptance_criteria_baseline). Every subsequent verdict — the first-stage evaluator, the MAV acceptance judge — reads only this frozen baseline, never a field that could have been edited afterward. This blocks contract drift in both directions: the AI employee can’t loosen its own acceptance bar mid-task, and an operator can’t mistake editing a dashboard field for actually changing what the judges check against.
- The AI employee can’t change it: if an agent-identity call to the MCP
tasks_updatetool tries to changeacceptance_criteriaon its own goal task, the whole call is rejected, and an audit entry is logged (reason:goal_contract_frozen). The goal’stitleanddescriptionare what the judge reads as the goal, so an AI employee’stasks_updatethat touches either of them on a goal task is refused the same way. The control tags on the task (outcome:…,grant:…,auto-research) cannot be added, removed or reordered by an AI employee either (reason:reserved_tag_change). - The operator can edit the display copy, but that never changes what gets judged: the dashboard’s
tasks.updatecan still edit theacceptance_criteriashown on a task (for example, adding a note for a human reader), but the frozen baseline value does not move with it. The judge and the evaluator keep grading against the standard set when the task was created. Changing the acceptance criteria for real means creating a new goal. - Older tasks with no frozen baseline (created before this mechanism shipped) fall back to reading the mutable field, unchanged from before.
Guidance when no acceptance criteria are given
Section titled “Guidance when no acceptance criteria are given”/goal <goal description> (with no ||) still creates the task as usual, using the goal description itself as the acceptance basis, but the confirmation message adds a hint to help you think through what to specify next time:
💡 No separate acceptance criteria this time — next time these four things will make it easier to pin down: • Goal: what needs to be accomplished • Input: what data or material is needed • Output format: what the result should look like (e.g. Word / Excel / a block of text / an image) • Constraints: style, deadlines, and other limits, plus "what counts as done" Try 3-5 concrete, checkable acceptance criteria (e.g. "the report includes a monthly revenue chart" rather than "make it good"), then resubmit with /goal description || criteria.This hint doesn’t appear when || acceptance criteria are given explicitly. Subtask decomposition under planner_enabled follows the same discipline: each acceptance criterion states the result, not the method (freezing HOW would fail a correct output that took a different but equally valid path); keep it to 3-5 criteria (a heavier contract makes subtasks harder to converge); anything out of scope goes under Non-goals instead of being crammed into acceptance criteria.
Acceptance ledger (per-criterion)
Section titled “Acceptance ledger (per-criterion)”The frozen acceptance criteria are one block of text. The judge is told to check them item by item, but nothing a program could read showed that every item had actually been handled. The acceptance ledger fixes that.
When it is created. A goal created through the dashboard or the MCP tasks_create kind="goal" tool gets a ledger at creation, from the frozen baseline: one criterion per non-empty line, numbered C1, C2, … in order. At most 20 criteria are tracked; lines beyond the 20th are folded into C20 with a note, never dropped. The system also derives a stable id per criterion (canonical_id(task id, "criterion", index, fingerprint of the text)), so the model only ever echoes a short handle and never invents an id. Goals created before this feature, or while the mode is off, have no ledger and run exactly as before. Goals started from the chat /goal command and from a confirmed goal suggestion (both “立為目標任務” and “想一想”) get a ledger the same way. Goals created by autopilot rules and the sub-tasks of a planner-decomposed goal have no ledger.
What the worker sees. Every dispatch round carries a ## 驗收帳本 section right after the <state> block, one line per criterion with its current status, and the reporting instruction. The worker ends its tasks_complete result summary with a tag:
<criteria_status>[{"id": "C1", "status": "covered", "evidence": ["wrote reports/summary.md"], "unresolved": []}, {"id": "C2", "status": "blocked", "evidence": [], "unresolved": ["no permission to send mail"]}]</criteria_status>The tag body must be exactly one JSON array (no prose around it, no extra fields), every handle exactly once. status is covered (evidence required, unresolved empty), blocked (unresolved required) or candidate (thinks it is done, wants confirmation; evidence required). Each evidence or unresolved entry is cut at 500 characters, at most 8 entries per field. A report that breaks any rule is discarded as a whole: the ledger stays as it was, its invalid_reports counter goes up, and one criteria_status_invalid event goes to security_audit.jsonl with the violation and at most 200 masked characters of the tag body. A round without the tag changes nothing and is not counted. The tag is removed from the text the judge reads and from the ✅ message pushed back to the chat.
Modes (config.toml [goal_loop] criteria_ledger, read at every round, no restart):
| Mode | Ledger | Judge |
|---|---|---|
off |
Not created, not injected, not parsed | Unchanged |
report (default) |
Created, injected, parsed, shown to people | Gets the worker’s self-reported ledger as a reference block marked “self-report, not evidence”; its reply format is unchanged |
enforce |
Same as report |
The panel must also return criteria: [{"id": "C1", "pass": true, "reason": "..."}] with every handle exactly once. A missing or duplicated handle is a FAIL for that criterion, and correctness passes only if every criterion passes and the aspect itself passes |
An unknown value reads as report. Moving to enforce follows the same observation discipline as strict_reply_parsing: switch it yourself after watching real judge rounds. Under enforce an external judge (judge = "external") is not asked for per-criterion verdicts; its own pass/fail stands.
What you see. tasks.timeline returns criteria_ledger (null when the goal has none) with mode (the mode in effect now), units[] (id, handle, text, status, evidence, unresolved, updated_round), last_report_round and invalid_reports. A needs_human card for a goal with a ledger adds one line such as 驗收帳本:1/3 條已回報達成;C2 受阻(沒有寄信權限);C3 尚未回報. Each round’s ledger is also kept on its iteration row for later review.
needs_human’s pause_reason and the ledger answer different questions (why the loop stopped versus which criterion is still open); they are not merged.
Outer-loop progress board
Section titled “Outer-loop progress board”Every state transition on a goal task pushes a short (one-to-three-line) progress message back to the source conversation:
- Started / retrying (round N of the cap)
- In review
- Rejected → revising and retrying (with a summary of the judge’s feedback)
- Done ✅ (with a result summary)
- Stuck → needs your decision (also pushes approval buttons, tagged with a pause-reason category — see “needs_human button semantics” below)
The same task won’t push the same state twice. If the source conversation no longer exists, it falls back to the AI employee’s [proactive] channel; if neither exists, it’s written only to the dashboard’s Activity Feed, without interrupting you.
Stall-timeout progress reports
Section titled “Stall-timeout progress reports”Once a task is claimed (in_progress), if [goal_loop] progress_report_minutes (10 minutes by default) passes with no observable progress signal (measured against Activity Feed events — the updated_at field is refreshed periodically by the lease renewer and can’t be used as evidence of “something is happening”), the driver pushes a single “still running after X minutes with no progress reported” notice (Activity Feed + source conversation), at most once per round.
This is purely a report, not an intervention: it doesn’t redispatch, escalate, or cancel the task. The guardrails that actually act remain stalled_secs (redispatch), iteration_cap (escalate to a human), and wall_clock_hours (escalate to a human). Setting it to 0 (or any negative number) turns the whole feature off, with zero extra queries when disabled.
Tool streak advisory
Section titled “Tool streak advisory”In the autonomous loop, an AI employee sometimes gets stuck calling the same tool with the same arguments over and over — it already has the result and just hasn’t noticed it’s repeating itself. This is different from “Stall detection” below: stall detection needs two full rounds (dispatch → review) to notice a stuck task, while the tool streak advisory watches the tool-call sequence within a single round, so it can warn before the judge even gets involved.
The trigger is a run of consecutive calls to the same tool with the same masked arguments (reusing the secret-masking the audit trail already does), and the warning escalates by threshold:
| Consecutive calls | Warning |
|---|---|
| 3 | Suggests re-reading the last result to check whether the needed information is already there, before wasting another round repeating the call |
| 5 | Notes the current approach may not be making progress, and suggests a different method or angle |
| 8 | Strongly suggests stopping the repetition: wrap up whatever result is already available and report it, or call tasks_block to explain what’s blocking and ask for help |
The reminder is injected into the next round’s <state> block, at zero LLM cost and purely advisory: it never blocks, retries, or vetoes a tool call or a dispatch round — whether to act on it is left to the AI employee. The reminder itself is deliberately excluded from state_hash, so it doesn’t interfere with the existing (state, action) oscillation detection.
config.toml [goal_loop] tool_streak_advisory (true by default) turns this off entirely when disabled.
AutonomyLevel: five levels of autonomy
Section titled “AutonomyLevel: five levels of autonomy”Each AI employee’s autonomy level is controlled by a single dial: agent.toml [capabilities] autonomy_level. Unset or unparseable defaults to Approver (conservative: only asks you when it’s stuck or needs a human).
| Level | Behavior |
|---|---|
operator |
The loop never drives itself; a task sits idle after creation and a human advances it manually. |
collaborator |
Needs human approval before the first dispatch (kickoff approval); once approved, it retries on its own until done. |
consultant |
Same kickoff approval as collaborator. |
approver |
Default. No kickoff gate; escalates to a human only when stuck or genuinely needing one. |
observer |
Fully automatic; when it needs a human it only notifies, never waits (the task ends on its own). |
[capabilities]autonomy_level = "approver"Task-scoped tool grants (scoped_tools, v1.41)
Section titled “Task-scoped tool grants (scoped_tools, v1.41)”High-risk tools can be declared “usable only with a grant”: a tool listed under scoped_tools is refused until the AI employee holds a valid grant, and a grant lives only for the lifetime of a single task — it’s revoked automatically the moment that task ends (accepted, rejected, escalated to a human, or cancelled), never carrying over to the next one.
[capabilities]scoped_tools = ["shared_wiki_delete", "odoo_execute"] # these tools need a per-task grantgrant_ttl_secs = 3600 # hard cap on how long a grant can live (seconds), default 3600Two ways to get a grant:
- The AI employee requests one itself: calling the MCP tool
capability_request { tool, reason, task_id? }turns into an approval request (through the same notification/dashboard flow as other approvals); once you approve it the grant takes effect, and an approval left undecided past its deadline counts as denied. - Granted at goal kickoff: adding a
grant:<tool name>tag to a goal task grants it atomically when the kickoff approval (collaborator/consultant level) is approved, and it’s reclaimed automatically when the task ends.
The check is always fail-closed: if the grant store can’t be read, that’s treated as no grant. Tools not listed under scoped_tools are completely unaffected.
Configuration keys
Section titled “Configuration keys”config.toml (global)
Section titled “config.toml (global)”[dispatch]enabled = true # Enable the autonomous dispatch engine (includes the goal loop driver). Default falsepolicy = "fixed_hierarchy" # Dispatch policy (which AI employee picks up a task). See "Dispatch policy" below. Default fixed_hierarchygrounding_precheck_enabled = true # Grounding precheck before acceptance (see "Grounding precheck"). Default truetwo_stage_judge = true # Run a cheap first-stage evaluation before acceptance (see "Two-stage acceptance judging"). Default truestrict_reply_parsing = "shadow" # Judge replies under the strict JSON contract: off / shadow / enforce (see "Strict reply contract"). Default shadowjudge = "mav" # Who makes the acceptance call (see "Swapping the acceptance judge"). mav / external (evaluator_only / human_only were removed in v1.69.0). Default mavjudge_provider = "antigravity" # Optional: run the judge on another runtime (see "Running the judge on a different model"). Unset ⇒ the default utility runtimejudge_model = "gemini-3-pro-preview" # Optional: judge model id within that runtime. Unset ⇒ the default utility modeladmission = "queue" # What happens when ephemeral spawns hit the concurrency cap, "queue" or "fail". Default queue (see "Ephemeral spawn admission queueing" below)
[task_forward_model] # Task-level forward model (see the section of the same name). On by default since v1.54enabled = true
[goal_loop]iteration_cap = 5 # Hard dispatch cap for hard goals; escalates to a human past this. Default 5iteration_cap_simple = 3 # Dispatch cap for simple goals (dynamic judge depth). Default 3wall_clock_hours = 24 # Wall-clock budget from creation (hours); escalates to a human past this. Default 24max_concurrent = 3 # Cap on goal tasks in flight at once (prevents a spawn storm). Default 3tick_secs = 30 # Driver polling interval (seconds). Default 30stalled_secs = 600 # Seconds after dispatch with no claim before a task counts as stalled and can be redispatched. Default 600planner_enabled = false # When on, allows splitting a goal into a dependency DAG of subtasks (see "Parallel subtasks"). Default falseresume_on_restart = "pause" # What happens to in-flight goal tasks on gateway restart, "auto" or "pause" (see "Restart behavior"). Default pause, switchable under Settings → Automation on the dashboardprogress_report_minutes = 10 # How long a claimed task can go without a progress signal before one report fires; `0` disables it (see "Stall-timeout progress reports"). Default 10tool_streak_advisory = true # Whether to inject a reminder at 3/5/8 consecutive calls to the same tool with the same arguments (see "Tool streak advisory"). Default truecriteria_ledger = "report" # Per-criterion acceptance ledger: off / report / enforce (see "Acceptance ledger"). Default reportsteering_enabled = false # Directions for the next round from the task page (see "Directions and stopping on the task page"). Default false
[dispatch_guard] # Feedback-path circuit breaker (guards against self-reinforcing loops)window_secs = 60 # Sliding window length (seconds). Default 60max_in_window = 20 # Dispatches allowed within one window before tripping. Default 20cooldown_secs = 60 # Cooldown seconds after tripping, during which dispatch is refused. Default 60max_hop_depth = 5 # Cap on cross-process re-spawn depth along a delegation chain. Default 5All blocks are optional; anything missing or partially set falls back to the built-in defaults above. An unrecognized policy value always falls back to fixed_hierarchy, logging a warning.
Parallel subtasks (dependency DAG)
Section titled “Parallel subtasks (dependency DAG)”With [goal_loop] planner_enabled = true, creating a goal first lets the AI employee “try” splitting it into a set of subtasks annotated with dependencies (for example: query two data sources independently, then merge). The resulting subtasks each land on the Task Board, and any subtask whose depends_on list is fully satisfied runs in parallel, each checked for acceptance independently. Parallelism is still bounded by max_concurrent and the dispatch_guard circuit breaker — it never bypasses them.
- Not mandatory: when the model decides a split isn’t needed (or its reply can’t be parsed), it falls back to a single task, behaving exactly as if the setting were off.
- Cycle protection: if the resulting plan has a dependency cycle (or an out-of-range index), the whole plan is discarded, falls back to a single task, and a warning is logged. A broken DAG is never allowed to land.
- A stuck upstream never orphans its downstream: if a subtask’s upstream dependency lands in
failed/cancelled/needs_human(or the dependency doesn’t exist), the downstream inherits the escalation and also moves to needs-human, so you see the whole stuck branch at once. If the upstream is simply still running, the downstream is frozen for that round and reconsidered the next one.
The expected payoff is greatest for “multi-source lookup” style goals; an independent re-test measured roughly a 1.25x speedup (not the 3.7x a paper self-reported) — treat eval measurements as the source of truth before generalizing.
Dispatch policy (DispatchPolicy)
Section titled “Dispatch policy (DispatchPolicy)”[dispatch] policy decides which AI employee picks up a goal task. The default fixed_hierarchy behaves exactly as before (dispatches to whoever the task was already assigned to).
| Policy | Behavior |
|---|---|
fixed_hierarchy |
Default. Dispatches to the task’s existing assigned_to, unchanged. Zero LLM cost, fully deterministic. |
round_robin |
Rotates through the roster by “task category” (the first tag if any, otherwise priority). State lives in memory only, resetting on restart. |
llm_select |
An LLM picks the best-fitting AI employee from the roster via a tool call. Fails closed: if the output isn’t on the roster, or parsing/the LLM call fails, it always falls back to the fixed_hierarchy result, never dispatching to a made-up AI employee. No model name is hardcoded; it uses whatever utility runtime is configured. |
role_team |
Selection still yields the AI employee, exactly as fixed_hierarchy does. The roles live inside that employee — see “Team rounds” below. The setting exists so “this deployment composes teams” is visible in logs and telemetry; it does not change who a task is assigned to. |
Roster = the AI employee directories under <home>/agents/. When the roster is empty, both round_robin and llm_select fall back to the original assignment (never orphaning a task). A reassignment writes back to the task’s assigned_to, so heartbeat pulls and the activity log stay consistent.
Team rounds (Team-as-Agent)
Section titled “Team rounds (Team-as-Agent)”An AI employee can work a goal round as a small team inside itself: 規劃 → 執行 → 審核, each role on its own vendor’s model. Since v1.66 [team] enabled defaults to true — but what actually forms a team is naming a second vendor under [team.roles], not the flag. Write no roles and the executor and verifier both cascade onto the employee’s own model, share a model family, and the decorrelation rule refuses the spec: the task runs Solo, quietly, with no audit row. See 56-team-as-agent.md for the full picture; this section is what it means for the goal loop specifically.
The whole path — planner, executors, verifier and the
team_handofftool that carries packets between them — has landed and has been driven through a live round the judge accepted; that is one integration result, not a measured win over Solo. Configure[team.roles]only on a deployment you are watching.enabled = false(fleet-wide or per employee) andgate = "always_solo"are both one-line reverts.
What you see
Section titled “What you see”Nothing new in the conversation. The employee answers with one voice, progress arrives on the same board, and a task that needs you still carries one of the six pause classes. Under the hood the round runs as three stages instead of one wake-up message, and only the executor’s final product reaches the acceptance judge — which is the same two-stage evaluator and three-aspect panel as always. A team changes who does the work, not who decides it is done.
Two new reasons a team task can land in your lap:
| Pause | What happened |
|---|---|
需要決策 (blocked_needs_decision) |
The planning stage handed back no sub-tasks. Rather than guessing a breakdown from its prose, the loop asks you whether the goal is actually splittable — or whether it should just run as a single employee. |
預算用盡 (budget_exhausted) |
The task ran out of role-member spawns after spending some. It hands over the best round it managed, same as any other budget escalation. A task whose budget could never have paid for even one degraded round runs Solo instead — being parked over work that never started would be a worse answer than doing it the ordinary way. |
The gate
Section titled “The gate”Before each round, a zero-LLM rule set (crates/duduclaw-core/src/team_gate.rs) decides Solo or Team. It leans Solo on purpose.
- Hard Solo: the employee has
[container] sandbox_enabled = true(checked before every mode, includingalways_team, so a sandboxed employee never forms a team), a live channel turn, a plan-first goal still awaiting your approval, an irreversible action in the plan, or fewer than three rounds of budget left. - Count four signals: bulk (at least four independent work items and no dependency hubs), context overflow (the estimated input is larger than the model’s context window), capability gap (the role-matrix gap is at or above the matrix’s declared MDE) and long horizon (at least three acceptance criteria and the task produces artifacts). A signal it cannot measure never fires.
- Decide: three or more signals form the team; zero or one runs Solo; exactly two is the grey band.
In the grey band the loop runs the planning stage once (a call the task needed anyway) and applies the same rule again, this time with the bulk signal measured from the sub-task packets the planner wrote. Before a plan exists the bulk signal has nothing to measure, so the grey band always means two of the other three signals fired, and the second pass forms the team only when the plan holds four or more sub-tasks with no dependency hubs between them. Otherwise the round falls back to the ordinary single-employee round. A team spec with no planner role cannot take the second pass and runs Solo.
Every verdict is written to the audit log as team_gate_decision with the signals that fired, so a deployment’s Solo/Team split is measurable rather than anecdotal.
The spec is frozen per task
Section titled “The spec is frozen per task”The team a task runs with is decided once, at creation, and stored on the task. Later changes to [team] affect the next task. A spec that fails validation — most often a verifier sharing the executor’s model family — forms no team at all: the task runs Solo and there is no partial team.
Whether that refusal is loud depends on who asked for the team. An operator who wrote enabled = true, or who configured roles that then fail validation, gets the refusal audited as team_refused with the role and the reason. A deployment that never touched [team] at all gets a debug! line and nothing in the audit log — that refusal is the default state of every install since the v1.66 flip, and stamping a row on every goal task everywhere is how a real refusal stops being findable.
A role whose model is unset can also be filled from a measured capability matrix (role_model_matrix.toml in <DUDUCLAW_HOME>, written by duduclaw eval --matrix) before it falls through to the employee’s [model] preferred. Only resolved cells on the role’s own runtime count, ties select nothing, and an explicitly configured model always wins.
Budget and degrade chain
Section titled “Budget and degrade chain”[dispatch.team_budget]max_spawns_per_task = 12 # 4 roles x 3 rounds; clamped to a floor of 3max_turns_per_role = 3degrade_order = ["utility", "verifier_second_pass", "executor_replica"]One round is charged for planner? + executors + verifier + repair? — the verifier is a utility call rather than a scaffold, but it writes its own ledger row and the budget counts exactly the rows that carry a member id, so it costs a slot like every other stage. The floor of 3 is planner + executor + verifier.
As the spawn budget shrinks, capabilities are surrendered in that order — 合成 first, then the verifier’s one repair pass, then fan-out collapses to a single executor — before the task escalates as budget_exhausted. An unrecognised entry is dropped with a warning; a completely unusable list keeps the default chain.
Before its first team round the loop also checks, once, that [dispatch] ephemeral_max_active can hold the concurrent role members this configuration can need (max_concurrent × iteration_cap × roles — 45 against a default of 32) and warns with the number to raise. Nothing is enforced: past the ceiling, role spawns queue and can expire, which would otherwise surface much later as a round mysteriously missing its verifier.
Role-member spawns run on their own circuit-breaker bucket and budget ([dispatch_guard] role_team_max_in_window, default 60), so a composing team can never trip the breaker that guards an employee’s own sub-agent spawns.
Evidence on a team round: the employee and its members
Section titled “Evidence on a team round: the employee and its members”Everything downstream of a round — the team’s own verifier, the grounding precheck, and the MAV panel’s <tool_activity> digest — reads the tool-call audit trail for the task’s claim→review window. On a Solo round that trail belongs to one agent id. On a team round it does not: the work is done by ephemeral role members under their own ids, and the employee’s own window may hold nothing but the bookkeeping call that opened the round.
Read that way, a team round looks exactly like a task that claimed to do work and did nothing — which is what happened in live round 3, where the verifier and the settle evaluator both rejected honest work with “no tool activity supports file creation”. So the evidence set for a team round is the employee ∪ that round’s role members, taken from role_turns.jsonl, on both paths:
- the team verifier’s
<tool_activity>block, and - the settle path’s grounding precheck plus the judge digest.
Duplicate ids are collapsed, and a task with no role members adds no ids — a Solo task sees byte-identical evidence to before. The time window is unchanged, so members from other rounds contribute nothing even when the round-number lookup falls back to the whole task (the loop’s in-flight iteration counter and the settle path’s revision counter are separate and can diverge across a gateway restart).
Role members also work in the employee’s workspace rather than in their own throwaway scaffold, so files a member writes are still there when the verifier asks about them. Roles that write files are limited to claude and codex today; see 56-team-as-agent.md.
Union-ing the ids was necessary but not sufficient. Two further evidence sources landed after live round 8, both shared by the team verifier and the settle path so neither ever judges on a different account of the round:
- Native tool work is persisted. A member that works through native tools (a codex
shell, a ClaudeWrite) makes no MCP call, so before this its work left nothing in the audit trail at all — round 8 rejected three files that demonstrably existed. Each native tool event is now written as atool_calls.jsonlrow under the member’s id, carrying the masked call input and result text plussource = "native"and the runtime/model that produced it. Self-echo tools keep their output suppressed, exactly as the MCP writer does, so a role can never ground a claim on its own echoed packet. <artifact_receipts>— everyartifacts[].patha packet declares is stat’d and hashed against the employee’s workspace, and the result (<path> <bytes>B sha256=<hex> exists|missing|mismatch) becomes both a prompt block and anartifact_receiptaudit row. A declared-but-absent hash is filled in from the bytes; a declared hash that disagrees is kept as declared and recorded asmismatchwith ateam_packet_artifact_mismatchevent, so a swap stays visible rather than being quietly corrected. Onlyexistscounts as confirmation.
A task that declares no artefacts adds no block, and a Solo round with no native-tool members sees exactly what it saw before.
Audit and per-role records
Section titled “Audit and per-role records”| Event | Meaning |
|---|---|
team_gate_decision |
Solo / Team / grey band, with the signals that fired |
team_refused |
An enabled spec failed validation; the task runs Solo |
team_round_started |
A three-stage round began, with the roles and any degrade steps |
team_member_spawned |
One role member was created (role, runtime, model) |
team_stage_failed |
A stage did not produce what the next one needs, or could not run at all — carries role, runtime, model and the error |
team_handoff |
A role filed a packet (leg, round, size, audience, file) |
team_packet_skipped |
A packet file was unreadable, mislabelled or invalid and was ignored |
team_packet_artifact_refused |
A packet declared an artifact path resolving outside the employee’s workspace |
team_packet_fidelity_corrected |
A packet’s self-declared evidence grade disagreed with what was observed, and was overwritten |
Where the packets live
Section titled “Where the packets live”Each stage hands over by calling team_handoff, which derives the file path rather than taking one:
~/.duduclaw/team_packets/<task_id>/r<round>/planner-to-executor.json~/.duduclaw/team_packets/<task_id>/r<round>/planner-to-executor.01.json … up to .99A planner that breaks a goal into four sub-tasks writes four packets on the same leg, so a new packet takes the next numbered slot while re-filing the same packet_id overwrites its own file (a retry after a timeout does not duplicate the sub-task). The next stage reads those slots back in numeric order — canonical file first, then .01, .02, … A packet whose own from_role / to_role / task / round disagree with the file it sits in, or that fails validation, is skipped and audited as team_packet_skipped; one bad file does not cost the stage the rest of its packets.
Each stage also appends a row to <home>/role_turns.jsonl — role, runtime, model, effort, the packet it produced, the evidence grade of its observations, how it ended. Same permissions, locking and rotation as tool_calls.jsonl.
Ephemeral spawn admission queueing
Section titled “Ephemeral spawn admission queueing”Goal decomposition, delegation, and similar paths sometimes need to spin up a short-lived sub-agent (an ephemeral spawn). When that hits the concurrency cap (ephemeral_max_active, 32 by default), config.toml [dispatch] admission decides how to handle the request that pushed it over the limit:
| Value | Behavior |
|---|---|
queue |
Default. Bounded FIFO queueing: a request never simply vanishes, it runs once a slot frees up, in order. Each queued ticket carries a TTL (queue_item_ttl_secs, 600 seconds by default); past that it’s dropped and an audit entry is logged. The queue itself has a depth cap (queue_max_depth, 64 by default), and a full queue rejects outright (so an unbounded queue can’t itself become a new runaway risk). When the turn/session that made the request ends, its queued tickets are voided along with it, so a process that has already finished can’t suddenly spawn a late sub-agent. |
fail |
The old behavior: refuse outright past the cap, no queueing. |
ephemeral_max_active itself follows “adjustable, but never zero”: setting it to 0 clamps it to 1 and logs a warning. The concurrency cap can never be turned fully off.
[dispatch]admission = "queue" # "queue" (default) or "fail"queue_max_depth = 64 # queue depth cap, rejects past thisqueue_item_ttl_secs = 600 # seconds a queued ticket survives before being dropped and auditedephemeral_max_active = 32 # concurrency cap, 0 gets clamped to 1Structured outcome verification (outcome schema, WP2.4)
Section titled “Structured outcome verification (outcome schema, WP2.4)”/goal … || outcome:<spec> lets you add a machine-checkable deliverable contract on top of free-text acceptance criteria. When the AI employee reports it’s done and the task enters review, this contract runs a deterministic, zero-LLM-cost check before the acceptance judge:
- Check fails → the task goes straight back to
revisingwith feedback naming the specific gap (which field is missing, which file is missing), without calling the judge at all. This is a defense against judge false positives: an output with an obvious structural defect never gets waved through by an over-lenient judge, and no judge call is wasted on it. - Check passes → only then does it reach the judge, and the judge’s prompt gets a note that “structured outcome verification already passed its deterministic check,” so the judge can focus on quality.
Three spec types:
| Spec | Meaning |
|---|---|
outcome:text |
Default. No structured contract; behaves exactly as if no outcome were attached (not persisted, no check runs before the judge). |
outcome:json:<JSON Schema> |
A subset of JSON Schema (object / array / string / number / integer / boolean, supporting properties / required / items). Checks the ```json block in the AI employee’s final reply (falling back to parsing the whole reply if no fenced block is found). Missing fields or type mismatches are all listed as specific defects. |
outcome:files:<glob,glob> |
Asserts that a deliverable file matching each glob (*/? supported) exists under the AI employee’s working directory. Example: outcome:files:report.docx, out/*.pdf. |
Example