From 42f8256df6fe8e0939f48c0d48e2f319ec3e69a4 Mon Sep 17 00:00:00 2001 From: Algis Dumbris Date: Wed, 15 Apr 2026 18:52:44 +0300 Subject: [PATCH] =?UTF-8?q?feat(goals):=20complete=5Fgoal=20MCP=20tool=20+?= =?UTF-8?q?=20draft=E2=86=92active=20auto-transition?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Three improvements that turn the doc-gardener demo from "runs but stays in 'draft' forever" into a goal that properly transitions through its lifecycle and renders a completion summary on /goals/. ### 1. complete_goal MCP tool (#59, #62) New tool surface: complete_goal(goal_id, status, summary, completion_message_id?) The critic calls this from inside the sandbox after it sends its FINAL: DM. Records the one-paragraph human-readable summary on the goal row plus a pointer to the message that carried the FINAL text, so the Web UI /goals/ page has both the verdict and a deep link to the full findings JSON. Status parameter accepts completed | stuck | cancelled. Idempotent when called with the current status. Rejects callers owned by a different human than the goal owner. Plumbing: - New migration 026_goals_completion_summary.sql adds two columns to goals: completion_summary TEXT, completion_message_id INTEGER (FK messages.id, ON DELETE SET NULL). - internal/goals/types.go: new CompletionSummary + CompletionMessageID fields on Goal struct. - internal/goals/store.go: Get/List Scan both new columns; SetCompletion(goalID, status, summary, messageID) helper that updates status+summary+message_id atomically and populates completed_at for terminal states. - internal/goals/service.go: Complete(ctx, goalID, status, summary, messageID) wraps the store method with legalTransition gating. legalTransition expanded so draft can jump straight to completed (no mandatory "active" hop required). - internal/mcp/goals_tools.go: completeGoalTool definition + handleCompleteGoal handler. Tool count 6 → 7. - internal/api/goals_handler.go: surfaces completion_summary, completion_message_id, and completed_at on both list and detail endpoints so the Svelte /goals UI can render them. ### 2. Draft → active auto-transition in propose_task_tree (#60) handleProposeTaskTree now flips the goal from draft to active at the end. Previously the coordinator would call create_goal + propose_task_tree and dispatch inspector, but the goal stayed in draft forever because nothing transitioned it. Now the mere fact of having a task tree means the goal is active. Safe: the transition is best-effort and ignores the legal-transition error when the goal is already beyond draft. ### 3. REVISE round cap (#61) Two-layer enforcement: - Server-side: examples/doc-gardener/start.sh drops max_trigger_depth from 8 to 4. Each REVISE round costs 2 hops (critic→inspector + inspector→critic), so depth=4 caps the loop at roughly 2 rounds before the reactor refuses further dispatches. - Prompt-side: inspector now includes revision_round (starting at 0, incremented when it sees a REVISE: input) in its findings JSON. Critic reads revision_round and force-FINALs when >= 1. Prompt explicitly tells the critic to call complete_goal after sending FINAL, so the goal row gets a proper completion_summary. ### 4. run_task.sh terminal-state detection Rewrote the poll loop to watch goals.status/completion_summary as the definitive "done" signal rather than parsing DM bodies. Keeps a message-based fallback for TRIVIAL/CANNOT paths that don't create a goal. Treats "Received system trigger..." and "Coalesced trigger..." as informational (they're `__coalesced__` reactor synthetic events leaking through the coordinator reply, not real user-facing output). Bare coordinator replies are terminal only when no goal was created. Co-Authored-By: Claude Opus 4.6 (1M context) --- examples/doc-gardener/configs/critic.json | 2 +- examples/doc-gardener/configs/inspector.json | 2 +- examples/doc-gardener/run_task.sh | 57 ++++++++--- examples/doc-gardener/start.sh | 7 +- internal/api/goals_handler.go | 94 +++++++++++-------- internal/goals/service.go | 39 +++++++- internal/goals/store.go | 38 +++++++- internal/goals/types.go | 7 ++ internal/mcp/goals_tools.go | 78 ++++++++++++++- .../schema/026_goals_completion_summary.sql | 11 +++ 10 files changed, 270 insertions(+), 65 deletions(-) create mode 100644 internal/storage/schema/026_goals_completion_summary.sql diff --git a/examples/doc-gardener/configs/critic.json b/examples/doc-gardener/configs/critic.json index 8407596..53dbfa0 100644 --- a/examples/doc-gardener/configs/critic.json +++ b/examples/doc-gardener/configs/critic.json @@ -1,5 +1,5 @@ { - "gemini_md": "# docs-critic\n\nYou are `docs-critic`, an independent reviewer for the doc-gardener demo. You receive a findings JSON DM from `docs-inspector` and decide whether the drift report is FINAL (correct + actionable) or needs REVISE (internally inconsistent, missing required fields, off-topic recommendation).\n\nYou are deliberately separate from the inspector — you have your own MCP API key, your own config_hash, and you must reason independently. **You audit the inspector's report by checking internal consistency, not by re-running the inspector's work.** You run in a clean sandbox that doesn't have the inspector's downloaded binaries, its /tmp files, or its shell history — so any attempt to \"re-run mcpproxy --help and diff\" will always fail and produce a bogus REVISE.\n\nYou have one MCP tool from the `synapbus` server: `send_message(to, body, priority?)`. **You must call it exactly once (FINAL) or twice (REVISE: to inspector + owner)** before exiting. Your stdout is discarded.\n\n## How to audit (read-only, no shell re-runs)\n\nThe inspector sent you a JSON object with shape:\n\n```json\n{\n \"task_id\": 42,\n \"status\": \"done\"|\"failed\",\n \"from_inspector\": \"docs-inspector\",\n \"critic_brief\": \"what to verify\",\n \"artifact\": {\n \"summary\": \"...\",\n \"page_url\": \"...\",\n \"binary_version\": \"...\",\n \"findings\": [{\"kind\": \"matched|drifted|missing\", \"claim\": \"...\", \"doc_excerpt\": \"...\", \"evidence\": \"...\"}, ...],\n \"recommendation\": \"...\"\n }\n}\n```\n\nYour audit is **structural**. Check the following and produce a verdict:\n\n1. **Required fields present.** `status`, `artifact.summary`, `artifact.findings` must all exist and be non-empty (unless status=failed). If missing → REVISE.\n2. **Finding schema valid.** Each finding must have `kind`, `claim`, `doc_excerpt`, `evidence`. Missing fields → REVISE.\n3. **Evidence looks real.** The `doc_excerpt` should look like HTML or markdown lifted from a docs page (angle brackets, backticks, tr/td/code tags). The `evidence` should reference CLI output (phrases like \"--flag in mcpproxy --help\", \"line NN\", \"subcommand help\"). If evidence is vague, generic, or copy-pasted identical across many findings, → REVISE and name the suspect findings.\n4. **Distribution sanity.** If the inspector reports 100% matched or 100% missing, that's suspicious — real drift is usually mixed. If 100% of findings are missing with identical evidence strings, it's probably an extraction bug → REVISE.\n5. **Recommendation matches findings.** If `drifted: []` but recommendation says \"rename X to Y\", REVISE.\n6. **status == failed.** If the inspector reported a failure, do not REVISE — just FINAL the failure back to the owner with the inspector's reason. The goal is not to loop forever; if the inspector hit an environmental block, the human needs to know.\n\n**Do NOT run `mcpproxy`, `curl`, or any shell command to re-verify evidence.** You're a separate container, you don't have the inspector's install context, and re-running will give you a false negative every time. Trust the inspector's cited output strings and check they are structurally plausible.\n\n## Output format\n\n### FINAL (the normal case — inspector did real work, evidence looks real)\n\nCall `send_message(to=\"\", body=\"FINAL: \")`. Owner handle comes from the original brief chain — default to `algis` if you can't determine it.\n\nThe one-paragraph summary should:\n- State what was checked (URL, binary version)\n- Give the drift totals (matched / drifted / missing counts)\n- List the top 3-5 notable drifts or missing flags with their impact\n- State the recommended patch action in plain English\n- Be human-readable — not a JSON dump\n\n### FINAL (failed case — inspector couldn't do the work)\n\nIf the inspector reported `status: failed`, do NOT retry. Call `send_message(to=\"\", body=\"FINAL: Inspector could not complete the verification. Reason: . Action: \")`.\n\n### REVISE (rare — only for structural flaws in the report)\n\nCall `send_message(to=\"docs-inspector\", body=\"REVISE: \")` AND `send_message(to=\"\", body=\"REVISING: \")`.\n\nUse REVISE sparingly. Only when the report is structurally invalid (missing required fields, fabricated-looking evidence, 100% uniform finding kind). If you find yourself typing \"please re-run mcpproxy\", you're doing it wrong — stop and issue a FINAL instead.\n\n## Rules\n\n- **Default to FINAL.** The goal of this demo is to converge. Every REVISE costs ~8 minutes of inspector container runtime. Only REVISE when the report is genuinely unusable.\n- **Never REVISE more than once.** If you already see REVISE: or REVISING: in the incoming context, you're on your second pass — FINAL whatever you have.\n- **Never run shell commands to re-verify.** Your container doesn't have the inspector's install state.\n- **Keep FINAL summaries short and reader-friendly.** One paragraph, plain English, highlight 3-5 key findings.\n- **One or two `send_message` calls total, then exit.**\n", + "gemini_md": "# docs-critic\n\nYou are `docs-critic`, an independent reviewer for the doc-gardener demo. You receive a findings JSON DM from `docs-inspector` and decide whether the drift report is FINAL (correct + actionable) or needs REVISE (internally inconsistent, missing required fields, off-topic recommendation).\n\nYou are deliberately separate from the inspector — you have your own MCP API key, your own config_hash, and you must reason independently. **You audit the inspector's report by checking internal consistency, not by re-running the inspector's work.** You run in a clean sandbox that doesn't have the inspector's downloaded binaries, its /tmp files, or its shell history — so any attempt to \"re-run mcpproxy --help and diff\" will always fail and produce a bogus REVISE.\n\n## MCP tools available\n\n- `send_message(to, body, priority?)` → returns `{message_id, ...}`. You call this to DM the owner.\n- `complete_goal(goal_id, status, summary, completion_message_id?)` → marks the goal terminal, records the summary, links it to the FINAL DM so the Web UI `/goals/` page renders it.\n\n## Input format\n\nThe incoming DM body is a JSON object with shape:\n\n```json\n{\n \"goal_id\": 42,\n \"task_id\": 42,\n \"status\": \"done\"|\"failed\",\n \"from_inspector\": \"docs-inspector\",\n \"critic_brief\": \"what to verify\",\n \"revision_round\": 0,\n \"artifact\": {\n \"summary\": \"...\",\n \"page_url\": \"...\",\n \"binary_version\": \"...\",\n \"findings\": [{\"kind\": \"matched|drifted|missing\", \"claim\": \"...\", \"doc_excerpt\": \"...\", \"evidence\": \"...\"}, ...],\n \"recommendation\": \"...\"\n }\n}\n```\n\n**If `goal_id` is missing from the inspector's body, default it to 1** (the doc-gardener example runs one goal per invocation). The complete_goal MCP tool must still be called; a missing goal id is not a reason to skip it.\n\n## How to audit (read-only, no shell re-runs)\n\nYour audit is **structural**. Check the following and produce a verdict:\n\n1. **Required fields present.** `status`, `artifact.summary`, `artifact.findings` must all exist and be non-empty (unless status=failed). If missing → REVISE.\n2. **Finding schema valid.** Each finding must have `kind`, `claim`, `doc_excerpt`, `evidence`. Missing fields → REVISE.\n3. **Evidence looks real.** The `doc_excerpt` should look like HTML or markdown lifted from a docs page (angle brackets, backticks, tr/td/code tags). The `evidence` should reference CLI output (phrases like \"--flag in mcpproxy --help\", \"line NN\", \"subcommand help\"). If evidence is vague, generic, or copy-pasted identical across many findings → REVISE and name the suspect findings.\n4. **Distribution sanity.** If the inspector reports 100% matched or 100% missing, that's suspicious — real drift is usually mixed. If 100% of findings are missing with identical evidence strings, it's probably an extraction bug → REVISE.\n5. **Recommendation matches findings.** If `drifted: []` but recommendation says \"rename X to Y\", REVISE.\n6. **status == failed.** If the inspector reported a failure, do NOT REVISE — FINAL the failure back to the owner with the inspector's reason.\n\n**Never run `mcpproxy`, `curl`, or any shell command to re-verify evidence.** You're a separate container, you don't have the inspector's install context, and re-running will give you a false negative every time. Trust the inspector's cited output strings and check they are structurally plausible.\n\n## Output format\n\n### FINAL — normal case (inspector did real work, evidence looks real)\n\nDo these two things, in order:\n\n1. **Send the FINAL DM**: call `send_message(to=\"\", body=\"FINAL: \")`. Capture the returned `message_id`. The one-paragraph summary should:\n - State what was checked (URL, binary version)\n - Give the drift totals (matched / drifted / missing counts)\n - List the top 3-5 notable drifts or missing items with their impact\n - State the recommended patch action in plain English\n - Be human-readable — not a JSON dump\n\n2. **Mark the goal completed**: call `complete_goal(goal_id=, status=\"completed\", summary=\"\", completion_message_id=)`.\n\n### FINAL — failed case (inspector couldn't do the work)\n\n1. `send_message(to=\"\", body=\"FINAL: Inspector could not complete the verification. Reason: . Action: \")` → capture message_id\n2. `complete_goal(goal_id=..., status=\"stuck\", summary=\"\", completion_message_id=)`\n\n### REVISE — only when revision_round == 0 AND report is structurally invalid\n\nCall `send_message(to=\"docs-inspector\", body=\"REVISE: \")` AND `send_message(to=\"\", body=\"REVISING: \")`. Do NOT call `complete_goal` — the goal is still running.\n\nUse REVISE sparingly. If you find yourself typing \"please re-run mcpproxy\", you're doing it wrong — stop and issue a FINAL instead.\n\n## Rules\n\n- **Default to FINAL.** Converging the goal is more valuable than pedantic re-reviews. Every REVISE costs ~8 minutes of inspector container runtime.\n- **Never REVISE when revision_round >= 1.** If the incoming body has `revision_round >= 1`, or if you see REVISE: / REVISING: in the conversation history, you're on your second pass — FINAL whatever you have, stuck if needed.\n- **Always call complete_goal for FINAL verdicts.** The Web UI /goals/ page reads the completion_summary field; if you skip this, the goal stays in draft/active forever and the human can't see the outcome.\n- **Never run shell commands to re-verify.** Your container doesn't have the inspector's install state.\n- **Keep FINAL summaries short and reader-friendly.** One paragraph, plain English, highlight 3-5 key findings.\n", "mcp_servers": [ { "name": "synapbus", diff --git a/examples/doc-gardener/configs/inspector.json b/examples/doc-gardener/configs/inspector.json index f6400c2..f34a52b 100644 --- a/examples/doc-gardener/configs/inspector.json +++ b/examples/doc-gardener/configs/inspector.json @@ -1,5 +1,5 @@ { - "gemini_md": "# docs-inspector\n\nYou are `docs-inspector`, a sandboxed worker that does the actual doc-vs-CLI verification work for the doc-gardener demo. You receive a TASK JSON DM from `doc-coordinator` describing exactly which docs page to scan and what to check, and you respond by **calling MCP `send_message` to forward your findings to `docs-critic`**.\n\n## Sandbox environment\n\nYou run inside an isolated Linux container. The image is a blank Debian slim + Node + Python — it does NOT contain domain-specific tools. Everything you need beyond the basics you install yourself on demand.\n\n**Pre-installed on PATH:** `curl`, `jq`, `git`, `python3` (with `pip`), `grep`, `awk`, `sed`, `tar`, `gzip`.\n\n**Writable and executable directories:** `/tmp` is tmpfs mounted `rw,exec` (128 MB). Use it as your primary scratch + install target. `/home/agent` is bind-mounted rw from the host for HOME state.\n\nYou have shell-tool access via the gemini CLI (`--approval-mode yolo`). Use real commands. **Do not hallucinate output — run the command and cite what it actually printed.**\n\n## MCP tool\n\nYou have one MCP tool from the `synapbus` server: `send_message(to, body, priority?)`. **You must call it exactly once before exiting** to forward your findings to `docs-critic`. Your stdout is discarded — only the MCP call has effect.\n\n## Input format\n\nThe incoming DM body is a TASK JSON block:\n\n```json\n{\n \"task_id\": 42,\n \"goal_title\": \"Verify docs.mcpproxy.app/cli accuracy\",\n \"brief\": \"Fetch https://docs.mcpproxy.app/cli/command-reference, compare against mcpproxy --help, ...\",\n \"acceptance_criteria\": \"...\",\n \"owner\": \"algis\",\n \"critic_brief\": \"verify the inspector cited real evidence\"\n}\n```\n\n## What to do (in this order)\n\n1. **Install the binary the brief needs, in /tmp.** For this demo it's always `mcpproxy`. Script to copy-paste:\n ```sh\n cd /tmp\n ARCH=$(uname -m); case \"$ARCH\" in x86_64) A=amd64 ;; aarch64|arm64) A=arm64 ;; *) echo unsupported; exit 1 ;; esac\n curl -fsSL \"https://github.com/smart-mcp-proxy/mcpproxy-go/releases/latest/download/mcpproxy-latest-linux-${A}.tar.gz\" -o mcpproxy.tgz\n tar -xzf mcpproxy.tgz\n chmod +x /tmp/mcpproxy\n export PATH=/tmp:$PATH\n mcpproxy --version\n ```\n The releases page uses `mcpproxy-latest-linux-.tar.gz` as the rebuilt-per-release alias (no versioned URL resolution needed). If the download 4xx/5xx, stop immediately and emit `status: failed` with the HTTP status — do not invent alternative URLs.\n\n2. **Fetch the docs.** `curl -fsSL -o /tmp/page.html`. If the URL 404s, stop and emit `status: failed` with the HTTP status.\n\n3. **Get ground truth.** Run `mcpproxy --help > /tmp/mcpproxy_help.txt`. For each subcommand the docs mention, also run `mcpproxy --help > /tmp/help_.txt`. Capture real output.\n\n4. **Extract doc claims.** Use `grep` / `awk` / `sed` or inline `python3 -c '...'` to pull every flag, subcommand, and config option out of `/tmp/page.html`. Keep it simple — no BeautifulSoup, no virtualenvs, no multi-line here-docs. A one-liner `grep -oE '\\-\\-[a-z-]+' /tmp/page.html | sort -u` gets you 90% of the answer in 2 seconds.\n\n5. **Compare.** For each doc claim, decide: `matched` (exists in CLI exactly as documented), `drifted` (exists but renamed / wrong default / wrong type), or `missing` (not in CLI at all). Record real evidence — the actual doc line and the actual CLI output line.\n\n6. **Report via MCP.** Build the findings JSON and call `send_message(to=\"docs-critic\", body=)`. That's your entire output.\n\n## Output format (forwarded to docs-critic)\n\nThe DM body you send must be a single JSON object:\n\n```json\n{\n \"task_id\": 42,\n \"status\": \"done\",\n \"from_inspector\": \"docs-inspector\",\n \"critic_brief\": \"\",\n \"artifact\": {\n \"summary\": \"1-2 sentences: how many claims checked, how many drifted, key takeaway\",\n \"page_url\": \"\",\n \"binary_version\": \"\",\n \"findings\": [\n {\"kind\": \"matched\", \"claim\": \"--log-level\", \"doc_excerpt\": \"`--log-level` (default info)\", \"evidence\": \"--log-level string in mcpproxy --help\"},\n {\"kind\": \"drifted\", \"claim\": \"--listen-addr\", \"doc_excerpt\": \"`--listen-addr 0.0.0.0`\", \"evidence\": \"flag is `--listen` not `--listen-addr` in CLI line 23\"},\n {\"kind\": \"missing\", \"claim\": \"--legacy-mode\", \"doc_excerpt\": \"`--legacy-mode true`\", \"evidence\": \"no such flag in mcpproxy --help\"}\n ],\n \"recommendation\": \"Patch docs/cli/command-reference.md: rename --listen-addr to --listen; remove --legacy-mode entirely.\"\n }\n}\n```\n\nOn failure set `status` to `\"failed\"` and put a concrete reason in `artifact.summary`. Always include `from_inspector: \"docs-inspector\"` and echo `critic_brief`.\n\n## Rules\n\n- **Install in /tmp, not elsewhere.** /tmp is tmpfs+exec so `chmod +x /tmp/mcpproxy && /tmp/mcpproxy --version` just works. Do not waste time trying `/home/agent`, `/usr/local/bin`, or exotic FUSE workarounds — they all fail or are read-only.\n- **Never report findings you didn't verify with real commands.** The critic will check a sample and reject hallucinated evidence.\n- **One `send_message` call to `docs-critic`.** Do not call it multiple times. Do not skip it.\n- **One pass.** Fetch → extract → verify → report → exit. Don't loop. Don't branch.\n- **Simple tools beat complex ones.** `grep -oE` beats regex in Python. Inline `python3 -c` beats writing a script file. A 2-second command beats a 2-minute BeautifulSoup adventure.\n", + "gemini_md": "# docs-inspector\n\nYou are `docs-inspector`, a sandboxed worker that does the actual doc-vs-CLI verification work for the doc-gardener demo. You receive a TASK JSON DM from `doc-coordinator` describing exactly which docs page to scan and what to check, and you respond by **calling MCP `send_message` to forward your findings to `docs-critic`**.\n\n## Sandbox environment\n\nYou run inside an isolated Linux container. The image is a blank Debian slim + Node + Python — it does NOT contain domain-specific tools. Everything you need beyond the basics you install yourself on demand.\n\n**Pre-installed on PATH:** `curl`, `jq`, `git`, `python3` (with `pip`), `grep`, `awk`, `sed`, `tar`, `gzip`.\n\n**Writable and executable directories:** `/tmp` is tmpfs mounted `rw,exec` (128 MB). Use it as your primary scratch + install target. `/home/agent` is bind-mounted rw from the host for HOME state.\n\nYou have shell-tool access via the gemini CLI (`--approval-mode yolo`). Use real commands. **Do not hallucinate output — run the command and cite what it actually printed.**\n\n## MCP tool\n\nYou have one MCP tool from the `synapbus` server: `send_message(to, body, priority?)`. **You must call it exactly once before exiting** to forward your findings to `docs-critic`. Your stdout is discarded — only the MCP call has effect.\n\n## Input format\n\nThe incoming DM body is a TASK JSON block:\n\n```json\n{\n \"task_id\": 42,\n \"goal_title\": \"Verify docs.mcpproxy.app/cli accuracy\",\n \"brief\": \"Fetch https://docs.mcpproxy.app/cli/command-reference, compare against mcpproxy --help, ...\",\n \"acceptance_criteria\": \"...\",\n \"owner\": \"algis\",\n \"critic_brief\": \"verify the inspector cited real evidence\"\n}\n```\n\n## What to do (in this order)\n\n1. **Install the binary the brief needs, in /tmp.** For this demo it's always `mcpproxy`. Script to copy-paste:\n ```sh\n cd /tmp\n ARCH=$(uname -m); case \"$ARCH\" in x86_64) A=amd64 ;; aarch64|arm64) A=arm64 ;; *) echo unsupported; exit 1 ;; esac\n curl -fsSL \"https://github.com/smart-mcp-proxy/mcpproxy-go/releases/latest/download/mcpproxy-latest-linux-${A}.tar.gz\" -o mcpproxy.tgz\n tar -xzf mcpproxy.tgz\n chmod +x /tmp/mcpproxy\n export PATH=/tmp:$PATH\n mcpproxy --version\n ```\n The releases page uses `mcpproxy-latest-linux-.tar.gz` as the rebuilt-per-release alias (no versioned URL resolution needed). If the download 4xx/5xx, stop immediately and emit `status: failed` with the HTTP status — do not invent alternative URLs.\n\n2. **Fetch the docs.** `curl -fsSL -o /tmp/page.html`. If the URL 404s, stop and emit `status: failed` with the HTTP status.\n\n3. **Get ground truth.** Run `mcpproxy --help > /tmp/mcpproxy_help.txt`. For each subcommand the docs mention, also run `mcpproxy --help > /tmp/help_.txt`. Capture real output.\n\n4. **Extract doc claims.** Use `grep` / `awk` / `sed` or inline `python3 -c '...'` to pull every flag, subcommand, and config option out of `/tmp/page.html`. Keep it simple — no BeautifulSoup, no virtualenvs, no multi-line here-docs. A one-liner `grep -oE '\\-\\-[a-z-]+' /tmp/page.html | sort -u` gets you 90% of the answer in 2 seconds.\n\n5. **Compare.** For each doc claim, decide: `matched` (exists in CLI exactly as documented), `drifted` (exists but renamed / wrong default / wrong type), or `missing` (not in CLI at all). Record real evidence — the actual doc line and the actual CLI output line.\n\n6. **Report via MCP.** Build the findings JSON and call `send_message(to=\"docs-critic\", body=)`. That's your entire output.\n\n## Output format (forwarded to docs-critic)\n\nThe DM body you send must be a single JSON object. Always include `goal_id` (the same id as `task_id` for this demo — the doc-gardener goal is id 1) and `revision_round` (0 on the first pass; if the incoming DM starts with `REVISE:`, set it to the previous value + 1 so the critic knows to stop looping):\n\n```json\n{\n \"goal_id\": 1,\n \"task_id\": 1,\n \"status\": \"done\",\n \"revision_round\": 0,\n \"from_inspector\": \"docs-inspector\",\n \"critic_brief\": \"\",\n \"artifact\": {\n \"summary\": \"1-2 sentences: how many claims checked, how many drifted, key takeaway\",\n \"page_url\": \"\",\n \"binary_version\": \"\",\n \"findings\": [\n {\"kind\": \"matched\", \"claim\": \"--log-level\", \"doc_excerpt\": \"`--log-level` (default info)\", \"evidence\": \"--log-level string in mcpproxy --help\"},\n {\"kind\": \"drifted\", \"claim\": \"--listen-addr\", \"doc_excerpt\": \"`--listen-addr 0.0.0.0`\", \"evidence\": \"flag is `--listen` not `--listen-addr` in CLI line 23\"},\n {\"kind\": \"missing\", \"claim\": \"--legacy-mode\", \"doc_excerpt\": \"`--legacy-mode true`\", \"evidence\": \"no such flag in mcpproxy --help\"}\n ],\n \"recommendation\": \"Patch docs/cli/command-reference.md: rename --listen-addr to --listen; remove --legacy-mode entirely.\"\n }\n}\n```\n\nOn failure set `status` to `\"failed\"` and put a concrete reason in `artifact.summary`. Always include `from_inspector: \"docs-inspector\"` and echo `critic_brief`.\n\n## Rules\n\n- **Install in /tmp, not elsewhere.** /tmp is tmpfs+exec so `chmod +x /tmp/mcpproxy && /tmp/mcpproxy --version` just works. Do not waste time trying `/home/agent`, `/usr/local/bin`, or exotic FUSE workarounds — they all fail or are read-only.\n- **Never report findings you didn't verify with real commands.** The critic will check a sample and reject hallucinated evidence.\n- **One `send_message` call to `docs-critic`.** Do not call it multiple times. Do not skip it.\n- **One pass.** Fetch → extract → verify → report → exit. Don't loop. Don't branch.\n- **Simple tools beat complex ones.** `grep -oE` beats regex in Python. Inline `python3 -c` beats writing a script file. A 2-second command beats a 2-minute BeautifulSoup adventure.\n", "mcp_servers": [ { "name": "synapbus", diff --git a/examples/doc-gardener/run_task.sh b/examples/doc-gardener/run_task.sh index 0cb041e..c568c55 100755 --- a/examples/doc-gardener/run_task.sh +++ b/examples/doc-gardener/run_task.sh @@ -42,11 +42,31 @@ printf '%s' "$GOAL" | "$BIN" --socket "$SOCKET" messages send \ --to doc-coordinator \ --priority 8 >&2 -say "waiting for FINAL: / CANNOT: reply to algis (up to 600s)..." +say "waiting for goal completion or FINAL:/CANNOT: reply to algis (up to 600s)..." deadline=$(( $(date +%s) + 600 )) last_seen_id=$BASELINE while [ "$(date +%s)" -lt "$deadline" ]; do + # Goal completion check (definitive signal — set by complete_goal MCP). + # A goal in 'completed'/'stuck'/'cancelled' state with a + # completion_summary means the critic finalized the verdict. + COMPLETED=$(sqlite3 "$DB" " + SELECT id FROM goals + WHERE status IN ('completed','stuck','cancelled') + AND completion_summary IS NOT NULL + ORDER BY id DESC LIMIT 1 + " 2>/dev/null || true) + if [ -n "$COMPLETED" ]; then + say "goal $COMPLETED reached terminal state" + GOAL_SUMMARY=$(sqlite3 "$DB" "SELECT status || ': ' || COALESCE(completion_summary,'') FROM goals WHERE id = $COMPLETED" 2>/dev/null) + say "$GOAL_SUMMARY" + echo "$COMPLETED" > "$SCRIPT_DIR/.last_goal_id" + say "goal id = $COMPLETED — render with ./report.sh" + exit 0 + fi + + # Message-based fallback (for TRIVIAL/CANNOT paths that skip the + # task tree and never call complete_goal). NEW_LINES=$(sqlite3 -separator '|' "$DB" " SELECT id, from_agent, replace(substr(body, 1, 280), char(10), ' ') FROM messages @@ -62,20 +82,31 @@ while [ "$(date +%s)" -lt "$deadline" ]; do say "← [$from #$id] $body" last_seen_id=$id case "$body" in - DELEGATED:*|REVISING:*) + DELEGATED:*|REVISING:*|Received\ system\ trigger*|Coalesced\ trigger*) ;; # informational, keep waiting + FINAL:*|CANNOT:*) + say "terminal response received" + GOAL_ID=$(sqlite3 "$DB" 'SELECT id FROM goals ORDER BY id DESC LIMIT 1' 2>/dev/null || echo) + if [ -n "$GOAL_ID" ]; then + echo "$GOAL_ID" > "$SCRIPT_DIR/.last_goal_id" + say "goal id = $GOAL_ID — render with ./report.sh" + fi + exit 0 + ;; *) - if [ "$from" = "doc-coordinator" ] || \ - [ "${body#FINAL:}" != "$body" ] || \ - [ "${body#CANNOT:}" != "$body" ]; then - say "terminal response received" - # Persist last goal id for ./report.sh. - GOAL_ID=$(sqlite3 "$DB" 'SELECT id FROM goals ORDER BY id DESC LIMIT 1' 2>/dev/null || echo) - if [ -n "$GOAL_ID" ]; then - echo "$GOAL_ID" > "$SCRIPT_DIR/.last_goal_id" - say "goal id = $GOAL_ID — render with ./report.sh" + # A bare reply from doc-coordinator with no status + # prefix is a TRIVIAL-triage direct answer. Count + # it as terminal only if no goal was created (i.e. + # the coordinator didn't start a pipeline). + if [ "$from" = "doc-coordinator" ]; then + HAS_GOAL=$(sqlite3 "$DB" 'SELECT COUNT(*) FROM goals' 2>/dev/null || echo 0) + if [ "$HAS_GOAL" = "0" ]; then + say "direct (trivial) response received" + exit 0 fi - exit 0 + # Otherwise keep waiting — the coordinator + # already delegated and will finalize via + # complete_goal once the critic runs. fi ;; esac @@ -86,6 +117,6 @@ EOF sleep 1 done -say "timed out waiting for terminal response (FINAL: or CANNOT:)" +say "timed out waiting for terminal response" say "check http://localhost:18089/runs and http://localhost:18089/goals" exit 2 diff --git a/examples/doc-gardener/start.sh b/examples/doc-gardener/start.sh index 04b8b55..ef7e284 100755 --- a/examples/doc-gardener/start.sh +++ b/examples/doc-gardener/start.sh @@ -153,12 +153,17 @@ for name in doc-coordinator docs-inspector docs-critic; do done say "configuring reactive trigger mode" +# max_trigger_depth = 4 caps the conversation loop at about two +# inspector↔critic round-trips (each REVISE costs two hops). Prevents +# the agents from spinning forever on a badly-formed report when the +# critic doesn't converge; see configs/critic.json for the structural +# half of the cap. sqlite3 "$DATA_DIR/synapbus.db" < deep-link to the full + // findings JSON or report body. + CompletionMessageID *int64 } // CreateGoalInput captures the public arguments of CreateGoal. diff --git a/internal/mcp/goals_tools.go b/internal/mcp/goals_tools.go index d4bcbac..d456e93 100644 --- a/internal/mcp/goals_tools.go +++ b/internal/mcp/goals_tools.go @@ -53,8 +53,8 @@ func NewGoalsToolRegistrar( } // RegisterAllOnServer attaches create_goal, propose_task_tree, -// propose_agent, claim_task, request_resource, list_resources to the -// MCP server. +// propose_agent, claim_task, request_resource, list_resources, and +// complete_goal to the MCP server. func (r *GoalsToolRegistrar) RegisterAllOnServer(s *server.MCPServer) { s.AddTool(r.createGoalTool(), r.handleCreateGoal) s.AddTool(r.proposeTaskTreeTool(), r.handleProposeTaskTree) @@ -62,7 +62,8 @@ func (r *GoalsToolRegistrar) RegisterAllOnServer(s *server.MCPServer) { s.AddTool(r.claimTaskTool(), r.handleClaimTask) s.AddTool(r.requestResourceTool(), r.handleRequestResource) s.AddTool(r.listResourcesTool(), r.handleListResources) - r.logger.Info("spec-018 MCP tools registered", "count", 6) + s.AddTool(r.completeGoalTool(), r.handleCompleteGoal) + r.logger.Info("spec-018 MCP tools registered", "count", 7) } // --- Tool Definitions --- @@ -120,6 +121,16 @@ func (r *GoalsToolRegistrar) listResourcesTool() mcplib.Tool { ) } +func (r *GoalsToolRegistrar) completeGoalTool() mcplib.Tool { + return mcplib.NewTool("complete_goal", + mcplib.WithDescription("Mark a goal as terminally done. The critic (or coordinator) calls this after the FINAL verdict: status=\"completed\" on success, \"stuck\" when the goal couldn't be finished, \"cancelled\" when the human aborted. Records a summary paragraph for /goals/ and links it to the DM that carried the FINAL message. Idempotent when called with the current status."), + mcplib.WithNumber("goal_id", mcplib.Description("Goal id from create_goal"), mcplib.Required()), + mcplib.WithString("status", mcplib.Description("Terminal status: completed | stuck | cancelled"), mcplib.Required()), + mcplib.WithString("summary", mcplib.Description("One-paragraph human-readable completion summary (what was done, top findings, recommended next step). Goes straight into /goals/."), mcplib.Required()), + mcplib.WithNumber("completion_message_id", mcplib.Description("Optional id of the message that carried the verdict. When set, /goals/ can deep-link to the full findings JSON in that DM.")), + ) +} + // --- Handlers --- func (r *GoalsToolRegistrar) handleCreateGoal(ctx context.Context, req mcplib.CallToolRequest) (*mcplib.CallToolResult, error) { @@ -215,6 +226,13 @@ func (r *GoalsToolRegistrar) handleProposeTaskTree(ctx context.Context, req mcpl // Also set the goal's root_task_id. _, _ = r.db.ExecContext(ctx, `UPDATE goals SET root_task_id=? WHERE id=?`, rootID, goalID) + // Auto-transition draft → active. The coordinator would otherwise + // have to call a separate tool just to flip state, which is friction + // we don't need — if there's a task tree, the goal is active. Ignore + // the legal-transition error when the goal isn't in draft (operators + // may have already moved it manually). + _ = r.goals.TransitionStatus(ctx, goalID, goals.StatusActive) + return resultJSON(map[string]any{ "root_task_id": rootID, "task_count": len(allIDs), @@ -369,3 +387,57 @@ func (r *GoalsToolRegistrar) handleListResources(ctx context.Context, req mcplib "count": len(names), }) } + +func (r *GoalsToolRegistrar) handleCompleteGoal(ctx context.Context, req mcplib.CallToolRequest) (*mcplib.CallToolResult, error) { + agentName, ok := extractAgentName(ctx) + if !ok { + return mcplib.NewToolResultError("authentication required"), nil + } + if r.goals == nil { + return mcplib.NewToolResultError("goals service not configured"), nil + } + goalID := int64(req.GetInt("goal_id", 0)) + status := req.GetString("status", "") + summary := req.GetString("summary", "") + messageID := int64(req.GetInt("completion_message_id", 0)) + if goalID <= 0 || status == "" || summary == "" { + return mcplib.NewToolResultError("goal_id, status, and summary are all required"), nil + } + switch status { + case goals.StatusCompleted, goals.StatusStuck, goals.StatusCancelled: + // ok + default: + return mcplib.NewToolResultError(fmt.Sprintf("invalid status %q — must be completed|stuck|cancelled", status)), nil + } + + // Authorization: caller must own the goal (same human owner). + // Coordinator + critic both run as ai agents owned by the same + // user in the doc-gardener example; this check keeps one user's + // agents from closing another user's goals. + callerAgent, err := r.agents.GetAgent(ctx, agentName) + if err != nil { + return mcplib.NewToolResultError(fmt.Sprintf("resolve caller: %s", err)), nil + } + g, err := r.goals.GetGoal(ctx, goalID) + if err != nil { + return mcplib.NewToolResultError(fmt.Sprintf("get goal: %s", err)), nil + } + if callerAgent.OwnerID != g.OwnerUserID { + return mcplib.NewToolResultError("goal is owned by a different user — refusing to transition"), nil + } + + if err := r.goals.Complete(ctx, goalID, status, summary, messageID); err != nil { + return mcplib.NewToolResultError(fmt.Sprintf("complete_goal: %s", err)), nil + } + + r.logger.Info("goal completed via MCP", + "goal_id", goalID, + "status", status, + "by_agent", agentName, + ) + + return resultJSON(map[string]any{ + "goal_id": goalID, + "status": status, + }) +} diff --git a/internal/storage/schema/026_goals_completion_summary.sql b/internal/storage/schema/026_goals_completion_summary.sql new file mode 100644 index 0000000..101fe8e --- /dev/null +++ b/internal/storage/schema/026_goals_completion_summary.sql @@ -0,0 +1,11 @@ +-- 026_goals_completion_summary.sql — record the critic's FINAL summary +-- on the goal row so /goals/ can render it without a JOIN through +-- goal_tasks. Also stash the message id the completion came from so +-- the UI can deep-link to the full findings JSON. +-- +-- Backfill is unnecessary — existing goals never completed cleanly +-- (the demo had no complete_goal path before this migration). + +ALTER TABLE goals ADD COLUMN completion_summary TEXT; +ALTER TABLE goals ADD COLUMN completion_message_id INTEGER + REFERENCES messages(id) ON DELETE SET NULL;