Files
synapbus/examples/doc-gardener/run_task.sh
T
Algis DumbrisandClaude Opus 4.6 f1e2b1fa38 feat(doc-gardener): MCP-native + docker-isolated multi-agent demo
Replace the legacy cmd/docgardener orchestration (~2400 LOC of Go
spawning subprocess workers via local_command + admin socket) with
three Docker-isolated agents that all reach SynapBus through MCP:

  doc-coordinator   — Gemini Pro, triages goal, calls create_goal +
                      propose_task_tree + send_message via MCP
  docs-inspector    — Gemini Flash, fetches docs, installs mcpproxy,
                      shells out to verify, reports findings via MCP
  docs-critic       — Gemini Flash, independent reviewer with its
                      own MCP API key + config_hash, audits the
                      inspector's evidence and DMs the owner

Every agent runs inside synapbus-agent:latest with --cap-drop=ALL,
--security-opt=no-new-privileges, --read-only root + tmpfs /tmp,
--pids-limit, memory + CPU quotas. The container reaches the
SynapBus MCP server on the host at host.docker.internal:18089
because the docker harness rewrites .gemini/settings.json URLs
from 127.0.0.1 automatically.

Wrapper baked into the image at /usr/local/bin/synapbus-agent-wrapper.sh
so configs don't need to mount or template a per-example wrapper.
The harness's default no longer overrides docker CMD — the image's
baked entry script is used unless docker.command is set explicitly.

start.sh changes:
  - Preflight: docker daemon, GEMINI_API_KEY (or ~/.gemini/oauth_creds.json)
  - Builds synapbus-agent image lazily on first run
  - Mints one MCP API key per agent via `agent revoke-key`
  - Templates each config with __PORT__, __*_APIKEY__, __MODEL__,
    __GEMINI_API_KEY__, __EXTRA_MOUNTS__
  - With OAuth fallback: copies host ~/.gemini → data/agent-home/.gemini
    once and bind-mounts the whole agent-home rw at /home/agent so
    in-container gemini has a writable HOME without polluting the host
  - SYNAPBUS_KEEP_WORKDIR=1 preserves per-run docker workdirs for
    debugging
  - Sets harness_name=docker explicitly so the resolver picks the
    right backend even with empty local_command

stop.sh: best-effort cleanup of lingering synapbus-* containers so a
killed parent doesn't leave bind-mount holders that block the next
start.sh from re-mounting the same paths.

run_task.sh: snapshot-baseline pattern (only watches replies newer
than the max msg id at send time), 600s deadline, treats any reply
from doc-coordinator that isn't DELEGATED:/REVISING: as terminal,
plus FINAL:/CANNOT: from any sender.

cmd/docgardener slimmed from 7 files / 2580 LOC to 3 files / ~370 LOC.
The remaining binary only renders the HTML report (queries goals +
goal_tasks + traces + harness_runs from the SynapBus DB read-only).
agent.go, channels.go, flow.go, gemini_tree.go all deleted.

Verified end-to-end against gemini-2.5-pro coordinator + gemini-2.5-flash
workers (with OAuth fallback mount):

  ./run_task.sh "what does this demo do?"
    → coordinator TRIVIAL: replies directly via MCP send_message

  ./run_task.sh "Verify the CLI commands on docs.mcpproxy.app/cli/command-reference"
    → coordinator calls create_goal (slug verify-mcpproxy-cli-...),
      propose_task_tree (3-node tree: coordinator/plan,
      doc-gardener/scan, doc-gardener/audit) and send_message to
      docs-inspector
    → inspector container runs ~10 minutes inside the sandbox:
      installs mcpproxy from real release URL (linux-arm64), curls
      the docs page, falls back from BeautifulSoup → grep when
      python3-venv is missing, debugs its own f-string syntax, writes
      extract_flags.py, runs `mcpproxy --help` for ground truth
    → real multi-agent iteration loop: critic REVISE: → inspector
      retry → critic REVISE: with new feedback

The agents discovered real environment quirks (tmpfs noexec on /tmp,
externally-managed Python, missing python3-venv) and worked around
them inside the sandbox without touching the host.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-15 10:28:05 +03:00

92 lines
3.7 KiB
Bash
Executable File

#!/bin/bash
# run_task.sh — send a doc-verification goal DM from algis to
# doc-coordinator and wait for the FINAL: reply that flows back from
# docs-critic. The whole flow is driven by MCP tool calls inside three
# Docker-isolated agent containers — nothing here writes to the DB
# directly.
#
# Usage:
# ./run_task.sh # default doc-gardener brief
# ./run_task.sh "your custom goal here"
#
# The default brief asks the inspector to verify mcpproxy CLI flag
# documentation against the actual binary. Override with any free-form
# brief — the coordinator triages it.
set -euo pipefail
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
BIN="$SCRIPT_DIR/bin/synapbus"
SOCKET="$SCRIPT_DIR/data/synapbus.sock"
DEFAULT_GOAL='Verify the CLI commands listed on https://docs.mcpproxy.app/cli/command-reference still exist in the current mcpproxy binary. Install mcpproxy in the sandbox first (releases at https://github.com/smart-mcp-proxy/mcpproxy-go/releases — pick the linux-arm64 or linux-amd64 variant matching `uname -m`). For each documented command, check whether `mcpproxy --help` and `mcpproxy <command> --help` show it; flag any drift, missing commands, or doc claims that no longer match. Produce a patch suggestion list.'
GOAL="${1:-$DEFAULT_GOAL}"
say() { printf '\033[1;36m[run]\033[0m %s\n' "$*"; }
die() { printf '\033[1;31m[run][FAIL]\033[0m %s\n' "$*" >&2; exit 1; }
[ -x "$BIN" ] || die "synapbus binary not found at $BIN — run ./start.sh first"
[ -S "$SOCKET" ] || die "admin socket missing — is synapbus running?"
cd "$SCRIPT_DIR"
DB="$SCRIPT_DIR/data/synapbus.db"
# Snapshot the current max message id so we only look at replies from
# THIS run, not stale replies left from previous invocations.
BASELINE=$(sqlite3 "$DB" "SELECT COALESCE(MAX(id), 0) FROM messages" 2>/dev/null || echo 0)
say "sending goal DM: algis → doc-coordinator (baseline msg_id=$BASELINE)"
printf '%s' "$GOAL" | "$BIN" --socket "$SOCKET" messages send \
--from algis \
--to doc-coordinator \
--priority 8 >&2
say "waiting for FINAL: / CANNOT: reply to algis (up to 600s)..."
deadline=$(( $(date +%s) + 600 ))
last_seen_id=$BASELINE
while [ "$(date +%s)" -lt "$deadline" ]; do
NEW_LINES=$(sqlite3 -separator '|' "$DB" "
SELECT id, from_agent, replace(substr(body, 1, 280), char(10), ' ')
FROM messages
WHERE to_agent = 'algis'
AND from_agent != 'algis'
AND id > $last_seen_id
ORDER BY id ASC
" 2>/dev/null || true)
if [ -n "$NEW_LINES" ]; then
while IFS='|' read -r id from body; do
[ -z "$id" ] && continue
say "← [$from #$id] $body"
last_seen_id=$id
case "$body" in
DELEGATED:*|REVISING:*)
;; # informational, keep waiting
*)
if [ "$from" = "doc-coordinator" ] || \
[ "${body#FINAL:}" != "$body" ] || \
[ "${body#CANNOT:}" != "$body" ]; then
say "terminal response received"
# Persist last goal id for ./report.sh.
GOAL_ID=$(sqlite3 "$DB" 'SELECT id FROM goals ORDER BY id DESC LIMIT 1' 2>/dev/null || echo)
if [ -n "$GOAL_ID" ]; then
echo "$GOAL_ID" > "$SCRIPT_DIR/.last_goal_id"
say "goal id = $GOAL_ID — render with ./report.sh"
fi
exit 0
fi
;;
esac
done <<EOF
$NEW_LINES
EOF
fi
sleep 1
done
say "timed out waiting for terminal response (FINAL: or CANNOT:)"
say "check http://localhost:18089/runs and http://localhost:18089/goals"
exit 2