Replace the legacy cmd/docgardener orchestration (~2400 LOC of Go
spawning subprocess workers via local_command + admin socket) with
three Docker-isolated agents that all reach SynapBus through MCP:
doc-coordinator — Gemini Pro, triages goal, calls create_goal +
propose_task_tree + send_message via MCP
docs-inspector — Gemini Flash, fetches docs, installs mcpproxy,
shells out to verify, reports findings via MCP
docs-critic — Gemini Flash, independent reviewer with its
own MCP API key + config_hash, audits the
inspector's evidence and DMs the owner
Every agent runs inside synapbus-agent:latest with --cap-drop=ALL,
--security-opt=no-new-privileges, --read-only root + tmpfs /tmp,
--pids-limit, memory + CPU quotas. The container reaches the
SynapBus MCP server on the host at host.docker.internal:18089
because the docker harness rewrites .gemini/settings.json URLs
from 127.0.0.1 automatically.
Wrapper baked into the image at /usr/local/bin/synapbus-agent-wrapper.sh
so configs don't need to mount or template a per-example wrapper.
The harness's default no longer overrides docker CMD — the image's
baked entry script is used unless docker.command is set explicitly.
start.sh changes:
- Preflight: docker daemon, GEMINI_API_KEY (or ~/.gemini/oauth_creds.json)
- Builds synapbus-agent image lazily on first run
- Mints one MCP API key per agent via `agent revoke-key`
- Templates each config with __PORT__, __*_APIKEY__, __MODEL__,
__GEMINI_API_KEY__, __EXTRA_MOUNTS__
- With OAuth fallback: copies host ~/.gemini → data/agent-home/.gemini
once and bind-mounts the whole agent-home rw at /home/agent so
in-container gemini has a writable HOME without polluting the host
- SYNAPBUS_KEEP_WORKDIR=1 preserves per-run docker workdirs for
debugging
- Sets harness_name=docker explicitly so the resolver picks the
right backend even with empty local_command
stop.sh: best-effort cleanup of lingering synapbus-* containers so a
killed parent doesn't leave bind-mount holders that block the next
start.sh from re-mounting the same paths.
run_task.sh: snapshot-baseline pattern (only watches replies newer
than the max msg id at send time), 600s deadline, treats any reply
from doc-coordinator that isn't DELEGATED:/REVISING: as terminal,
plus FINAL:/CANNOT: from any sender.
cmd/docgardener slimmed from 7 files / 2580 LOC to 3 files / ~370 LOC.
The remaining binary only renders the HTML report (queries goals +
goal_tasks + traces + harness_runs from the SynapBus DB read-only).
agent.go, channels.go, flow.go, gemini_tree.go all deleted.
Verified end-to-end against gemini-2.5-pro coordinator + gemini-2.5-flash
workers (with OAuth fallback mount):
./run_task.sh "what does this demo do?"
→ coordinator TRIVIAL: replies directly via MCP send_message
./run_task.sh "Verify the CLI commands on docs.mcpproxy.app/cli/command-reference"
→ coordinator calls create_goal (slug verify-mcpproxy-cli-...),
propose_task_tree (3-node tree: coordinator/plan,
doc-gardener/scan, doc-gardener/audit) and send_message to
docs-inspector
→ inspector container runs ~10 minutes inside the sandbox:
installs mcpproxy from real release URL (linux-arm64), curls
the docs page, falls back from BeautifulSoup → grep when
python3-venv is missing, debugs its own f-string syntax, writes
extract_flags.py, runs `mcpproxy --help` for ground truth
→ real multi-agent iteration loop: critic REVISE: → inspector
retry → critic REVISE: with new feedback
The agents discovered real environment quirks (tmpfs noexec on /tmp,
externally-managed Python, missing python3-venv) and worked around
them inside the sandbox without touching the host.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
92 lines
3.7 KiB
Bash
Executable File
92 lines
3.7 KiB
Bash
Executable File
#!/bin/bash
|
|
# run_task.sh — send a doc-verification goal DM from algis to
|
|
# doc-coordinator and wait for the FINAL: reply that flows back from
|
|
# docs-critic. The whole flow is driven by MCP tool calls inside three
|
|
# Docker-isolated agent containers — nothing here writes to the DB
|
|
# directly.
|
|
#
|
|
# Usage:
|
|
# ./run_task.sh # default doc-gardener brief
|
|
# ./run_task.sh "your custom goal here"
|
|
#
|
|
# The default brief asks the inspector to verify mcpproxy CLI flag
|
|
# documentation against the actual binary. Override with any free-form
|
|
# brief — the coordinator triages it.
|
|
|
|
set -euo pipefail
|
|
|
|
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
|
BIN="$SCRIPT_DIR/bin/synapbus"
|
|
SOCKET="$SCRIPT_DIR/data/synapbus.sock"
|
|
|
|
DEFAULT_GOAL='Verify the CLI commands listed on https://docs.mcpproxy.app/cli/command-reference still exist in the current mcpproxy binary. Install mcpproxy in the sandbox first (releases at https://github.com/smart-mcp-proxy/mcpproxy-go/releases — pick the linux-arm64 or linux-amd64 variant matching `uname -m`). For each documented command, check whether `mcpproxy --help` and `mcpproxy <command> --help` show it; flag any drift, missing commands, or doc claims that no longer match. Produce a patch suggestion list.'
|
|
|
|
GOAL="${1:-$DEFAULT_GOAL}"
|
|
|
|
say() { printf '\033[1;36m[run]\033[0m %s\n' "$*"; }
|
|
die() { printf '\033[1;31m[run][FAIL]\033[0m %s\n' "$*" >&2; exit 1; }
|
|
|
|
[ -x "$BIN" ] || die "synapbus binary not found at $BIN — run ./start.sh first"
|
|
[ -S "$SOCKET" ] || die "admin socket missing — is synapbus running?"
|
|
|
|
cd "$SCRIPT_DIR"
|
|
DB="$SCRIPT_DIR/data/synapbus.db"
|
|
|
|
# Snapshot the current max message id so we only look at replies from
|
|
# THIS run, not stale replies left from previous invocations.
|
|
BASELINE=$(sqlite3 "$DB" "SELECT COALESCE(MAX(id), 0) FROM messages" 2>/dev/null || echo 0)
|
|
|
|
say "sending goal DM: algis → doc-coordinator (baseline msg_id=$BASELINE)"
|
|
printf '%s' "$GOAL" | "$BIN" --socket "$SOCKET" messages send \
|
|
--from algis \
|
|
--to doc-coordinator \
|
|
--priority 8 >&2
|
|
|
|
say "waiting for FINAL: / CANNOT: reply to algis (up to 600s)..."
|
|
deadline=$(( $(date +%s) + 600 ))
|
|
last_seen_id=$BASELINE
|
|
|
|
while [ "$(date +%s)" -lt "$deadline" ]; do
|
|
NEW_LINES=$(sqlite3 -separator '|' "$DB" "
|
|
SELECT id, from_agent, replace(substr(body, 1, 280), char(10), ' ')
|
|
FROM messages
|
|
WHERE to_agent = 'algis'
|
|
AND from_agent != 'algis'
|
|
AND id > $last_seen_id
|
|
ORDER BY id ASC
|
|
" 2>/dev/null || true)
|
|
|
|
if [ -n "$NEW_LINES" ]; then
|
|
while IFS='|' read -r id from body; do
|
|
[ -z "$id" ] && continue
|
|
say "← [$from #$id] $body"
|
|
last_seen_id=$id
|
|
case "$body" in
|
|
DELEGATED:*|REVISING:*)
|
|
;; # informational, keep waiting
|
|
*)
|
|
if [ "$from" = "doc-coordinator" ] || \
|
|
[ "${body#FINAL:}" != "$body" ] || \
|
|
[ "${body#CANNOT:}" != "$body" ]; then
|
|
say "terminal response received"
|
|
# Persist last goal id for ./report.sh.
|
|
GOAL_ID=$(sqlite3 "$DB" 'SELECT id FROM goals ORDER BY id DESC LIMIT 1' 2>/dev/null || echo)
|
|
if [ -n "$GOAL_ID" ]; then
|
|
echo "$GOAL_ID" > "$SCRIPT_DIR/.last_goal_id"
|
|
say "goal id = $GOAL_ID — render with ./report.sh"
|
|
fi
|
|
exit 0
|
|
fi
|
|
;;
|
|
esac
|
|
done <<EOF
|
|
$NEW_LINES
|
|
EOF
|
|
fi
|
|
sleep 1
|
|
done
|
|
|
|
say "timed out waiting for terminal response (FINAL: or CANNOT:)"
|
|
say "check http://localhost:18089/runs and http://localhost:18089/goals"
|
|
exit 2
|