4 Commits
Author SHA1 Message Date
Algis DumbrisandClaude Opus 4.6 42f8256df6 feat(goals): complete_goal MCP tool + draft→active auto-transition
Three improvements that turn the doc-gardener demo from "runs but
stays in 'draft' forever" into a goal that properly transitions
through its lifecycle and renders a completion summary on /goals/<id>.

### 1. complete_goal MCP tool (#59, #62)

New tool surface: complete_goal(goal_id, status, summary, completion_message_id?)

The critic calls this from inside the sandbox after it sends its FINAL:
DM. Records the one-paragraph human-readable summary on the goal row
plus a pointer to the message that carried the FINAL text, so the Web
UI /goals/<id> page has both the verdict and a deep link to the full
findings JSON.

Status parameter accepts completed | stuck | cancelled. Idempotent
when called with the current status. Rejects callers owned by a
different human than the goal owner.

Plumbing:
- New migration 026_goals_completion_summary.sql adds two columns
  to goals: completion_summary TEXT, completion_message_id INTEGER
  (FK messages.id, ON DELETE SET NULL).
- internal/goals/types.go: new CompletionSummary + CompletionMessageID
  fields on Goal struct.
- internal/goals/store.go: Get/List Scan both new columns;
  SetCompletion(goalID, status, summary, messageID) helper that
  updates status+summary+message_id atomically and populates
  completed_at for terminal states.
- internal/goals/service.go: Complete(ctx, goalID, status, summary,
  messageID) wraps the store method with legalTransition gating.
  legalTransition expanded so draft can jump straight to completed
  (no mandatory "active" hop required).
- internal/mcp/goals_tools.go: completeGoalTool definition +
  handleCompleteGoal handler. Tool count 6 → 7.
- internal/api/goals_handler.go: surfaces completion_summary,
  completion_message_id, and completed_at on both list and detail
  endpoints so the Svelte /goals UI can render them.

### 2. Draft → active auto-transition in propose_task_tree (#60)

handleProposeTaskTree now flips the goal from draft to active at the
end. Previously the coordinator would call create_goal +
propose_task_tree and dispatch inspector, but the goal stayed in
draft forever because nothing transitioned it. Now the mere fact
of having a task tree means the goal is active.

Safe: the transition is best-effort and ignores the legal-transition
error when the goal is already beyond draft.

### 3. REVISE round cap (#61)

Two-layer enforcement:

- Server-side: examples/doc-gardener/start.sh drops max_trigger_depth
  from 8 to 4. Each REVISE round costs 2 hops (critic→inspector +
  inspector→critic), so depth=4 caps the loop at roughly 2 rounds
  before the reactor refuses further dispatches.

- Prompt-side: inspector now includes revision_round (starting at
  0, incremented when it sees a REVISE: input) in its findings JSON.
  Critic reads revision_round and force-FINALs when >= 1. Prompt
  explicitly tells the critic to call complete_goal after sending
  FINAL, so the goal row gets a proper completion_summary.

### 4. run_task.sh terminal-state detection

Rewrote the poll loop to watch goals.status/completion_summary as
the definitive "done" signal rather than parsing DM bodies. Keeps
a message-based fallback for TRIVIAL/CANNOT paths that don't create
a goal. Treats "Received system trigger..." and "Coalesced
trigger..." as informational (they're `__coalesced__` reactor
synthetic events leaking through the coordinator reply, not real
user-facing output). Bare coordinator replies are terminal only
when no goal was created.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-15 18:52:44 +03:00
Algis DumbrisandClaude Opus 4.6 f1e2b1fa38 feat(doc-gardener): MCP-native + docker-isolated multi-agent demo
Replace the legacy cmd/docgardener orchestration (~2400 LOC of Go
spawning subprocess workers via local_command + admin socket) with
three Docker-isolated agents that all reach SynapBus through MCP:

  doc-coordinator   — Gemini Pro, triages goal, calls create_goal +
                      propose_task_tree + send_message via MCP
  docs-inspector    — Gemini Flash, fetches docs, installs mcpproxy,
                      shells out to verify, reports findings via MCP
  docs-critic       — Gemini Flash, independent reviewer with its
                      own MCP API key + config_hash, audits the
                      inspector's evidence and DMs the owner

Every agent runs inside synapbus-agent:latest with --cap-drop=ALL,
--security-opt=no-new-privileges, --read-only root + tmpfs /tmp,
--pids-limit, memory + CPU quotas. The container reaches the
SynapBus MCP server on the host at host.docker.internal:18089
because the docker harness rewrites .gemini/settings.json URLs
from 127.0.0.1 automatically.

Wrapper baked into the image at /usr/local/bin/synapbus-agent-wrapper.sh
so configs don't need to mount or template a per-example wrapper.
The harness's default no longer overrides docker CMD — the image's
baked entry script is used unless docker.command is set explicitly.

start.sh changes:
  - Preflight: docker daemon, GEMINI_API_KEY (or ~/.gemini/oauth_creds.json)
  - Builds synapbus-agent image lazily on first run
  - Mints one MCP API key per agent via `agent revoke-key`
  - Templates each config with __PORT__, __*_APIKEY__, __MODEL__,
    __GEMINI_API_KEY__, __EXTRA_MOUNTS__
  - With OAuth fallback: copies host ~/.gemini → data/agent-home/.gemini
    once and bind-mounts the whole agent-home rw at /home/agent so
    in-container gemini has a writable HOME without polluting the host
  - SYNAPBUS_KEEP_WORKDIR=1 preserves per-run docker workdirs for
    debugging
  - Sets harness_name=docker explicitly so the resolver picks the
    right backend even with empty local_command

stop.sh: best-effort cleanup of lingering synapbus-* containers so a
killed parent doesn't leave bind-mount holders that block the next
start.sh from re-mounting the same paths.

run_task.sh: snapshot-baseline pattern (only watches replies newer
than the max msg id at send time), 600s deadline, treats any reply
from doc-coordinator that isn't DELEGATED:/REVISING: as terminal,
plus FINAL:/CANNOT: from any sender.

cmd/docgardener slimmed from 7 files / 2580 LOC to 3 files / ~370 LOC.
The remaining binary only renders the HTML report (queries goals +
goal_tasks + traces + harness_runs from the SynapBus DB read-only).
agent.go, channels.go, flow.go, gemini_tree.go all deleted.

Verified end-to-end against gemini-2.5-pro coordinator + gemini-2.5-flash
workers (with OAuth fallback mount):

  ./run_task.sh "what does this demo do?"
    → coordinator TRIVIAL: replies directly via MCP send_message

  ./run_task.sh "Verify the CLI commands on docs.mcpproxy.app/cli/command-reference"
    → coordinator calls create_goal (slug verify-mcpproxy-cli-...),
      propose_task_tree (3-node tree: coordinator/plan,
      doc-gardener/scan, doc-gardener/audit) and send_message to
      docs-inspector
    → inspector container runs ~10 minutes inside the sandbox:
      installs mcpproxy from real release URL (linux-arm64), curls
      the docs page, falls back from BeautifulSoup → grep when
      python3-venv is missing, debugs its own f-string syntax, writes
      extract_flags.py, runs `mcpproxy --help` for ground truth
    → real multi-agent iteration loop: critic REVISE: → inspector
      retry → critic REVISE: with new feedback

The agents discovered real environment quirks (tmpfs noexec on /tmp,
externally-managed Python, missing python3-venv) and worked around
them inside the sandbox without touching the host.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-15 10:28:05 +03:00
Algis DumbrisandClaude Opus 4.6 3b94fab226 feat(018): real reactor-driven multi-agent doc-gardener flow
Until now the doc-gardener example was a single monolithic
orchestrator binary writing synthetic messages directly to SQLite.
That's now obsolete: the feature runs as a true multi-agent flow
where the SynapBus reactor fires subprocess runs for every DM, each
agent is its own reactive subprocess invocation, and follow-up DMs
go through the real MessagingService.Send → dispatcher path so the
reactor picks them up.

Changes:

- cmd/synapbus/main.go: gate the three legacy background workers
  (expiry, retention, stalemate) behind SYNAPBUS_DISABLE_*_WORKER env
  flags. These workers manage the legacy channel task-auction /
  message retention features the doc-gardener demo doesn't use, but
  they held the single-connection write pool long enough to wedge
  the whole server for interactive sessions. All three are disabled
  in the example's start.sh.

- cmd/docgardener/agent.go (new): the per-agent subprocess entry the
  reactor harness invokes for every reactive trigger. Reads
  message.json from the workdir, routes by SYNAPBUS_AGENT to either
  coordinator-kickoff, coordinator-completion, or specialist-work
  logic. Writes prompt.txt + response.txt for harness capture. Uses
  the admin socket (`synapbus messages send`) for follow-up DMs so
  the real MessagingService.Send path fires the dispatcher.

- cmd/docgardener/main.go: adds `docgardener agent` subcommand, plus
  helpers freshAPIKey / bcryptHash / absPath / selfPath used by the
  spawn flow.

- examples/doc-gardener/start.sh: provisions user + coordinator
  agent + algis human agent + approvals/requests channels; the
  coordinator is created with trigger_mode=reactive,
  harness_name=subprocess, local_command pointing to docgardener
  agent, and harness_config_json.env carrying SYNAPBUS_AGENT,
  SYNAPBUS_BIN, SYNAPBUS_SOCKET. Specialists are spawned
  dynamically by the coordinator at runtime (not pre-registered),
  so the demo exercises dynamic agent spawning end-to-end.

- examples/doc-gardener/run_task.sh: collapsed to a 3-line kickoff
  that just DMs the coordinator and polls algis's inbox for the
  coordinator's FINAL: reply. Everything else happens via the
  reactor.

Verified end-to-end in Chrome on a fresh instance:
- 4 agents registered (coordinator + 3 specialists dynamically
  spawned by the coordinator on receipt of the first DM)
- 7 reactive_runs + 6 harness_runs across the goal lifecycle:
    algis → coordinator (kickoff, 624ms, builds goal+tree+spawns)
    coordinator → docs-scanner (claim task 2)
    coordinator → cli-verifier (claim task 3)
    coordinator → drift-reporter (claim task 4)
    docs-scanner → coordinator (DONE task=2)
    cli-verifier → coordinator (DONE task=3)
    drift-reporter → coordinator (DONE task=4, coalesced)
- Web UI Agent Runs page shows all 7 runs with the real
  "DM from X" trigger lines and correct sender/receiver chain
- Goal ends at status=completed with all 3 leaf tasks at status=done
- Each specialist run posts a real subprocess artifact to the
  goal channel (#finding, #verified, #summary) and appends a real
  reputation_evidence row keyed by config_hash.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 17:09:08 +03:00
Algis DumbrisandClaude Opus 4.6 ff5d0c49f4 feat(018): dynamic agent spawning — primitives + doc-gardener demo
Ships the MVP slice of spec 018 (dynamic agent spawning):

- 5 new SQLite migrations (021-025): goals + goal_tasks + agent_proposals
  + reputation_evidence + secrets + harness_runs.task_id. The legacy
  `tasks` table (channel auctions) and `agent_trust` table (reactions
  workflow) are left untouched — the new schema coexists.

- 4 new internal packages, fully tested:
  - internal/goals: Goal struct + store + service, slug collision dedup,
    backing-channel auto-create via ChannelCreator adapter
  - internal/goaltasks: goal_tasks table with denormalized 16 KB
    ancestry snapshots, single-statement optimistic-lock atomic claim,
    recursive-CTE cost rollup, state machine, per-billing-code rollup
  - internal/secrets: NaCl-secretbox encrypted blobs, user/agent/task
    scope precedence, sanitized env injection, master-key bootstrap
  - internal/trust additions: ConfigHash (deterministic SHA-256 of
    model + prompt + tools + skills + mcp + subagents, sorted),
    DelegationCap (tier + tool-scope + budget + depth enforcement),
    append-only Ledger with exponential time-decay rolling score and
    70%-of-parent child seeding. Existing trust package unchanged.

- Critical invariants under test:
  - 50-goroutine concurrent claim race → exactly one winner per round
  - ConfigHash stable under shuffled array inputs, sensitive to
    capability changes
  - DelegationCap full tier × tool-scope matrix
  - Ledger time-decay + parent seed at 70 % ± 1 %
  - Secret name sanitization, scope precedence, plaintext never
    returned via MCP-equivalent paths

- internal/agents/types.go extended with dynamic-spawning columns
  (config_hash, parent_agent_id, spawn_depth, system_prompt,
  autonomy_tier, tool_scope_json, quarantined_at). Existing tests
  still pass.

- cmd/docgardener: self-contained demo binary driving the end-to-end
  flow. `docgardener run` creates a goal, builds a task tree with
  denormalized ancestry, spawns 3 specialists (each going through
  real delegation-cap validation and config-hash computation and
  70 %-of-parent reputation seeding), claims tasks atomically, runs
  them through the state machine, records reputation evidence.
  `docgardener report` queries all of that back out and renders a
  rich dark-mode HTML report (header, spend metrics, task tree,
  spawned-agent cards with reputation bars, cost breakdown, artifacts,
  timeline).

- examples/doc-gardener: start.sh / run_task.sh / report.sh / stop.sh
  mirroring the cold-topic-explainer pattern. Launches an isolated
  synapbus instance on port 18089, drives the demo, renders
  report.html, cleans up. Full README documenting what's real vs
  deferred, plus examples/README.md listing both examples.

- specs/018: tasks.md updated with MVP completion status; legacy tasks
  naming collision noted.

Deferred (marked explicitly in example README):
- Real LLM-driven coordinator (needs MCP tool wiring + prompt
  iteration)
- Real subprocess runs (needs reactor integration with task_id on
  ExecRequest)
- Full MCP tool surface (contracts are written at
  specs/018-dynamic-agent-spawning/contracts/mcp-tools.md)
- Svelte /goals UI (REST endpoints remain a follow-up)
- Full budget race + quarantine auto-trigger wiring
- Full resource-request → secrets fulfill reaction-workflow path

Cross-compiles clean for linux/amd64 and darwin/arm64 with no CGO
(SC-010). All new package tests pass (SC-004, SC-005, SC-007).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 15:29:21 +03:00