main
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
42f8256df6 |
feat(goals): complete_goal MCP tool + draft→active auto-transition
Three improvements that turn the doc-gardener demo from "runs but stays in 'draft' forever" into a goal that properly transitions through its lifecycle and renders a completion summary on /goals/<id>. ### 1. complete_goal MCP tool (#59, #62) New tool surface: complete_goal(goal_id, status, summary, completion_message_id?) The critic calls this from inside the sandbox after it sends its FINAL: DM. Records the one-paragraph human-readable summary on the goal row plus a pointer to the message that carried the FINAL text, so the Web UI /goals/<id> page has both the verdict and a deep link to the full findings JSON. Status parameter accepts completed | stuck | cancelled. Idempotent when called with the current status. Rejects callers owned by a different human than the goal owner. Plumbing: - New migration 026_goals_completion_summary.sql adds two columns to goals: completion_summary TEXT, completion_message_id INTEGER (FK messages.id, ON DELETE SET NULL). - internal/goals/types.go: new CompletionSummary + CompletionMessageID fields on Goal struct. - internal/goals/store.go: Get/List Scan both new columns; SetCompletion(goalID, status, summary, messageID) helper that updates status+summary+message_id atomically and populates completed_at for terminal states. - internal/goals/service.go: Complete(ctx, goalID, status, summary, messageID) wraps the store method with legalTransition gating. legalTransition expanded so draft can jump straight to completed (no mandatory "active" hop required). - internal/mcp/goals_tools.go: completeGoalTool definition + handleCompleteGoal handler. Tool count 6 → 7. - internal/api/goals_handler.go: surfaces completion_summary, completion_message_id, and completed_at on both list and detail endpoints so the Svelte /goals UI can render them. ### 2. Draft → active auto-transition in propose_task_tree (#60) handleProposeTaskTree now flips the goal from draft to active at the end. Previously the coordinator would call create_goal + propose_task_tree and dispatch inspector, but the goal stayed in draft forever because nothing transitioned it. Now the mere fact of having a task tree means the goal is active. Safe: the transition is best-effort and ignores the legal-transition error when the goal is already beyond draft. ### 3. REVISE round cap (#61) Two-layer enforcement: - Server-side: examples/doc-gardener/start.sh drops max_trigger_depth from 8 to 4. Each REVISE round costs 2 hops (critic→inspector + inspector→critic), so depth=4 caps the loop at roughly 2 rounds before the reactor refuses further dispatches. - Prompt-side: inspector now includes revision_round (starting at 0, incremented when it sees a REVISE: input) in its findings JSON. Critic reads revision_round and force-FINALs when >= 1. Prompt explicitly tells the critic to call complete_goal after sending FINAL, so the goal row gets a proper completion_summary. ### 4. run_task.sh terminal-state detection Rewrote the poll loop to watch goals.status/completion_summary as the definitive "done" signal rather than parsing DM bodies. Keeps a message-based fallback for TRIVIAL/CANNOT paths that don't create a goal. Treats "Received system trigger..." and "Coalesced trigger..." as informational (they're `__coalesced__` reactor synthetic events leaking through the coordinator reply, not real user-facing output). Bare coordinator replies are terminal only when no goal was created. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
f1e2b1fa38 |
feat(doc-gardener): MCP-native + docker-isolated multi-agent demo
Replace the legacy cmd/docgardener orchestration (~2400 LOC of Go
spawning subprocess workers via local_command + admin socket) with
three Docker-isolated agents that all reach SynapBus through MCP:
doc-coordinator — Gemini Pro, triages goal, calls create_goal +
propose_task_tree + send_message via MCP
docs-inspector — Gemini Flash, fetches docs, installs mcpproxy,
shells out to verify, reports findings via MCP
docs-critic — Gemini Flash, independent reviewer with its
own MCP API key + config_hash, audits the
inspector's evidence and DMs the owner
Every agent runs inside synapbus-agent:latest with --cap-drop=ALL,
--security-opt=no-new-privileges, --read-only root + tmpfs /tmp,
--pids-limit, memory + CPU quotas. The container reaches the
SynapBus MCP server on the host at host.docker.internal:18089
because the docker harness rewrites .gemini/settings.json URLs
from 127.0.0.1 automatically.
Wrapper baked into the image at /usr/local/bin/synapbus-agent-wrapper.sh
so configs don't need to mount or template a per-example wrapper.
The harness's default no longer overrides docker CMD — the image's
baked entry script is used unless docker.command is set explicitly.
start.sh changes:
- Preflight: docker daemon, GEMINI_API_KEY (or ~/.gemini/oauth_creds.json)
- Builds synapbus-agent image lazily on first run
- Mints one MCP API key per agent via `agent revoke-key`
- Templates each config with __PORT__, __*_APIKEY__, __MODEL__,
__GEMINI_API_KEY__, __EXTRA_MOUNTS__
- With OAuth fallback: copies host ~/.gemini → data/agent-home/.gemini
once and bind-mounts the whole agent-home rw at /home/agent so
in-container gemini has a writable HOME without polluting the host
- SYNAPBUS_KEEP_WORKDIR=1 preserves per-run docker workdirs for
debugging
- Sets harness_name=docker explicitly so the resolver picks the
right backend even with empty local_command
stop.sh: best-effort cleanup of lingering synapbus-* containers so a
killed parent doesn't leave bind-mount holders that block the next
start.sh from re-mounting the same paths.
run_task.sh: snapshot-baseline pattern (only watches replies newer
than the max msg id at send time), 600s deadline, treats any reply
from doc-coordinator that isn't DELEGATED:/REVISING: as terminal,
plus FINAL:/CANNOT: from any sender.
cmd/docgardener slimmed from 7 files / 2580 LOC to 3 files / ~370 LOC.
The remaining binary only renders the HTML report (queries goals +
goal_tasks + traces + harness_runs from the SynapBus DB read-only).
agent.go, channels.go, flow.go, gemini_tree.go all deleted.
Verified end-to-end against gemini-2.5-pro coordinator + gemini-2.5-flash
workers (with OAuth fallback mount):
./run_task.sh "what does this demo do?"
→ coordinator TRIVIAL: replies directly via MCP send_message
./run_task.sh "Verify the CLI commands on docs.mcpproxy.app/cli/command-reference"
→ coordinator calls create_goal (slug verify-mcpproxy-cli-...),
propose_task_tree (3-node tree: coordinator/plan,
doc-gardener/scan, doc-gardener/audit) and send_message to
docs-inspector
→ inspector container runs ~10 minutes inside the sandbox:
installs mcpproxy from real release URL (linux-arm64), curls
the docs page, falls back from BeautifulSoup → grep when
python3-venv is missing, debugs its own f-string syntax, writes
extract_flags.py, runs `mcpproxy --help` for ground truth
→ real multi-agent iteration loop: critic REVISE: → inspector
retry → critic REVISE: with new feedback
The agents discovered real environment quirks (tmpfs noexec on /tmp,
externally-managed Python, missing python3-venv) and worked around
them inside the sandbox without touching the host.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
3b94fab226 |
feat(018): real reactor-driven multi-agent doc-gardener flow
Until now the doc-gardener example was a single monolithic
orchestrator binary writing synthetic messages directly to SQLite.
That's now obsolete: the feature runs as a true multi-agent flow
where the SynapBus reactor fires subprocess runs for every DM, each
agent is its own reactive subprocess invocation, and follow-up DMs
go through the real MessagingService.Send → dispatcher path so the
reactor picks them up.
Changes:
- cmd/synapbus/main.go: gate the three legacy background workers
(expiry, retention, stalemate) behind SYNAPBUS_DISABLE_*_WORKER env
flags. These workers manage the legacy channel task-auction /
message retention features the doc-gardener demo doesn't use, but
they held the single-connection write pool long enough to wedge
the whole server for interactive sessions. All three are disabled
in the example's start.sh.
- cmd/docgardener/agent.go (new): the per-agent subprocess entry the
reactor harness invokes for every reactive trigger. Reads
message.json from the workdir, routes by SYNAPBUS_AGENT to either
coordinator-kickoff, coordinator-completion, or specialist-work
logic. Writes prompt.txt + response.txt for harness capture. Uses
the admin socket (`synapbus messages send`) for follow-up DMs so
the real MessagingService.Send path fires the dispatcher.
- cmd/docgardener/main.go: adds `docgardener agent` subcommand, plus
helpers freshAPIKey / bcryptHash / absPath / selfPath used by the
spawn flow.
- examples/doc-gardener/start.sh: provisions user + coordinator
agent + algis human agent + approvals/requests channels; the
coordinator is created with trigger_mode=reactive,
harness_name=subprocess, local_command pointing to docgardener
agent, and harness_config_json.env carrying SYNAPBUS_AGENT,
SYNAPBUS_BIN, SYNAPBUS_SOCKET. Specialists are spawned
dynamically by the coordinator at runtime (not pre-registered),
so the demo exercises dynamic agent spawning end-to-end.
- examples/doc-gardener/run_task.sh: collapsed to a 3-line kickoff
that just DMs the coordinator and polls algis's inbox for the
coordinator's FINAL: reply. Everything else happens via the
reactor.
Verified end-to-end in Chrome on a fresh instance:
- 4 agents registered (coordinator + 3 specialists dynamically
spawned by the coordinator on receipt of the first DM)
- 7 reactive_runs + 6 harness_runs across the goal lifecycle:
algis → coordinator (kickoff, 624ms, builds goal+tree+spawns)
coordinator → docs-scanner (claim task 2)
coordinator → cli-verifier (claim task 3)
coordinator → drift-reporter (claim task 4)
docs-scanner → coordinator (DONE task=2)
cli-verifier → coordinator (DONE task=3)
drift-reporter → coordinator (DONE task=4, coalesced)
- Web UI Agent Runs page shows all 7 runs with the real
"DM from X" trigger lines and correct sender/receiver chain
- Goal ends at status=completed with all 3 leaf tasks at status=done
- Each specialist run posts a real subprocess artifact to the
goal channel (#finding, #verified, #summary) and appends a real
reputation_evidence row keyed by config_hash.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
ff5d0c49f4 |
feat(018): dynamic agent spawning — primitives + doc-gardener demo
Ships the MVP slice of spec 018 (dynamic agent spawning):
- 5 new SQLite migrations (021-025): goals + goal_tasks + agent_proposals
+ reputation_evidence + secrets + harness_runs.task_id. The legacy
`tasks` table (channel auctions) and `agent_trust` table (reactions
workflow) are left untouched — the new schema coexists.
- 4 new internal packages, fully tested:
- internal/goals: Goal struct + store + service, slug collision dedup,
backing-channel auto-create via ChannelCreator adapter
- internal/goaltasks: goal_tasks table with denormalized 16 KB
ancestry snapshots, single-statement optimistic-lock atomic claim,
recursive-CTE cost rollup, state machine, per-billing-code rollup
- internal/secrets: NaCl-secretbox encrypted blobs, user/agent/task
scope precedence, sanitized env injection, master-key bootstrap
- internal/trust additions: ConfigHash (deterministic SHA-256 of
model + prompt + tools + skills + mcp + subagents, sorted),
DelegationCap (tier + tool-scope + budget + depth enforcement),
append-only Ledger with exponential time-decay rolling score and
70%-of-parent child seeding. Existing trust package unchanged.
- Critical invariants under test:
- 50-goroutine concurrent claim race → exactly one winner per round
- ConfigHash stable under shuffled array inputs, sensitive to
capability changes
- DelegationCap full tier × tool-scope matrix
- Ledger time-decay + parent seed at 70 % ± 1 %
- Secret name sanitization, scope precedence, plaintext never
returned via MCP-equivalent paths
- internal/agents/types.go extended with dynamic-spawning columns
(config_hash, parent_agent_id, spawn_depth, system_prompt,
autonomy_tier, tool_scope_json, quarantined_at). Existing tests
still pass.
- cmd/docgardener: self-contained demo binary driving the end-to-end
flow. `docgardener run` creates a goal, builds a task tree with
denormalized ancestry, spawns 3 specialists (each going through
real delegation-cap validation and config-hash computation and
70 %-of-parent reputation seeding), claims tasks atomically, runs
them through the state machine, records reputation evidence.
`docgardener report` queries all of that back out and renders a
rich dark-mode HTML report (header, spend metrics, task tree,
spawned-agent cards with reputation bars, cost breakdown, artifacts,
timeline).
- examples/doc-gardener: start.sh / run_task.sh / report.sh / stop.sh
mirroring the cold-topic-explainer pattern. Launches an isolated
synapbus instance on port 18089, drives the demo, renders
report.html, cleans up. Full README documenting what's real vs
deferred, plus examples/README.md listing both examples.
- specs/018: tasks.md updated with MVP completion status; legacy tasks
naming collision noted.
Deferred (marked explicitly in example README):
- Real LLM-driven coordinator (needs MCP tool wiring + prompt
iteration)
- Real subprocess runs (needs reactor integration with task_id on
ExecRequest)
- Full MCP tool surface (contracts are written at
specs/018-dynamic-agent-spawning/contracts/mcp-tools.md)
- Svelte /goals UI (REST endpoints remain a follow-up)
- Full budget race + quarantine auto-trigger wiring
- Full resource-request → secrets fulfill reaction-workflow path
Cross-compiles clean for linux/amd64 and darwin/arm64 with no CGO
(SC-010). All new package tests pass (SC-004, SC-005, SC-007).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|