bfb2551b451a2b909d3e028b040ac6a41eceff0b
204
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bfb2551b45 |
feat(020): configurable dream parallelism + 3 bug fixes from kubic drain
Drain-on-demand: SYNAPBUS_DREAM_PARALLEL (default 1) and `synapbus memory dream-run --parallel N` fan out N concurrent dream-agent k8s Jobs per (owner, job_type) in one shot. Set high (e.g. 8) to drain backlog quickly, then back to 1 for normal hourly operation. Schema: - migration 030_dream_parallelism: adds slot INTEGER NOT NULL DEFAULT 0 to memory_consolidation_jobs. Drops + recreates the partial unique in-flight index as (owner, job_type, slot) so slots 0..N-1 each hold one in-flight job independently. Stores: - JobsStore.CreateOnSlot + CreateNextAvailableSlot. - ConsolidatorWorker.ForceRunN dispatches N parallel jobs through the existing launchOne path (extracted from ForceRun). - core_rewrite coerces to N=1 regardless of the knob — per-(owner, agent) blob is wholesale-replace and concurrent rewrites would race. Three bug fixes discovered while bringing the parallel path up on kubic: 1. k8s Job names collided on rapid relaunch because runner.go used "synapbus-<agent>-<msg_id>", and dream dispatches have msg_id=0. Now appends a unique (timestamp%1e6, 4-byte random) suffix when msg_id is zero; historical "synapbus-<agent>-<id>" prefix preserved. 2. memory_list_unprocessed didn't actually exclude already-refined messages — the contract said it should, the implementation returned the same oldest-50 every cycle. The dream agent kept re-refining the same set: 221 refines links touched only 55 unique dst messages, so progress flat-lined. Added the NOT IN (refines/duplicate_of/superseded_by) filter and a from_agent NOT LIKE 'dream:%' clause so the agent never refines its own reflections. 3. The k8sjob harness was constructed with nil Waiter in main.go, so every dream dispatch failed instantly with "k8sjob: no Waiter configured". Now builds a ClientsetWaiter from the in-cluster clientset. Plus admin/server.go gets DreamRunN closure + DefaultDreamParallel (sourced from MemoryConfig.DreamParallel). admin/socket.go handleMemoryDreamRun accepts `parallel` arg and returns job_ids[]. CLI admin command grows --parallel N flag. Live evidence from kubic (image v0.21.0-amd64): 1 CLI call with --parallel 8 produced 8 job rows on slots 0..7, spawned 8 distinct k8s Jobs with unique suffixes, retired ~86 unprocessed messages in <1 min (vs ~10/cycle for the buggy serial version pre-fix-2). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
1d894ec5aa |
fix(020): wire dispatch-token + claude config bridge
Two short fixes discovered while bringing the dream-claude agent live on kubic against a Max20 subscription: 1. internal/mcp/server.go: HTTPContextFunc now reads X-Synapbus-Dispatch-Token from request headers and stuffs it into ctx via WithDispatchToken. The memory_* tools already expected it in context; the bridge was missing on the HTTP boundary. Without this, every memory_* call returned dispatch_token_missing — which is the failure the live dream-agent hit on first run. 2. Followed searcher's proven pattern for Max20 OAuth: the k8sjob harness already auto-mounts /home/user/.claude → /app/.claude; the agent record just needs CLAUDE_CONFIG_DIR=/app/.claude in k8s_env_json. Documented for future agents in the k8s-job template, no code change needed here. Live evidence (kubic, image v0.20.7-amd64): Job 2015 reflection: status=succeeded, 3 actions, 2 new reflection memories (ids 32111, 32112) on reflections-mcpproxy, each synthesizing 14 source memories. End-to-end functional. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
069a985af5 |
feat(020): 14d window + token-budget circuit breaker + dream-agent + dashboard
Backend (Go, in this commit):
- migration 029_memory_dream_usage: per (date, owner) counters for
tokens_in/out, jobs_started/succeeded/failed/circuit_broken
- DreamUsageStore + UsageGate (internal/messaging/dream_usage.go).
Gate inspects today's usage against new env knobs:
- SYNAPBUS_DREAM_RECENT_WINDOW (default 336h / 14d)
- SYNAPBUS_DREAM_DAILY_TOKEN_LIMIT_IN (default 1M)
- SYNAPBUS_DREAM_DAILY_TOKEN_LIMIT_OUT (default 200k)
- SYNAPBUS_DREAM_DAILY_JOB_LIMIT (default 100)
- Consolidator now bounds watermarks + core_rewrite eligibility by the
recency window. core_rewrite skipped for owners with no in-window
activity. ForceRun honors the breaker.
- Recency fallback in BuildContextPacket + memory_list_unprocessed now
accept RecentWindowDays so injection and dream queries see the same
14d slice.
- Prometheus metrics registered (internal/metrics/metrics.go):
synapbus_dream_jobs_total{owner,job_type,status},
synapbus_dream_tokens_total{owner,direction},
synapbus_dream_job_duration_seconds{owner,job_type},
synapbus_dream_circuit_broken_total{owner,reason},
synapbus_injection_packets_total{tool},
synapbus_injection_memories_per_packet{tool},
synapbus_injection_packet_chars{tool},
synapbus_injection_skipped_total{tool,reason}.
- deploy/kubic/deployment.yaml: liveness/readiness timeoutSeconds: 1→5
(root-causes the "connection refused" mcpproxy errors at 13:02 today —
/readyz occasionally exceeded 1s under dream-worker tick load, so the
pod fell out of the Service endpoints intermittently).
Dream-claude agent (Python, in /dream-agent/):
- dream_runner.py uses claude-agent-sdk 0.1.48 to drive Claude Code
against SynapBus's MCP server. MCP transport carries
Authorization: Bearer <api_key> AND X-Synapbus-Dispatch-Token from env
via the SDK's McpHttpServerConfig.headers field — confirmed supported.
- Tools restricted via allowed_tools to mcp__synapbus__memory_*.
- Final JSON envelope reports tokens_in/out so harness.Usage stays
populated and the circuit breaker can count consumption.
- Dockerfile builds linux/amd64 at 189 MB, mirroring searcher's
agents/universal recipe.
- k8s-job-template.yaml: backoffLimit 0, ttl 600s, 512Mi/1CPU,
Anthropic credentials via secret-ref.
Grafana dashboard (deploy/kubic/grafana/):
- dream-dashboard.json — 14 panels across 5 rows (dream activity,
token usage vs limit, circuit breaker, injection layer, MCP
transport health), all templated to ${DS_PROMETHEUS}.
- import.sh: resolves the cluster's Prometheus DS uid and POSTs the
dashboard via Grafana API.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
5a5794beb1 |
fix(020): inject relevant_context on every eligible tool call
Two bugs discovered while verifying on kubic: 1. HTTP→MCP context propagation lost the full *agents.Agent struct (only the bare name was copied via ContextWithAgentName). The injection wrapper called agents.AgentFromContext and got nothing, silently skipping the packet. Now propagate both: full struct for middleware that needs the owner_id, name kept for backward-compat. 2. BuildContextPacket gated retrieval on `query != ""`, so my_status — the highest-value injection target — always returned a nil packet. FR-009 actually says "use recent owner activity as the implicit query" in that case. Added recentMemoriesForOwner: a direct SQL query over memory channels filtered by author owner, sorted by id DESC. New search_mode "recent" surfaces the fallback path in the packet so clients can tell it apart from semantic/fulltext. Live verification on kubic (image v0.20.3-amd64): - my_status as `claude-code` now returns relevant_context with 2 memories from algis-owned agents, packet sized exactly at the 500-token budget. - memory_injections audit ring captures each packet (research-mcpproxy → search → 1 item; claude-code → my_status → 2 items). - Cross-owner SC-008 holds: the recency query filters on agents.owner_id, so an unrelated owner's agent sees nothing. Dream worker autonomously fired 31 consolidation jobs (4 types × 2 owners × periodic ticks); all failed at the harness step with "k8sjob: no Waiter configured" — expected, the claude-code agent has no k8s_image set. Dispatch chain itself works end-to-end (job row → token → harness.Execute → audit). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
2044b199b8 |
feat(020): US3 — dream worker + 6 MCP consolidation tools
Background ConsolidatorWorker dispatches consolidation work to a Claude Code agent through harness.Harness.Execute (NOT via system DMs — per feedback_system_dm_no_trigger.md) with a one-time 15m dispatch token. The dispatched agent uses six new MCP tools, all token-gated and recording every action to memory_consolidation_jobs.actions JSON. Stores: - memory_links.go (+ test): typed edges with actor-prefix reserved-type guard; AddConsolidationLink bypass for memory_mark_duplicate / memory_supersede (their contractual writers). - memory_pins.go (+ test): owner pin overlay, bypasses score floor. - memory_status.go: queries the memory_status view. - consolidation_jobs.go: Create / Dispatch / Lease / AppendAction / Complete with ErrJobAlreadyInFlight via partial unique index. - auto_links.go: MessageListener generating mention / reply_to / channel_cooccurrence links automatically on send. Worker: - consolidator.go (+ test): ticker pattern modeled on StalemateWorker. Watermark trigger for link_gen / dedup_contradiction; daily 03:00 UTC for sleep-time core rewrite. Wallclock budget via harness Budget. Global semaphore gates concurrent owners. Mocked-harness test asserts no system DM is ever sent. - consolidator_prompts.go: four job-type prompts passed via env to the dispatched agent. MCP tools (internal/mcp/memory_tools.go + test): - memory_list_unprocessed, memory_write_reflection, memory_rewrite_core, memory_mark_duplicate, memory_supersede, memory_add_link. - Full error-code matrix tested per contracts/mcp-memory-tools.md. - Registered only when SYNAPBUS_DREAM_ENABLED=1. Injection extensions: - search/injection.go: pin overlay applied after retrieval; status filter drops soft_deleted / superseded unless pinned. New PinProvider, StatusProvider, MessageLookup hooks on InjectionOpts. Wiring: - cmd/synapbus/main.go: stores constructed, AutoLinkListener attached to MessagingService, mcpSrv.SetDream wired, ConsolidatorWorker start/stop, admin DreamRun closure. - cmd/synapbus/admin.go: synapbus memory dream-run --owner --job socket-RPC command (forces a single job bypassing trigger). Cycle workarounds (documented in code): - messaging.DreamAgent / HarnessDispatcher are local interfaces (the agents and harness packages import messaging, not the reverse). main.go wraps the real types via adapter structs. Stubbed: - Cron expression parsing (DreamDeepCron). Hardcoded daily 03:00 UTC. Adding robfig/cron deferred to keep no-new-deps. Pre-existing reactor test failures (5) are unchanged; confirmed pre-020 via stash check. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
a52d68ed88 |
feat(020): US2 — per-agent core memory blob
Letta-style identity blob, one per (owner, agent), always included in
session-start (my_status) responses. Replaces wholesale on rewrite; size
capped at SYNAPBUS_CORE_MEMORY_MAX_BYTES (default 2048); owner-scoped.
Components:
- internal/messaging/memory_core.go (+ test): CoreMemoryStore with
Get/Set/Delete/List, ErrCoreMemoryTooLarge, NewCoreProvider adapter
for search.CoreMemoryProvider.
- internal/mcp/server.go SetInjection: wires the core provider into
the my_status handler wrap.
- internal/mcp/injection_core_test.go: seed → wrapped my_status →
relevant_context.core_memory matches; missing row → no field.
- internal/api/memory_core.go + router: GET/PUT/DELETE
/api/owner/{ownerID}/agents/{agentName}/core-memory, session-auth,
413 on oversize.
- internal/admin/socket.go: memory.core.{get,set,delete} dispatch
handlers with username→user.id resolution.
- cmd/synapbus/admin.go: `synapbus memory core {get,set,delete}` cobra
subtree.
- cmd/synapbus/main.go: wires ParseMemoryConfig, CoreMemoryStore,
MemoryInjections; calls mcpSrv.SetInjection on startup.
Pin overlay still TODO (US3-T029).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
8a5d5e1f59 |
feat(020): US1 — proactive injection on MCP tool responses
Wraps eligible MCP tool handlers (my_status, send_message, search, execute; get_replies excluded as pure metadata) with a middleware that appends relevant_context to the JSON response. Retrieval reuses the existing search.Service hybrid pipeline; owner scoping filters out memories from other owners' agents (SC-008). Pin overlay is a marked TODO for US3. Components: - internal/search/injection.go (+ test): BuildContextPacket with token budget greedy fill, score floor, truncation flag, CoreMemoryProvider interface stubbed for US2. - internal/mcp/injection_wrap.go (+ test): WrapInjection middleware, registered via SetInjection on the existing handler. - internal/mcp/injection_e2e_test.go: adversarial cross-owner test asserts H1 cannot see H2's memories on any wrapped tool. - internal/messaging/memory_injections.go (+ test): 24h audit ring, hourly cleanup tick wired into stalemate worker. Discovery during impl: claim_messages/read_inbox/read_channel live as actions inside the execute bridge, not as registered top-level MCP tools. They inherit injection through the execute wrapper. This commit also bundles pre-existing working-tree changes for the 027 "remove approval noise" cleanup (migration 027, design doc, removal of reminder/escalate logic from stalemate worker, related trims in goals_tools.go and tools_hybrid.go). The two changes touch the same files (stalemate.go, tools_hybrid.go) and bundling them keeps history readable. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
da827c03b3 |
feat(020): foundational — migration 028, dispatch tokens, owner resolver
Adds the SQL substrate (6 tables + memory_status view) and the helpers every user story depends on: - migration 028_memory_consolidation.sql + smoke test - internal/messaging/memory_config.go (env-flag plumbing) - internal/messaging/dispatch_tokens.go (32-byte rand, 15m TTL, single-job-bound) - internal/messaging/memory_channels.go (open-brain / reflections-* / is_memory flag) - internal/agents/owner.go (OwnerFor with sentinel errors) Deviations from spec, all documented in code: - owner_id is stored as INTEGER FK to users; OwnerFor converts to the string scope-key the new tables use. - MemoryChannel is a local struct to avoid an import cycle between internal/channels and internal/messaging. - channels.metadata column does not exist yet; IsMemoryChannel honors it conditionally so MemoryChannelIDs can extend trivially when added. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
84973ad665 |
specs(020): proactive memory injection + dream worker
Owner-scoped memory the system pushes onto MCP tool responses, plus a background "dream" worker that dispatches consolidation jobs to a Claude Code agent through the existing harness seam. Memory pool reuses the messages table on memory-flagged channels; six new SQLite tables for core blob, links, audit, pins, dispatch tokens, and a 24h injection ring. Zero CGO, no new external deps. Includes: spec.md (4 user stories), plan.md, research.md (10 decisions), data-model.md, contracts/ (injection shape + 6 memory tools), quickstart.md, tasks.md (47 tasks across foundational + 3 stories + polish, with parallel-subagent cluster plan). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
a3d8681715 |
chore(deploy): replace stale helm chart with plain kubic manifests
The helm release went into failed state in March 2026 after an out-of-band
kubectl set image broke server-side-apply ownership; every deploy since has
been a direct kubectl set image, leaving the chart values drifting against
live state.
Drop deploy/helm/ entirely. Add deploy/kubic/{namespace,pvc,service,
deployment,secret.example}.yaml mirroring what's actually running, plus
scripts/deploy-kubic.sh encoding the build → docker save → scp → microk8s
ctr image import → kubectl set image flow used for v0.13.x-reactive through
v0.17.0. README documents why no helm and how to back up /data before
schema-touching versions.
kubectl diff -f deploy/kubic/ is empty against the live cluster.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
ed8a6da8b2 |
merge: plugin system framework (019)
Release / Build darwin/amd64 (push) Canceled after 0s
Release / Build linux/amd64 (push) Canceled after 0s
Release / Build darwin/arm64 (push) Canceled after 0s
Release / Build linux/arm64 (push) Canceled after 0s
Release / Generate Homebrew Formula (push) Canceled after 0s
Release / GitHub Release (push) Canceled after 0s
Release / Docker Image (push) Canceled after 0s
Release / Publish to MCP Registry (push) Canceled after 0s
Adds internal/plugin runtime, plugintest harness, demo plugin, and 103-task spec under specs/019-plugin-system. ~5k LOC, no overlap with messaging core. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>v0.17.0 |
||
|
|
8814e87e16 |
fix(messaging): non-destructive read_inbox + stalemate UPDATE race guard
read_inbox now requires explicit MarkRead (default false). Worker-queue callers opt in. Resolves bugs-synapbus #30674 where consecutive identical calls returned 0 the second time and produced inconsistent views with the claim/process/done loop and StalemateWorker. failTimedOutProcessing UPDATE now re-checks claimed_at < cutoff so a fresh re-claim between SELECT and UPDATE can't be stomped to failed. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
494e8f83e9 |
feat(plugin): plugin framework + plugintest + demo plugin + integration tests
Implements the compile-in plugin system designed in spec 019:
- internal/plugin/ (~1,300 LOC):
* Tiny Plugin interface + 10 optional HasX capability sub-interfaces
* Host struct with Logger / DB / Messenger / Channels / Attachments /
Search / Secrets / Events / Config / DataDir / Tracer / Metrics /
DefaultOwner / BaseURL
* Registry with panic-on-duplicate, stability-level tracking, capability
indexing (MCP tools, actions, panels, channel types, routes, event
subscribers, CLI commands)
* Migrator: per-plugin SHA-256-checksum'd migration chain, namespaced
plugin_<name>_* table enforcement, idempotent re-apply
* Three-phase lifecycle (Migrate → Init → Start) with panic-safe
wrappers around every plugin call; failure per plugin isolated,
core continues
* YAML config loader that preserves unknown top-level keys on round-trip
* Status store exposing /api/plugins/status JSON
* Restart helpers (Noop + SignalRestarter); graceful reload is
in-process for the demo
- internal/plugin/plugintest/ (~345 LOC):
* NopHost(t) with in-memory modernc.org/sqlite
* Run(t, plugin) full-lifecycle smoke helper
* Assertions: HasTool, HasAction, HasPanel, HasChannelType,
HasMigration, PluginStarted, PluginFailed
* ScopedSecrets that returns ErrSecretNotFound for cross-plugin
reads (satisfies SC-006)
- internal/plugins/demo/ (canonical showcase):
* Plugin that exercises every HasX capability (migrations, actions,
HTTP routes, web panel, lifecycle, config schema, stability)
* Own SQL migration creating plugin_demo_notes
* Embedded HTML panel that fetches notes via JS
* 4 unit tests covering smoke, full capability registration, action
handlers, and config-driven max_notes limit
- cmd/plugindemo/ (~290 LOC):
* Demo HTTP server wiring registry to chi
* Mounts /api/plugins/status, /api/admin/plugins/{name}/{enable,disable},
/api/actions/{name}, /api/plugins/<name>/* (per-plugin REST),
/ui/plugins/<name>/ (per-plugin UI)
* SIGHUP-triggered config reload + registry rebuild + mux swap
* SIGTERM/SIGINT graceful shutdown
- test/integration/ (~357 LOC, build-tag "integration"):
* 6 end-to-end tests against a spawned plugindemo binary
* Enable/disable round-trip with data preservation
* SIGHUP reload timing (measured 41 ms — SC-008 target is 2 s)
* Action-404 on disabled plugin, panel-404 on disabled plugin
* REST endpoints + UI panel reachable
Contract deviation: admin toggle endpoints moved from
/api/plugins/{name}/{enable,disable} to /api/admin/plugins/{name}/{...}
to avoid URL collision with chi per-plugin route mounts. rest.md updated.
Scope deferred to next session (mechanical follow-ups):
- Port internal/wiki/ to internal/plugins/wiki/
- Squash 26 migrations to schema/000_initial.sql
- Backup scripts for live kubic instance
- Remaining 9 plugin extractions
- Boundary-lint static analyzer
- Wire into cmd/synapbus/main.go
All unit + integration tests green. Chrome UI smoke test passes.
autonomous_summary.md carries the full verification record.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
89aaa51f71 |
tasks(019): 103-task execution plan organized by user story
Setup (3) + Foundational (15) + US4 backup (5) + US1 toggle (5) + US3 wiki extraction (12) + US5 failure isolation (4) + US2 author docs (6) + Verification (9) + Polish (5). MVP = Setup + Foundational + US3 + US1. Parallel opportunities marked [P] within each story. All tasks follow checklist format with concrete file paths. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
69cd13dce1 |
plan(019): plugin system plan + research + data-model + contracts + quickstart
Phase 0 research resolves all 12 open decisions (interface shape, registration, host API, dynamic toggle, migrations, UI panel integration, config format, testing, boundary enforcement, squash, failure notification, integration test). Phase 1 artifacts: data-model.md (Plugin, Registry, Migration, Host, Status, Backup), contracts/plugin.md (Plugin + HasX interfaces), contracts/host.md (Host struct + plugintest constructor), contracts/rest.md (/api/plugins/*), quickstart.md (end-to-end "hello" plugin in 8 steps). All ten constitution gates pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
b021768a9a |
spec(019): plugin system for SynapBus core
Compile-in plugin framework with tiny Plugin interface + HasX capability sub-interfaces, typed Host struct, explicit registration, SIGHUP graceful restart, three-phase boot, per-plugin failure isolation, plugintest helpers. Scope: Phase 0 backup+squash, Phase 1 framework plumbing, Phase 2 extract wiki as canonical pilot plugin. Remaining 9 extractions are follow-up. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
121d05f875 |
feat(docker): auto-stage host OAuth credentials into agent containers
The docker harness now detects and stages host CLI auth files (~/.gemini/oauth_creds.json, ~/.claude/.credentials.json) into a writable agent-home directory mounted at /home/agent. This lets containerized agents reuse the host's Gemini Pro / Claude Pro OAuth sessions without manual secret management or API keys. Only auth files are copied — not the host's settings.json or MCP configs (which contain stale localhost URLs that would hang Gemini CLI inside containers). The staged dir is writable so CLIs can create projects.json, history, etc. alongside the auth files. Also sets GEMINI_DEFAULT_AUTH_TYPE=oauth-personal and GEMINI_CLI_NO_RELAUNCH=true when OAuth creds are detected, writes Claude's hasCompletedOnboarding flag, and simplifies the doc-gardener example to use the harness-level credential staging instead of manual HOME directory seeding. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
bd54aac82f |
fix(synapbus-agent): early-exit on __coalesced__ synthetic triggers
When the reactor sets pending_work while an agent run is in progress,
checkPendingWork fires a synthetic follow-up event with FromAgent=
__coalesced__ and body="Coalesced trigger: process all pending
messages." That event gets delivered to the container as a
message.json with that placeholder body, and the wrapper happily
feeds it to gemini — which then produces a spurious "What does this
demo do?" reply because the only thing the model sees is a generic
filler body.
The doc-gardener coordinator kept emitting stray replies to algis
between real DELEGATED: messages because of this. The reactor would
fail the coalesced run ("Received system trigger..."), the critic
would get confused by intermediate traffic, and the /goals panel
would accumulate garbage.
Wrapper now checks $FROM at the top of main. If it's __coalesced__
we log it and exit 0 without invoking the CLI. The reactor marks
the run succeeded, no tokens burned, no spurious DMs produced. Any
real pending work re-triggers naturally when the next actual
message arrives.
Smoke-verified:
docker run --rm -v /tmp/test:/workspace synapbus-agent:latest \
/usr/local/bin/synapbus-agent-wrapper.sh
[wrapper test] cli=gemini from=__coalesced__ body_bytes=9
[wrapper test] synthetic coalesced trigger — skipping CLI invocation
EXIT=0
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
42f8256df6 |
feat(goals): complete_goal MCP tool + draft→active auto-transition
Three improvements that turn the doc-gardener demo from "runs but stays in 'draft' forever" into a goal that properly transitions through its lifecycle and renders a completion summary on /goals/<id>. ### 1. complete_goal MCP tool (#59, #62) New tool surface: complete_goal(goal_id, status, summary, completion_message_id?) The critic calls this from inside the sandbox after it sends its FINAL: DM. Records the one-paragraph human-readable summary on the goal row plus a pointer to the message that carried the FINAL text, so the Web UI /goals/<id> page has both the verdict and a deep link to the full findings JSON. Status parameter accepts completed | stuck | cancelled. Idempotent when called with the current status. Rejects callers owned by a different human than the goal owner. Plumbing: - New migration 026_goals_completion_summary.sql adds two columns to goals: completion_summary TEXT, completion_message_id INTEGER (FK messages.id, ON DELETE SET NULL). - internal/goals/types.go: new CompletionSummary + CompletionMessageID fields on Goal struct. - internal/goals/store.go: Get/List Scan both new columns; SetCompletion(goalID, status, summary, messageID) helper that updates status+summary+message_id atomically and populates completed_at for terminal states. - internal/goals/service.go: Complete(ctx, goalID, status, summary, messageID) wraps the store method with legalTransition gating. legalTransition expanded so draft can jump straight to completed (no mandatory "active" hop required). - internal/mcp/goals_tools.go: completeGoalTool definition + handleCompleteGoal handler. Tool count 6 → 7. - internal/api/goals_handler.go: surfaces completion_summary, completion_message_id, and completed_at on both list and detail endpoints so the Svelte /goals UI can render them. ### 2. Draft → active auto-transition in propose_task_tree (#60) handleProposeTaskTree now flips the goal from draft to active at the end. Previously the coordinator would call create_goal + propose_task_tree and dispatch inspector, but the goal stayed in draft forever because nothing transitioned it. Now the mere fact of having a task tree means the goal is active. Safe: the transition is best-effort and ignores the legal-transition error when the goal is already beyond draft. ### 3. REVISE round cap (#61) Two-layer enforcement: - Server-side: examples/doc-gardener/start.sh drops max_trigger_depth from 8 to 4. Each REVISE round costs 2 hops (critic→inspector + inspector→critic), so depth=4 caps the loop at roughly 2 rounds before the reactor refuses further dispatches. - Prompt-side: inspector now includes revision_round (starting at 0, incremented when it sees a REVISE: input) in its findings JSON. Critic reads revision_round and force-FINALs when >= 1. Prompt explicitly tells the critic to call complete_goal after sending FINAL, so the goal row gets a proper completion_summary. ### 4. run_task.sh terminal-state detection Rewrote the poll loop to watch goals.status/completion_summary as the definitive "done" signal rather than parsing DM bodies. Keeps a message-based fallback for TRIVIAL/CANNOT paths that don't create a goal. Treats "Received system trigger..." and "Coalesced trigger..." as informational (they're `__coalesced__` reactor synthetic events leaking through the coordinator reply, not real user-facing output). Bare coordinator replies are terminal only when no goal was created. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
b5b5d6d26c |
fix(reactor): don't truncate body for subprocess + docker dispatch
dispatchHarness was capping event.Body at 4096 bytes before handing it to the subprocess/docker harness. Those backends write the body to a message.json file in the per-run workdir (bind-mounted into the container) and have no shell/env-var size limits, so silent truncation was hostile. The doc-gardener inspector routinely produces 10-20 KiB findings JSON (drift report with per-flag evidence). Truncation cut off the trailing artifact.findings entries + artifact.recommendation, making the report look incomplete to the critic — which then spuriously REVISE'd, blowing the 600s deadline. The K8s job path still truncates in createJob() because Kubernetes imposes a 1 MiB env-var cap and most shells misbehave past a few KiB. That's a separate code path, untouched. Also: critic prompt rewrite (examples/doc-gardener/configs/critic.json). The old critic spec told the critic to "spot-check evidence by re-running the inspector's commands". That's structurally wrong: the critic runs in a fresh container with no install state, so re-running mcpproxy --help always fails and produces a false REVISE. New prompt says: audit by structural consistency only, never run shell commands to re-verify, default to FINAL, never REVISE more than once, and FINAL the failure summary back to the owner when the inspector reports status: failed. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
ddbb1bd329 |
fix(docker): tmpfs /tmp mounted rw,exec,size=128m
Docker Desktop's default noexec on tmpfs broke the "download a CLI to /tmp, chmod +x, run it" workflow — exactly what the doc-gardener inspector needs to verify docs.mcpproxy.app against the real mcpproxy binary. Previously the agent spent ~10 minutes in a self- debug loop discovering the noexec, falling back to /home/agent, running into externally-managed Python, missing python3-venv, etc. With /tmp exec, the inspector's own install pipeline works on the first try: curl | tar | chmod | run. First real run produced a 72-claim drift report (21 matched / 1 drifted / 50 missing) against mcpproxy v0.24.4 in ~8 minutes, no REVISE loop. The 64m → 128m bump gives breathing room for curl'd tarballs that need a temp extraction directory alongside the final binary. Inspector prompt updated to tell the agent about the /tmp install path explicitly and forbid the previous /home/agent detours. Also updated the coordinator brief template to match. Note the image itself is UNCHANGED — we deliberately do NOT bake mcpproxy (or any other domain-specific tool) into synapbus-agent. The image stays a blank Linux shell with Node + Python + core tools, and each example's prompt teaches its agent how to install whatever it needs. This keeps the gardener universal: swap in any other docs domain and the inspector figures out what to install on demand. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
bdec096190 |
feat(doc-gardener): default coordinator to gemini-3.1-pro-preview
Align with the goal-coordinator example — both now use Gemini 3 Pro as the default for triage/coordination. Workers stay on 2.5 Flash. Override via SYNAPBUS_COORDINATOR_MODEL when the preview model is rate-limited. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
f1e2b1fa38 |
feat(doc-gardener): MCP-native + docker-isolated multi-agent demo
Replace the legacy cmd/docgardener orchestration (~2400 LOC of Go
spawning subprocess workers via local_command + admin socket) with
three Docker-isolated agents that all reach SynapBus through MCP:
doc-coordinator — Gemini Pro, triages goal, calls create_goal +
propose_task_tree + send_message via MCP
docs-inspector — Gemini Flash, fetches docs, installs mcpproxy,
shells out to verify, reports findings via MCP
docs-critic — Gemini Flash, independent reviewer with its
own MCP API key + config_hash, audits the
inspector's evidence and DMs the owner
Every agent runs inside synapbus-agent:latest with --cap-drop=ALL,
--security-opt=no-new-privileges, --read-only root + tmpfs /tmp,
--pids-limit, memory + CPU quotas. The container reaches the
SynapBus MCP server on the host at host.docker.internal:18089
because the docker harness rewrites .gemini/settings.json URLs
from 127.0.0.1 automatically.
Wrapper baked into the image at /usr/local/bin/synapbus-agent-wrapper.sh
so configs don't need to mount or template a per-example wrapper.
The harness's default no longer overrides docker CMD — the image's
baked entry script is used unless docker.command is set explicitly.
start.sh changes:
- Preflight: docker daemon, GEMINI_API_KEY (or ~/.gemini/oauth_creds.json)
- Builds synapbus-agent image lazily on first run
- Mints one MCP API key per agent via `agent revoke-key`
- Templates each config with __PORT__, __*_APIKEY__, __MODEL__,
__GEMINI_API_KEY__, __EXTRA_MOUNTS__
- With OAuth fallback: copies host ~/.gemini → data/agent-home/.gemini
once and bind-mounts the whole agent-home rw at /home/agent so
in-container gemini has a writable HOME without polluting the host
- SYNAPBUS_KEEP_WORKDIR=1 preserves per-run docker workdirs for
debugging
- Sets harness_name=docker explicitly so the resolver picks the
right backend even with empty local_command
stop.sh: best-effort cleanup of lingering synapbus-* containers so a
killed parent doesn't leave bind-mount holders that block the next
start.sh from re-mounting the same paths.
run_task.sh: snapshot-baseline pattern (only watches replies newer
than the max msg id at send time), 600s deadline, treats any reply
from doc-coordinator that isn't DELEGATED:/REVISING: as terminal,
plus FINAL:/CANNOT: from any sender.
cmd/docgardener slimmed from 7 files / 2580 LOC to 3 files / ~370 LOC.
The remaining binary only renders the HTML report (queries goals +
goal_tasks + traces + harness_runs from the SynapBus DB read-only).
agent.go, channels.go, flow.go, gemini_tree.go all deleted.
Verified end-to-end against gemini-2.5-pro coordinator + gemini-2.5-flash
workers (with OAuth fallback mount):
./run_task.sh "what does this demo do?"
→ coordinator TRIVIAL: replies directly via MCP send_message
./run_task.sh "Verify the CLI commands on docs.mcpproxy.app/cli/command-reference"
→ coordinator calls create_goal (slug verify-mcpproxy-cli-...),
propose_task_tree (3-node tree: coordinator/plan,
doc-gardener/scan, doc-gardener/audit) and send_message to
docs-inspector
→ inspector container runs ~10 minutes inside the sandbox:
installs mcpproxy from real release URL (linux-arm64), curls
the docs page, falls back from BeautifulSoup → grep when
python3-venv is missing, debugs its own f-string syntax, writes
extract_flags.py, runs `mcpproxy --help` for ground truth
→ real multi-agent iteration loop: critic REVISE: → inspector
retry → critic REVISE: with new feedback
The agents discovered real environment quirks (tmpfs noexec on /tmp,
externally-managed Python, missing python3-venv) and worked around
them inside the sandbox without touching the host.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
560d9d4125 |
feat(harness): docker isolation backend + canonical synapbus-agent image
New `internal/harness/docker` package: per-run ephemeral container
backend that runs each agent in `docker run --rm`, bind-mounts the
materialized workdir at /workspace, and captures stdout/stderr/exit
code/result.json the same way the subprocess backend does.
Inspired by scion's pkg/runtime/docker.go: shell out to the docker
CLI (zero new Go deps, zero CGO), per-task ephemeral containers with
no warm pool, host-side scratch dir bind-mounted in.
Default security posture (overridable per agent):
--rm
--cap-drop=ALL
--security-opt=no-new-privileges
--read-only with tmpfs /tmp
--pids-limit=512
--user=<host uid:gid>
--network=bridge (configurable; --network=none for air-gap)
--memory / --cpus from agent config
--add-host host.docker.internal:host-gateway on Linux
The backend reuses subprocess.AgentConfig for gemini_md/claude_md/
mcp_servers/skills materialization so existing example configs work
unchanged. Per-agent docker tunables go under a new `docker` block
in harness_config_json: image, memory, cpus, network, extra_mounts,
cap_add, read_only_root, user, entrypoint, command, extra_args.
MCP host rewrite: `.gemini/settings.json` URLs of the form
http://127.0.0.1:<port>/mcp are rewritten to
http://host.docker.internal:<port>/mcp at materialization time so the
in-container Gemini CLI can reach the SynapBus MCP server on the host
without code changes in the example wrappers.
Wired into the reactor and Registry resolver:
- Registry.Resolve picks "docker" when harness_config_json contains a
`"docker"` block, taking precedence over local_command so explicit
isolation never silently downgrades.
- reactor.agentBackendKind() returns backendDocker for the same case.
- evaluateTrigger's harness-backend gate accepts backendDocker
alongside subprocess + webhook.
- main.go registers docker.Harness with the harness registry, passing
the SynapBus listen port so the URL rewrite uses the correct host
port.
Smoke tests in docker_test.go (skipped when no docker daemon):
- TestExecute_Hello: env injection + bind-mount writeback + message.json
+ result.json + stdout capture using alpine:3.20
- TestExecute_NoImage: rejects agents missing docker.image
- TestExecute_TimeoutCancel: wall-clock budget kills the container
New canonical agent image at image-build/synapbus-agent/:
- Debian bookworm-slim base
- Node 22 + @google/gemini-cli + @anthropic-ai/claude-code
- jq, sqlite3, curl, git, python3, tini (PID 1 for signal forwarding)
- Non-root agent user uid/gid 1000
- ENTRYPOINT tini, CMD /workspace/wrapper.sh
No SynapBus binary inside the image — agents reach the host MCP server
over the network at host.docker.internal:<port>.
Pre-existing reactor test failures (TestReactorNoK8sImage,
TestReactorDepthExceeded, TestReactorBudgetExhausted,
TestReactorCooldownSkipped, TestReactorSequentialExecution) verified
to exist on
|
||
|
|
f319290ef9 |
feat(goal-coordinator): native MCP tool surface via Gemini session
The coordinator now reaches SynapBus's MCP endpoint directly from inside the Gemini session. wrapper.sh's coordinator branch is a pure pass-through — no more JSON-plan parsing. When the coordinator runs, Gemini connects to /mcp with the coordinator's own Bearer API key and calls `create_goal`, `propose_task_tree`, and `send_message` as native tools. Goal rows, task trees, and DMs all land in the DB in one in-session flow. - start.sh mints a fresh API key for goal-coordinator via `agent revoke-key` and substitutes it into configs/coordinator.json (plus the port) at apply_config time. - coordinator.json declares the synapbus MCP server in mcp_servers; the subprocess harness already writes .gemini/settings.json from that array, so gemini picks it up automatically. - GEMINI.md rewritten to instruct the model to call MCP tools instead of emitting a JSON action blob. Stdout is explicitly discarded; every reply goes through send_message. - wrapper.sh coordinator branch is ~15 lines: invoke gemini, log, exit. Inspector + critic keep the legacy JSON-plan pattern since they're workers with fixed contracts. - SYNAPBUS_KEEP_WORKDIR=1 preserves per-run workdirs for debugging MCP traces, gemini output, and materialized configs. - Reactor checkPendingWork now fires after subprocess run completion (previously only K8s poller hit this path). The synthetic coalesced trigger uses a `__coalesced__` sentinel instead of `system` so it bypasses the FromAgent=="system" dispatch guard. Verified e2e (with rate-limit-induced retries): - TRIVIAL: "what is 2+2?" → coordinator send_message(algis, "4") - INFEASIBLE: "Transfer \$50…" → coordinator send_message(algis, "CANNOT: …") - SINGLE-STEP: 3-node task tree materialized in goal_tasks, TASK JSON forwarded to generic-inspector → critic-auditor chain. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
40706a22ce |
feat(example): universal goal-coordinator with triage + delegation
New examples/goal-coordinator/ demonstrates an LLM-driven coordinator that classifies any natural-language goal into one of four paths: - TRIVIAL → coordinator answers directly (no delegation) - INFEASIBLE → coordinator refuses with a concrete reason - SINGLE-STEP → delegate to one inspector + one critic (the default) - MULTI-STEP → multi-phase plan (rare) Architecture: - goal-coordinator (Gemini 3.1 Pro) triages and delegates - generic-inspector (Gemini 2.5 Flash) does scan+verify+report in one pass (shared context, no artificial splitting) - critic-auditor (Gemini 2.5 Flash) reviews the artifact with its own config_hash → independent reputation, no shared reasoning trace → can't rationalize the worker's mistakes Harness-agnostic via the existing subprocess harness + a wrapper.sh that calls `gemini`. Swapping to claude / codex is a 3-line change in the call block — nothing in SynapBus itself is tied to a CLI. Verified e2e on gemini-3.1-pro-preview: - "what is 2+2?" → TRIVIAL, direct "4" reply - "check Go version >= 1.23" → SINGLE-STEP, 3 runs, FINAL: "The installed Go version (1.25.1) meets the specified requirement" - "transfer $50 from my bank account" → INFEASIBLE, CANNOT: refusal citing missing credentials 7 runs visible in /runs with captured prompt/response per run. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
4546bef955 |
fix(auth): split session/user reads onto read pool
Root cause of the Web UI wedge on /conversations/1 (and every other authenticated page when the reactor is busy): SQLiteSessionStore and SQLiteUserStore routed every read through the single-connection write pool. On every authenticated request RequireSession does a GetSession + GetUserByID — both hit the write pool, so each one queues behind every reactor / tracer / messaging write. Observed /api/conversations/1 returning 401 after 113 seconds and login POST timing out for 15+ seconds. - SQLiteSessionStore: new NewSQLiteSessionStoreWithRead that takes separate write + read handles. GetSession routes SELECTs through readDB; the last_active_at bump and expired-session cleanup now fire-and-forget on a background goroutine so HTTP handlers never wait on the write pool for a non-critical liveness poke. - SQLiteUserStore: same split. GetUserByID / GetUserByEmail / GetUserByUsername go through readDB. - main.go: wires db.QueryDB() (the query_only=ON read pool) into both stores via the new constructors. Verified: /api/conversations/1 now returns 200 in <2ms even while the coordinator subprocess is blocking on a long Gemini call. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
190df8119f |
fix(example): rebuild web dist in start.sh when missing
The binary embeds internal/web/dist via go:embed, but .gitignore only tracks index.html. A fresh clone has an empty dist, so the binary serves only the shell HTML + no _app JS — the channel page renders its skeleton placeholder forever because the SPA never loads. start.sh now detects an empty dist, runs `npm run build`, and copies web/build into internal/web/dist before `go build`. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
4e6f6cb539 |
feat(018): Gemini LLM coordinator + spec-018 MCP tool surface
- cmd/docgardener: coordinator calls the `gemini` CLI with the goal
brief when SYNAPBUS_GEMINI_MODEL is set, parses the returned JSON
into a goaltasks.TreeNode, and aligns leaf billing codes so the
fixed dispatch table still routes specialists correctly. Falls
back to the hardcoded template on any failure (missing CLI, non-
zero exit, bad JSON) so the demo still works offline.
- internal/mcp: new GoalsToolRegistrar exposing 6 spec-018 tools —
create_goal, propose_task_tree, propose_agent, claim_task,
request_resource, list_resources. All require an authenticated
agent context; wire-only changes on the MCP server side.
- main.go: builds + attaches the new registrar after the hybrid
tool registrar, logs the 6 tools at startup.
Verified e2e: demo run with Gemini produces an LLM-generated root
task title ("Verify and patch mcpproxy documentation drift"), all
3 specialists dispatched and completed, $1.05 cost rollup on the
/goals/1 page, and the MCP server registers 11 tools total (5
hybrid + 6 spec-018) at boot.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
f9d8f1a1c9 |
feat(018): /goals page, budget cascade, quarantine, secrets loop
- New /api/goals + /api/goals/{id} endpoints serving list + task tree
+ cost rollup + billing breakdown + spawned agents + timeline.
- New Svelte /goals and /goals/[id] pages with sidebar link.
- goals.Service.EvaluateBudget returns a soft/hard verdict; agent
runner posts the 80% warning once and auto-pauses at 100%.
- Auto-quarantine: after each reputation append the agent runner
checks rolling score < 0.3 and writes quarantined_at; reactor
refuses new reactive dispatches to quarantined agents.
- Reactor exposes SetSecretProvider; main.go wires secrets.Store
so reactive subprocess runs inherit user/agent-scoped env vars.
- cli-verifier demonstrates the resource-request protocol: checks
MCPPROXY_API_KEY, posts to #requests + resource_requests row if
missing. New `synapbus secrets set/list` CLI (direct-DB) closes
the loop.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
3b94fab226 |
feat(018): real reactor-driven multi-agent doc-gardener flow
Until now the doc-gardener example was a single monolithic
orchestrator binary writing synthetic messages directly to SQLite.
That's now obsolete: the feature runs as a true multi-agent flow
where the SynapBus reactor fires subprocess runs for every DM, each
agent is its own reactive subprocess invocation, and follow-up DMs
go through the real MessagingService.Send → dispatcher path so the
reactor picks them up.
Changes:
- cmd/synapbus/main.go: gate the three legacy background workers
(expiry, retention, stalemate) behind SYNAPBUS_DISABLE_*_WORKER env
flags. These workers manage the legacy channel task-auction /
message retention features the doc-gardener demo doesn't use, but
they held the single-connection write pool long enough to wedge
the whole server for interactive sessions. All three are disabled
in the example's start.sh.
- cmd/docgardener/agent.go (new): the per-agent subprocess entry the
reactor harness invokes for every reactive trigger. Reads
message.json from the workdir, routes by SYNAPBUS_AGENT to either
coordinator-kickoff, coordinator-completion, or specialist-work
logic. Writes prompt.txt + response.txt for harness capture. Uses
the admin socket (`synapbus messages send`) for follow-up DMs so
the real MessagingService.Send path fires the dispatcher.
- cmd/docgardener/main.go: adds `docgardener agent` subcommand, plus
helpers freshAPIKey / bcryptHash / absPath / selfPath used by the
spawn flow.
- examples/doc-gardener/start.sh: provisions user + coordinator
agent + algis human agent + approvals/requests channels; the
coordinator is created with trigger_mode=reactive,
harness_name=subprocess, local_command pointing to docgardener
agent, and harness_config_json.env carrying SYNAPBUS_AGENT,
SYNAPBUS_BIN, SYNAPBUS_SOCKET. Specialists are spawned
dynamically by the coordinator at runtime (not pre-registered),
so the demo exercises dynamic agent spawning end-to-end.
- examples/doc-gardener/run_task.sh: collapsed to a 3-line kickoff
that just DMs the coordinator and polls algis's inbox for the
coordinator's FINAL: reply. Everything else happens via the
reactor.
Verified end-to-end in Chrome on a fresh instance:
- 4 agents registered (coordinator + 3 specialists dynamically
spawned by the coordinator on receipt of the first DM)
- 7 reactive_runs + 6 harness_runs across the goal lifecycle:
algis → coordinator (kickoff, 624ms, builds goal+tree+spawns)
coordinator → docs-scanner (claim task 2)
coordinator → cli-verifier (claim task 3)
coordinator → drift-reporter (claim task 4)
docs-scanner → coordinator (DONE task=2)
cli-verifier → coordinator (DONE task=3)
drift-reporter → coordinator (DONE task=4, coalesced)
- Web UI Agent Runs page shows all 7 runs with the real
"DM from X" trigger lines and correct sender/receiver chain
- Goal ends at status=completed with all 3 leaf tasks at status=done
- Each specialist run posts a real subprocess artifact to the
goal channel (#finding, #verified, #summary) and appends a real
reputation_evidence row keyed by config_hash.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
2ba0f95666 |
feat(018): real subprocess runs in docgardener + agent trust UI
docgardener: each leaf task now launches a real subprocess via exec.CommandContext and records a full reactive_runs + harness_runs row chain with task_id populated, captured prompt, captured response, exit code, duration, tokens, cost. The Agent Runs page and /runs/:id detail page now show real data for the doc-gardener demo — including "What the model saw" and "What the model said" panels — without needing the coordinator LLM loop. agents store: agentSelectSQL and both scanAgent functions extended to read the feature-018 columns (config_hash, parent_agent_id, spawn_depth, system_prompt, autonomy_tier, tool_scope_json, quarantined_at, quarantine_reason). /api/agents and /api/agents/:name now return these fields end-to-end. Web UI agent detail (web/src/routes/agents/[name]/+page.svelte): adds a Trust & Spawn section (config_hash, autonomy tier, spawn depth, parent agent, tool scope chips) and a full-height System Prompt pre block. Rebuilt internal/web/dist/. Verified in Chrome against a fresh ./start.sh && ./run_task.sh run: - Agent Runs page lists 3 completed runs (docs-scanner, cli-verifier, drift-reporter) with task.claim event and non-zero durations - /runs/1 detail page renders captured prompt + structured #finding output with 12 flags - /agents/docs-scanner shows config_hash=a0b5c6538b2d…, parent=#1, depth=1, tool-scope chips, and the 170-char system prompt - #goal-... channel loads all 12 messages (no "Joining..." hang) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
ff5d0c49f4 |
feat(018): dynamic agent spawning — primitives + doc-gardener demo
Ships the MVP slice of spec 018 (dynamic agent spawning):
- 5 new SQLite migrations (021-025): goals + goal_tasks + agent_proposals
+ reputation_evidence + secrets + harness_runs.task_id. The legacy
`tasks` table (channel auctions) and `agent_trust` table (reactions
workflow) are left untouched — the new schema coexists.
- 4 new internal packages, fully tested:
- internal/goals: Goal struct + store + service, slug collision dedup,
backing-channel auto-create via ChannelCreator adapter
- internal/goaltasks: goal_tasks table with denormalized 16 KB
ancestry snapshots, single-statement optimistic-lock atomic claim,
recursive-CTE cost rollup, state machine, per-billing-code rollup
- internal/secrets: NaCl-secretbox encrypted blobs, user/agent/task
scope precedence, sanitized env injection, master-key bootstrap
- internal/trust additions: ConfigHash (deterministic SHA-256 of
model + prompt + tools + skills + mcp + subagents, sorted),
DelegationCap (tier + tool-scope + budget + depth enforcement),
append-only Ledger with exponential time-decay rolling score and
70%-of-parent child seeding. Existing trust package unchanged.
- Critical invariants under test:
- 50-goroutine concurrent claim race → exactly one winner per round
- ConfigHash stable under shuffled array inputs, sensitive to
capability changes
- DelegationCap full tier × tool-scope matrix
- Ledger time-decay + parent seed at 70 % ± 1 %
- Secret name sanitization, scope precedence, plaintext never
returned via MCP-equivalent paths
- internal/agents/types.go extended with dynamic-spawning columns
(config_hash, parent_agent_id, spawn_depth, system_prompt,
autonomy_tier, tool_scope_json, quarantined_at). Existing tests
still pass.
- cmd/docgardener: self-contained demo binary driving the end-to-end
flow. `docgardener run` creates a goal, builds a task tree with
denormalized ancestry, spawns 3 specialists (each going through
real delegation-cap validation and config-hash computation and
70 %-of-parent reputation seeding), claims tasks atomically, runs
them through the state machine, records reputation evidence.
`docgardener report` queries all of that back out and renders a
rich dark-mode HTML report (header, spend metrics, task tree,
spawned-agent cards with reputation bars, cost breakdown, artifacts,
timeline).
- examples/doc-gardener: start.sh / run_task.sh / report.sh / stop.sh
mirroring the cold-topic-explainer pattern. Launches an isolated
synapbus instance on port 18089, drives the demo, renders
report.html, cleans up. Full README documenting what's real vs
deferred, plus examples/README.md listing both examples.
- specs/018: tasks.md updated with MVP completion status; legacy tasks
naming collision noted.
Deferred (marked explicitly in example README):
- Real LLM-driven coordinator (needs MCP tool wiring + prompt
iteration)
- Real subprocess runs (needs reactor integration with task_id on
ExecRequest)
- Full MCP tool surface (contracts are written at
specs/018-dynamic-agent-spawning/contracts/mcp-tools.md)
- Svelte /goals UI (REST endpoints remain a follow-up)
- Full budget race + quarantine auto-trigger wiring
- Full resource-request → secrets fulfill reaction-workflow path
Cross-compiles clean for linux/amd64 and darwin/arm64 with no CGO
(SC-010). All new package tests pass (SC-004, SC-005, SC-007).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
d9d1b7fee2 |
spec(018): dynamic agent spawning — full design
Complete speckit spec for the feature: a coordinator-driven system where a human types a goal, a meta-agent decomposes into a task tree, proposes spawning specialist sub-agents, runs them on heartbeats, verifies their outputs, and iterates. Includes: - spec.md (9 user stories, 46 FRs, 12 SCs) - plan.md (constitution check PASS) - research.md (17 design decisions documented) - data-model.md (5 migrations) - contracts/mcp-tools.md (9 new MCP tools + 5 REST endpoints) - quickstart.md (10-minute runbook) - tasks.md (135 tasks, 12 phases, MVP at phase 8) - checklists/requirements.md (quality gates) Doc-gardener example is the end-to-end acceptance test. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
fee73e33a0 |
feat(ux): run detail page + reaction pills + captured prompt/response
Makes the Web UI reflect what agents are actually doing: reactions
on DMs that trigger a subprocess run, a per-run detail page that
shows the exact prompt the model received and the raw response, and
cross-linked reactive_runs ↔ harness_runs data for a single composite
API call.
Migration 020 (internal/storage/schema/020_harness_run_detail.sql):
ALTER TABLE harness_runs ADD COLUMN reactive_run_id INTEGER;
ALTER TABLE harness_runs ADD COLUMN prompt TEXT;
ALTER TABLE harness_runs ADD COLUMN response TEXT;
CREATE INDEX idx_harness_runs_reactive ON harness_runs(reactive_run_id);
internal/harness:
* ExecRequest.ReactiveRunID — reactor pins the reactive_runs row id
so the observer can JOIN the two tables.
* ExecResult.Prompt / Response — the subprocess harness reads
prompt.txt / response.txt that wrappers write into the workdir,
and runs.Store persists them (capped at 32 KiB each).
* runs.Run struct now has JSON tags — previously the API returned
PascalCase field names that didn't match the Web UI's snake_case
TypeScript types.
* New runs.Store.GetByReactiveRunID for the composite API endpoint.
* Test schema updated to include the new columns.
internal/reactor:
* New ReactionNotifier interface + SetReactionNotifier.
* dispatchHarness now reacts `in_progress` on the triggering DM
before spawning the goroutine.
* runHarness reacts `done` on success, `reject` on failure. The
existing reactionPriority ordering means the terminal reaction
wins for badge display — no need to remove in_progress first.
* dispatchHarness sets ExecRequest.ReactiveRunID.
cmd/synapbus/main.go:
* reactorReactionAdapter: adapts reactions.Service.Toggle to the
reactor's one-shot AddReaction signature.
* HarnessRunsStore wired into the API router config.
internal/api/runs_handler.go — GetRun composite endpoint:
The GET /api/runs/{id} response now returns everything the Web UI
needs to render the run detail page in one call:
{
"run": <reactive_runs row>,
"harness_run": <linked harness_runs row with prompt/response>,
"agent": <current agent snapshot with harness_config_json>,
"trigger_message": <DM that started the run>,
"outgoing_message": <first DM the agent produced after startedAt>
}
The outgoing-message lookup wraps both sides of the created_at
comparison in datetime() so SQLite parses the stored 'YYYY-MM-DD
HH:MM:SS' and the Go-emitted RFC3339 into the same canonical form
before comparing — a raw string compare was silently returning no
rows.
internal/api/router.go: HarnessRunsStore field in RouterConfig, wired
through to NewRunsHandler.
examples/cold-topic-explainer/wrapper.sh:
Writes prompt.txt and response.txt alongside gemini.stdout.raw so
the subprocess harness can capture "what the model saw" and "what
the model said" post-hoc.
web/src/lib/components/MessageList.svelte:
New ReactionPills render below each message body when the message
carries a `reactions` array (already populated by
EnrichMessages/ReactionEnricher on the server side). Makes the
👀 in_progress / ✔ done / ❌ reject lifecycle visible in every DM
view and conversation.
web/src/routes/runs/[id]/+page.svelte (NEW):
New run detail page at /runs/:id with sections:
1. Header strip — agent, status pill, backend badge, trigger
info, duration, tokens in/out, cost, exit code, trace id.
2. Triggering message — body + sender.
3. What the model saw — GEMINI.md / CLAUDE.md from agent snapshot
+ the captured rendered prompt (byte count on each summary
bar, collapsible details).
4. What the model said — captured response, falling back to
logs_excerpt or error_log when unavailable.
5. Outgoing message — body + recipient + status.
6. Metadata — reactive_run.id, harness_run.run_id, backend,
session_id, tokens_cached, k8s_job, agent trigger config.
Styled against the existing dark tailwind system — no design
overhaul, fits the current aesthetic (editorial sectioning,
monospace for code-like content, accent-blue for links,
accent-purple for system-instructions, accent-green for model
output, accent-red for errors).
web/src/routes/runs/+page.svelte: the inline expand panel now has
a "View full details →" link next to the Retry button.
E2E VERIFIED on a live subprocess run:
* Topic: "why does the subprocess harness materialise GEMINI.md
alongside .gemini/settings.json in the per-run workdir?"
* 3 subprocess runs + 3 reactive_runs + 3 harness_runs, all linked.
* message_reactions: 6 rows — in_progress + done for each hop.
* GET /api/runs/1 returns a composite with
harness_run.prompt=883 bytes, harness_run.response=550 bytes,
reactive_run_id=1, trigger_message populated, outgoing_message
populated (decomposer-pro → writer-flash), agent.gemini_md=747
bytes. All keys are snake_case as the Svelte types expect.
Full go test ./... green. `vite build` green. Demo instance still
running on port 18088 for browser verification.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
b140879bc3 |
fix(examples): own demo agents by the algis user, not admin
start.sh was passing --owner 1 to every `agent create` call, under the
assumption that the freshly created `algis` user would be user id 1.
It isn't — the `admin` user is auto-seeded at id=1, so `algis` comes
in at id=2. Result: the 3 AI agents AND the `algis` human agent were
owned by `admin`, and when the user logged into the Web UI as `algis`
the message handler called GetHumanAgentForUser(2) which returned a
different auto-created `algis-human` agent (id=5, owned by user 2).
The critic DM'd the name `algis` → landed on agent id=1 (admin-owned),
but the UI listed DMs for `algis-human` → panel showed
"No conversations" despite reactive_runs clearly showing the chain
succeeded.
Fix: after `user create`, query sqlite for the algis user id and use
that value as --owner for every subsequent agent create. Bails with a
clear error if the id lookup fails or returns 1 (sanity check that
admin/algis aren't conflated).
Verified end-to-end:
1. ./stop.sh && ./start.sh — new instance, ownership correct from
the first `agent create` call.
2. ./run_task.sh with a fresh topic — 3 subprocess runs succeeded,
messages #1 (algis → decomposer-pro) and #4 (critic-lite → algis)
now belong to an agent owned by user 2, so
GetHumanAgentForUser(2) returns the same agent the messages are
addressed to.
3. sqlite3 agents table shows all four agents with owner_id=2.
4. Chrome navigation to http://localhost:18088/ reaches the login
form with no errors — once the user logs in as
algis / algis-demo-pw the Direct Messages panel will show the
four-message conversation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
95491402a9 |
fix(otel,examples): schema-URL merge + stale embedded web dist
Two independent fixes hit while running the cold-topic-explainer demo
end-to-end.
1. observability/otel.go — schema URL conflict on Init
When SYNAPBUS_OTEL_ENABLED=1, Init() failed with:
observability: build resource: conflicting Schema URL:
https://opentelemetry.io/schemas/1.26.0 and
https://opentelemetry.io/schemas/1.21.0
resource.Default() ships with schema 1.26.0 (newer otel/sdk) but I
was passing semconv.SchemaURL from v1.21 into a NewWithAttributes
call. resource.Merge rejects that.
Fix: use resource.NewSchemaless for the service.* attributes so our
side of the merge has no schema URL and slots cleanly into whatever
Default provides. ServiceVersion is now only attached when non-empty
(avoids a stray service.version="" attribute).
Two new regression tests:
TestInit_EnabledSucceeds — Enabled=true with all fields set
TestInit_EnabledWithNoVersion — Enabled=true with empty version
Both point at an unroutable endpoint so the batcher never actually
exports; the bug reproduced during Init(), which is all we need.
2. examples/cold-topic-explainer/start.sh — rebuild embedded SPA
The Svelte Web UI loaded blank because internal/web/dist/ had a
mismatched index.html + stale _app/immutable/entry/ assets (a build
had updated index.html but not the chunks, so every asset URL fell
through to the SPA HTML fallback and the browser tried to execute
HTML as JavaScript).
The canonical path is `make web`, but start.sh never ran it, so a
working demo depended on the developer having run `make web` first.
Fix: start.sh now rebuilds the SPA when web/src is newer than the
embedded dist/index.html, using the already-installed
web/node_modules (no reinstall). Falls back with a "run make web
once" hint when node_modules isn't present. This keeps the fast
path fast (~2s vite build after cache warm) and eliminates the
silent-stale-dist trap.
E2E verified after both fixes:
* SYNAPBUS_OTEL_ENABLED=1 start.sh no longer crashes.
* `curl /_app/immutable/entry/start.*.js` returns real JavaScript
(Content-Type: text/javascript) instead of the index.html
fallback.
* Chrome-in-MCP navigation to http://localhost:18088/ renders the
login form with no SynapBus-originated console errors.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
04e4e3ca37 |
feat(harness): GEMINI.md support + cold-topic-explainer example
Adds everything needed to run a real multi-Gemini-model reactive agent
loop end-to-end on SynapBus.
internal/harness/subprocess/config.go:
* AgentConfig.GeminiMD — content of workdir/GEMINI.md
* MaterialiseAgentConfig writes GEMINI.md AND workdir/.gemini/settings.json
when gemini_md is set. The settings file carries the same mcp_servers
list as .mcp.json (so a Gemini child running from the workdir gets
the exact MCP surface the operator configured, not the user's
~/.gemini/settings.json).
* 2 new config_test cases: GEMINI.md + .gemini/settings.json round
trip, GEMINI.md with empty mcp_servers still writes the settings
file (explicitly clearing any inherited home config).
internal/harness/registry.go — BUG FIX:
Resolve() now honours agent.HarnessName (explicit selection) BEFORE
the inference chain, matching the reactor's own agentBackendKind
policy. Previously, when multiple backends were registered,
Resolve would pick "webhook" for every non-K8s agent — even when the
agent's HarnessName was "subprocess" — because the original fallback
chain put webhook first. This is why the first cold-topic-explainer
run failed with "webhook: agent has no webhook config". Discovered
during e2e testing.
internal/admin/socket.go + cmd/synapbus/admin.go:
New `messages.send` admin command (socket + CLI). Sends a DM as any
agent through the messaging service, bypassing the REST/MCP auth
layers. Local-only via the admin Unix socket, so the threat model is
"whoever can reach the socket is already admin".
CLI:
synapbus messages send --from X --to Y --body "..." [--priority N]
synapbus messages send --from X --to Y --body-file path
echo "..." | synapbus messages send --from X --to Y
Used by the harness shell wrappers (so Gemini subprocess agents can
DM each other) and by run_task.sh (to kick off a chain as a human
user without implementing the REST session flow).
examples/cold-topic-explainer/ (NEW):
Runnable 3-agent Gemini demo that exercises the subprocess harness,
reactive triggers, recursive update, and all the preconditions (depth,
budget, cooldown) end-to-end on a separate isolated synapbus instance.
Layout:
README.md — usage + troubleshooting + cost notes
start.sh — builds synapbus, launches on port 18088 with
./data, creates user + agents + harness configs,
marks agents reactive via sqlite3
run_task.sh — sends initial DM algis → decomposer-pro, polls
reactive_runs + messages for the FINAL: reply,
prints the result or dumps reactive_runs on
timeout for debugging
stop.sh — SIGTERM + 5s grace + SIGKILL fallback
wrapper.sh — shared subprocess local_command: reads
message.json + GEMINI.md, calls gemini headless
with --approval-mode yolo, strips the
"MCP issues detected" noise prefix, routes the
cleaned response to the next agent via
`synapbus messages send` over the admin socket
configs/
decomposer-pro.json — gemini-3.1-pro-preview
(gemini-2.5-pro is currently capacity-
exhausted on Google's side)
writer-flash.json — gemini-2.5-flash
critic-lite.json — gemini-2.5-flash-lite
.gitignore — data/, bin/, synapbus.log, .synapbus.pid
The wrapper does NOT rely on gemini's MCP tool-calling (which was
unreliable in testing). Gemini is used as a pure text generator; the
shell decides routing based on AGENT_ROLE:
- decomposer → NEXT_AGENT (writer)
- writer → NEXT_AGENT (critic)
- critic → OWNER_AGENT if response starts with FINAL:,
REVISE_AGENT otherwise
E2E VERIFICATION (real run, real Gemini, not a mock):
Topic: "how does SynapBus unify message delivery, reactive agent
triggers, and harness runs on a single SQLite database?"
Result (from data/synapbus.db after one successful run):
harness_runs:
#1 decomposer-pro subprocess success 106s
#2 writer-flash subprocess success 155s
#3 critic-lite subprocess success 10s
reactive_runs: 3 rows, all succeeded, trigger_from chain:
algis → decomposer-pro → writer-flash → critic-lite
messages:
#1 algis → decomposer-pro (topic)
#2 decomposer-pro → writer-flash (Q1/Q2/Q3 breakdown)
#3 writer-flash → critic-lite (3-paragraph draft)
#4 critic-lite → algis (FINAL: + polished 3-paragraph explainer)
Critic converged in one pass (all scores ≥ 8), so the writer↔critic
refinement loop didn't need to recurse — but the plumbing for it
(REVISE: branch in wrapper.sh, depth limit in reactor) is wired and
ready. Flipping the critic's acceptance bar exercises the recursion.
Full `go test ./...` remained green through all changes.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
b8a70bfe66 |
feat(harness): Option C — subprocess agent config (CLAUDE.md, MCP, skills)
Makes the subprocess backend fully self-contained: each agent carries
its instructions, MCP servers, skills, and subagents in its
harness_config_json column, viewable in the Web UI, editable via CLI.
internal/harness/subprocess/config.go (NEW):
AgentConfig struct with optional fields:
- claude_md → workdir/CLAUDE.md
- agents_md → workdir/AGENTS.md
- mcp_servers → workdir/.mcp.json (Claude Code format)
- skills → workdir/.claude/skills/<name>/SKILL.md
- subagents → workdir/.claude/agents/<name>.md
- env → layered into child env (after k8s_env_json,
before caller overrides)
ParseAgentConfig tolerates empty / returns error on invalid JSON.
MaterialiseAgentConfig writes all artifacts into the workdir with
path-traversal sanitisation on skill/subagent names.
subprocess.Harness.Execute now calls Parse + Materialise before exec,
so an agent's declarative config is on disk by the time the child
CLI's cwd lookup fires. buildEnv takes the parsed config and overlays
cfg.Env on top of k8s_env_json.
Tests:
config_test.go — 6 cases: empty, invalid JSON, full round-trip,
materialise writes all artefacts, empty is a no-op, skill names
are sanitised against "../escape" / "/etc/passwd", mcp entries
without a name are dropped.
subprocess_test.go — 2 new e2e cases: agent with CLAUDE.md + mcp
servers + skills + env sees all of them from inside the child via
cat/echo; invalid harness_config_json surfaces as Execute error.
internal/agents/store.go:
AgentStore gains UpdateHarnessConfig(ctx, name, harnessName,
localCommand, harnessConfigJSON). Empty strings leave a field
unchanged; literal "-" clears (sets to NULL). Returns sql.ErrNoRows
on missing agent. AgentService exposes Store() so admin handlers
can reach it without adding a full service method for a
config-set-style operation.
store_test.go: 6-subcase test covers set-all, partial update, clear,
unknown agent, and no-field no-op.
internal/admin/socket.go:
Two new admin commands:
harness.config_get {agent_name} → {harness_name, local_command,
harness_config_json, harness_config (parsed), parse_error?}
harness.config_set {agent_name, harness_name?, local_command?,
harness_config_json?} → updated fields
config_set validates JSON shape before calling the store; null /
"-" literals clear the column.
cmd/synapbus/admin.go:
New top-level `harness config` command group:
synapbus harness config get --agent <name> [--raw]
synapbus harness config set --agent <name>
[--harness-name subprocess]
[--local-command '["claude","--print"]']
[--file config.json] # or pipe from stdin
[--clear]
synapbus harness config edit --agent <name>
# fetches current config, opens $VISUAL/$EDITOR/vi,
# validates JSON on save, writes back via config_set
web/src/routes/agents/[name]/+page.svelte:
New read-only "Harness" panel on the agent detail page:
- Resolved backend badge (explicit or inferred from k8s_image /
local_command / harness_config_json.url)
- Grid summary: CLAUDE.md size, AGENTS.md size, MCP server count,
skills count
- Collapsible details for CLAUDE.md, AGENTS.md, each MCP server
(name / type / url|command / header count), skill filenames,
subagent filenames, env vars
- Footer hint showing the CLI edit command
No edit controls — editing is CLI-only by design (safer, fits an
ops-heavy workflow).
Verified: full project test suite (40+ packages including integration
tests) plus `vite build` of the Svelte app all green; `go vet ./...`
clean; the existing TestSubprocess_Execute_MaterialisesHarnessConfig
e2e test proves the round-trip from harness_config_json → workdir →
child process works end-to-end.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
8ddccd2452 |
feat(reactor): route non-K8s reactive runs through harness registry
Phase 7: the reactor now branches on agent backend kind.
Reactor changes (internal/reactor/reactor.go):
* Adds `registry *harness.Registry` field + `SetHarnessRegistry`.
* `agentBackendKind()` picks k8s | subprocess | webhook | none from
the agent's `HarnessName`, `K8sImage`, `LocalCommand`, and
`HarnessConfigJSON` fields. Explicit `HarnessName` wins.
* `evaluateTrigger` applies the same preconditions (depth, daily
budget, cooldown, already-running, pending_work coalescing) to
every backend — a subprocess agent mentioned in a channel now
goes through the exact same rate limits a K8s agent does.
* K8s agents keep the existing `createJob` fast-return path with
the async poller for restart safety. Non-K8s agents use a new
`dispatchHarness` that inserts the reactive_runs row, spawns a
detached goroutine, blocks on `Registry.Execute`, and writes the
terminal status / error_log / metrics / failure DM on return.
* Import `harness`, `messaging`, `google/uuid` for building the
ExecRequest.
main.go wiring:
* Build one `harness.Registry` with all three real backends:
`k8sjob.New(k8sRunner, …)`, `subprocess.New(Config{BaseDir:
dataDir/harness/subprocess}, …)`, `webhook.New(Config{}, …)`.
* Attach a `runs.Store` as the registry Observer so every dispatch
writes a harness_runs row — no per-caller code required.
* Hand the registry to the reactor via `SetHarnessRegistry`.
* Log the registered backend names at startup.
Tests (internal/reactor/reactor_test.go):
* New `insertSubprocessAgent`, `newHarnessReactor`, `waitForRun`,
and `fakeNotifier` helpers.
* Seven new tests that register a stub harness under "subprocess"
and verify: success from @mention, failure recorded + DM sent,
depth-exceeded skipped, budget-exhausted skipped, cooldown
skipped, already-running queued, no-backend fails cleanly. Each
checks the harness stub is NOT called when a precondition skips.
* Existing `TestReactorNoK8sImage` keeps working — the old
k8s-specific error message is replaced with the backend-agnostic
"no backend configured" phrasing.
* `setupTestDB` now pins `SetMaxOpenConns(1)`: modernc.org/sqlite
in-memory DBs give each pool connection a fresh empty database,
which races the new dispatchHarness goroutine and main-test
goroutine. Pinning is the standard workaround.
The overall behaviour: `@local-agent` in a channel message now starts
the configured subprocess/webhook under the same depth/budget/cooldown
rate limits as a K8s agent, tracked in reactive_runs and harness_runs,
instrumented with an OTel span, with trace context propagated into the
child via env vars. Failure DMs go to the human owner, as with K8s.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
574897df4d |
feat: harness-agnostic wrappers + OpenTelemetry integration
Introduces internal/harness — a minimal Harness interface inspired by
GoogleCloudPlatform/scion — plus four backends (k8sjob, subprocess,
webhook, stub) and an OTel-traced Registry that spans every dispatch
and injects W3C trace context into child processes via env vars.
Phases landed together on this branch:
1. internal/harness scaffold: Harness/Capabilities/ExecRequest/
ExecResult/Budget/Usage types, Registry with Resolve/Execute,
in-memory stub backend.
2. internal/harness/k8sjob: wraps existing k8s.JobRunner behind the
Harness interface with a Waiter abstraction (real clientset +
test fake). BuildHandler exports the per-agent config logic.
3. internal/harness/subprocess: os/exec-based backend (Mac+Linux),
per-run workdir, result.json handoff, bounded log capture,
Budget-driven wall-clock timeout.
4. internal/harness/webhook: synchronous HTTP POST with HMAC
signing via internal/webhooks.ComputeHMACSignature, per-agent
URL/secret/timeout read from harness_config_json.
5. internal/observability: OTel tracer init via OTLP HTTP (opt-in
via SYNAPBUS_OTEL_ENABLED), W3C propagator always installed;
Registry.Execute starts a harness.execute span per dispatch and
calls InjectTraceContext into req.Env so children inherit it.
6. internal/harness/runs: SQLite-backed Observer that persists a
harness_runs row per dispatch with status, usage, cost, duration,
trace_id, session_id, and a bounded logs excerpt.
Schema: new migration 019_harness.sql adds agents.harness_name /
local_command / harness_config_json columns and the backend-agnostic
harness_runs table with indices on (agent, created_at), (status),
(trace_id), (run_id). internal/reactor/reactor_test.go inline schema
updated to match.
Deployment: deploy/kubic/otel-collector.yaml stands up an otel-collector
Deployment + ConfigMap + ClusterIP Service in the synapbus namespace on
kubic, receiving OTLP gRPC (4317) and HTTP (4318) and exporting debug
output until a Tempo/Jaeger backend lands.
Docs: docs/harness-otel-research.html compares scion and paperclip
side-by-side and maps the current synapbus executor surface; its
companion docs/harness-otel-design.md carries the phase plan, span
taxonomy, and migration schema verbatim.
The reactor currently still calls k8s.JobRunner directly — rewiring it
through the Registry is a follow-up, intentionally out of scope for
this branch to keep the refactor reversible. The new packages are
independently tested (~78 new tests across 7 packages) and the full
project test suite passes.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
0e25fbcccb |
feat: sdk_backend + autonomous run integration
Adds benchmark/sdk_backend.py that routes model calls through either the anthropic SDK (preferred, requires ANTHROPIC_API_KEY) or the claude-agent-sdk as a Claude Code session fallback. agents.py and baseline.py now go through this unified backend instead of calling anthropic directly. Ran benchmark/run.py --mode single-shot --question q1 end-to-end with real Claude API calls via claude-agent-sdk. Real numbers: - Marketplace (Haiku 4.5): 3314 tokens, F1 1.000 (exact match) - Baseline (Sonnet 4.6): 697 tokens, F1 0.857 (penalized for "1783") - Pareto verdict: FAIL (not strictly NW; marketplace wins quality, loses cost — informative failure per spec design). Added autonomous_report.html (rich narrative with Pareto chart) and autonomous_summary.md. All 34 Go packages still green. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
8fd42cb957 | merge: 017-musique-benchmark MVP (Python harness + trio + Pareto report) | ||
|
|
dc1ae4da60 | merge: 016-agent-marketplace MVP (auction + manifests + reputation) | ||
|
|
cda3365863 |
feat(016): agent marketplace MVP — capability manifests, auction channels, reputation ledger
Implements US1, US2, and US3 of spec 016 by layering a marketplace service on top of existing primitives rather than reinventing them: - Capability manifests (US2) reuse the wiki subsystem. Each agent publishes a per-agent article at slug "agent-<name>" and gets versioning, revision history, and FTS search for free. - Auction channels (US1) reuse the existing auction channel type, swarm service, and task/bid store. post_auction / bid / award wrap post_task / bid_task / accept_bid and attach marketplace metadata (max_budget_tokens, domains, estimated_tokens, confidence, approach) in the task.requirements and bid.capabilities JSON blobs. Award converts the auction into a claim by DM'ing the winner at priority 8 with task_id metadata, so the existing claim/process/done lifecycle takes over with zero new machinery. - Reputation ledger (US3) adds migration 018_agent_marketplace.sql with a new agent_reputation table keyed by (agent_name, domain). mark_task_done completes the task via the swarm service and writes one ledger row per declared domain using the reported actual_tokens and success_score. query_reputation returns a rolled-up summary plus recent entries for a given (agent, domain) pair — reputation is always a vector, never a global score (FR-013). Also: - Adds the "awarded" reaction type (FR-008) alongside existing approve/ reject/in_progress/done/published. Migration 018 widens the reactions CHECK constraint via a table rebuild. - 6 new actions added to the action registry (post_auction, bid, award, mark_task_done, read_skill_card, query_reputation) so the search tool can discover them and the execute tool can dispatch them. - New internal/marketplace package (store.go + service.go). - New internal/mcp/marketplace.go bridge handlers. - New internal/mcp/marketplace_test.go covers the full auction lifecycle, capability manifest publish/read/update, self-bid rejection, non-auction channel rejection, and reputation summary aggregation. Out of scope for MVP (deferred per spec prompt): US4 reflection loop, tombstoning FR-020a/b, multi-owner quorums, auto-escalation on zero bids, bootstrap exploration credit, epsilon-greedy selection, and the hard-stop budget enforcement daemon (only soft recording of estimated vs actual is included). All existing tests pass; new marketplace tests pass. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
02b8548eac |
feat(017): MuSiQue benchmark harness — marketplace stub, mixed-tier agents, Pareto scoring, HTML report
MVP implementation of the MuSiQue multi-agent benchmark (spec 017): - benchmark/setup.py: downloads musique_v1.0.zip from the canonical Google Drive source (mirrors upstream download_data.sh). Idempotent. - benchmark/curate.py: deterministic selection of 3 4-hop questions from the dev set sharing a US pivot entity; writes benchmark/trio.jsonl. - benchmark/marketplace.py: in-process 016-marketplace stub with post_auction / bid / award / mark_done / query_reputation and a domain-scoped reputation ledger. Designed for mechanical swap to real SynapBus MCP tools. - benchmark/agents.py: HaikuAgent + SonnetAgent, using the official anthropic SDK (no Claude Agent SDK, no subprocesses). Models pinned to claude-haiku-4-5-20251001 and claude-sonnet-4-6. - benchmark/baseline.py: single Sonnet call with all 20 distractors plus chain-of-thought. - benchmark/score.py: SQuAD-style normalized F1 + Pareto verdict (strictly northwest = PASS). - benchmark/run.py: main entry. --mode single-shot, --question, --dry-run. - benchmark/report.py: self-contained HTML with inline SVG scatter plot. - benchmark/trio.jsonl: curated reproducible trio (all three converge on "Treaty of Paris" US territory cession). Verified with benchmark/run.py --dry-run end-to-end; all 8 files py_compile clean. Real-token execution is deferred to the user's main session. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
e77fd7afdf |
spec(017): MuSiQue MAS benchmark harness
Feature spec for Python benchmark that integration-tests 016 marketplace against a real multi-hop reasoning task. 4 prioritized user stories: P1 single-shot Pareto verification, P2 curated trio with dedup, P3 learning tier, P1 rich HTML report. 23 FRs, 7 success criteria. Also: brainstorming design doc at docs/superpowers/specs/ capturing the 6 clarifying questions and chosen decisions (mixed-tier pool, curated trio, tiered run modes, wait-for-016 execution strategy). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
96db7c06a4 |
spec(016): agent marketplace spec + research reports
- specs/016-agent-marketplace: self-organizing marketplace spec with capability manifests, auction channels, domain-scoped reputation, and reflection loop. Four user stories (P1: auction + manifests, P2: reputation + reflection). 27 FRs, 10 success criteria, checklist. - multiagent_systems_report.html: landscape of OSS MAS frameworks, coordination patterns (blackboard/stigmergy/contract-net/gossip), problem classes, toy benchmarks. - agent_marketplace_guide_ru.html: Russian technical guide with terminology dictionary, Fermi walkthrough, Voyager lessons. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
660da6d646 |
feat: wiki export/import CLI for backup and restore
Usage: synapbus wiki export --data ./data --output ./wiki-export synapbus wiki import --data ./data --input ./wiki-export Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
75fb483ae9 |
docs: add spec and plan for 013-agent-wiki feature
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |