Makes the subprocess backend fully self-contained: each agent carries
its instructions, MCP servers, skills, and subagents in its
harness_config_json column, viewable in the Web UI, editable via CLI.
internal/harness/subprocess/config.go (NEW):
AgentConfig struct with optional fields:
- claude_md → workdir/CLAUDE.md
- agents_md → workdir/AGENTS.md
- mcp_servers → workdir/.mcp.json (Claude Code format)
- skills → workdir/.claude/skills/<name>/SKILL.md
- subagents → workdir/.claude/agents/<name>.md
- env → layered into child env (after k8s_env_json,
before caller overrides)
ParseAgentConfig tolerates empty / returns error on invalid JSON.
MaterialiseAgentConfig writes all artifacts into the workdir with
path-traversal sanitisation on skill/subagent names.
subprocess.Harness.Execute now calls Parse + Materialise before exec,
so an agent's declarative config is on disk by the time the child
CLI's cwd lookup fires. buildEnv takes the parsed config and overlays
cfg.Env on top of k8s_env_json.
Tests:
config_test.go — 6 cases: empty, invalid JSON, full round-trip,
materialise writes all artefacts, empty is a no-op, skill names
are sanitised against "../escape" / "/etc/passwd", mcp entries
without a name are dropped.
subprocess_test.go — 2 new e2e cases: agent with CLAUDE.md + mcp
servers + skills + env sees all of them from inside the child via
cat/echo; invalid harness_config_json surfaces as Execute error.
internal/agents/store.go:
AgentStore gains UpdateHarnessConfig(ctx, name, harnessName,
localCommand, harnessConfigJSON). Empty strings leave a field
unchanged; literal "-" clears (sets to NULL). Returns sql.ErrNoRows
on missing agent. AgentService exposes Store() so admin handlers
can reach it without adding a full service method for a
config-set-style operation.
store_test.go: 6-subcase test covers set-all, partial update, clear,
unknown agent, and no-field no-op.
internal/admin/socket.go:
Two new admin commands:
harness.config_get {agent_name} → {harness_name, local_command,
harness_config_json, harness_config (parsed), parse_error?}
harness.config_set {agent_name, harness_name?, local_command?,
harness_config_json?} → updated fields
config_set validates JSON shape before calling the store; null /
"-" literals clear the column.
cmd/synapbus/admin.go:
New top-level `harness config` command group:
synapbus harness config get --agent <name> [--raw]
synapbus harness config set --agent <name>
[--harness-name subprocess]
[--local-command '["claude","--print"]']
[--file config.json] # or pipe from stdin
[--clear]
synapbus harness config edit --agent <name>
# fetches current config, opens $VISUAL/$EDITOR/vi,
# validates JSON on save, writes back via config_set
web/src/routes/agents/[name]/+page.svelte:
New read-only "Harness" panel on the agent detail page:
- Resolved backend badge (explicit or inferred from k8s_image /
local_command / harness_config_json.url)
- Grid summary: CLAUDE.md size, AGENTS.md size, MCP server count,
skills count
- Collapsible details for CLAUDE.md, AGENTS.md, each MCP server
(name / type / url|command / header count), skill filenames,
subagent filenames, env vars
- Footer hint showing the CLI edit command
No edit controls — editing is CLI-only by design (safer, fits an
ops-heavy workflow).
Verified: full project test suite (40+ packages including integration
tests) plus `vite build` of the Svelte app all green; `go vet ./...`
clean; the existing TestSubprocess_Execute_MaterialisesHarnessConfig
e2e test proves the round-trip from harness_config_json → workdir →
child process works end-to-end.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Phase 7: the reactor now branches on agent backend kind.
Reactor changes (internal/reactor/reactor.go):
* Adds `registry *harness.Registry` field + `SetHarnessRegistry`.
* `agentBackendKind()` picks k8s | subprocess | webhook | none from
the agent's `HarnessName`, `K8sImage`, `LocalCommand`, and
`HarnessConfigJSON` fields. Explicit `HarnessName` wins.
* `evaluateTrigger` applies the same preconditions (depth, daily
budget, cooldown, already-running, pending_work coalescing) to
every backend — a subprocess agent mentioned in a channel now
goes through the exact same rate limits a K8s agent does.
* K8s agents keep the existing `createJob` fast-return path with
the async poller for restart safety. Non-K8s agents use a new
`dispatchHarness` that inserts the reactive_runs row, spawns a
detached goroutine, blocks on `Registry.Execute`, and writes the
terminal status / error_log / metrics / failure DM on return.
* Import `harness`, `messaging`, `google/uuid` for building the
ExecRequest.
main.go wiring:
* Build one `harness.Registry` with all three real backends:
`k8sjob.New(k8sRunner, …)`, `subprocess.New(Config{BaseDir:
dataDir/harness/subprocess}, …)`, `webhook.New(Config{}, …)`.
* Attach a `runs.Store` as the registry Observer so every dispatch
writes a harness_runs row — no per-caller code required.
* Hand the registry to the reactor via `SetHarnessRegistry`.
* Log the registered backend names at startup.
Tests (internal/reactor/reactor_test.go):
* New `insertSubprocessAgent`, `newHarnessReactor`, `waitForRun`,
and `fakeNotifier` helpers.
* Seven new tests that register a stub harness under "subprocess"
and verify: success from @mention, failure recorded + DM sent,
depth-exceeded skipped, budget-exhausted skipped, cooldown
skipped, already-running queued, no-backend fails cleanly. Each
checks the harness stub is NOT called when a precondition skips.
* Existing `TestReactorNoK8sImage` keeps working — the old
k8s-specific error message is replaced with the backend-agnostic
"no backend configured" phrasing.
* `setupTestDB` now pins `SetMaxOpenConns(1)`: modernc.org/sqlite
in-memory DBs give each pool connection a fresh empty database,
which races the new dispatchHarness goroutine and main-test
goroutine. Pinning is the standard workaround.
The overall behaviour: `@local-agent` in a channel message now starts
the configured subprocess/webhook under the same depth/budget/cooldown
rate limits as a K8s agent, tracked in reactive_runs and harness_runs,
instrumented with an OTel span, with trace context propagated into the
child via env vars. Failure DMs go to the human owner, as with K8s.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Introduces internal/harness — a minimal Harness interface inspired by
GoogleCloudPlatform/scion — plus four backends (k8sjob, subprocess,
webhook, stub) and an OTel-traced Registry that spans every dispatch
and injects W3C trace context into child processes via env vars.
Phases landed together on this branch:
1. internal/harness scaffold: Harness/Capabilities/ExecRequest/
ExecResult/Budget/Usage types, Registry with Resolve/Execute,
in-memory stub backend.
2. internal/harness/k8sjob: wraps existing k8s.JobRunner behind the
Harness interface with a Waiter abstraction (real clientset +
test fake). BuildHandler exports the per-agent config logic.
3. internal/harness/subprocess: os/exec-based backend (Mac+Linux),
per-run workdir, result.json handoff, bounded log capture,
Budget-driven wall-clock timeout.
4. internal/harness/webhook: synchronous HTTP POST with HMAC
signing via internal/webhooks.ComputeHMACSignature, per-agent
URL/secret/timeout read from harness_config_json.
5. internal/observability: OTel tracer init via OTLP HTTP (opt-in
via SYNAPBUS_OTEL_ENABLED), W3C propagator always installed;
Registry.Execute starts a harness.execute span per dispatch and
calls InjectTraceContext into req.Env so children inherit it.
6. internal/harness/runs: SQLite-backed Observer that persists a
harness_runs row per dispatch with status, usage, cost, duration,
trace_id, session_id, and a bounded logs excerpt.
Schema: new migration 019_harness.sql adds agents.harness_name /
local_command / harness_config_json columns and the backend-agnostic
harness_runs table with indices on (agent, created_at), (status),
(trace_id), (run_id). internal/reactor/reactor_test.go inline schema
updated to match.
Deployment: deploy/kubic/otel-collector.yaml stands up an otel-collector
Deployment + ConfigMap + ClusterIP Service in the synapbus namespace on
kubic, receiving OTLP gRPC (4317) and HTTP (4318) and exporting debug
output until a Tempo/Jaeger backend lands.
Docs: docs/harness-otel-research.html compares scion and paperclip
side-by-side and maps the current synapbus executor surface; its
companion docs/harness-otel-design.md carries the phase plan, span
taxonomy, and migration schema verbatim.
The reactor currently still calls k8s.JobRunner directly — rewiring it
through the Registry is a follow-up, intentionally out of scope for
this branch to keep the refactor reversible. The new packages are
independently tested (~78 new tests across 7 packages) and the full
project test suite passes.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Adds benchmark/sdk_backend.py that routes model calls through either
the anthropic SDK (preferred, requires ANTHROPIC_API_KEY) or the
claude-agent-sdk as a Claude Code session fallback. agents.py and
baseline.py now go through this unified backend instead of calling
anthropic directly.
Ran benchmark/run.py --mode single-shot --question q1 end-to-end
with real Claude API calls via claude-agent-sdk. Real numbers:
- Marketplace (Haiku 4.5): 3314 tokens, F1 1.000 (exact match)
- Baseline (Sonnet 4.6): 697 tokens, F1 0.857 (penalized for "1783")
- Pareto verdict: FAIL (not strictly NW; marketplace wins quality,
loses cost — informative failure per spec design).
Added autonomous_report.html (rich narrative with Pareto chart)
and autonomous_summary.md. All 34 Go packages still green.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Implements US1, US2, and US3 of spec 016 by layering a marketplace service
on top of existing primitives rather than reinventing them:
- Capability manifests (US2) reuse the wiki subsystem. Each agent publishes
a per-agent article at slug "agent-<name>" and gets versioning, revision
history, and FTS search for free.
- Auction channels (US1) reuse the existing auction channel type, swarm
service, and task/bid store. post_auction / bid / award wrap post_task /
bid_task / accept_bid and attach marketplace metadata (max_budget_tokens,
domains, estimated_tokens, confidence, approach) in the task.requirements
and bid.capabilities JSON blobs. Award converts the auction into a claim
by DM'ing the winner at priority 8 with task_id metadata, so the existing
claim/process/done lifecycle takes over with zero new machinery.
- Reputation ledger (US3) adds migration 018_agent_marketplace.sql with a
new agent_reputation table keyed by (agent_name, domain). mark_task_done
completes the task via the swarm service and writes one ledger row per
declared domain using the reported actual_tokens and success_score.
query_reputation returns a rolled-up summary plus recent entries for a
given (agent, domain) pair — reputation is always a vector, never a
global score (FR-013).
Also:
- Adds the "awarded" reaction type (FR-008) alongside existing approve/
reject/in_progress/done/published. Migration 018 widens the reactions
CHECK constraint via a table rebuild.
- 6 new actions added to the action registry (post_auction, bid, award,
mark_task_done, read_skill_card, query_reputation) so the search tool
can discover them and the execute tool can dispatch them.
- New internal/marketplace package (store.go + service.go).
- New internal/mcp/marketplace.go bridge handlers.
- New internal/mcp/marketplace_test.go covers the full auction lifecycle,
capability manifest publish/read/update, self-bid rejection, non-auction
channel rejection, and reputation summary aggregation.
Out of scope for MVP (deferred per spec prompt): US4 reflection loop,
tombstoning FR-020a/b, multi-owner quorums, auto-escalation on zero bids,
bootstrap exploration credit, epsilon-greedy selection, and the hard-stop
budget enforcement daemon (only soft recording of estimated vs actual is
included).
All existing tests pass; new marketplace tests pass.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
MVP implementation of the MuSiQue multi-agent benchmark (spec 017):
- benchmark/setup.py: downloads musique_v1.0.zip from the canonical Google
Drive source (mirrors upstream download_data.sh). Idempotent.
- benchmark/curate.py: deterministic selection of 3 4-hop questions from
the dev set sharing a US pivot entity; writes benchmark/trio.jsonl.
- benchmark/marketplace.py: in-process 016-marketplace stub with
post_auction / bid / award / mark_done / query_reputation and a
domain-scoped reputation ledger. Designed for mechanical swap to real
SynapBus MCP tools.
- benchmark/agents.py: HaikuAgent + SonnetAgent, using the official
anthropic SDK (no Claude Agent SDK, no subprocesses). Models pinned
to claude-haiku-4-5-20251001 and claude-sonnet-4-6.
- benchmark/baseline.py: single Sonnet call with all 20 distractors
plus chain-of-thought.
- benchmark/score.py: SQuAD-style normalized F1 + Pareto verdict
(strictly northwest = PASS).
- benchmark/run.py: main entry. --mode single-shot, --question, --dry-run.
- benchmark/report.py: self-contained HTML with inline SVG scatter plot.
- benchmark/trio.jsonl: curated reproducible trio (all three converge on
"Treaty of Paris" US territory cession).
Verified with benchmark/run.py --dry-run end-to-end; all 8 files
py_compile clean. Real-token execution is deferred to the user's main
session.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Feature spec for Python benchmark that integration-tests 016 marketplace
against a real multi-hop reasoning task. 4 prioritized user stories:
P1 single-shot Pareto verification, P2 curated trio with dedup,
P3 learning tier, P1 rich HTML report. 23 FRs, 7 success criteria.
Also: brainstorming design doc at docs/superpowers/specs/ capturing
the 6 clarifying questions and chosen decisions (mixed-tier pool,
curated trio, tiered run modes, wait-for-016 execution strategy).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Auto mode now runs both semantic and fulltext searches, merging
results using Reciprocal Rank Fusion (RRF, k=60) for best of both
- New min_similarity parameter (default 0.25) filters semantic noise
- Results that match both sources are marked as "hybrid" match_type
- New getFloat bridge helper for MCP min_similarity parameter
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Words like "to", "from", "and", "or", "not", "near" are FTS5 operators
and caused SQL errors (e.g. "no such column: to") when passed as search
queries. sanitizeFTS5Query() wraps each token in double quotes so they
are treated as literal phrase tokens by SQLite's FTS5 MATCH operator.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
CVE: GHSA-go-jose DoS in parsing (GO-2025-3485)
The vulnerability allowed crafted JWS tokens to cause excessive
memory allocation. Affects our OAuth token validation path.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Wrong password now shows "Invalid username or password" (was "Session expired")
- Client API differentiates 401 on login page vs elsewhere
- Added per-IP login rate limiter: 3 failures → blocked 1 minute
- 429 status code returned with remaining seconds in message
- Rate limit cleared on successful login
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Removed the 'peer NOT IN (owned)' filter that excluded all owned agents.
Since all agents (algis, research-*, social-commenter) are owned by the
same user, the filter was hiding all inter-agent DMs. Now shows all
unique DM partners regardless of ownership.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The previous implementation only scanned inbox (pending messages),
so historical conversations with read/done messages were invisible.
New GetDMPartners() does a direct SQL query with window functions
to find all unique DM partners with most recent message preview
and unread count. Historical conversations now always show.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- New API: GET /api/dm/partners — returns DM conversation partners
ordered by most recent message, with unread counts
- Sidebar DM section now shows actual conversation partners (agents
you've exchanged messages with) instead of owned agents
- Each partner shows name, unread badge, clickable to /dm/{name}
- Fixes issue where all DMs were shown mixed in one view
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The StaleWorker sends DMs from 'system' to all channel members when
workflow messages are stuck in 'proposed' state. These DMs were
triggering reactive agent runs, which couldn't action the stale
messages, burning daily budget on wasted K8s Jobs.
Now: reactor silently ignores all messages from 'system' sender.
System notifications are for human review, not agent action.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The onMount + async pattern wasn't triggering Svelte 5 reactivity
properly. Switched to $effect with $user dependency (same pattern
used by Sidebar and other components). Also waits for auth before
loading data.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
New agents now learn about the query action during onboarding:
tables (my_messages, my_channels, channel_messages), examples,
and limitations (100 rows, SELECT only, 5s timeout).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The reactor was inserting the run record first, then creating the K8s
Job, then updating the record with the job name. If the update failed
(SQLITE_BUSY), the run would be stuck in 'running' with no job name,
making it invisible to the poller.
Now: create K8s Job first, then insert the run record with job name
already set in a single atomic write.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Increase busy_timeout from 5s to 15s
- Set synchronous=NORMAL (safe with WAL, reduces fsync)
- Limit MaxOpenConns to 4 to reduce write lock contention
- Explicit wal_autocheckpoint=1000
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- New /runs route with agent summary cards, run list, filtering
- Agent cards show budget usage, cooldown status, current state
- Expandable run rows with error logs and retry button
- API client: runs.list, runs.get, runs.retry, runs.reactiveAgents
- Sidebar navigation updated with "Agent Runs" link
- Rebuilt web dist
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Specifies the reactive agent system: DM/@mention triggers K8s Jobs
with reactor decision engine, cooldown/budget/depth rate limiting,
sequential execution with coalescing, Web UI Agent Runs panel,
failure notifications, and admin CLI.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Show attachment previews on DM messages (was missing, only channels had it)
- Add file upload button to DM compose bar with paperclip icon
- Enrich messages with attachment data in all MCP bridge functions
(read_inbox, claim_messages, search, channel_messages, list_by_state)
- Remove file type restrictions — allow any file type, keep 50MB size limit
- Rebuild web dist
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Prevents 181K+ responses when channels have many messages with long
bodies. New params: limit (default 20, max 100), offset (default 0),
max_body_length (default 500 chars when include_messages=true).
Response now includes total count alongside paginated results.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
DM messages from agents were cut off at 300 characters in the
MessageList view. Increased to 800 to show more context while
still keeping long messages manageable.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
6 demo scenarios from single agent to 4-agent outreach pipeline.
SynapBus as agent memory (channels + semantic search). Three-stage
progression (experiment → stabilize → scale). Identified gaps in
code, website, and documentation. Website restructure proposal.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
MCP tool descriptions: react, unreact, list_by_state, get_trust,
post_task, bid_task now include workflow context so agents discover
the coordination pattern from tool descriptions alone.
Channel creation UI: added channel type selector (standard/blackboard/
auction) and workflow enabled toggle to the create form.
Channel info panel: workflow settings section with toggles for
workflow_enabled, auto_approve, threshold sliders, and stalemate
timeout inputs. Changes apply via PUT /api/channels/{name}/settings.
Agent skill docs: created stigmergy-workflow.md and task-auction.md
reference skills for agent workspaces.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
StalemateWorker: new Phase 2 scans workflow-enabled channels for stale
messages in non-terminal states. Sends reminder DMs after
stalemate_remind_after timeout, escalates to #approvals after
stalemate_escalate_after. Deduplication prevents repeat notifications.
7 new tests.
Website: blog post "SynapBus v0.10: Trust Scores, Reactions, and the
Agent Platform Vision". Updated features page with reactions, trust,
and archetypes sections.
Searcher: all 4 agent AGENT.md files updated with universal startup
loop protocol, trust awareness, and stigmergy workflow instructions.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>