Commit Graph
176 Commits
Author SHA1 Message Date
Algis DumbrisandClaude Opus 4.6 4e6f6cb539 feat(018): Gemini LLM coordinator + spec-018 MCP tool surface
- cmd/docgardener: coordinator calls the `gemini` CLI with the goal
  brief when SYNAPBUS_GEMINI_MODEL is set, parses the returned JSON
  into a goaltasks.TreeNode, and aligns leaf billing codes so the
  fixed dispatch table still routes specialists correctly. Falls
  back to the hardcoded template on any failure (missing CLI, non-
  zero exit, bad JSON) so the demo still works offline.
- internal/mcp: new GoalsToolRegistrar exposing 6 spec-018 tools —
  create_goal, propose_task_tree, propose_agent, claim_task,
  request_resource, list_resources. All require an authenticated
  agent context; wire-only changes on the MCP server side.
- main.go: builds + attaches the new registrar after the hybrid
  tool registrar, logs the 6 tools at startup.

Verified e2e: demo run with Gemini produces an LLM-generated root
task title ("Verify and patch mcpproxy documentation drift"), all
3 specialists dispatched and completed, $1.05 cost rollup on the
/goals/1 page, and the MCP server registers 11 tools total (5
hybrid + 6 spec-018) at boot.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 20:37:07 +03:00
Algis DumbrisandClaude Opus 4.6 f9d8f1a1c9 feat(018): /goals page, budget cascade, quarantine, secrets loop
- New /api/goals + /api/goals/{id} endpoints serving list + task tree
  + cost rollup + billing breakdown + spawned agents + timeline.
- New Svelte /goals and /goals/[id] pages with sidebar link.
- goals.Service.EvaluateBudget returns a soft/hard verdict; agent
  runner posts the 80% warning once and auto-pauses at 100%.
- Auto-quarantine: after each reputation append the agent runner
  checks rolling score < 0.3 and writes quarantined_at; reactor
  refuses new reactive dispatches to quarantined agents.
- Reactor exposes SetSecretProvider; main.go wires secrets.Store
  so reactive subprocess runs inherit user/agent-scoped env vars.
- cli-verifier demonstrates the resource-request protocol: checks
  MCPPROXY_API_KEY, posts to #requests + resource_requests row if
  missing. New `synapbus secrets set/list` CLI (direct-DB) closes
  the loop.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 20:27:25 +03:00
Algis DumbrisandClaude Opus 4.6 3b94fab226 feat(018): real reactor-driven multi-agent doc-gardener flow
Until now the doc-gardener example was a single monolithic
orchestrator binary writing synthetic messages directly to SQLite.
That's now obsolete: the feature runs as a true multi-agent flow
where the SynapBus reactor fires subprocess runs for every DM, each
agent is its own reactive subprocess invocation, and follow-up DMs
go through the real MessagingService.Send → dispatcher path so the
reactor picks them up.

Changes:

- cmd/synapbus/main.go: gate the three legacy background workers
  (expiry, retention, stalemate) behind SYNAPBUS_DISABLE_*_WORKER env
  flags. These workers manage the legacy channel task-auction /
  message retention features the doc-gardener demo doesn't use, but
  they held the single-connection write pool long enough to wedge
  the whole server for interactive sessions. All three are disabled
  in the example's start.sh.

- cmd/docgardener/agent.go (new): the per-agent subprocess entry the
  reactor harness invokes for every reactive trigger. Reads
  message.json from the workdir, routes by SYNAPBUS_AGENT to either
  coordinator-kickoff, coordinator-completion, or specialist-work
  logic. Writes prompt.txt + response.txt for harness capture. Uses
  the admin socket (`synapbus messages send`) for follow-up DMs so
  the real MessagingService.Send path fires the dispatcher.

- cmd/docgardener/main.go: adds `docgardener agent` subcommand, plus
  helpers freshAPIKey / bcryptHash / absPath / selfPath used by the
  spawn flow.

- examples/doc-gardener/start.sh: provisions user + coordinator
  agent + algis human agent + approvals/requests channels; the
  coordinator is created with trigger_mode=reactive,
  harness_name=subprocess, local_command pointing to docgardener
  agent, and harness_config_json.env carrying SYNAPBUS_AGENT,
  SYNAPBUS_BIN, SYNAPBUS_SOCKET. Specialists are spawned
  dynamically by the coordinator at runtime (not pre-registered),
  so the demo exercises dynamic agent spawning end-to-end.

- examples/doc-gardener/run_task.sh: collapsed to a 3-line kickoff
  that just DMs the coordinator and polls algis's inbox for the
  coordinator's FINAL: reply. Everything else happens via the
  reactor.

Verified end-to-end in Chrome on a fresh instance:
- 4 agents registered (coordinator + 3 specialists dynamically
  spawned by the coordinator on receipt of the first DM)
- 7 reactive_runs + 6 harness_runs across the goal lifecycle:
    algis → coordinator (kickoff, 624ms, builds goal+tree+spawns)
    coordinator → docs-scanner (claim task 2)
    coordinator → cli-verifier (claim task 3)
    coordinator → drift-reporter (claim task 4)
    docs-scanner → coordinator (DONE task=2)
    cli-verifier → coordinator (DONE task=3)
    drift-reporter → coordinator (DONE task=4, coalesced)
- Web UI Agent Runs page shows all 7 runs with the real
  "DM from X" trigger lines and correct sender/receiver chain
- Goal ends at status=completed with all 3 leaf tasks at status=done
- Each specialist run posts a real subprocess artifact to the
  goal channel (#finding, #verified, #summary) and appends a real
  reputation_evidence row keyed by config_hash.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 17:09:08 +03:00
Algis DumbrisandClaude Opus 4.6 2ba0f95666 feat(018): real subprocess runs in docgardener + agent trust UI
docgardener: each leaf task now launches a real subprocess via
exec.CommandContext and records a full reactive_runs + harness_runs
row chain with task_id populated, captured prompt, captured response,
exit code, duration, tokens, cost. The Agent Runs page and
/runs/:id detail page now show real data for the doc-gardener demo
— including "What the model saw" and "What the model said" panels —
without needing the coordinator LLM loop.

agents store: agentSelectSQL and both scanAgent functions extended
to read the feature-018 columns (config_hash, parent_agent_id,
spawn_depth, system_prompt, autonomy_tier, tool_scope_json,
quarantined_at, quarantine_reason). /api/agents and
/api/agents/:name now return these fields end-to-end.

Web UI agent detail (web/src/routes/agents/[name]/+page.svelte):
adds a Trust & Spawn section (config_hash, autonomy tier, spawn
depth, parent agent, tool scope chips) and a full-height System
Prompt pre block. Rebuilt internal/web/dist/.

Verified in Chrome against a fresh ./start.sh && ./run_task.sh run:
- Agent Runs page lists 3 completed runs (docs-scanner, cli-verifier,
  drift-reporter) with task.claim event and non-zero durations
- /runs/1 detail page renders captured prompt + structured #finding
  output with 12 flags
- /agents/docs-scanner shows config_hash=a0b5c6538b2d…, parent=#1,
  depth=1, tool-scope chips, and the 170-char system prompt
- #goal-... channel loads all 12 messages (no "Joining..." hang)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 16:10:08 +03:00
Algis DumbrisandClaude Opus 4.6 ff5d0c49f4 feat(018): dynamic agent spawning — primitives + doc-gardener demo
Ships the MVP slice of spec 018 (dynamic agent spawning):

- 5 new SQLite migrations (021-025): goals + goal_tasks + agent_proposals
  + reputation_evidence + secrets + harness_runs.task_id. The legacy
  `tasks` table (channel auctions) and `agent_trust` table (reactions
  workflow) are left untouched — the new schema coexists.

- 4 new internal packages, fully tested:
  - internal/goals: Goal struct + store + service, slug collision dedup,
    backing-channel auto-create via ChannelCreator adapter
  - internal/goaltasks: goal_tasks table with denormalized 16 KB
    ancestry snapshots, single-statement optimistic-lock atomic claim,
    recursive-CTE cost rollup, state machine, per-billing-code rollup
  - internal/secrets: NaCl-secretbox encrypted blobs, user/agent/task
    scope precedence, sanitized env injection, master-key bootstrap
  - internal/trust additions: ConfigHash (deterministic SHA-256 of
    model + prompt + tools + skills + mcp + subagents, sorted),
    DelegationCap (tier + tool-scope + budget + depth enforcement),
    append-only Ledger with exponential time-decay rolling score and
    70%-of-parent child seeding. Existing trust package unchanged.

- Critical invariants under test:
  - 50-goroutine concurrent claim race → exactly one winner per round
  - ConfigHash stable under shuffled array inputs, sensitive to
    capability changes
  - DelegationCap full tier × tool-scope matrix
  - Ledger time-decay + parent seed at 70 % ± 1 %
  - Secret name sanitization, scope precedence, plaintext never
    returned via MCP-equivalent paths

- internal/agents/types.go extended with dynamic-spawning columns
  (config_hash, parent_agent_id, spawn_depth, system_prompt,
  autonomy_tier, tool_scope_json, quarantined_at). Existing tests
  still pass.

- cmd/docgardener: self-contained demo binary driving the end-to-end
  flow. `docgardener run` creates a goal, builds a task tree with
  denormalized ancestry, spawns 3 specialists (each going through
  real delegation-cap validation and config-hash computation and
  70 %-of-parent reputation seeding), claims tasks atomically, runs
  them through the state machine, records reputation evidence.
  `docgardener report` queries all of that back out and renders a
  rich dark-mode HTML report (header, spend metrics, task tree,
  spawned-agent cards with reputation bars, cost breakdown, artifacts,
  timeline).

- examples/doc-gardener: start.sh / run_task.sh / report.sh / stop.sh
  mirroring the cold-topic-explainer pattern. Launches an isolated
  synapbus instance on port 18089, drives the demo, renders
  report.html, cleans up. Full README documenting what's real vs
  deferred, plus examples/README.md listing both examples.

- specs/018: tasks.md updated with MVP completion status; legacy tasks
  naming collision noted.

Deferred (marked explicitly in example README):
- Real LLM-driven coordinator (needs MCP tool wiring + prompt
  iteration)
- Real subprocess runs (needs reactor integration with task_id on
  ExecRequest)
- Full MCP tool surface (contracts are written at
  specs/018-dynamic-agent-spawning/contracts/mcp-tools.md)
- Svelte /goals UI (REST endpoints remain a follow-up)
- Full budget race + quarantine auto-trigger wiring
- Full resource-request → secrets fulfill reaction-workflow path

Cross-compiles clean for linux/amd64 and darwin/arm64 with no CGO
(SC-010). All new package tests pass (SC-004, SC-005, SC-007).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 15:29:21 +03:00
Algis DumbrisandClaude Opus 4.6 d9d1b7fee2 spec(018): dynamic agent spawning — full design
Complete speckit spec for the feature: a coordinator-driven system where
a human types a goal, a meta-agent decomposes into a task tree, proposes
spawning specialist sub-agents, runs them on heartbeats, verifies their
outputs, and iterates.

Includes:
- spec.md (9 user stories, 46 FRs, 12 SCs)
- plan.md (constitution check PASS)
- research.md (17 design decisions documented)
- data-model.md (5 migrations)
- contracts/mcp-tools.md (9 new MCP tools + 5 REST endpoints)
- quickstart.md (10-minute runbook)
- tasks.md (135 tasks, 12 phases, MVP at phase 8)
- checklists/requirements.md (quality gates)

Doc-gardener example is the end-to-end acceptance test.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 15:00:25 +03:00
Algis DumbrisandClaude Opus 4.6 fee73e33a0 feat(ux): run detail page + reaction pills + captured prompt/response
Makes the Web UI reflect what agents are actually doing: reactions
on DMs that trigger a subprocess run, a per-run detail page that
shows the exact prompt the model received and the raw response, and
cross-linked reactive_runs ↔ harness_runs data for a single composite
API call.

Migration 020 (internal/storage/schema/020_harness_run_detail.sql):

  ALTER TABLE harness_runs ADD COLUMN reactive_run_id INTEGER;
  ALTER TABLE harness_runs ADD COLUMN prompt          TEXT;
  ALTER TABLE harness_runs ADD COLUMN response        TEXT;
  CREATE INDEX idx_harness_runs_reactive ON harness_runs(reactive_run_id);

internal/harness:

  * ExecRequest.ReactiveRunID — reactor pins the reactive_runs row id
    so the observer can JOIN the two tables.
  * ExecResult.Prompt / Response — the subprocess harness reads
    prompt.txt / response.txt that wrappers write into the workdir,
    and runs.Store persists them (capped at 32 KiB each).
  * runs.Run struct now has JSON tags — previously the API returned
    PascalCase field names that didn't match the Web UI's snake_case
    TypeScript types.
  * New runs.Store.GetByReactiveRunID for the composite API endpoint.
  * Test schema updated to include the new columns.

internal/reactor:

  * New ReactionNotifier interface + SetReactionNotifier.
  * dispatchHarness now reacts `in_progress` on the triggering DM
    before spawning the goroutine.
  * runHarness reacts `done` on success, `reject` on failure. The
    existing reactionPriority ordering means the terminal reaction
    wins for badge display — no need to remove in_progress first.
  * dispatchHarness sets ExecRequest.ReactiveRunID.

cmd/synapbus/main.go:

  * reactorReactionAdapter: adapts reactions.Service.Toggle to the
    reactor's one-shot AddReaction signature.
  * HarnessRunsStore wired into the API router config.

internal/api/runs_handler.go — GetRun composite endpoint:

  The GET /api/runs/{id} response now returns everything the Web UI
  needs to render the run detail page in one call:

    {
      "run":              <reactive_runs row>,
      "harness_run":      <linked harness_runs row with prompt/response>,
      "agent":            <current agent snapshot with harness_config_json>,
      "trigger_message":  <DM that started the run>,
      "outgoing_message": <first DM the agent produced after startedAt>
    }

  The outgoing-message lookup wraps both sides of the created_at
  comparison in datetime() so SQLite parses the stored 'YYYY-MM-DD
  HH:MM:SS' and the Go-emitted RFC3339 into the same canonical form
  before comparing — a raw string compare was silently returning no
  rows.

internal/api/router.go: HarnessRunsStore field in RouterConfig, wired
through to NewRunsHandler.

examples/cold-topic-explainer/wrapper.sh:

  Writes prompt.txt and response.txt alongside gemini.stdout.raw so
  the subprocess harness can capture "what the model saw" and "what
  the model said" post-hoc.

web/src/lib/components/MessageList.svelte:

  New ReactionPills render below each message body when the message
  carries a `reactions` array (already populated by
  EnrichMessages/ReactionEnricher on the server side). Makes the
  👀 in_progress / ✔ done / ❌ reject lifecycle visible in every DM
  view and conversation.

web/src/routes/runs/[id]/+page.svelte (NEW):

  New run detail page at /runs/:id with sections:

    1. Header strip — agent, status pill, backend badge, trigger
       info, duration, tokens in/out, cost, exit code, trace id.
    2. Triggering message — body + sender.
    3. What the model saw — GEMINI.md / CLAUDE.md from agent snapshot
       + the captured rendered prompt (byte count on each summary
       bar, collapsible details).
    4. What the model said — captured response, falling back to
       logs_excerpt or error_log when unavailable.
    5. Outgoing message — body + recipient + status.
    6. Metadata — reactive_run.id, harness_run.run_id, backend,
       session_id, tokens_cached, k8s_job, agent trigger config.

  Styled against the existing dark tailwind system — no design
  overhaul, fits the current aesthetic (editorial sectioning,
  monospace for code-like content, accent-blue for links,
  accent-purple for system-instructions, accent-green for model
  output, accent-red for errors).

web/src/routes/runs/+page.svelte: the inline expand panel now has
a "View full details →" link next to the Retry button.

E2E VERIFIED on a live subprocess run:

  * Topic: "why does the subprocess harness materialise GEMINI.md
           alongside .gemini/settings.json in the per-run workdir?"
  * 3 subprocess runs + 3 reactive_runs + 3 harness_runs, all linked.
  * message_reactions: 6 rows — in_progress + done for each hop.
  * GET /api/runs/1 returns a composite with
    harness_run.prompt=883 bytes, harness_run.response=550 bytes,
    reactive_run_id=1, trigger_message populated, outgoing_message
    populated (decomposer-pro → writer-flash), agent.gemini_md=747
    bytes. All keys are snake_case as the Svelte types expect.

  Full go test ./... green. `vite build` green. Demo instance still
  running on port 18088 for browser verification.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 09:26:18 +03:00
Algis DumbrisandClaude Opus 4.6 b140879bc3 fix(examples): own demo agents by the algis user, not admin
start.sh was passing --owner 1 to every `agent create` call, under the
assumption that the freshly created `algis` user would be user id 1.
It isn't — the `admin` user is auto-seeded at id=1, so `algis` comes
in at id=2. Result: the 3 AI agents AND the `algis` human agent were
owned by `admin`, and when the user logged into the Web UI as `algis`
the message handler called GetHumanAgentForUser(2) which returned a
different auto-created `algis-human` agent (id=5, owned by user 2).
The critic DM'd the name `algis` → landed on agent id=1 (admin-owned),
but the UI listed DMs for `algis-human` → panel showed
"No conversations" despite reactive_runs clearly showing the chain
succeeded.

Fix: after `user create`, query sqlite for the algis user id and use
that value as --owner for every subsequent agent create. Bails with a
clear error if the id lookup fails or returns 1 (sanity check that
admin/algis aren't conflated).

Verified end-to-end:
  1. ./stop.sh && ./start.sh — new instance, ownership correct from
     the first `agent create` call.
  2. ./run_task.sh with a fresh topic — 3 subprocess runs succeeded,
     messages #1 (algis → decomposer-pro) and #4 (critic-lite → algis)
     now belong to an agent owned by user 2, so
     GetHumanAgentForUser(2) returns the same agent the messages are
     addressed to.
  3. sqlite3 agents table shows all four agents with owner_id=2.
  4. Chrome navigation to http://localhost:18088/ reaches the login
     form with no errors — once the user logs in as
     algis / algis-demo-pw the Direct Messages panel will show the
     four-message conversation.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 08:54:57 +03:00
Algis DumbrisandClaude Opus 4.6 95491402a9 fix(otel,examples): schema-URL merge + stale embedded web dist
Two independent fixes hit while running the cold-topic-explainer demo
end-to-end.

1. observability/otel.go — schema URL conflict on Init

  When SYNAPBUS_OTEL_ENABLED=1, Init() failed with:

    observability: build resource: conflicting Schema URL:
      https://opentelemetry.io/schemas/1.26.0 and
      https://opentelemetry.io/schemas/1.21.0

  resource.Default() ships with schema 1.26.0 (newer otel/sdk) but I
  was passing semconv.SchemaURL from v1.21 into a NewWithAttributes
  call. resource.Merge rejects that.

  Fix: use resource.NewSchemaless for the service.* attributes so our
  side of the merge has no schema URL and slots cleanly into whatever
  Default provides. ServiceVersion is now only attached when non-empty
  (avoids a stray service.version="" attribute).

  Two new regression tests:
    TestInit_EnabledSucceeds       — Enabled=true with all fields set
    TestInit_EnabledWithNoVersion  — Enabled=true with empty version
  Both point at an unroutable endpoint so the batcher never actually
  exports; the bug reproduced during Init(), which is all we need.

2. examples/cold-topic-explainer/start.sh — rebuild embedded SPA

  The Svelte Web UI loaded blank because internal/web/dist/ had a
  mismatched index.html + stale _app/immutable/entry/ assets (a build
  had updated index.html but not the chunks, so every asset URL fell
  through to the SPA HTML fallback and the browser tried to execute
  HTML as JavaScript).

  The canonical path is `make web`, but start.sh never ran it, so a
  working demo depended on the developer having run `make web` first.

  Fix: start.sh now rebuilds the SPA when web/src is newer than the
  embedded dist/index.html, using the already-installed
  web/node_modules (no reinstall). Falls back with a "run make web
  once" hint when node_modules isn't present. This keeps the fast
  path fast (~2s vite build after cache warm) and eliminates the
  silent-stale-dist trap.

E2E verified after both fixes:
  * SYNAPBUS_OTEL_ENABLED=1 start.sh no longer crashes.
  * `curl /_app/immutable/entry/start.*.js` returns real JavaScript
    (Content-Type: text/javascript) instead of the index.html
    fallback.
  * Chrome-in-MCP navigation to http://localhost:18088/ renders the
    login form with no SynapBus-originated console errors.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 08:38:40 +03:00
Algis DumbrisandClaude Opus 4.6 04e4e3ca37 feat(harness): GEMINI.md support + cold-topic-explainer example
Adds everything needed to run a real multi-Gemini-model reactive agent
loop end-to-end on SynapBus.

internal/harness/subprocess/config.go:
  * AgentConfig.GeminiMD — content of workdir/GEMINI.md
  * MaterialiseAgentConfig writes GEMINI.md AND workdir/.gemini/settings.json
    when gemini_md is set. The settings file carries the same mcp_servers
    list as .mcp.json (so a Gemini child running from the workdir gets
    the exact MCP surface the operator configured, not the user's
    ~/.gemini/settings.json).
  * 2 new config_test cases: GEMINI.md + .gemini/settings.json round
    trip, GEMINI.md with empty mcp_servers still writes the settings
    file (explicitly clearing any inherited home config).

internal/harness/registry.go — BUG FIX:
  Resolve() now honours agent.HarnessName (explicit selection) BEFORE
  the inference chain, matching the reactor's own agentBackendKind
  policy. Previously, when multiple backends were registered,
  Resolve would pick "webhook" for every non-K8s agent — even when the
  agent's HarnessName was "subprocess" — because the original fallback
  chain put webhook first. This is why the first cold-topic-explainer
  run failed with "webhook: agent has no webhook config". Discovered
  during e2e testing.

internal/admin/socket.go + cmd/synapbus/admin.go:
  New `messages.send` admin command (socket + CLI). Sends a DM as any
  agent through the messaging service, bypassing the REST/MCP auth
  layers. Local-only via the admin Unix socket, so the threat model is
  "whoever can reach the socket is already admin".

  CLI:
    synapbus messages send --from X --to Y --body "..." [--priority N]
    synapbus messages send --from X --to Y --body-file path
    echo "..." | synapbus messages send --from X --to Y

  Used by the harness shell wrappers (so Gemini subprocess agents can
  DM each other) and by run_task.sh (to kick off a chain as a human
  user without implementing the REST session flow).

examples/cold-topic-explainer/ (NEW):
  Runnable 3-agent Gemini demo that exercises the subprocess harness,
  reactive triggers, recursive update, and all the preconditions (depth,
  budget, cooldown) end-to-end on a separate isolated synapbus instance.

  Layout:
    README.md         — usage + troubleshooting + cost notes
    start.sh          — builds synapbus, launches on port 18088 with
                        ./data, creates user + agents + harness configs,
                        marks agents reactive via sqlite3
    run_task.sh       — sends initial DM algis → decomposer-pro, polls
                        reactive_runs + messages for the FINAL: reply,
                        prints the result or dumps reactive_runs on
                        timeout for debugging
    stop.sh           — SIGTERM + 5s grace + SIGKILL fallback
    wrapper.sh        — shared subprocess local_command: reads
                        message.json + GEMINI.md, calls gemini headless
                        with --approval-mode yolo, strips the
                        "MCP issues detected" noise prefix, routes the
                        cleaned response to the next agent via
                        `synapbus messages send` over the admin socket
    configs/
      decomposer-pro.json  — gemini-3.1-pro-preview
                              (gemini-2.5-pro is currently capacity-
                              exhausted on Google's side)
      writer-flash.json     — gemini-2.5-flash
      critic-lite.json      — gemini-2.5-flash-lite
    .gitignore        — data/, bin/, synapbus.log, .synapbus.pid

  The wrapper does NOT rely on gemini's MCP tool-calling (which was
  unreliable in testing). Gemini is used as a pure text generator; the
  shell decides routing based on AGENT_ROLE:
    - decomposer → NEXT_AGENT (writer)
    - writer     → NEXT_AGENT (critic)
    - critic     → OWNER_AGENT if response starts with FINAL:,
                   REVISE_AGENT otherwise

E2E VERIFICATION (real run, real Gemini, not a mock):

  Topic: "how does SynapBus unify message delivery, reactive agent
          triggers, and harness runs on a single SQLite database?"

  Result (from data/synapbus.db after one successful run):
    harness_runs:
      #1 decomposer-pro subprocess success 106s
      #2 writer-flash    subprocess success 155s
      #3 critic-lite     subprocess success  10s
    reactive_runs: 3 rows, all succeeded, trigger_from chain:
      algis → decomposer-pro → writer-flash → critic-lite
    messages:
      #1 algis → decomposer-pro (topic)
      #2 decomposer-pro → writer-flash (Q1/Q2/Q3 breakdown)
      #3 writer-flash → critic-lite (3-paragraph draft)
      #4 critic-lite → algis (FINAL: + polished 3-paragraph explainer)

  Critic converged in one pass (all scores ≥ 8), so the writer↔critic
  refinement loop didn't need to recurse — but the plumbing for it
  (REVISE: branch in wrapper.sh, depth limit in reactor) is wired and
  ready. Flipping the critic's acceptance bar exercises the recursion.

  Full `go test ./...` remained green through all changes.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-13 22:01:36 +03:00
Algis DumbrisandClaude Opus 4.6 b8a70bfe66 feat(harness): Option C — subprocess agent config (CLAUDE.md, MCP, skills)
Makes the subprocess backend fully self-contained: each agent carries
its instructions, MCP servers, skills, and subagents in its
harness_config_json column, viewable in the Web UI, editable via CLI.

internal/harness/subprocess/config.go (NEW):

  AgentConfig struct with optional fields:
    - claude_md       → workdir/CLAUDE.md
    - agents_md       → workdir/AGENTS.md
    - mcp_servers     → workdir/.mcp.json (Claude Code format)
    - skills          → workdir/.claude/skills/<name>/SKILL.md
    - subagents       → workdir/.claude/agents/<name>.md
    - env             → layered into child env (after k8s_env_json,
                        before caller overrides)

  ParseAgentConfig  tolerates empty / returns error on invalid JSON.
  MaterialiseAgentConfig writes all artifacts into the workdir with
  path-traversal sanitisation on skill/subagent names.

subprocess.Harness.Execute now calls Parse + Materialise before exec,
so an agent's declarative config is on disk by the time the child
CLI's cwd lookup fires. buildEnv takes the parsed config and overlays
cfg.Env on top of k8s_env_json.

Tests:
  config_test.go — 6 cases: empty, invalid JSON, full round-trip,
    materialise writes all artefacts, empty is a no-op, skill names
    are sanitised against "../escape" / "/etc/passwd", mcp entries
    without a name are dropped.
  subprocess_test.go — 2 new e2e cases: agent with CLAUDE.md + mcp
    servers + skills + env sees all of them from inside the child via
    cat/echo; invalid harness_config_json surfaces as Execute error.

internal/agents/store.go:

  AgentStore gains UpdateHarnessConfig(ctx, name, harnessName,
  localCommand, harnessConfigJSON). Empty strings leave a field
  unchanged; literal "-" clears (sets to NULL). Returns sql.ErrNoRows
  on missing agent. AgentService exposes Store() so admin handlers
  can reach it without adding a full service method for a
  config-set-style operation.

  store_test.go: 6-subcase test covers set-all, partial update, clear,
  unknown agent, and no-field no-op.

internal/admin/socket.go:

  Two new admin commands:
    harness.config_get {agent_name} → {harness_name, local_command,
        harness_config_json, harness_config (parsed), parse_error?}
    harness.config_set {agent_name, harness_name?, local_command?,
        harness_config_json?} → updated fields
  config_set validates JSON shape before calling the store; null /
  "-" literals clear the column.

cmd/synapbus/admin.go:

  New top-level `harness config` command group:
    synapbus harness config get --agent <name> [--raw]
    synapbus harness config set --agent <name>
        [--harness-name subprocess]
        [--local-command '["claude","--print"]']
        [--file config.json]    # or pipe from stdin
        [--clear]
    synapbus harness config edit --agent <name>
        # fetches current config, opens $VISUAL/$EDITOR/vi,
        # validates JSON on save, writes back via config_set

web/src/routes/agents/[name]/+page.svelte:

  New read-only "Harness" panel on the agent detail page:
    - Resolved backend badge (explicit or inferred from k8s_image /
      local_command / harness_config_json.url)
    - Grid summary: CLAUDE.md size, AGENTS.md size, MCP server count,
      skills count
    - Collapsible details for CLAUDE.md, AGENTS.md, each MCP server
      (name / type / url|command / header count), skill filenames,
      subagent filenames, env vars
    - Footer hint showing the CLI edit command
  No edit controls — editing is CLI-only by design (safer, fits an
  ops-heavy workflow).

Verified: full project test suite (40+ packages including integration
tests) plus `vite build` of the Svelte app all green; `go vet ./...`
clean; the existing TestSubprocess_Execute_MaterialisesHarnessConfig
e2e test proves the round-trip from harness_config_json → workdir →
child process works end-to-end.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-13 20:57:28 +03:00
Algis DumbrisandClaude Opus 4.6 8ddccd2452 feat(reactor): route non-K8s reactive runs through harness registry
Phase 7: the reactor now branches on agent backend kind.

Reactor changes (internal/reactor/reactor.go):

  * Adds `registry *harness.Registry` field + `SetHarnessRegistry`.
  * `agentBackendKind()` picks k8s | subprocess | webhook | none from
    the agent's `HarnessName`, `K8sImage`, `LocalCommand`, and
    `HarnessConfigJSON` fields. Explicit `HarnessName` wins.
  * `evaluateTrigger` applies the same preconditions (depth, daily
    budget, cooldown, already-running, pending_work coalescing) to
    every backend — a subprocess agent mentioned in a channel now
    goes through the exact same rate limits a K8s agent does.
  * K8s agents keep the existing `createJob` fast-return path with
    the async poller for restart safety. Non-K8s agents use a new
    `dispatchHarness` that inserts the reactive_runs row, spawns a
    detached goroutine, blocks on `Registry.Execute`, and writes the
    terminal status / error_log / metrics / failure DM on return.
  * Import `harness`, `messaging`, `google/uuid` for building the
    ExecRequest.

main.go wiring:

  * Build one `harness.Registry` with all three real backends:
    `k8sjob.New(k8sRunner, …)`, `subprocess.New(Config{BaseDir:
    dataDir/harness/subprocess}, …)`, `webhook.New(Config{}, …)`.
  * Attach a `runs.Store` as the registry Observer so every dispatch
    writes a harness_runs row — no per-caller code required.
  * Hand the registry to the reactor via `SetHarnessRegistry`.
  * Log the registered backend names at startup.

Tests (internal/reactor/reactor_test.go):

  * New `insertSubprocessAgent`, `newHarnessReactor`, `waitForRun`,
    and `fakeNotifier` helpers.
  * Seven new tests that register a stub harness under "subprocess"
    and verify: success from @mention, failure recorded + DM sent,
    depth-exceeded skipped, budget-exhausted skipped, cooldown
    skipped, already-running queued, no-backend fails cleanly. Each
    checks the harness stub is NOT called when a precondition skips.
  * Existing `TestReactorNoK8sImage` keeps working — the old
    k8s-specific error message is replaced with the backend-agnostic
    "no backend configured" phrasing.
  * `setupTestDB` now pins `SetMaxOpenConns(1)`: modernc.org/sqlite
    in-memory DBs give each pool connection a fresh empty database,
    which races the new dispatchHarness goroutine and main-test
    goroutine. Pinning is the standard workaround.

The overall behaviour: `@local-agent` in a channel message now starts
the configured subprocess/webhook under the same depth/budget/cooldown
rate limits as a K8s agent, tracked in reactive_runs and harness_runs,
instrumented with an OTel span, with trace context propagated into the
child via env vars. Failure DMs go to the human owner, as with K8s.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-13 20:38:10 +03:00
Algis DumbrisandClaude Opus 4.6 574897df4d feat: harness-agnostic wrappers + OpenTelemetry integration
Introduces internal/harness — a minimal Harness interface inspired by
GoogleCloudPlatform/scion — plus four backends (k8sjob, subprocess,
webhook, stub) and an OTel-traced Registry that spans every dispatch
and injects W3C trace context into child processes via env vars.

Phases landed together on this branch:

  1. internal/harness scaffold: Harness/Capabilities/ExecRequest/
     ExecResult/Budget/Usage types, Registry with Resolve/Execute,
     in-memory stub backend.
  2. internal/harness/k8sjob: wraps existing k8s.JobRunner behind the
     Harness interface with a Waiter abstraction (real clientset +
     test fake). BuildHandler exports the per-agent config logic.
  3. internal/harness/subprocess: os/exec-based backend (Mac+Linux),
     per-run workdir, result.json handoff, bounded log capture,
     Budget-driven wall-clock timeout.
  4. internal/harness/webhook: synchronous HTTP POST with HMAC
     signing via internal/webhooks.ComputeHMACSignature, per-agent
     URL/secret/timeout read from harness_config_json.
  5. internal/observability: OTel tracer init via OTLP HTTP (opt-in
     via SYNAPBUS_OTEL_ENABLED), W3C propagator always installed;
     Registry.Execute starts a harness.execute span per dispatch and
     calls InjectTraceContext into req.Env so children inherit it.
  6. internal/harness/runs: SQLite-backed Observer that persists a
     harness_runs row per dispatch with status, usage, cost, duration,
     trace_id, session_id, and a bounded logs excerpt.

Schema: new migration 019_harness.sql adds agents.harness_name /
local_command / harness_config_json columns and the backend-agnostic
harness_runs table with indices on (agent, created_at), (status),
(trace_id), (run_id). internal/reactor/reactor_test.go inline schema
updated to match.

Deployment: deploy/kubic/otel-collector.yaml stands up an otel-collector
Deployment + ConfigMap + ClusterIP Service in the synapbus namespace on
kubic, receiving OTLP gRPC (4317) and HTTP (4318) and exporting debug
output until a Tempo/Jaeger backend lands.

Docs: docs/harness-otel-research.html compares scion and paperclip
side-by-side and maps the current synapbus executor surface; its
companion docs/harness-otel-design.md carries the phase plan, span
taxonomy, and migration schema verbatim.

The reactor currently still calls k8s.JobRunner directly — rewiring it
through the Registry is a follow-up, intentionally out of scope for
this branch to keep the refactor reversible. The new packages are
independently tested (~78 new tests across 7 packages) and the full
project test suite passes.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-13 20:23:04 +03:00
Algis DumbrisandClaude Opus 4.6 0e25fbcccb feat: sdk_backend + autonomous run integration
Adds benchmark/sdk_backend.py that routes model calls through either
the anthropic SDK (preferred, requires ANTHROPIC_API_KEY) or the
claude-agent-sdk as a Claude Code session fallback. agents.py and
baseline.py now go through this unified backend instead of calling
anthropic directly.

Ran benchmark/run.py --mode single-shot --question q1 end-to-end
with real Claude API calls via claude-agent-sdk. Real numbers:
- Marketplace (Haiku 4.5): 3314 tokens, F1 1.000 (exact match)
- Baseline (Sonnet 4.6): 697 tokens, F1 0.857 (penalized for "1783")
- Pareto verdict: FAIL (not strictly NW; marketplace wins quality,
  loses cost — informative failure per spec design).

Added autonomous_report.html (rich narrative with Pareto chart)
and autonomous_summary.md. All 34 Go packages still green.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 15:29:56 +03:00
Algis Dumbris 8fd42cb957 merge: 017-musique-benchmark MVP (Python harness + trio + Pareto report) 2026-04-11 15:18:49 +03:00
Algis Dumbris dc1ae4da60 merge: 016-agent-marketplace MVP (auction + manifests + reputation) 2026-04-11 15:18:46 +03:00
Algis DumbrisandClaude Opus 4.6 cda3365863 feat(016): agent marketplace MVP — capability manifests, auction channels, reputation ledger
Implements US1, US2, and US3 of spec 016 by layering a marketplace service
on top of existing primitives rather than reinventing them:

- Capability manifests (US2) reuse the wiki subsystem. Each agent publishes
  a per-agent article at slug "agent-<name>" and gets versioning, revision
  history, and FTS search for free.
- Auction channels (US1) reuse the existing auction channel type, swarm
  service, and task/bid store. post_auction / bid / award wrap post_task /
  bid_task / accept_bid and attach marketplace metadata (max_budget_tokens,
  domains, estimated_tokens, confidence, approach) in the task.requirements
  and bid.capabilities JSON blobs. Award converts the auction into a claim
  by DM'ing the winner at priority 8 with task_id metadata, so the existing
  claim/process/done lifecycle takes over with zero new machinery.
- Reputation ledger (US3) adds migration 018_agent_marketplace.sql with a
  new agent_reputation table keyed by (agent_name, domain). mark_task_done
  completes the task via the swarm service and writes one ledger row per
  declared domain using the reported actual_tokens and success_score.
  query_reputation returns a rolled-up summary plus recent entries for a
  given (agent, domain) pair — reputation is always a vector, never a
  global score (FR-013).

Also:
- Adds the "awarded" reaction type (FR-008) alongside existing approve/
  reject/in_progress/done/published. Migration 018 widens the reactions
  CHECK constraint via a table rebuild.
- 6 new actions added to the action registry (post_auction, bid, award,
  mark_task_done, read_skill_card, query_reputation) so the search tool
  can discover them and the execute tool can dispatch them.
- New internal/marketplace package (store.go + service.go).
- New internal/mcp/marketplace.go bridge handlers.
- New internal/mcp/marketplace_test.go covers the full auction lifecycle,
  capability manifest publish/read/update, self-bid rejection, non-auction
  channel rejection, and reputation summary aggregation.

Out of scope for MVP (deferred per spec prompt): US4 reflection loop,
tombstoning FR-020a/b, multi-owner quorums, auto-escalation on zero bids,
bootstrap exploration credit, epsilon-greedy selection, and the hard-stop
budget enforcement daemon (only soft recording of estimated vs actual is
included).

All existing tests pass; new marketplace tests pass.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 15:17:30 +03:00
Algis DumbrisandClaude Opus 4.6 02b8548eac feat(017): MuSiQue benchmark harness — marketplace stub, mixed-tier agents, Pareto scoring, HTML report
MVP implementation of the MuSiQue multi-agent benchmark (spec 017):

- benchmark/setup.py: downloads musique_v1.0.zip from the canonical Google
  Drive source (mirrors upstream download_data.sh). Idempotent.
- benchmark/curate.py: deterministic selection of 3 4-hop questions from
  the dev set sharing a US pivot entity; writes benchmark/trio.jsonl.
- benchmark/marketplace.py: in-process 016-marketplace stub with
  post_auction / bid / award / mark_done / query_reputation and a
  domain-scoped reputation ledger. Designed for mechanical swap to real
  SynapBus MCP tools.
- benchmark/agents.py: HaikuAgent + SonnetAgent, using the official
  anthropic SDK (no Claude Agent SDK, no subprocesses). Models pinned
  to claude-haiku-4-5-20251001 and claude-sonnet-4-6.
- benchmark/baseline.py: single Sonnet call with all 20 distractors
  plus chain-of-thought.
- benchmark/score.py: SQuAD-style normalized F1 + Pareto verdict
  (strictly northwest = PASS).
- benchmark/run.py: main entry. --mode single-shot, --question, --dry-run.
- benchmark/report.py: self-contained HTML with inline SVG scatter plot.
- benchmark/trio.jsonl: curated reproducible trio (all three converge on
  "Treaty of Paris" US territory cession).

Verified with benchmark/run.py --dry-run end-to-end; all 8 files
py_compile clean. Real-token execution is deferred to the user's main
session.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 15:12:19 +03:00
Algis DumbrisandClaude Opus 4.6 e77fd7afdf spec(017): MuSiQue MAS benchmark harness
Feature spec for Python benchmark that integration-tests 016 marketplace
against a real multi-hop reasoning task. 4 prioritized user stories:
P1 single-shot Pareto verification, P2 curated trio with dedup,
P3 learning tier, P1 rich HTML report. 23 FRs, 7 success criteria.

Also: brainstorming design doc at docs/superpowers/specs/ capturing
the 6 clarifying questions and chosen decisions (mixed-tier pool,
curated trio, tiered run modes, wait-for-016 execution strategy).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 15:02:59 +03:00
Algis DumbrisandClaude Opus 4.6 96db7c06a4 spec(016): agent marketplace spec + research reports
- specs/016-agent-marketplace: self-organizing marketplace spec with
  capability manifests, auction channels, domain-scoped reputation,
  and reflection loop. Four user stories (P1: auction + manifests,
  P2: reputation + reflection). 27 FRs, 10 success criteria, checklist.
- multiagent_systems_report.html: landscape of OSS MAS frameworks,
  coordination patterns (blackboard/stigmergy/contract-net/gossip),
  problem classes, toy benchmarks.
- agent_marketplace_guide_ru.html: Russian technical guide with
  terminology dictionary, Fermi walkthrough, Voyager lessons.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 14:59:51 +03:00
Algis DumbrisandClaude Opus 4.6 660da6d646 feat: wiki export/import CLI for backup and restore
Usage:
  synapbus wiki export --data ./data --output ./wiki-export
  synapbus wiki import --data ./data --input ./wiki-export

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-05 20:04:06 +03:00
Algis DumbrisandClaude Opus 4.6 75fb483ae9 docs: add spec and plan for 013-agent-wiki feature
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-05 19:14:46 +03:00
Algis DumbrisandClaude Opus 4.6 ea1554fd6e feat: wiki web UI — map of content, article view, revision history
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-05 19:06:51 +03:00
Algis DumbrisandClaude Opus 4.6 eea6176ca9 feat: agent wiki — articles, revisions, backlinks, FTS search, REST API
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-05 19:06:36 +03:00
Algis DumbrisandClaude Opus 4.6 53c8ba5bb0 feat: hybrid search (RRF fusion) + minimum similarity threshold
- Auto mode now runs both semantic and fulltext searches, merging
  results using Reciprocal Rank Fusion (RRF, k=60) for best of both
- New min_similarity parameter (default 0.25) filters semantic noise
- Results that match both sources are marked as "hybrid" match_type
- New getFloat bridge helper for MCP min_similarity parameter

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 16:56:26 +03:00
Algis DumbrisandClaude Opus 4.6 ed33da1093 docs: add spec 012-search-quality-platform improvements
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 11:54:58 +03:00
Algis DumbrisandClaude Opus 4.6 36750cc1d0 fix: escape FTS5 reserved words in fulltext search queries
Words like "to", "from", "and", "or", "not", "near" are FTS5 operators
and caused SQL errors (e.g. "no such column: to") when passed as search
queries. sanitizeFTS5Query() wraps each token in double quotes so they
are treated as literal phrase tokens by SQLite's FTS5 MATCH operator.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 11:53:14 +03:00
Algis DumbrisandClaude Opus 4.6 02d76babff sec: upgrade go-jose v3.0.3→v3.0.4 (fixes DoS via JWS parsing)
CVE: GHSA-go-jose DoS in parsing (GO-2025-3485)
The vulnerability allowed crafted JWS tokens to cause excessive
memory allocation. Affects our OAuth token validation path.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-31 14:22:45 +03:00
Algis DumbrisandClaude Opus 4.6 6b65e0130f fix: login shows correct error + brute-force protection
Release / Build darwin/amd64 (push) Canceled after 0s
Release / Build linux/amd64 (push) Canceled after 0s
Release / Build darwin/arm64 (push) Canceled after 0s
Release / Build linux/arm64 (push) Canceled after 0s
Release / Generate Homebrew Formula (push) Canceled after 0s
Release / GitHub Release (push) Canceled after 0s
Release / Docker Image (push) Canceled after 0s
Release / Publish to MCP Registry (push) Canceled after 0s
- Wrong password now shows "Invalid username or password" (was "Session expired")
- Client API differentiates 401 on login page vs elsewhere
- Added per-IP login rate limiter: 3 failures → blocked 1 minute
- 429 status code returned with remaining seconds in message
- Rate limit cleared on successful login

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
v0.14.6
2026-03-27 06:28:19 +02:00
Algis DumbrisandClaude Opus 4.6 726479a57f fix: DM partners query includes owned agents as conversation partners
Release / Build darwin/amd64 (push) Canceled after 0s
Release / Build linux/amd64 (push) Canceled after 0s
Release / Build darwin/arm64 (push) Canceled after 0s
Release / Build linux/arm64 (push) Canceled after 0s
Release / Generate Homebrew Formula (push) Canceled after 0s
Release / GitHub Release (push) Canceled after 0s
Release / Docker Image (push) Canceled after 0s
Release / Publish to MCP Registry (push) Canceled after 0s
Removed the 'peer NOT IN (owned)' filter that excluded all owned agents.
Since all agents (algis, research-*, social-commenter) are owned by the
same user, the filter was hiding all inter-agent DMs. Now shows all
unique DM partners regardless of ownership.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
v0.14.5
2026-03-26 09:35:46 +02:00
Algis DumbrisandClaude Opus 4.6 5dd5a6f23c fix: DM partners query uses SQL over all messages, not just inbox
Release / Build darwin/amd64 (push) Canceled after 0s
Release / Build linux/amd64 (push) Canceled after 0s
Release / Build darwin/arm64 (push) Canceled after 0s
Release / Build linux/arm64 (push) Canceled after 0s
Release / Generate Homebrew Formula (push) Canceled after 0s
Release / GitHub Release (push) Canceled after 0s
Release / Docker Image (push) Canceled after 0s
Release / Publish to MCP Registry (push) Canceled after 0s
The previous implementation only scanned inbox (pending messages),
so historical conversations with read/done messages were invisible.
New GetDMPartners() does a direct SQL query with window functions
to find all unique DM partners with most recent message preview
and unread count. Historical conversations now always show.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
v0.14.4
2026-03-26 09:30:10 +02:00
Algis DumbrisandClaude Opus 4.6 3b97429f20 fix: DM sidebar shows conversation partners, not owned agents
Release / Build darwin/amd64 (push) Canceled after 0s
Release / Build linux/amd64 (push) Canceled after 0s
Release / Build darwin/arm64 (push) Canceled after 0s
Release / Build linux/arm64 (push) Canceled after 0s
Release / Generate Homebrew Formula (push) Canceled after 0s
Release / GitHub Release (push) Canceled after 0s
Release / Docker Image (push) Canceled after 0s
Release / Publish to MCP Registry (push) Canceled after 0s
- New API: GET /api/dm/partners — returns DM conversation partners
  ordered by most recent message, with unread counts
- Sidebar DM section now shows actual conversation partners (agents
  you've exchanged messages with) instead of owned agents
- Each partner shows name, unread badge, clickable to /dm/{name}
- Fixes issue where all DMs were shown mixed in one view

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
v0.14.3
2026-03-26 09:19:22 +02:00
Algis DumbrisandClaude Opus 4.6 7848911a5f fix: ignore system DMs in reactor — prevent stalemate notification cascade
Release / Build darwin/amd64 (push) Canceled after 0s
Release / Build linux/amd64 (push) Canceled after 0s
Release / Build darwin/arm64 (push) Canceled after 0s
Release / Build linux/arm64 (push) Canceled after 0s
Release / Generate Homebrew Formula (push) Canceled after 0s
Release / GitHub Release (push) Canceled after 0s
Release / Docker Image (push) Canceled after 0s
Release / Publish to MCP Registry (push) Canceled after 0s
The StaleWorker sends DMs from 'system' to all channel members when
workflow messages are stuck in 'proposed' state. These DMs were
triggering reactive agent runs, which couldn't action the stale
messages, burning daily budget on wasted K8s Jobs.

Now: reactor silently ignores all messages from 'system' sender.
System notifications are for human review, not agent action.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
v0.14.2
2026-03-26 08:56:37 +02:00
Algis DumbrisandClaude Opus 4.6 4b8c574096 fix: Agent Runs page stuck on Loading — use $effect instead of onMount
Release / Build darwin/amd64 (push) Canceled after 0s
Release / Build linux/amd64 (push) Canceled after 0s
Release / Build darwin/arm64 (push) Canceled after 0s
Release / Build linux/arm64 (push) Canceled after 0s
Release / Generate Homebrew Formula (push) Canceled after 0s
Release / GitHub Release (push) Canceled after 0s
Release / Docker Image (push) Canceled after 0s
Release / Publish to MCP Registry (push) Canceled after 0s
The onMount + async pattern wasn't triggering Svelte 5 reactivity
properly. Switched to $effect with $user dependency (same pattern
used by Sidebar and other components). Also waits for auth before
loading data.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
v0.14.1
2026-03-26 08:30:36 +02:00
Algis Dumbris aed7cb5e98 Merge features 014+015: Reactive Agent Triggers + SQL Query Interface
Release / Build darwin/amd64 (push) Canceled after 0s
Release / Build linux/amd64 (push) Canceled after 0s
Release / Build darwin/arm64 (push) Canceled after 0s
Release / Build linux/arm64 (push) Canceled after 0s
Release / Generate Homebrew Formula (push) Canceled after 0s
Release / GitHub Release (push) Canceled after 0s
Release / Docker Image (push) Canceled after 0s
Release / Publish to MCP Registry (push) Canceled after 0s
v0.14.0
2026-03-26 07:41:37 +02:00
Algis DumbrisandClaude Opus 4.6 107b5e930d docs: add SQL query action to CLAUDE.md onboarding template
New agents now learn about the query action during onboarding:
tables (my_messages, my_channels, channel_messages), examples,
and limitations (100 rows, SELECT only, 5s timeout).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-26 07:26:46 +02:00
Algis DumbrisandClaude Opus 4.6 e5ee8d16e4 fix(015): remove SQL LIMIT injection — enforce in Go only
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-26 07:19:38 +02:00
Algis DumbrisandClaude Opus 4.6 bd1bccc692 feat(015): SQL query interface for agents + split read/write pools
Split Connection Pools:
- writeDB: MaxOpenConns=1, serializes all writes (no SQLITE_BUSY)
- readDB: MaxOpenConns=8, query_only=ON, for all SELECTs
- QueryDB() helper returns read pool when available

SQL Query Interface:
- New 'query' action via execute MCP tool
- Read-only enforcement (PRAGMA query_only=ON + SQL validation)
- Curated views: my_messages, my_channels, channel_messages
- Per-agent access control via CTE injection
- Auto LIMIT 100, 5s timeout, SELECT-only validation
- Blocks: INSERT, UPDATE, DELETE, DROP, PRAGMA, etc.
- 12 new tests (access control, validation, limits, CTEs)

Migration 016: agent query views (v_agent_messages, etc.)
Action registry: 30 actions (was 29, added 'query')
All 29 test packages pass.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-26 07:17:23 +02:00
Algis DumbrisandClaude Opus 4.6 b6fc298595 feat(015): add spec for SQL query interface + split connection pools
Two features:
1. SQL query action for agents via execute MCP tool — read-only,
   curated views, LIMIT/timeout, SELECT-only validation
2. Split read/write SQLite connection pools — writeDB (1 conn)
   + readDB (8 conns) to eliminate SQLITE_BUSY

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-26 07:04:55 +02:00
Algis DumbrisandClaude Opus 4.6 64c68c22be fix(014): prevent stuck runs by creating K8s Job before DB insert
The reactor was inserting the run record first, then creating the K8s
Job, then updating the record with the job name. If the update failed
(SQLITE_BUSY), the run would be stuck in 'running' with no job name,
making it invisible to the poller.

Now: create K8s Job first, then insert the run record with job name
already set in a single atomic write.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-26 06:09:16 +02:00
Algis DumbrisandClaude Opus 4.6 cf6066229f feat(014): add Prometheus metrics, Grafana dashboard, volume mounts, resource tuning
- Reactor Prometheus metrics: triggers_total, run_duration_seconds, agent_running, budget_used_today
- Integrated promauto metrics into hand-rolled WritePrometheus endpoint
- K8s runner: ImagePullPolicy=IfNotPresent, volume mounts, CLI args support
- Reactor: 2Gi/500m default resources (agent SDK needs it), 1h timeout
- Grafana dashboard "SynapBus Reactive Agents" with 8 panels:
  triggers by status, agent state, budget gauge, run duration,
  agent turns from Loki, reactor events log, agent container logs
- SQLite: busy_timeout=15s, synchronous=NORMAL, MaxOpenConns=4

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 22:15:54 +02:00
Algis DumbrisandClaude Opus 4.6 012b7f6fba fix: reduce SQLITE_BUSY errors under concurrent load
- Increase busy_timeout from 5s to 15s
- Set synchronous=NORMAL (safe with WAL, reduces fsync)
- Limit MaxOpenConns to 4 to reduce write lock contention
- Explicit wal_autocheckpoint=1000

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 18:48:44 +02:00
Algis DumbrisandClaude Opus 4.6 c6c96f64be feat(014): add Web UI Agent Runs page
- New /runs route with agent summary cards, run list, filtering
- Agent cards show budget usage, cooldown status, current state
- Expandable run rows with error logs and retry button
- API client: runs.list, runs.get, runs.retry, runs.reactiveAgents
- Sidebar navigation updated with "Agent Runs" link
- Rebuilt web dist

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 17:45:57 +02:00
Algis DumbrisandClaude Opus 4.6 6afe1853ad feat(014): implement reactive agent triggering engine
- Migration 015: extends agents with trigger config, adds reactive_runs table
- Reactor engine: decision chain (mode, depth, budget, cooldown, sequential)
- Reactor store: SQLite persistence for runs with RFC3339 timestamps
- Reactor poller: K8s Job status polling (15s interval)
- Failure notifier: system DM to owner on job failure
- REST API: /api/runs, /api/runs/:id, /api/runs/:id/retry, /api/agents/reactive
- Agent model: trigger_mode, cooldown, budget, depth, k8s_image, pending_work
- K8s runner: GetClientset() for poller
- All 28 test packages pass (8 new reactor tests)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 17:42:48 +02:00
Algis DumbrisandClaude Opus 4.6 68f356b5e3 feat(014): add implementation plan, research, data model, and contracts
Phase 0: research.md — 7 decisions on polling, coalescing, depth, cooldown
Phase 1: data-model.md — schema for reactive_runs + agent extensions
Phase 1: contracts — REST API, MCP tools, CLI commands
Phase 1: quickstart.md — developer onboarding guide

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 17:23:48 +02:00
Algis DumbrisandClaude Opus 4.6 ea256ed526 feat(014): add reactive agent triggering spec
Specifies the reactive agent system: DM/@mention triggers K8s Jobs
with reactor decision engine, cooldown/budget/depth rate limiting,
sequential execution with coalescing, Web UI Agent Runs panel,
failure notifications, and admin CLI.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 17:20:52 +02:00
Algis DumbrisandClaude Opus 4.6 0e28c0b45e feat: fix attachment handling — display in DMs, enrich in MCP, allow all file types
- Show attachment previews on DM messages (was missing, only channels had it)
- Add file upload button to DM compose bar with paperclip icon
- Enrich messages with attachment data in all MCP bridge functions
  (read_inbox, claim_messages, search, channel_messages, list_by_state)
- Remove file type restrictions — allow any file type, keep 50MB size limit
- Rebuild web dist

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-24 13:00:28 +02:00
Algis DumbrisandClaude Opus 4.6 8134a7eef5 chore: rebuild web dist with v0.12.2
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 21:11:39 +02:00
Algis DumbrisandClaude Opus 4.6 91a1f2adcb feat: add pagination + body truncation to list_by_state
Prevents 181K+ responses when channels have many messages with long
bodies. New params: limit (default 20, max 100), offset (default 0),
max_body_length (default 500 chars when include_messages=true).

Response now includes total count alongside paginated results.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 20:09:38 +02:00
Algis DumbrisandClaude Opus 4.6 faab0f7f17 chore: rebuild web dist with truncation fix, update agent context
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 20:06:53 +02:00