Root cause
----------
ExpireTasks in internal/channels/task_store.go ran a single unbounded
UPDATE filtered on (status='open' AND deadline IS NOT NULL AND deadline < now).
The only indexes on tasks were idx_tasks_status(status) and
idx_tasks_channel(channel_id). With status cardinality of ~4 and a growing
auction-tasks table on kubic, the planner used idx_tasks_status to enumerate
all open rows then evaluated deadline per row, holding a SQLite write
transaction the whole time. Under WAL contention with concurrent writers
(message inserts, consolidator) the worker's 30s context regularly expired,
producing the recurring expiry-worker log line.
Fix
---
1. New migration 031_tasks_expiry_index.sql: partial composite index
idx_tasks_expiry(status, deadline) WHERE status='open' AND deadline IS NOT NULL.
This is the exact predicate ExpireTasks uses, so the planner now seeks
straight to eligible rows. The partial form keeps the index empty for the
steady-state majority of rows (completed/cancelled), so writes elsewhere
aren't penalized.
2. Batch the UPDATE in chunks of 500 (rowid IN subquery; UPDATE ... LIMIT
isn't compiled into modernc.org/sqlite by default). Bounded write
transactions stop the worker from starving other writers and let it
observe context cancellation between batches.
Perf
----
New test exercises 2400 mixed rows (1200 expirable). With the index +
batching, expiry finishes in ~3ms inside a 5s context; without the index a
regression to full status-scan would be measurably worse and is also
guarded by an EXPLAIN QUERY PLAN test.
Operational notes
-----------------
- Migration is additive and idempotent (CREATE INDEX IF NOT EXISTS). No
backfill needed; it will apply on next pod startup.
- After rollout, expiry-worker error logs should clear within one tick
(default 1m).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Drain-on-demand: SYNAPBUS_DREAM_PARALLEL (default 1) and
`synapbus memory dream-run --parallel N` fan out N concurrent
dream-agent k8s Jobs per (owner, job_type) in one shot. Set
high (e.g. 8) to drain backlog quickly, then back to 1 for normal
hourly operation.
Schema:
- migration 030_dream_parallelism: adds slot INTEGER NOT NULL DEFAULT 0
to memory_consolidation_jobs. Drops + recreates the partial unique
in-flight index as (owner, job_type, slot) so slots 0..N-1 each hold
one in-flight job independently.
Stores:
- JobsStore.CreateOnSlot + CreateNextAvailableSlot.
- ConsolidatorWorker.ForceRunN dispatches N parallel jobs through the
existing launchOne path (extracted from ForceRun).
- core_rewrite coerces to N=1 regardless of the knob — per-(owner,
agent) blob is wholesale-replace and concurrent rewrites would race.
Three bug fixes discovered while bringing the parallel path up on
kubic:
1. k8s Job names collided on rapid relaunch because runner.go used
"synapbus-<agent>-<msg_id>", and dream dispatches have msg_id=0.
Now appends a unique (timestamp%1e6, 4-byte random) suffix when
msg_id is zero; historical "synapbus-<agent>-<id>" prefix preserved.
2. memory_list_unprocessed didn't actually exclude already-refined
messages — the contract said it should, the implementation
returned the same oldest-50 every cycle. The dream agent kept
re-refining the same set: 221 refines links touched only 55
unique dst messages, so progress flat-lined. Added the
NOT IN (refines/duplicate_of/superseded_by) filter and a
from_agent NOT LIKE 'dream:%' clause so the agent never refines
its own reflections.
3. The k8sjob harness was constructed with nil Waiter in main.go,
so every dream dispatch failed instantly with "k8sjob: no Waiter
configured". Now builds a ClientsetWaiter from the in-cluster
clientset.
Plus admin/server.go gets DreamRunN closure + DefaultDreamParallel
(sourced from MemoryConfig.DreamParallel). admin/socket.go
handleMemoryDreamRun accepts `parallel` arg and returns job_ids[].
CLI admin command grows --parallel N flag.
Live evidence from kubic (image v0.21.0-amd64):
1 CLI call with --parallel 8 produced 8 job rows on slots 0..7,
spawned 8 distinct k8s Jobs with unique suffixes, retired ~86
unprocessed messages in <1 min (vs ~10/cycle for the buggy
serial version pre-fix-2).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Backend (Go, in this commit):
- migration 029_memory_dream_usage: per (date, owner) counters for
tokens_in/out, jobs_started/succeeded/failed/circuit_broken
- DreamUsageStore + UsageGate (internal/messaging/dream_usage.go).
Gate inspects today's usage against new env knobs:
- SYNAPBUS_DREAM_RECENT_WINDOW (default 336h / 14d)
- SYNAPBUS_DREAM_DAILY_TOKEN_LIMIT_IN (default 1M)
- SYNAPBUS_DREAM_DAILY_TOKEN_LIMIT_OUT (default 200k)
- SYNAPBUS_DREAM_DAILY_JOB_LIMIT (default 100)
- Consolidator now bounds watermarks + core_rewrite eligibility by the
recency window. core_rewrite skipped for owners with no in-window
activity. ForceRun honors the breaker.
- Recency fallback in BuildContextPacket + memory_list_unprocessed now
accept RecentWindowDays so injection and dream queries see the same
14d slice.
- Prometheus metrics registered (internal/metrics/metrics.go):
synapbus_dream_jobs_total{owner,job_type,status},
synapbus_dream_tokens_total{owner,direction},
synapbus_dream_job_duration_seconds{owner,job_type},
synapbus_dream_circuit_broken_total{owner,reason},
synapbus_injection_packets_total{tool},
synapbus_injection_memories_per_packet{tool},
synapbus_injection_packet_chars{tool},
synapbus_injection_skipped_total{tool,reason}.
- deploy/kubic/deployment.yaml: liveness/readiness timeoutSeconds: 1→5
(root-causes the "connection refused" mcpproxy errors at 13:02 today —
/readyz occasionally exceeded 1s under dream-worker tick load, so the
pod fell out of the Service endpoints intermittently).
Dream-claude agent (Python, in /dream-agent/):
- dream_runner.py uses claude-agent-sdk 0.1.48 to drive Claude Code
against SynapBus's MCP server. MCP transport carries
Authorization: Bearer <api_key> AND X-Synapbus-Dispatch-Token from env
via the SDK's McpHttpServerConfig.headers field — confirmed supported.
- Tools restricted via allowed_tools to mcp__synapbus__memory_*.
- Final JSON envelope reports tokens_in/out so harness.Usage stays
populated and the circuit breaker can count consumption.
- Dockerfile builds linux/amd64 at 189 MB, mirroring searcher's
agents/universal recipe.
- k8s-job-template.yaml: backoffLimit 0, ttl 600s, 512Mi/1CPU,
Anthropic credentials via secret-ref.
Grafana dashboard (deploy/kubic/grafana/):
- dream-dashboard.json — 14 panels across 5 rows (dream activity,
token usage vs limit, circuit breaker, injection layer, MCP
transport health), all templated to ${DS_PROMETHEUS}.
- import.sh: resolves the cluster's Prometheus DS uid and POSTs the
dashboard via Grafana API.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Background ConsolidatorWorker dispatches consolidation work to a Claude
Code agent through harness.Harness.Execute (NOT via system DMs — per
feedback_system_dm_no_trigger.md) with a one-time 15m dispatch token.
The dispatched agent uses six new MCP tools, all token-gated and
recording every action to memory_consolidation_jobs.actions JSON.
Stores:
- memory_links.go (+ test): typed edges with actor-prefix reserved-type
guard; AddConsolidationLink bypass for memory_mark_duplicate /
memory_supersede (their contractual writers).
- memory_pins.go (+ test): owner pin overlay, bypasses score floor.
- memory_status.go: queries the memory_status view.
- consolidation_jobs.go: Create / Dispatch / Lease / AppendAction /
Complete with ErrJobAlreadyInFlight via partial unique index.
- auto_links.go: MessageListener generating mention / reply_to /
channel_cooccurrence links automatically on send.
Worker:
- consolidator.go (+ test): ticker pattern modeled on StalemateWorker.
Watermark trigger for link_gen / dedup_contradiction; daily 03:00
UTC for sleep-time core rewrite. Wallclock budget via harness Budget.
Global semaphore gates concurrent owners. Mocked-harness test asserts
no system DM is ever sent.
- consolidator_prompts.go: four job-type prompts passed via env to the
dispatched agent.
MCP tools (internal/mcp/memory_tools.go + test):
- memory_list_unprocessed, memory_write_reflection, memory_rewrite_core,
memory_mark_duplicate, memory_supersede, memory_add_link.
- Full error-code matrix tested per contracts/mcp-memory-tools.md.
- Registered only when SYNAPBUS_DREAM_ENABLED=1.
Injection extensions:
- search/injection.go: pin overlay applied after retrieval; status
filter drops soft_deleted / superseded unless pinned. New
PinProvider, StatusProvider, MessageLookup hooks on InjectionOpts.
Wiring:
- cmd/synapbus/main.go: stores constructed, AutoLinkListener attached
to MessagingService, mcpSrv.SetDream wired, ConsolidatorWorker
start/stop, admin DreamRun closure.
- cmd/synapbus/admin.go: synapbus memory dream-run --owner --job
socket-RPC command (forces a single job bypassing trigger).
Cycle workarounds (documented in code):
- messaging.DreamAgent / HarnessDispatcher are local interfaces (the
agents and harness packages import messaging, not the reverse).
main.go wraps the real types via adapter structs.
Stubbed:
- Cron expression parsing (DreamDeepCron). Hardcoded daily 03:00 UTC.
Adding robfig/cron deferred to keep no-new-deps.
Pre-existing reactor test failures (5) are unchanged; confirmed
pre-020 via stash check.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Wraps eligible MCP tool handlers (my_status, send_message, search,
execute; get_replies excluded as pure metadata) with a middleware that
appends relevant_context to the JSON response. Retrieval reuses the
existing search.Service hybrid pipeline; owner scoping filters out
memories from other owners' agents (SC-008). Pin overlay is a marked
TODO for US3.
Components:
- internal/search/injection.go (+ test): BuildContextPacket with token
budget greedy fill, score floor, truncation flag, CoreMemoryProvider
interface stubbed for US2.
- internal/mcp/injection_wrap.go (+ test): WrapInjection middleware,
registered via SetInjection on the existing handler.
- internal/mcp/injection_e2e_test.go: adversarial cross-owner test
asserts H1 cannot see H2's memories on any wrapped tool.
- internal/messaging/memory_injections.go (+ test): 24h audit ring,
hourly cleanup tick wired into stalemate worker.
Discovery during impl: claim_messages/read_inbox/read_channel live as
actions inside the execute bridge, not as registered top-level MCP
tools. They inherit injection through the execute wrapper.
This commit also bundles pre-existing working-tree changes for the
027 "remove approval noise" cleanup (migration 027, design doc,
removal of reminder/escalate logic from stalemate worker, related
trims in goals_tools.go and tools_hybrid.go). The two changes touch
the same files (stalemate.go, tools_hybrid.go) and bundling them
keeps history readable.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds the SQL substrate (6 tables + memory_status view) and the helpers
every user story depends on:
- migration 028_memory_consolidation.sql + smoke test
- internal/messaging/memory_config.go (env-flag plumbing)
- internal/messaging/dispatch_tokens.go (32-byte rand, 15m TTL, single-job-bound)
- internal/messaging/memory_channels.go (open-brain / reflections-* / is_memory flag)
- internal/agents/owner.go (OwnerFor with sentinel errors)
Deviations from spec, all documented in code:
- owner_id is stored as INTEGER FK to users; OwnerFor converts to the
string scope-key the new tables use.
- MemoryChannel is a local struct to avoid an import cycle between
internal/channels and internal/messaging.
- channels.metadata column does not exist yet; IsMemoryChannel honors
it conditionally so MemoryChannelIDs can extend trivially when added.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Three improvements that turn the doc-gardener demo from "runs but
stays in 'draft' forever" into a goal that properly transitions
through its lifecycle and renders a completion summary on /goals/<id>.
### 1. complete_goal MCP tool (#59, #62)
New tool surface: complete_goal(goal_id, status, summary, completion_message_id?)
The critic calls this from inside the sandbox after it sends its FINAL:
DM. Records the one-paragraph human-readable summary on the goal row
plus a pointer to the message that carried the FINAL text, so the Web
UI /goals/<id> page has both the verdict and a deep link to the full
findings JSON.
Status parameter accepts completed | stuck | cancelled. Idempotent
when called with the current status. Rejects callers owned by a
different human than the goal owner.
Plumbing:
- New migration 026_goals_completion_summary.sql adds two columns
to goals: completion_summary TEXT, completion_message_id INTEGER
(FK messages.id, ON DELETE SET NULL).
- internal/goals/types.go: new CompletionSummary + CompletionMessageID
fields on Goal struct.
- internal/goals/store.go: Get/List Scan both new columns;
SetCompletion(goalID, status, summary, messageID) helper that
updates status+summary+message_id atomically and populates
completed_at for terminal states.
- internal/goals/service.go: Complete(ctx, goalID, status, summary,
messageID) wraps the store method with legalTransition gating.
legalTransition expanded so draft can jump straight to completed
(no mandatory "active" hop required).
- internal/mcp/goals_tools.go: completeGoalTool definition +
handleCompleteGoal handler. Tool count 6 → 7.
- internal/api/goals_handler.go: surfaces completion_summary,
completion_message_id, and completed_at on both list and detail
endpoints so the Svelte /goals UI can render them.
### 2. Draft → active auto-transition in propose_task_tree (#60)
handleProposeTaskTree now flips the goal from draft to active at the
end. Previously the coordinator would call create_goal +
propose_task_tree and dispatch inspector, but the goal stayed in
draft forever because nothing transitioned it. Now the mere fact
of having a task tree means the goal is active.
Safe: the transition is best-effort and ignores the legal-transition
error when the goal is already beyond draft.
### 3. REVISE round cap (#61)
Two-layer enforcement:
- Server-side: examples/doc-gardener/start.sh drops max_trigger_depth
from 8 to 4. Each REVISE round costs 2 hops (critic→inspector +
inspector→critic), so depth=4 caps the loop at roughly 2 rounds
before the reactor refuses further dispatches.
- Prompt-side: inspector now includes revision_round (starting at
0, incremented when it sees a REVISE: input) in its findings JSON.
Critic reads revision_round and force-FINALs when >= 1. Prompt
explicitly tells the critic to call complete_goal after sending
FINAL, so the goal row gets a proper completion_summary.
### 4. run_task.sh terminal-state detection
Rewrote the poll loop to watch goals.status/completion_summary as
the definitive "done" signal rather than parsing DM bodies. Keeps
a message-based fallback for TRIVIAL/CANNOT paths that don't create
a goal. Treats "Received system trigger..." and "Coalesced
trigger..." as informational (they're `__coalesced__` reactor
synthetic events leaking through the coordinator reply, not real
user-facing output). Bare coordinator replies are terminal only
when no goal was created.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Makes the Web UI reflect what agents are actually doing: reactions
on DMs that trigger a subprocess run, a per-run detail page that
shows the exact prompt the model received and the raw response, and
cross-linked reactive_runs ↔ harness_runs data for a single composite
API call.
Migration 020 (internal/storage/schema/020_harness_run_detail.sql):
ALTER TABLE harness_runs ADD COLUMN reactive_run_id INTEGER;
ALTER TABLE harness_runs ADD COLUMN prompt TEXT;
ALTER TABLE harness_runs ADD COLUMN response TEXT;
CREATE INDEX idx_harness_runs_reactive ON harness_runs(reactive_run_id);
internal/harness:
* ExecRequest.ReactiveRunID — reactor pins the reactive_runs row id
so the observer can JOIN the two tables.
* ExecResult.Prompt / Response — the subprocess harness reads
prompt.txt / response.txt that wrappers write into the workdir,
and runs.Store persists them (capped at 32 KiB each).
* runs.Run struct now has JSON tags — previously the API returned
PascalCase field names that didn't match the Web UI's snake_case
TypeScript types.
* New runs.Store.GetByReactiveRunID for the composite API endpoint.
* Test schema updated to include the new columns.
internal/reactor:
* New ReactionNotifier interface + SetReactionNotifier.
* dispatchHarness now reacts `in_progress` on the triggering DM
before spawning the goroutine.
* runHarness reacts `done` on success, `reject` on failure. The
existing reactionPriority ordering means the terminal reaction
wins for badge display — no need to remove in_progress first.
* dispatchHarness sets ExecRequest.ReactiveRunID.
cmd/synapbus/main.go:
* reactorReactionAdapter: adapts reactions.Service.Toggle to the
reactor's one-shot AddReaction signature.
* HarnessRunsStore wired into the API router config.
internal/api/runs_handler.go — GetRun composite endpoint:
The GET /api/runs/{id} response now returns everything the Web UI
needs to render the run detail page in one call:
{
"run": <reactive_runs row>,
"harness_run": <linked harness_runs row with prompt/response>,
"agent": <current agent snapshot with harness_config_json>,
"trigger_message": <DM that started the run>,
"outgoing_message": <first DM the agent produced after startedAt>
}
The outgoing-message lookup wraps both sides of the created_at
comparison in datetime() so SQLite parses the stored 'YYYY-MM-DD
HH:MM:SS' and the Go-emitted RFC3339 into the same canonical form
before comparing — a raw string compare was silently returning no
rows.
internal/api/router.go: HarnessRunsStore field in RouterConfig, wired
through to NewRunsHandler.
examples/cold-topic-explainer/wrapper.sh:
Writes prompt.txt and response.txt alongside gemini.stdout.raw so
the subprocess harness can capture "what the model saw" and "what
the model said" post-hoc.
web/src/lib/components/MessageList.svelte:
New ReactionPills render below each message body when the message
carries a `reactions` array (already populated by
EnrichMessages/ReactionEnricher on the server side). Makes the
👀 in_progress / ✔ done / ❌ reject lifecycle visible in every DM
view and conversation.
web/src/routes/runs/[id]/+page.svelte (NEW):
New run detail page at /runs/:id with sections:
1. Header strip — agent, status pill, backend badge, trigger
info, duration, tokens in/out, cost, exit code, trace id.
2. Triggering message — body + sender.
3. What the model saw — GEMINI.md / CLAUDE.md from agent snapshot
+ the captured rendered prompt (byte count on each summary
bar, collapsible details).
4. What the model said — captured response, falling back to
logs_excerpt or error_log when unavailable.
5. Outgoing message — body + recipient + status.
6. Metadata — reactive_run.id, harness_run.run_id, backend,
session_id, tokens_cached, k8s_job, agent trigger config.
Styled against the existing dark tailwind system — no design
overhaul, fits the current aesthetic (editorial sectioning,
monospace for code-like content, accent-blue for links,
accent-purple for system-instructions, accent-green for model
output, accent-red for errors).
web/src/routes/runs/+page.svelte: the inline expand panel now has
a "View full details →" link next to the Retry button.
E2E VERIFIED on a live subprocess run:
* Topic: "why does the subprocess harness materialise GEMINI.md
alongside .gemini/settings.json in the per-run workdir?"
* 3 subprocess runs + 3 reactive_runs + 3 harness_runs, all linked.
* message_reactions: 6 rows — in_progress + done for each hop.
* GET /api/runs/1 returns a composite with
harness_run.prompt=883 bytes, harness_run.response=550 bytes,
reactive_run_id=1, trigger_message populated, outgoing_message
populated (decomposer-pro → writer-flash), agent.gemini_md=747
bytes. All keys are snake_case as the Svelte types expect.
Full go test ./... green. `vite build` green. Demo instance still
running on port 18088 for browser verification.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Introduces internal/harness — a minimal Harness interface inspired by
GoogleCloudPlatform/scion — plus four backends (k8sjob, subprocess,
webhook, stub) and an OTel-traced Registry that spans every dispatch
and injects W3C trace context into child processes via env vars.
Phases landed together on this branch:
1. internal/harness scaffold: Harness/Capabilities/ExecRequest/
ExecResult/Budget/Usage types, Registry with Resolve/Execute,
in-memory stub backend.
2. internal/harness/k8sjob: wraps existing k8s.JobRunner behind the
Harness interface with a Waiter abstraction (real clientset +
test fake). BuildHandler exports the per-agent config logic.
3. internal/harness/subprocess: os/exec-based backend (Mac+Linux),
per-run workdir, result.json handoff, bounded log capture,
Budget-driven wall-clock timeout.
4. internal/harness/webhook: synchronous HTTP POST with HMAC
signing via internal/webhooks.ComputeHMACSignature, per-agent
URL/secret/timeout read from harness_config_json.
5. internal/observability: OTel tracer init via OTLP HTTP (opt-in
via SYNAPBUS_OTEL_ENABLED), W3C propagator always installed;
Registry.Execute starts a harness.execute span per dispatch and
calls InjectTraceContext into req.Env so children inherit it.
6. internal/harness/runs: SQLite-backed Observer that persists a
harness_runs row per dispatch with status, usage, cost, duration,
trace_id, session_id, and a bounded logs excerpt.
Schema: new migration 019_harness.sql adds agents.harness_name /
local_command / harness_config_json columns and the backend-agnostic
harness_runs table with indices on (agent, created_at), (status),
(trace_id), (run_id). internal/reactor/reactor_test.go inline schema
updated to match.
Deployment: deploy/kubic/otel-collector.yaml stands up an otel-collector
Deployment + ConfigMap + ClusterIP Service in the synapbus namespace on
kubic, receiving OTLP gRPC (4317) and HTTP (4318) and exporting debug
output until a Tempo/Jaeger backend lands.
Docs: docs/harness-otel-research.html compares scion and paperclip
side-by-side and maps the current synapbus executor surface; its
companion docs/harness-otel-design.md carries the phase plan, span
taxonomy, and migration schema verbatim.
The reactor currently still calls k8s.JobRunner directly — rewiring it
through the Registry is a follow-up, intentionally out of scope for
this branch to keep the refactor reversible. The new packages are
independently tested (~78 new tests across 7 packages) and the full
project test suite passes.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Implements US1, US2, and US3 of spec 016 by layering a marketplace service
on top of existing primitives rather than reinventing them:
- Capability manifests (US2) reuse the wiki subsystem. Each agent publishes
a per-agent article at slug "agent-<name>" and gets versioning, revision
history, and FTS search for free.
- Auction channels (US1) reuse the existing auction channel type, swarm
service, and task/bid store. post_auction / bid / award wrap post_task /
bid_task / accept_bid and attach marketplace metadata (max_budget_tokens,
domains, estimated_tokens, confidence, approach) in the task.requirements
and bid.capabilities JSON blobs. Award converts the auction into a claim
by DM'ing the winner at priority 8 with task_id metadata, so the existing
claim/process/done lifecycle takes over with zero new machinery.
- Reputation ledger (US3) adds migration 018_agent_marketplace.sql with a
new agent_reputation table keyed by (agent_name, domain). mark_task_done
completes the task via the swarm service and writes one ledger row per
declared domain using the reported actual_tokens and success_score.
query_reputation returns a rolled-up summary plus recent entries for a
given (agent, domain) pair — reputation is always a vector, never a
global score (FR-013).
Also:
- Adds the "awarded" reaction type (FR-008) alongside existing approve/
reject/in_progress/done/published. Migration 018 widens the reactions
CHECK constraint via a table rebuild.
- 6 new actions added to the action registry (post_auction, bid, award,
mark_task_done, read_skill_card, query_reputation) so the search tool
can discover them and the execute tool can dispatch them.
- New internal/marketplace package (store.go + service.go).
- New internal/mcp/marketplace.go bridge handlers.
- New internal/mcp/marketplace_test.go covers the full auction lifecycle,
capability manifest publish/read/update, self-bid rejection, non-auction
channel rejection, and reputation summary aggregation.
Out of scope for MVP (deferred per spec prompt): US4 reflection loop,
tombstoning FR-020a/b, multi-owner quorums, auto-escalation on zero bids,
bootstrap exploration credit, epsilon-greedy selection, and the hard-stop
budget enforcement daemon (only soft recording of estimated vs actual is
included).
All existing tests pass; new marketplace tests pass.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Increase busy_timeout from 5s to 15s
- Set synchronous=NORMAL (safe with WAL, reduces fsync)
- Limit MaxOpenConns to 4 to reduce write lock contention
- Explicit wal_autocheckpoint=1000
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Trust scores: per (agent, action_type) pair, stored in agent_trust
table. Auto-adjusts when human reacts to AI agent messages (approve
+0.05, reject -0.1). Scores clamped [0.0, 1.0]. MCP get_trust action
+ REST API /api/trust/{agent}. Web UI shows trust progress bars on
agent detail pages.
Claim semantics: only one in_progress reaction per message enforced.
First agent to claim wins, duplicates rejected with clear error.
State-change webhooks: StateChangeNotifier interface fires
workflow.state_changed events through existing webhook infrastructure
when reactions change a message's derived workflow state.
Channel thresholds: publish_threshold and approve_threshold fields
on channels for configuring autonomy gates.
Migration 014_trust_claims.sql. 17 new test cases across trust
model + store.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Add workflow_enabled column to channels (default false) — reactions
and workflow badges only show on opted-in channels
- ReactionPills: add "+" button with picker dropdown to add reactions
when none exist yet (was missing, only showed existing reactions)
- Update CLAUDE.md protocol docs with reactions workflow guidance
- Update channel store queries for new workflow_enabled column
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Add typed reactions (approve/reject/in_progress/done/published) with
toggle semantics. Workflow state derived from highest-priority reaction.
New reactions package with model, SQLite store, and service layer.
REST API: POST/GET/DELETE /api/messages/{id}/reactions for toggle/query,
PUT /api/channels/{name}/settings for workflow config, GET by-state
endpoint for listing messages by workflow state.
MCP: react/unreact/get_reactions/list_by_state actions via bridge.
Web UI: WorkflowBadge (colored state pills) and ReactionPills (toggle
pills with agent names) components integrated into channel view.
Channel settings: auto_approve, stalemate_remind_after,
stalemate_escalate_after columns. CLI: channels update command.
Migration 013_reactions.sql adds message_reactions table and channel
workflow columns. 29+ new test cases across model and store.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Analytics: time-series message graph with 5 time spans (1h/4h/24h/7d/30d),
top-5 agents and channels leaderboards, summary cards. 4 new REST endpoints.
PWA: web app manifest, service worker with cache-first static/network-only
API strategy, push notifications via Web Push API with VAPID keys, push
subscription management endpoints, SQLite migration for subscriptions.
UX fixes: auto-resize compose textarea (3-12 lines), inline editable agent
display name, editable human display name in settings, smart mention/channel
highlighting (existing→link, deleted→inactive badge, unknown→plain text),
font size -/+ preference (12-24px persisted in localStorage), version footer
with GitHub link.
MCP: 4 prompts — daily-digest, agent-health-check, channel-overview,
debug-agent. Registered with prompt capabilities enabled.
Code review fixes: scoped push unsubscribe to user, capped analytics limit
at 100, hardened HTML strip regex, bounded SW cache, backend push unsub on
disable.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Major feature additions across backend and frontend:
Backend:
- Unix domain socket admin server with JSON-RPC protocol
- CLI subcommands: user/agent management, audit, backup, messages, channels
- Managed API keys (sb_ prefix) with permissions, channel limits, expiry
- Thread/reply support in messaging core (reply_to column)
- Context-based trace owner_id propagation for proper audit filtering
- Two-step auth middleware supporting both agent keys and managed API keys
Frontend:
- Complete Slack-like dark theme redesign with custom CSS properties
- Sidebar with Channels, Direct Messages, and Admin sections
- Thread panel (slide-in) for viewing message replies
- API key management page with create form, key display, and
ready-to-use MCP/Claude Code config snippets with copy-to-clipboard
- All pages restyled: login, dashboard, agents, conversations, settings
E2E Tests:
- Fixed test runner binary path resolution
- Added pyproject.toml for test dependencies
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add HNSW-based vector search with configurable embedding providers
(OpenAI, Ollama) and automatic FTS5 fallback when no provider is
configured. Background pipeline embeds messages asynchronously on
ingest, stores vectors in a pure-Go HNSW index, and retries on
failure with exponential backoff. The search_messages MCP tool now
supports search_mode (auto/semantic/fulltext) and returns ranked
results with similarity scores. All existing tests continue to pass,
CGO_ENABLED=0 cross-compilation verified.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add complete auth subsystem with OAuth 2.1 authorization server using
ory/fosite, local user accounts with bcrypt password hashing, session
management, and HTTP handlers for the Web UI.
Components:
- User store with bcrypt hashing (configurable cost, default 12), CRUD,
validation (username 3-64 chars alphanumeric+underscore, password 8-72 bytes)
- Session store with secure random IDs, configurable lifetime (default 24h),
expiration cleanup, and per-user invalidation
- OAuth client store with client_id/secret generation and bcrypt verification
- Fosite storage adapter implementing CoreStorage, TokenRevocationStorage,
and PKCERequestStorage backed by SQLite
- OAuth provider configured with authorization code (PKCE S256 mandatory),
client credentials, refresh token rotation, and token introspection
- HTTP handlers: POST /auth/register, POST /auth/login, POST /auth/logout,
GET /auth/me, PUT /auth/password, GET /oauth/authorize, POST /oauth/token,
POST /oauth/introspect
- Middleware: RequireSession (cookie), RequireBearer (access token),
RequireAuth (either), RequireAdmin (role check)
- Structured auth event logging (login, token issuance, session lifecycle)
- Schema migration 002_auth.sql extending users, oauth_clients, oauth_tokens
tables and adding sessions, oauth_authorization_codes tables
- Initial admin user auto-created on first run with random password printed
to stdout
- All tests pass with CGO_ENABLED=0, zero external runtime dependencies
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add comprehensive trace logging and observability features:
- Enhanced trace store with owner-scoped queries, filtering (agent, action,
time range), pagination, streaming export, and retention cleanup
- REST API endpoints: GET /api/traces (list with filters), GET /api/traces/export
(streaming JSON/CSV), GET /api/traces/stats (action counts)
- Owner isolation enforced at every layer (store, API, tests)
- Hand-rolled Prometheus metrics (pure Go, zero CGO): traces_total,
traces_by_action, errors_total, active_agents at GET /metrics
- Configurable slog JSON handler with --log-level flag
- Request ID middleware for cross-referencing logs and traces
- Batch trace writing (64 entries or 100ms flush interval)
- Background retention cleanup via --trace-retention flag
- SQL migration 002 adds owner_id column and composite indexes
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Implements the foundational layer that all SynapBus features depend on:
- internal/storage: SQLite connection manager (WAL mode, busy_timeout,
foreign_keys) and embedded migration runner using modernc.org/sqlite
- internal/messaging: MessagingService with send, read inbox, claim,
mark done/failed, and FTS5 search. SQLite-backed MessageStore with
conversation auto-creation and read/unread tracking via inbox_state.
- internal/agents: AgentService with register, authenticate (bcrypt),
update, deregister, discover by capability. HTTP auth middleware.
- internal/mcp: MCP server using mark3labs/mcp-go with 9 registered
tools (send_message, read_inbox, claim_messages, mark_done,
search_messages, register_agent, discover_agents, update_agent,
deregister_agent). SSE transport, health endpoint, connection manager.
- internal/trace: Async trace recorder with buffered channel for
recording agent actions to SQLite traces table.
- cmd/synapbus: Updated main.go wiring storage, migrations, services,
MCP server, chi router, and graceful shutdown.
All code compiles with CGO_ENABLED=0. Full test suite passes.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>