Token-usage wiring (closes the 0-tokens gap in memory_dream_usage):
- internal/harness/k8sjob/k8sjob.go: after extractResultJSON, parse
tokens_in/tokens_out/tokens_cached/cost_usd from the final JSON
envelope and stash into ExecResult.Usage so the dream worker's
UsageGate circuit breaker actually counts consumption.
- dream-agent/dream_runner.py: Max20 OAuth sessions don't surface
per-call tokens through the SDK's ResultMessage.usage. Falls back
to a turn-based estimate so the gate has SOME signal:
tokens_in_est = turns * 5000 + tool_calls * 2000
tokens_out_est = turns * 300
Calibrated against observed reflection runs.
Watchdog (deploy/kubic/watchdog/):
- watchdog.yaml: in-cluster CronJob runs every hour at :05 past UTC,
with a dedicated ServiceAccount + Role granting (get/list/exec on
pods, patch+update on deployments/scale) inside the synapbus
namespace only.
- Health checks: pod readiness + restart count; last-1h job
succ/fail/in_flight counts; today's jobs_started + tokens_in +
circuit_broken.
- Red flags that auto-stop synapbus (scale to 0):
* pod restart count > 3
* failed dream jobs in last 1h > 20
* jobs_started today > 200 OR tokens_in > 30M
* circuit broke AND still firing (started >> completed)
- Dockerfile: slim alpine + kubectl v1.30.5 binary (synapbus-watchdog:v1).
Built locally and imported into kubic's containerd because the
public docker.io/bitnami/kubectl manifest was returning text/html
from kubic's network egress.
Replaces the schedule-skill remote-agent approach because Anthropic
cloud agents can't reach kubic.home.arpa (LAN-only) and can't call
kubectl scale. The k8s CronJob is the right primitive for an
in-cluster safety watchdog.
Live evidence: first manual run on kubic reported
pod=synapbus-... ready=true restarts=0
last_1h jobs total=18 succ=18 fail=0 in_flight=0
today: jobs_started=189 tokens_in=0 succeeded=169 failed=15
HEALTHY — no action
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Backend (Go, in this commit):
- migration 029_memory_dream_usage: per (date, owner) counters for
tokens_in/out, jobs_started/succeeded/failed/circuit_broken
- DreamUsageStore + UsageGate (internal/messaging/dream_usage.go).
Gate inspects today's usage against new env knobs:
- SYNAPBUS_DREAM_RECENT_WINDOW (default 336h / 14d)
- SYNAPBUS_DREAM_DAILY_TOKEN_LIMIT_IN (default 1M)
- SYNAPBUS_DREAM_DAILY_TOKEN_LIMIT_OUT (default 200k)
- SYNAPBUS_DREAM_DAILY_JOB_LIMIT (default 100)
- Consolidator now bounds watermarks + core_rewrite eligibility by the
recency window. core_rewrite skipped for owners with no in-window
activity. ForceRun honors the breaker.
- Recency fallback in BuildContextPacket + memory_list_unprocessed now
accept RecentWindowDays so injection and dream queries see the same
14d slice.
- Prometheus metrics registered (internal/metrics/metrics.go):
synapbus_dream_jobs_total{owner,job_type,status},
synapbus_dream_tokens_total{owner,direction},
synapbus_dream_job_duration_seconds{owner,job_type},
synapbus_dream_circuit_broken_total{owner,reason},
synapbus_injection_packets_total{tool},
synapbus_injection_memories_per_packet{tool},
synapbus_injection_packet_chars{tool},
synapbus_injection_skipped_total{tool,reason}.
- deploy/kubic/deployment.yaml: liveness/readiness timeoutSeconds: 1→5
(root-causes the "connection refused" mcpproxy errors at 13:02 today —
/readyz occasionally exceeded 1s under dream-worker tick load, so the
pod fell out of the Service endpoints intermittently).
Dream-claude agent (Python, in /dream-agent/):
- dream_runner.py uses claude-agent-sdk 0.1.48 to drive Claude Code
against SynapBus's MCP server. MCP transport carries
Authorization: Bearer <api_key> AND X-Synapbus-Dispatch-Token from env
via the SDK's McpHttpServerConfig.headers field — confirmed supported.
- Tools restricted via allowed_tools to mcp__synapbus__memory_*.
- Final JSON envelope reports tokens_in/out so harness.Usage stays
populated and the circuit breaker can count consumption.
- Dockerfile builds linux/amd64 at 189 MB, mirroring searcher's
agents/universal recipe.
- k8s-job-template.yaml: backoffLimit 0, ttl 600s, 512Mi/1CPU,
Anthropic credentials via secret-ref.
Grafana dashboard (deploy/kubic/grafana/):
- dream-dashboard.json — 14 panels across 5 rows (dream activity,
token usage vs limit, circuit breaker, injection layer, MCP
transport health), all templated to ${DS_PROMETHEUS}.
- import.sh: resolves the cluster's Prometheus DS uid and POSTs the
dashboard via Grafana API.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The helm release went into failed state in March 2026 after an out-of-band
kubectl set image broke server-side-apply ownership; every deploy since has
been a direct kubectl set image, leaving the chart values drifting against
live state.
Drop deploy/helm/ entirely. Add deploy/kubic/{namespace,pvc,service,
deployment,secret.example}.yaml mirroring what's actually running, plus
scripts/deploy-kubic.sh encoding the build → docker save → scp → microk8s
ctr image import → kubectl set image flow used for v0.13.x-reactive through
v0.17.0. README documents why no helm and how to back up /data before
schema-touching versions.
kubectl diff -f deploy/kubic/ is empty against the live cluster.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Introduces internal/harness — a minimal Harness interface inspired by
GoogleCloudPlatform/scion — plus four backends (k8sjob, subprocess,
webhook, stub) and an OTel-traced Registry that spans every dispatch
and injects W3C trace context into child processes via env vars.
Phases landed together on this branch:
1. internal/harness scaffold: Harness/Capabilities/ExecRequest/
ExecResult/Budget/Usage types, Registry with Resolve/Execute,
in-memory stub backend.
2. internal/harness/k8sjob: wraps existing k8s.JobRunner behind the
Harness interface with a Waiter abstraction (real clientset +
test fake). BuildHandler exports the per-agent config logic.
3. internal/harness/subprocess: os/exec-based backend (Mac+Linux),
per-run workdir, result.json handoff, bounded log capture,
Budget-driven wall-clock timeout.
4. internal/harness/webhook: synchronous HTTP POST with HMAC
signing via internal/webhooks.ComputeHMACSignature, per-agent
URL/secret/timeout read from harness_config_json.
5. internal/observability: OTel tracer init via OTLP HTTP (opt-in
via SYNAPBUS_OTEL_ENABLED), W3C propagator always installed;
Registry.Execute starts a harness.execute span per dispatch and
calls InjectTraceContext into req.Env so children inherit it.
6. internal/harness/runs: SQLite-backed Observer that persists a
harness_runs row per dispatch with status, usage, cost, duration,
trace_id, session_id, and a bounded logs excerpt.
Schema: new migration 019_harness.sql adds agents.harness_name /
local_command / harness_config_json columns and the backend-agnostic
harness_runs table with indices on (agent, created_at), (status),
(trace_id), (run_id). internal/reactor/reactor_test.go inline schema
updated to match.
Deployment: deploy/kubic/otel-collector.yaml stands up an otel-collector
Deployment + ConfigMap + ClusterIP Service in the synapbus namespace on
kubic, receiving OTLP gRPC (4317) and HTTP (4318) and exporting debug
output until a Tempo/Jaeger backend lands.
Docs: docs/harness-otel-research.html compares scion and paperclip
side-by-side and maps the current synapbus executor surface; its
companion docs/harness-otel-design.md carries the phase plan, span
taxonomy, and migration schema verbatim.
The reactor currently still calls k8s.JobRunner directly — rewiring it
through the Registry is a follow-up, intentionally out of scope for
this branch to keep the refactor reversible. The new packages are
independently tested (~78 new tests across 7 packages) and the full
project test suite passes.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Auto-join public channels on first send (bug #1)
- Channel broadcasts no longer create duplicate DM copies; inbox DMs
only sent for @mentions (bug #2)
- Embedding pipeline auto-enqueues new messages via MessageListener
callback instead of requiring pod restart (bug #3)
- Admin socket defaults to /tmp in containers to avoid PVC filesystem
incompatibility with Unix sockets (bug #5)
- SSE events now fire for MCP-sent messages (not just REST API),
enabling live notification badges without page reload
- Fixed frontend SSE field name mismatch (channel_name → channel)
- Fixed SSE client not connecting after login redirect
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Switch Docker runtime from scratch to alpine:3.19 so kubectl exec works
for admin CLI operations. Add `synapbus channels create` and
`synapbus channels join` CLI commands with corresponding admin socket
handlers. Change default socket path to /data/synapbus.sock (absolute).
Also add Helm envFrom support and NodePort configuration.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* docs: add specification for production readiness & website launch
Covers DevOps hooks, CI/CD, Prometheus observability, Docker/Helm
deployment, and synapbus.dev website with documentation and blog.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* docs: add implementation plan and research for production readiness
Covers 5 workstreams: git hooks, CI/CD, observability, deployment
artifacts, and website. All constitution gates pass.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* feat: add pre-commit and pre-push git hooks
Pre-commit runs go vet, golangci-lint (optional), and fast tests.
Pre-push runs full test suite and build verification.
Installable via `make hooks`.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* ci: add GitHub Actions for PR checks and releases
ci.yml: lint, test, build on PRs to main.
release.yml: multi-platform binaries + Docker image on version tags.
Targets: linux/amd64, linux/arm64, darwin/amd64, darwin/arm64, windows/amd64.
Docker pushed to ghcr.io/smart-mcp-proxy/synapbus.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* feat: add Dockerfile, docker-compose, and Helm chart
Multi-stage Docker build (node + golang + scratch), ~30MB image.
docker-compose.yml for local development with volume persistence.
Helm chart with configurable Deployment, Service, PVC, Ingress,
and Prometheus ServiceMonitor.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* feat: add Prometheus metrics and Kubernetes health endpoints
Add internal/metrics package with Prometheus collectors (HTTP requests,
duration, messages, agents, connections) and chi-compatible middleware.
Add internal/health package with /healthz (liveness) and /readyz
(readiness with DB ping) endpoints. Wire into main.go with promhttp
handler at /metrics.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: update Go version to 1.25, fix Docker build issues
- Update Dockerfile golang image from 1.23 to 1.25 (matches go.mod)
- Add tzdata package for timezone support in scratch image
- Use npm install --legacy-peer-deps for web frontend build
- Update CI/release workflows to use Go 1.25
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: resolve CI failures - golangci-lint v2 and npm peer deps
- Upgrade golangci-lint-action to v7 with v2.1 (supports Go 1.25)
- Use npm install --legacy-peer-deps instead of npm ci for web builds
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: replace golangci-lint with go vet (golangci-lint doesn't support Go 1.25 yet)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>