9 Commits
Author SHA1 Message Date
Algis DumbrisandClaude Opus 4.7 0d4a9539b5 fix(watchdog): add sqlite to runtime + aggregate today usage across owners
Two silent failures making the watchdog's circuit-breaker checks
toothless:

1) The Alpine runtime image had no sqlite3 binary, so every
   `kubectl exec -- sqlite3 …` call inside the watchdog returned empty.
   $JS/$TIN/$JOK/$JFL/$JCB all defaulted to 0 → every "soft cap" /
   "tokens > 30M" / "circuit broke but still firing" check trivially
   passed regardless of real state. apk add sqlite (≈700KB).

2) The "today usage" query filtered to owner_id='2' only, but dream
   jobs run for any owner (we just saw a clean owner_id=1 dispatch).
   Replace with SUM across all rows for date=date('now'); the caps
   are intentionally global, not per-owner.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-19 06:16:58 +03:00
Algis DumbrisandClaude Opus 4.7 73a1802155 feat(020): token accounting + watchdog CronJob
Token-usage wiring (closes the 0-tokens gap in memory_dream_usage):
- internal/harness/k8sjob/k8sjob.go: after extractResultJSON, parse
  tokens_in/tokens_out/tokens_cached/cost_usd from the final JSON
  envelope and stash into ExecResult.Usage so the dream worker's
  UsageGate circuit breaker actually counts consumption.
- dream-agent/dream_runner.py: Max20 OAuth sessions don't surface
  per-call tokens through the SDK's ResultMessage.usage. Falls back
  to a turn-based estimate so the gate has SOME signal:
    tokens_in_est = turns * 5000 + tool_calls * 2000
    tokens_out_est = turns * 300
  Calibrated against observed reflection runs.

Watchdog (deploy/kubic/watchdog/):
- watchdog.yaml: in-cluster CronJob runs every hour at :05 past UTC,
  with a dedicated ServiceAccount + Role granting (get/list/exec on
  pods, patch+update on deployments/scale) inside the synapbus
  namespace only.
- Health checks: pod readiness + restart count; last-1h job
  succ/fail/in_flight counts; today's jobs_started + tokens_in +
  circuit_broken.
- Red flags that auto-stop synapbus (scale to 0):
    * pod restart count > 3
    * failed dream jobs in last 1h > 20
    * jobs_started today > 200 OR tokens_in > 30M
    * circuit broke AND still firing (started >> completed)
- Dockerfile: slim alpine + kubectl v1.30.5 binary (synapbus-watchdog:v1).
  Built locally and imported into kubic's containerd because the
  public docker.io/bitnami/kubectl manifest was returning text/html
  from kubic's network egress.

Replaces the schedule-skill remote-agent approach because Anthropic
cloud agents can't reach kubic.home.arpa (LAN-only) and can't call
kubectl scale. The k8s CronJob is the right primitive for an
in-cluster safety watchdog.

Live evidence: first manual run on kubic reported
  pod=synapbus-... ready=true restarts=0
  last_1h jobs total=18 succ=18 fail=0 in_flight=0
  today: jobs_started=189 tokens_in=0 succeeded=169 failed=15
  HEALTHY — no action

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 22:22:13 +03:00
Algis DumbrisandClaude Opus 4.7 069a985af5 feat(020): 14d window + token-budget circuit breaker + dream-agent + dashboard
Backend (Go, in this commit):
- migration 029_memory_dream_usage: per (date, owner) counters for
  tokens_in/out, jobs_started/succeeded/failed/circuit_broken
- DreamUsageStore + UsageGate (internal/messaging/dream_usage.go).
  Gate inspects today's usage against new env knobs:
  - SYNAPBUS_DREAM_RECENT_WINDOW (default 336h / 14d)
  - SYNAPBUS_DREAM_DAILY_TOKEN_LIMIT_IN (default 1M)
  - SYNAPBUS_DREAM_DAILY_TOKEN_LIMIT_OUT (default 200k)
  - SYNAPBUS_DREAM_DAILY_JOB_LIMIT (default 100)
- Consolidator now bounds watermarks + core_rewrite eligibility by the
  recency window. core_rewrite skipped for owners with no in-window
  activity. ForceRun honors the breaker.
- Recency fallback in BuildContextPacket + memory_list_unprocessed now
  accept RecentWindowDays so injection and dream queries see the same
  14d slice.
- Prometheus metrics registered (internal/metrics/metrics.go):
  synapbus_dream_jobs_total{owner,job_type,status},
  synapbus_dream_tokens_total{owner,direction},
  synapbus_dream_job_duration_seconds{owner,job_type},
  synapbus_dream_circuit_broken_total{owner,reason},
  synapbus_injection_packets_total{tool},
  synapbus_injection_memories_per_packet{tool},
  synapbus_injection_packet_chars{tool},
  synapbus_injection_skipped_total{tool,reason}.
- deploy/kubic/deployment.yaml: liveness/readiness timeoutSeconds: 1→5
  (root-causes the "connection refused" mcpproxy errors at 13:02 today —
  /readyz occasionally exceeded 1s under dream-worker tick load, so the
  pod fell out of the Service endpoints intermittently).

Dream-claude agent (Python, in /dream-agent/):
- dream_runner.py uses claude-agent-sdk 0.1.48 to drive Claude Code
  against SynapBus's MCP server. MCP transport carries
  Authorization: Bearer <api_key> AND X-Synapbus-Dispatch-Token from env
  via the SDK's McpHttpServerConfig.headers field — confirmed supported.
- Tools restricted via allowed_tools to mcp__synapbus__memory_*.
- Final JSON envelope reports tokens_in/out so harness.Usage stays
  populated and the circuit breaker can count consumption.
- Dockerfile builds linux/amd64 at 189 MB, mirroring searcher's
  agents/universal recipe.
- k8s-job-template.yaml: backoffLimit 0, ttl 600s, 512Mi/1CPU,
  Anthropic credentials via secret-ref.

Grafana dashboard (deploy/kubic/grafana/):
- dream-dashboard.json — 14 panels across 5 rows (dream activity,
  token usage vs limit, circuit breaker, injection layer, MCP
  transport health), all templated to ${DS_PROMETHEUS}.
- import.sh: resolves the cluster's Prometheus DS uid and POSTs the
  dashboard via Grafana API.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 13:58:28 +03:00
Algis DumbrisandClaude Opus 4.7 a3d8681715 chore(deploy): replace stale helm chart with plain kubic manifests
The helm release went into failed state in March 2026 after an out-of-band
kubectl set image broke server-side-apply ownership; every deploy since has
been a direct kubectl set image, leaving the chart values drifting against
live state.

Drop deploy/helm/ entirely. Add deploy/kubic/{namespace,pvc,service,
deployment,secret.example}.yaml mirroring what's actually running, plus
scripts/deploy-kubic.sh encoding the build → docker save → scp → microk8s
ctr image import → kubectl set image flow used for v0.13.x-reactive through
v0.17.0. README documents why no helm and how to back up /data before
schema-touching versions.

kubectl diff -f deploy/kubic/ is empty against the live cluster.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-08 12:08:26 +03:00
Algis DumbrisandClaude Opus 4.6 574897df4d feat: harness-agnostic wrappers + OpenTelemetry integration
Introduces internal/harness — a minimal Harness interface inspired by
GoogleCloudPlatform/scion — plus four backends (k8sjob, subprocess,
webhook, stub) and an OTel-traced Registry that spans every dispatch
and injects W3C trace context into child processes via env vars.

Phases landed together on this branch:

  1. internal/harness scaffold: Harness/Capabilities/ExecRequest/
     ExecResult/Budget/Usage types, Registry with Resolve/Execute,
     in-memory stub backend.
  2. internal/harness/k8sjob: wraps existing k8s.JobRunner behind the
     Harness interface with a Waiter abstraction (real clientset +
     test fake). BuildHandler exports the per-agent config logic.
  3. internal/harness/subprocess: os/exec-based backend (Mac+Linux),
     per-run workdir, result.json handoff, bounded log capture,
     Budget-driven wall-clock timeout.
  4. internal/harness/webhook: synchronous HTTP POST with HMAC
     signing via internal/webhooks.ComputeHMACSignature, per-agent
     URL/secret/timeout read from harness_config_json.
  5. internal/observability: OTel tracer init via OTLP HTTP (opt-in
     via SYNAPBUS_OTEL_ENABLED), W3C propagator always installed;
     Registry.Execute starts a harness.execute span per dispatch and
     calls InjectTraceContext into req.Env so children inherit it.
  6. internal/harness/runs: SQLite-backed Observer that persists a
     harness_runs row per dispatch with status, usage, cost, duration,
     trace_id, session_id, and a bounded logs excerpt.

Schema: new migration 019_harness.sql adds agents.harness_name /
local_command / harness_config_json columns and the backend-agnostic
harness_runs table with indices on (agent, created_at), (status),
(trace_id), (run_id). internal/reactor/reactor_test.go inline schema
updated to match.

Deployment: deploy/kubic/otel-collector.yaml stands up an otel-collector
Deployment + ConfigMap + ClusterIP Service in the synapbus namespace on
kubic, receiving OTLP gRPC (4317) and HTTP (4318) and exporting debug
output until a Tempo/Jaeger backend lands.

Docs: docs/harness-otel-research.html compares scion and paperclip
side-by-side and maps the current synapbus executor surface; its
companion docs/harness-otel-design.md carries the phase plan, span
taxonomy, and migration schema verbatim.

The reactor currently still calls k8s.JobRunner directly — rewiring it
through the Registry is a follow-up, intentionally out of scope for
this branch to keep the refactor reversible. The new packages are
independently tested (~78 new tests across 7 packages) and the full
project test suite passes.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-13 20:23:04 +03:00
Algis DumbrisandClaude Opus 4.6 5938fcd555 fix: bugs #1-3-5 from #bugs-synapbus + live SSE notifications
- Auto-join public channels on first send (bug #1)
- Channel broadcasts no longer create duplicate DM copies; inbox DMs
  only sent for @mentions (bug #2)
- Embedding pipeline auto-enqueues new messages via MessageListener
  callback instead of requiring pod restart (bug #3)
- Admin socket defaults to /tmp in containers to avoid PVC filesystem
  incompatibility with Unix sockets (bug #5)
- SSE events now fire for MCP-sent messages (not just REST API),
  enabling live notification badges without page reload
- Fixed frontend SSE field name mismatch (channel_name → channel)
- Fixed SSE client not connecting after login redirect

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-15 18:56:30 +02:00
Algis DumbrisandClaude Opus 4.6 9dd27c4b9b feat: admin CLI & Docker fixes — alpine base, channels create/join, absolute socket path
Switch Docker runtime from scratch to alpine:3.19 so kubectl exec works
for admin CLI operations. Add `synapbus channels create` and
`synapbus channels join` CLI commands with corresponding admin socket
handlers. Change default socket path to /data/synapbus.sock (absolute).
Also add Helm envFrom support and NodePort configuration.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 14:54:13 +02:00
Algis DumbrisandClaude Opus 4.6 42775df65f feat: webhooks & Kubernetes Job runner for event-driven agents
- Webhook registration via MCP (register_webhook, list_webhooks, delete_webhook)
- HMAC-SHA256 payload signing, SSRF-safe HTTP client, loop detection (depth 5)
- 8-worker goroutine delivery pool with exponential backoff retry (1s/5s/30s)
- Dead letter queue with auto-purge, auto-disable after 50 consecutive failures
- Per-agent rate limiting (60 deliveries/min)
- K8s Job runner (register_k8s_handler, list_k8s_handlers, delete_k8s_handler)
- Auto-detect in-cluster via InClusterConfig, NoopRunner fallback
- REST API for webhook deliveries, dead letters, K8s job runs and logs
- Web UI: webhook management, K8s handler pages, dead letters view
- MultiDispatcher fan-out pattern for webhook + K8s event dispatch
- SQLite migration 009: webhooks, webhook_deliveries, k8s_handlers, k8s_job_runs
- 51 tests across 9 test packages, all passing

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 18:14:04 +02:00
ff81f2f218 feat: production readiness - metrics, health, CI/CD, Docker, Helm (#1)
* docs: add specification for production readiness & website launch

Covers DevOps hooks, CI/CD, Prometheus observability, Docker/Helm
deployment, and synapbus.dev website with documentation and blog.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs: add implementation plan and research for production readiness

Covers 5 workstreams: git hooks, CI/CD, observability, deployment
artifacts, and website. All constitution gates pass.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add pre-commit and pre-push git hooks

Pre-commit runs go vet, golangci-lint (optional), and fast tests.
Pre-push runs full test suite and build verification.
Installable via `make hooks`.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* ci: add GitHub Actions for PR checks and releases

ci.yml: lint, test, build on PRs to main.
release.yml: multi-platform binaries + Docker image on version tags.
Targets: linux/amd64, linux/arm64, darwin/amd64, darwin/arm64, windows/amd64.
Docker pushed to ghcr.io/smart-mcp-proxy/synapbus.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add Dockerfile, docker-compose, and Helm chart

Multi-stage Docker build (node + golang + scratch), ~30MB image.
docker-compose.yml for local development with volume persistence.
Helm chart with configurable Deployment, Service, PVC, Ingress,
and Prometheus ServiceMonitor.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add Prometheus metrics and Kubernetes health endpoints

Add internal/metrics package with Prometheus collectors (HTTP requests,
duration, messages, agents, connections) and chi-compatible middleware.
Add internal/health package with /healthz (liveness) and /readyz
(readiness with DB ping) endpoints. Wire into main.go with promhttp
handler at /metrics.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: update Go version to 1.25, fix Docker build issues

- Update Dockerfile golang image from 1.23 to 1.25 (matches go.mod)
- Add tzdata package for timezone support in scratch image
- Use npm install --legacy-peer-deps for web frontend build
- Update CI/release workflows to use Go 1.25

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: resolve CI failures - golangci-lint v2 and npm peer deps

- Upgrade golangci-lint-action to v7 with v2.1 (supports Go 1.25)
- Use npm install --legacy-peer-deps instead of npm ci for web builds

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: replace golangci-lint with go vet (golangci-lint doesn't support Go 1.25 yet)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-13 21:03:47 +02:00