5 Commits
Author SHA1 Message Date
Algis DumbrisandClaude Opus 4.7 0d4a9539b5 fix(watchdog): add sqlite to runtime + aggregate today usage across owners
Two silent failures making the watchdog's circuit-breaker checks
toothless:

1) The Alpine runtime image had no sqlite3 binary, so every
   `kubectl exec -- sqlite3 …` call inside the watchdog returned empty.
   $JS/$TIN/$JOK/$JFL/$JCB all defaulted to 0 → every "soft cap" /
   "tokens > 30M" / "circuit broke but still firing" check trivially
   passed regardless of real state. apk add sqlite (≈700KB).

2) The "today usage" query filtered to owner_id='2' only, but dream
   jobs run for any owner (we just saw a clean owner_id=1 dispatch).
   Replace with SUM across all rows for date=date('now'); the caps
   are intentionally global, not per-owner.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-19 06:16:58 +03:00
Algis DumbrisandClaude Opus 4.7 73a1802155 feat(020): token accounting + watchdog CronJob
Token-usage wiring (closes the 0-tokens gap in memory_dream_usage):
- internal/harness/k8sjob/k8sjob.go: after extractResultJSON, parse
  tokens_in/tokens_out/tokens_cached/cost_usd from the final JSON
  envelope and stash into ExecResult.Usage so the dream worker's
  UsageGate circuit breaker actually counts consumption.
- dream-agent/dream_runner.py: Max20 OAuth sessions don't surface
  per-call tokens through the SDK's ResultMessage.usage. Falls back
  to a turn-based estimate so the gate has SOME signal:
    tokens_in_est = turns * 5000 + tool_calls * 2000
    tokens_out_est = turns * 300
  Calibrated against observed reflection runs.

Watchdog (deploy/kubic/watchdog/):
- watchdog.yaml: in-cluster CronJob runs every hour at :05 past UTC,
  with a dedicated ServiceAccount + Role granting (get/list/exec on
  pods, patch+update on deployments/scale) inside the synapbus
  namespace only.
- Health checks: pod readiness + restart count; last-1h job
  succ/fail/in_flight counts; today's jobs_started + tokens_in +
  circuit_broken.
- Red flags that auto-stop synapbus (scale to 0):
    * pod restart count > 3
    * failed dream jobs in last 1h > 20
    * jobs_started today > 200 OR tokens_in > 30M
    * circuit broke AND still firing (started >> completed)
- Dockerfile: slim alpine + kubectl v1.30.5 binary (synapbus-watchdog:v1).
  Built locally and imported into kubic's containerd because the
  public docker.io/bitnami/kubectl manifest was returning text/html
  from kubic's network egress.

Replaces the schedule-skill remote-agent approach because Anthropic
cloud agents can't reach kubic.home.arpa (LAN-only) and can't call
kubectl scale. The k8s CronJob is the right primitive for an
in-cluster safety watchdog.

Live evidence: first manual run on kubic reported
  pod=synapbus-... ready=true restarts=0
  last_1h jobs total=18 succ=18 fail=0 in_flight=0
  today: jobs_started=189 tokens_in=0 succeeded=169 failed=15
  HEALTHY — no action

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 22:22:13 +03:00
Algis DumbrisandClaude Opus 4.7 069a985af5 feat(020): 14d window + token-budget circuit breaker + dream-agent + dashboard
Backend (Go, in this commit):
- migration 029_memory_dream_usage: per (date, owner) counters for
  tokens_in/out, jobs_started/succeeded/failed/circuit_broken
- DreamUsageStore + UsageGate (internal/messaging/dream_usage.go).
  Gate inspects today's usage against new env knobs:
  - SYNAPBUS_DREAM_RECENT_WINDOW (default 336h / 14d)
  - SYNAPBUS_DREAM_DAILY_TOKEN_LIMIT_IN (default 1M)
  - SYNAPBUS_DREAM_DAILY_TOKEN_LIMIT_OUT (default 200k)
  - SYNAPBUS_DREAM_DAILY_JOB_LIMIT (default 100)
- Consolidator now bounds watermarks + core_rewrite eligibility by the
  recency window. core_rewrite skipped for owners with no in-window
  activity. ForceRun honors the breaker.
- Recency fallback in BuildContextPacket + memory_list_unprocessed now
  accept RecentWindowDays so injection and dream queries see the same
  14d slice.
- Prometheus metrics registered (internal/metrics/metrics.go):
  synapbus_dream_jobs_total{owner,job_type,status},
  synapbus_dream_tokens_total{owner,direction},
  synapbus_dream_job_duration_seconds{owner,job_type},
  synapbus_dream_circuit_broken_total{owner,reason},
  synapbus_injection_packets_total{tool},
  synapbus_injection_memories_per_packet{tool},
  synapbus_injection_packet_chars{tool},
  synapbus_injection_skipped_total{tool,reason}.
- deploy/kubic/deployment.yaml: liveness/readiness timeoutSeconds: 1→5
  (root-causes the "connection refused" mcpproxy errors at 13:02 today —
  /readyz occasionally exceeded 1s under dream-worker tick load, so the
  pod fell out of the Service endpoints intermittently).

Dream-claude agent (Python, in /dream-agent/):
- dream_runner.py uses claude-agent-sdk 0.1.48 to drive Claude Code
  against SynapBus's MCP server. MCP transport carries
  Authorization: Bearer <api_key> AND X-Synapbus-Dispatch-Token from env
  via the SDK's McpHttpServerConfig.headers field — confirmed supported.
- Tools restricted via allowed_tools to mcp__synapbus__memory_*.
- Final JSON envelope reports tokens_in/out so harness.Usage stays
  populated and the circuit breaker can count consumption.
- Dockerfile builds linux/amd64 at 189 MB, mirroring searcher's
  agents/universal recipe.
- k8s-job-template.yaml: backoffLimit 0, ttl 600s, 512Mi/1CPU,
  Anthropic credentials via secret-ref.

Grafana dashboard (deploy/kubic/grafana/):
- dream-dashboard.json — 14 panels across 5 rows (dream activity,
  token usage vs limit, circuit breaker, injection layer, MCP
  transport health), all templated to ${DS_PROMETHEUS}.
- import.sh: resolves the cluster's Prometheus DS uid and POSTs the
  dashboard via Grafana API.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 13:58:28 +03:00
Algis DumbrisandClaude Opus 4.7 a3d8681715 chore(deploy): replace stale helm chart with plain kubic manifests
The helm release went into failed state in March 2026 after an out-of-band
kubectl set image broke server-side-apply ownership; every deploy since has
been a direct kubectl set image, leaving the chart values drifting against
live state.

Drop deploy/helm/ entirely. Add deploy/kubic/{namespace,pvc,service,
deployment,secret.example}.yaml mirroring what's actually running, plus
scripts/deploy-kubic.sh encoding the build → docker save → scp → microk8s
ctr image import → kubectl set image flow used for v0.13.x-reactive through
v0.17.0. README documents why no helm and how to back up /data before
schema-touching versions.

kubectl diff -f deploy/kubic/ is empty against the live cluster.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-08 12:08:26 +03:00
Algis DumbrisandClaude Opus 4.6 574897df4d feat: harness-agnostic wrappers + OpenTelemetry integration
Introduces internal/harness — a minimal Harness interface inspired by
GoogleCloudPlatform/scion — plus four backends (k8sjob, subprocess,
webhook, stub) and an OTel-traced Registry that spans every dispatch
and injects W3C trace context into child processes via env vars.

Phases landed together on this branch:

  1. internal/harness scaffold: Harness/Capabilities/ExecRequest/
     ExecResult/Budget/Usage types, Registry with Resolve/Execute,
     in-memory stub backend.
  2. internal/harness/k8sjob: wraps existing k8s.JobRunner behind the
     Harness interface with a Waiter abstraction (real clientset +
     test fake). BuildHandler exports the per-agent config logic.
  3. internal/harness/subprocess: os/exec-based backend (Mac+Linux),
     per-run workdir, result.json handoff, bounded log capture,
     Budget-driven wall-clock timeout.
  4. internal/harness/webhook: synchronous HTTP POST with HMAC
     signing via internal/webhooks.ComputeHMACSignature, per-agent
     URL/secret/timeout read from harness_config_json.
  5. internal/observability: OTel tracer init via OTLP HTTP (opt-in
     via SYNAPBUS_OTEL_ENABLED), W3C propagator always installed;
     Registry.Execute starts a harness.execute span per dispatch and
     calls InjectTraceContext into req.Env so children inherit it.
  6. internal/harness/runs: SQLite-backed Observer that persists a
     harness_runs row per dispatch with status, usage, cost, duration,
     trace_id, session_id, and a bounded logs excerpt.

Schema: new migration 019_harness.sql adds agents.harness_name /
local_command / harness_config_json columns and the backend-agnostic
harness_runs table with indices on (agent, created_at), (status),
(trace_id), (run_id). internal/reactor/reactor_test.go inline schema
updated to match.

Deployment: deploy/kubic/otel-collector.yaml stands up an otel-collector
Deployment + ConfigMap + ClusterIP Service in the synapbus namespace on
kubic, receiving OTLP gRPC (4317) and HTTP (4318) and exporting debug
output until a Tempo/Jaeger backend lands.

Docs: docs/harness-otel-research.html compares scion and paperclip
side-by-side and maps the current synapbus executor surface; its
companion docs/harness-otel-design.md carries the phase plan, span
taxonomy, and migration schema verbatim.

The reactor currently still calls k8s.JobRunner directly — rewiring it
through the Registry is a follow-up, intentionally out of scope for
this branch to keep the refactor reversible. The new packages are
independently tested (~78 new tests across 7 packages) and the full
project test suite passes.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-13 20:23:04 +03:00