Files
synapbus/deploy/kubic
Algis DumbrisandClaude Opus 4.7 73a1802155 feat(020): token accounting + watchdog CronJob
Token-usage wiring (closes the 0-tokens gap in memory_dream_usage):
- internal/harness/k8sjob/k8sjob.go: after extractResultJSON, parse
  tokens_in/tokens_out/tokens_cached/cost_usd from the final JSON
  envelope and stash into ExecResult.Usage so the dream worker's
  UsageGate circuit breaker actually counts consumption.
- dream-agent/dream_runner.py: Max20 OAuth sessions don't surface
  per-call tokens through the SDK's ResultMessage.usage. Falls back
  to a turn-based estimate so the gate has SOME signal:
    tokens_in_est = turns * 5000 + tool_calls * 2000
    tokens_out_est = turns * 300
  Calibrated against observed reflection runs.

Watchdog (deploy/kubic/watchdog/):
- watchdog.yaml: in-cluster CronJob runs every hour at :05 past UTC,
  with a dedicated ServiceAccount + Role granting (get/list/exec on
  pods, patch+update on deployments/scale) inside the synapbus
  namespace only.
- Health checks: pod readiness + restart count; last-1h job
  succ/fail/in_flight counts; today's jobs_started + tokens_in +
  circuit_broken.
- Red flags that auto-stop synapbus (scale to 0):
    * pod restart count > 3
    * failed dream jobs in last 1h > 20
    * jobs_started today > 200 OR tokens_in > 30M
    * circuit broke AND still firing (started >> completed)
- Dockerfile: slim alpine + kubectl v1.30.5 binary (synapbus-watchdog:v1).
  Built locally and imported into kubic's containerd because the
  public docker.io/bitnami/kubectl manifest was returning text/html
  from kubic's network egress.

Replaces the schedule-skill remote-agent approach because Anthropic
cloud agents can't reach kubic.home.arpa (LAN-only) and can't call
kubectl scale. The k8s CronJob is the right primitive for an
in-cluster safety watchdog.

Live evidence: first manual run on kubic reported
  pod=synapbus-... ready=true restarts=0
  last_1h jobs total=18 succ=18 fail=0 in_flight=0
  today: jobs_started=189 tokens_in=0 succeeded=169 failed=15
  HEALTHY — no action

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 22:22:13 +03:00
..

SynapBus on kubic

Plain Kubernetes manifests for the kubic single-node MicroK8s cluster (kubic.home.arpa). No Helm — the image is built locally, imported directly into MicroK8s containerd, and rolled with kubectl set image.

Files

File Purpose
namespace.yaml synapbus namespace
pvc.yaml 2 Gi PVC on microk8s-hostpath for /data (DB + WAL + attachments + HNSW index)
secret.example.yaml Template for synapbus-secrets (OpenAI/Gemini keys, mounted via envFrom)
deployment.yaml Single replica, docker.io/library/synapbus:vX.Y.Z-amd64, imagePullPolicy: IfNotPresent (image is pre-loaded into containerd)
service.yaml NodePort 30088 on port 8080
otel-collector.yaml OpenTelemetry collector for traces/metrics

Initial install

kubectl apply -f deploy/kubic/namespace.yaml
kubectl apply -f deploy/kubic/pvc.yaml
# Edit secret.example.yaml first — never commit real keys.
kubectl apply -f deploy/kubic/secret.example.yaml
kubectl apply -f deploy/kubic/service.yaml
kubectl apply -f deploy/kubic/deployment.yaml

Releasing a new version

scripts/deploy-kubic.sh v0.17.0

The script:

  1. docker buildx build --platform linux/amd64 with the version baked in.
  2. docker save to a tarball.
  3. scp to kubic.home.arpa.
  4. ssh kubic 'sudo microk8s ctr image import …' (loads the image into the in-cluster containerd registry — the image is not pushed to a remote registry).
  5. kubectl set image deploy/synapbus synapbus=docker.io/library/synapbus:vX.Y.Z-amd64.
  6. kubectl rollout status … and a /healthz smoke test.

The docker.io/library/ prefix is required because that's how containerd resolves image references that don't specify a registry — synapbus:v… written into the deployment is normalised to docker.io/library/synapbus:v… on the node.

Why no Helm?

The original chart under deploy/helm/ (since deleted) was used for the very first install (Mar 2026) and then went into a failed state when someone ran kubectl set image for a hotfix; subsequent helm upgrade attempts hit server-side-apply ownership conflicts. Rather than reconcile, we now own the manifests directly. The deploy flow is simple enough that templating buys nothing.

Backups

Before any version that touches schema, snapshot /data:

kubectl exec -n synapbus deploy/synapbus -- \
  tar -C /data -cf - synapbus.db synapbus.db-shm synapbus.db-wal vapid_keys.json \
  | tar -xf - -C "$HOME/synapbus-backups/$(date -u +%Y%m%dT%H%M%SZ)/"

Then sqlite3 synapbus.db 'PRAGMA wal_checkpoint(TRUNCATE); PRAGMA integrity_check;' to fold the WAL into the main file and verify integrity before archiving.