Token-usage wiring (closes the 0-tokens gap in memory_dream_usage):
- internal/harness/k8sjob/k8sjob.go: after extractResultJSON, parse
tokens_in/tokens_out/tokens_cached/cost_usd from the final JSON
envelope and stash into ExecResult.Usage so the dream worker's
UsageGate circuit breaker actually counts consumption.
- dream-agent/dream_runner.py: Max20 OAuth sessions don't surface
per-call tokens through the SDK's ResultMessage.usage. Falls back
to a turn-based estimate so the gate has SOME signal:
tokens_in_est = turns * 5000 + tool_calls * 2000
tokens_out_est = turns * 300
Calibrated against observed reflection runs.
Watchdog (deploy/kubic/watchdog/):
- watchdog.yaml: in-cluster CronJob runs every hour at :05 past UTC,
with a dedicated ServiceAccount + Role granting (get/list/exec on
pods, patch+update on deployments/scale) inside the synapbus
namespace only.
- Health checks: pod readiness + restart count; last-1h job
succ/fail/in_flight counts; today's jobs_started + tokens_in +
circuit_broken.
- Red flags that auto-stop synapbus (scale to 0):
* pod restart count > 3
* failed dream jobs in last 1h > 20
* jobs_started today > 200 OR tokens_in > 30M
* circuit broke AND still firing (started >> completed)
- Dockerfile: slim alpine + kubectl v1.30.5 binary (synapbus-watchdog:v1).
Built locally and imported into kubic's containerd because the
public docker.io/bitnami/kubectl manifest was returning text/html
from kubic's network egress.
Replaces the schedule-skill remote-agent approach because Anthropic
cloud agents can't reach kubic.home.arpa (LAN-only) and can't call
kubectl scale. The k8s CronJob is the right primitive for an
in-cluster safety watchdog.
Live evidence: first manual run on kubic reported
pod=synapbus-... ready=true restarts=0
last_1h jobs total=18 succ=18 fail=0 in_flight=0
today: jobs_started=189 tokens_in=0 succeeded=169 failed=15
HEALTHY — no action
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>