Repeated production crash: pod ran 27min–3.5h then went ready=false,
restarts=0 (process alive but readiness probe failing). Watchdog
correctly scaled deploy to 0 each time.
Root cause: /readyz calls db.PingContext() on the write pool, which
has MaxOpenConns=1 (serialized writes). The consolidator's dream-job
dispatch (introduced in 020) holds that single connection for 30s+
during one tick: it creates a K8s Job, writes the job row, issues a
dispatch token, all sequentially. /readyz blocks waiting for the
connection through the entire dispatch. With probe period=5s,
failureThreshold=3, the pod flips to NotReady after ~15s — long
before the dispatch finishes.
The new diagnostic: rebuilt v0.17.0 (pre-020) on kubic — runs 5h+
clean, memory flat at 134Mi. v0.21.2 (with 020) dies within hours.
The dispatch path is the only ~30s write holding the conn.
Fix: pass db.QueryDB() to health.NewChecker. QueryDB returns the
read pool (MaxOpenConns=8) when available, write pool when not, so
the readiness probe can run concurrently with any writer.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>