fix(health): route /readyz through read pool to survive long writers

Repeated production crash: pod ran 27min–3.5h then went ready=false,
restarts=0 (process alive but readiness probe failing). Watchdog
correctly scaled deploy to 0 each time.

Root cause: /readyz calls db.PingContext() on the write pool, which
has MaxOpenConns=1 (serialized writes). The consolidator's dream-job
dispatch (introduced in 020) holds that single connection for 30s+
during one tick: it creates a K8s Job, writes the job row, issues a
dispatch token, all sequentially. /readyz blocks waiting for the
connection through the entire dispatch. With probe period=5s,
failureThreshold=3, the pod flips to NotReady after ~15s — long
before the dispatch finishes.

The new diagnostic: rebuilt v0.17.0 (pre-020) on kubic — runs 5h+
clean, memory flat at 134Mi. v0.21.2 (with 020) dies within hours.
The dispatch path is the only ~30s write holding the conn.

Fix: pass db.QueryDB() to health.NewChecker. QueryDB returns the
read pool (MaxOpenConns=8) when available, write pool when not, so
the readiness probe can run concurrently with any writer.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
Algis Dumbris
2026-05-15 15:17:42 +03:00
co-authored by Claude Opus 4.7
parent c4d888082d
commit b03350ffc1
+7 -2
View File
@@ -734,8 +734,13 @@ func runServe(cmd *cobra.Command, args []string) error {
slog.Info("consolidator (dream) worker disabled (SYNAPBUS_DREAM_ENABLED=0)")
}
// Create health checker
healthChecker := health.NewChecker(db.DB, version)
// Create health checker. Use the read pool (8 concurrent conns) so /readyz
// can't be starved by a long-running writer holding the serialized write
// connection — most notably the consolidator's dream-job dispatch, which
// does a K8s Job create + DB writes that can run 30s+. With the write pool
// (MaxOpenConns=1) the probe blocked for the entire dispatch, the kubelet
// flipped the pod to not-ready, and the watchdog scaled the deploy to 0.
healthChecker := health.NewChecker(db.QueryDB(), version)
// Set up chi router
r := chi.NewRouter()