From b03350ffc1545ca168d6abcc6e5919383d6d88bd Mon Sep 17 00:00:00 2001 From: Algis Dumbris Date: Fri, 15 May 2026 15:17:42 +0300 Subject: [PATCH] fix(health): route /readyz through read pool to survive long writers MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Repeated production crash: pod ran 27min–3.5h then went ready=false, restarts=0 (process alive but readiness probe failing). Watchdog correctly scaled deploy to 0 each time. Root cause: /readyz calls db.PingContext() on the write pool, which has MaxOpenConns=1 (serialized writes). The consolidator's dream-job dispatch (introduced in 020) holds that single connection for 30s+ during one tick: it creates a K8s Job, writes the job row, issues a dispatch token, all sequentially. /readyz blocks waiting for the connection through the entire dispatch. With probe period=5s, failureThreshold=3, the pod flips to NotReady after ~15s β€” long before the dispatch finishes. The new diagnostic: rebuilt v0.17.0 (pre-020) on kubic β€” runs 5h+ clean, memory flat at 134Mi. v0.21.2 (with 020) dies within hours. The dispatch path is the only ~30s write holding the conn. Fix: pass db.QueryDB() to health.NewChecker. QueryDB returns the read pool (MaxOpenConns=8) when available, write pool when not, so the readiness probe can run concurrently with any writer. Co-Authored-By: Claude Opus 4.7 (1M context) --- cmd/synapbus/main.go | 9 +++++++-- 1 file changed, 7 insertions(+), 2 deletions(-) diff --git a/cmd/synapbus/main.go b/cmd/synapbus/main.go index e3e60c1..232d0f0 100644 --- a/cmd/synapbus/main.go +++ b/cmd/synapbus/main.go @@ -734,8 +734,13 @@ func runServe(cmd *cobra.Command, args []string) error { slog.Info("consolidator (dream) worker disabled (SYNAPBUS_DREAM_ENABLED=0)") } - // Create health checker - healthChecker := health.NewChecker(db.DB, version) + // Create health checker. Use the read pool (8 concurrent conns) so /readyz + // can't be starved by a long-running writer holding the serialized write + // connection β€” most notably the consolidator's dream-job dispatch, which + // does a K8s Job create + DB writes that can run 30s+. With the write pool + // (MaxOpenConns=1) the probe blocked for the entire dispatch, the kubelet + // flipped the pod to not-ready, and the watchdog scaled the deploy to 0. + healthChecker := health.NewChecker(db.QueryDB(), version) // Set up chi router r := chi.NewRouter()