Root cause: ConsolidatorWorker.tryDispatch / launchOne incremented
the per-(owner, day) `jobs_started` counter immediately after
JobsStore.Create, BEFORE the dispatch flip succeeded. When the
subsequent steps fail (token issue, agent lookup, dispatch flip
race) the job row is Completed as `failed` but `jobs_started`
remains incremented — and the error paths never call
RecordCompletion, so `jobs_failed` stays flat while `jobs_started`
drifts upward.
Over hours/days, owners whose dispatches fail consistently (e.g.
owner_id=2 in the kubic deployment, hitting one of the harness
failure modes from commit bfb2551) accumulate phantom
`jobs_started` until the default DreamDailyJobLimit=100 trips. From
that point every tick logs `circuit broken … reason=jobs_exceeded`
for all four job types, even though no real jobs ran — and the
counter never decays until midnight UTC.
Fix: move `usage.RecordStart(...)` to AFTER a successful
`jobs.Dispatch(...)` in both tryDispatch (consolidator.go:518)
and launchOne (consolidator.go:455). Now only dispatches that
actually transitioned a row to `dispatched` count against the
daily-job-limit gate.
Test: TestConsolidator_PreDispatchFailureDoesNotBurnJobsStarted
seeds DreamDailyJobLimit=2, makes the agent lookup fail, calls
ForceRun three times, asserts jobs_started stays 0 and the gate
still allows. Verified to fail without the fix
(jobs_started=2 / reason=jobs_exceeded) and pass with it.
Counterpart TestConsolidator_DispatchSuccessIncrementsJobsStarted
asserts jobs_started=1 on a real successful dispatch so the
counter still feeds the gate correctly.
Operational note: this prevents future inflation. Existing stuck
rows for owner_id=2 in today's `memory_dream_usage` bucket need a
one-shot SQL fix —
UPDATE memory_dream_usage
SET jobs_started = jobs_succeeded + jobs_failed
WHERE date = date('now')
AND owner_id = '2';
or simply wait for the next UTC-midnight reset.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>