Two independent fixes hit while running the cold-topic-explainer demo
end-to-end.
1. observability/otel.go — schema URL conflict on Init
When SYNAPBUS_OTEL_ENABLED=1, Init() failed with:
observability: build resource: conflicting Schema URL:
https://opentelemetry.io/schemas/1.26.0 and
https://opentelemetry.io/schemas/1.21.0
resource.Default() ships with schema 1.26.0 (newer otel/sdk) but I
was passing semconv.SchemaURL from v1.21 into a NewWithAttributes
call. resource.Merge rejects that.
Fix: use resource.NewSchemaless for the service.* attributes so our
side of the merge has no schema URL and slots cleanly into whatever
Default provides. ServiceVersion is now only attached when non-empty
(avoids a stray service.version="" attribute).
Two new regression tests:
TestInit_EnabledSucceeds — Enabled=true with all fields set
TestInit_EnabledWithNoVersion — Enabled=true with empty version
Both point at an unroutable endpoint so the batcher never actually
exports; the bug reproduced during Init(), which is all we need.
2. examples/cold-topic-explainer/start.sh — rebuild embedded SPA
The Svelte Web UI loaded blank because internal/web/dist/ had a
mismatched index.html + stale _app/immutable/entry/ assets (a build
had updated index.html but not the chunks, so every asset URL fell
through to the SPA HTML fallback and the browser tried to execute
HTML as JavaScript).
The canonical path is `make web`, but start.sh never ran it, so a
working demo depended on the developer having run `make web` first.
Fix: start.sh now rebuilds the SPA when web/src is newer than the
embedded dist/index.html, using the already-installed
web/node_modules (no reinstall). Falls back with a "run make web
once" hint when node_modules isn't present. This keeps the fast
path fast (~2s vite build after cache warm) and eliminates the
silent-stale-dist trap.
E2E verified after both fixes:
* SYNAPBUS_OTEL_ENABLED=1 start.sh no longer crashes.
* `curl /_app/immutable/entry/start.*.js` returns real JavaScript
(Content-Type: text/javascript) instead of the index.html
fallback.
* Chrome-in-MCP navigation to http://localhost:18088/ renders the
login form with no SynapBus-originated console errors.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
cold-topic-explainer
Toy multi-agent task that exercises the subprocess harness end-to-end. Three Gemini agents on different models collaborate via SynapBus DMs to produce a 3-paragraph explainer for a topic, with a writer ↔ critic refinement loop.
Roles
| Agent | Model | Job |
|---|---|---|
decomposer-pro |
gemini-2.5-pro |
Receives the topic, splits it into what / why / how, DMs writer-flash |
writer-flash |
gemini-2.5-flash |
Drafts (or revises) the 3-paragraph explainer, DMs critic-lite |
critic-lite |
gemini-2.5-flash-lite |
Rates each paragraph 1–10. Scores all ≥ 8 → DMs algis with FINAL:. Else DMs writer-flash with REVISE: and specific fixes |
This exercises:
- Decomposition —
decomposer-prosplits one request into 3 sub-questions - Delegation — each agent DMs the next, routed by the SynapBus reactor
- Recursive update — the writer↔critic loop runs until convergence or
max_trigger_depthfires (default 6, giving ~3 full refinement rounds)
Every hop is a subprocess reactive run, subject to the same depth /
budget / cooldown guards as a K8s reactive run. Each hop writes a
harness_runs row with usage, cost, duration, and trace id.
Prereqs
geminiCLI installed and authenticated (gemini auth logindone once)- Go 1.25+
jq,curl,sqlite3available on PATH- An unused TCP port (default 18088)
Run it
./start.sh
./run_task.sh "how does the SynapBus reactor's pending_work flag coalesce bursts of DMs?"
./stop.sh
What happens
start.shbuildssynapbusfrom the current checkout, launches a separate instance on port 18088 with a local./datadirectory, creates useralgis(passwordalgis), creates three AI agents, and configures each agent'sharness_config_jsonwith GEMINI.md, MCP pointer, role env, and the wrapper script invocation.run_task.shkicks off the chain by sending an initial DM fromalgistodecomposer-provia the admin socket, then polls for a DM toalgiswhose body starts withFINAL:. Prints the body when it arrives (or gives up after 4 min).stop.shsignals the synapbus PID and waits for it to exit cleanly.
View during the run
- Web UI: http://localhost:18088 — log in as
algis/algis-demo-pw - Agent detail (see Harness panel + traces):
- Live slog JSON:
tail -f synapbus.log | jq -c 'select(.component=="reactor" or .harness)' - All DMs in order:
./bin/synapbus --socket ./data/synapbus.sock messages list --limit 50 - Harness runs:
sqlite3 ./data/synapbus.db 'SELECT run_id, agent_name, backend, status, duration_ms, tokens_in, tokens_out, cost_usd FROM harness_runs ORDER BY id'
OpenTelemetry
Off by default. To ship spans to a collector while you run the task:
SYNAPBUS_OTEL_ENABLED=1 SYNAPBUS_OTEL_ENDPOINT=otel-collector.synapbus.svc.cluster.local:4318 ./start.sh
Or stand up a local collector first using deploy/kubic/otel-collector.yaml.
Without a collector, the same information is available in synapbus.log
as slog JSON and in the harness_runs table.
Cost
Rough cost per successful run, assuming 2 writer-critic iterations:
| Hops | Model | Cost |
|---|---|---|
| 1 | gemini-2.5-pro | ~$0.01 |
| 2 | gemini-2.5-flash | ~$0.01 |
| 3 | gemini-2.5-flash-lite | ~$0.002 |
| Total | ~$0.02 |
The daily trigger budget per agent is capped at 20 (see start.sh) so
this example cannot accidentally spend more than pennies per day even
if the reactor loops on a bug.
Files
start.sh— launch separate synapbus + configure agentsrun_task.sh— kickoff DM + poll for finalstop.sh— graceful shutdownwrapper.sh— shell wrapper used as the agents'local_command; readsmessage.json, callsgemini, routes the result back via the admin socketconfigs/*.json— per-agentharness_config_jsonblobs