Files
synapbus/examples/cold-topic-explainer/README.md
T
Algis DumbrisandClaude Opus 4.6 04e4e3ca37 feat(harness): GEMINI.md support + cold-topic-explainer example
Adds everything needed to run a real multi-Gemini-model reactive agent
loop end-to-end on SynapBus.

internal/harness/subprocess/config.go:
  * AgentConfig.GeminiMD — content of workdir/GEMINI.md
  * MaterialiseAgentConfig writes GEMINI.md AND workdir/.gemini/settings.json
    when gemini_md is set. The settings file carries the same mcp_servers
    list as .mcp.json (so a Gemini child running from the workdir gets
    the exact MCP surface the operator configured, not the user's
    ~/.gemini/settings.json).
  * 2 new config_test cases: GEMINI.md + .gemini/settings.json round
    trip, GEMINI.md with empty mcp_servers still writes the settings
    file (explicitly clearing any inherited home config).

internal/harness/registry.go — BUG FIX:
  Resolve() now honours agent.HarnessName (explicit selection) BEFORE
  the inference chain, matching the reactor's own agentBackendKind
  policy. Previously, when multiple backends were registered,
  Resolve would pick "webhook" for every non-K8s agent — even when the
  agent's HarnessName was "subprocess" — because the original fallback
  chain put webhook first. This is why the first cold-topic-explainer
  run failed with "webhook: agent has no webhook config". Discovered
  during e2e testing.

internal/admin/socket.go + cmd/synapbus/admin.go:
  New `messages.send` admin command (socket + CLI). Sends a DM as any
  agent through the messaging service, bypassing the REST/MCP auth
  layers. Local-only via the admin Unix socket, so the threat model is
  "whoever can reach the socket is already admin".

  CLI:
    synapbus messages send --from X --to Y --body "..." [--priority N]
    synapbus messages send --from X --to Y --body-file path
    echo "..." | synapbus messages send --from X --to Y

  Used by the harness shell wrappers (so Gemini subprocess agents can
  DM each other) and by run_task.sh (to kick off a chain as a human
  user without implementing the REST session flow).

examples/cold-topic-explainer/ (NEW):
  Runnable 3-agent Gemini demo that exercises the subprocess harness,
  reactive triggers, recursive update, and all the preconditions (depth,
  budget, cooldown) end-to-end on a separate isolated synapbus instance.

  Layout:
    README.md         — usage + troubleshooting + cost notes
    start.sh          — builds synapbus, launches on port 18088 with
                        ./data, creates user + agents + harness configs,
                        marks agents reactive via sqlite3
    run_task.sh       — sends initial DM algis → decomposer-pro, polls
                        reactive_runs + messages for the FINAL: reply,
                        prints the result or dumps reactive_runs on
                        timeout for debugging
    stop.sh           — SIGTERM + 5s grace + SIGKILL fallback
    wrapper.sh        — shared subprocess local_command: reads
                        message.json + GEMINI.md, calls gemini headless
                        with --approval-mode yolo, strips the
                        "MCP issues detected" noise prefix, routes the
                        cleaned response to the next agent via
                        `synapbus messages send` over the admin socket
    configs/
      decomposer-pro.json  — gemini-3.1-pro-preview
                              (gemini-2.5-pro is currently capacity-
                              exhausted on Google's side)
      writer-flash.json     — gemini-2.5-flash
      critic-lite.json      — gemini-2.5-flash-lite
    .gitignore        — data/, bin/, synapbus.log, .synapbus.pid

  The wrapper does NOT rely on gemini's MCP tool-calling (which was
  unreliable in testing). Gemini is used as a pure text generator; the
  shell decides routing based on AGENT_ROLE:
    - decomposer → NEXT_AGENT (writer)
    - writer     → NEXT_AGENT (critic)
    - critic     → OWNER_AGENT if response starts with FINAL:,
                   REVISE_AGENT otherwise

E2E VERIFICATION (real run, real Gemini, not a mock):

  Topic: "how does SynapBus unify message delivery, reactive agent
          triggers, and harness runs on a single SQLite database?"

  Result (from data/synapbus.db after one successful run):
    harness_runs:
      #1 decomposer-pro subprocess success 106s
      #2 writer-flash    subprocess success 155s
      #3 critic-lite     subprocess success  10s
    reactive_runs: 3 rows, all succeeded, trigger_from chain:
      algis → decomposer-pro → writer-flash → critic-lite
    messages:
      #1 algis → decomposer-pro (topic)
      #2 decomposer-pro → writer-flash (Q1/Q2/Q3 breakdown)
      #3 writer-flash → critic-lite (3-paragraph draft)
      #4 critic-lite → algis (FINAL: + polished 3-paragraph explainer)

  Critic converged in one pass (all scores ≥ 8), so the writer↔critic
  refinement loop didn't need to recurse — but the plumbing for it
  (REVISE: branch in wrapper.sh, depth limit in reactor) is wired and
  ready. Flipping the critic's acceptance bar exercises the recursion.

  Full `go test ./...` remained green through all changes.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-13 22:01:36 +03:00

103 lines
4.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# cold-topic-explainer
Toy multi-agent task that exercises the subprocess harness end-to-end.
Three Gemini agents on different models collaborate via SynapBus DMs
to produce a 3-paragraph explainer for a topic, with a
writer ↔ critic refinement loop.
## Roles
| Agent | Model | Job |
|-------------------|------------------------|---|
| `decomposer-pro` | `gemini-2.5-pro` | Receives the topic, splits it into what / why / how, DMs `writer-flash` |
| `writer-flash` | `gemini-2.5-flash` | Drafts (or revises) the 3-paragraph explainer, DMs `critic-lite` |
| `critic-lite` | `gemini-2.5-flash-lite` | Rates each paragraph 1–10. Scores all ≥ 8 → DMs `algis` with `FINAL:`. Else DMs `writer-flash` with `REVISE:` and specific fixes |
This exercises:
- **Decomposition** — `decomposer-pro` splits one request into 3 sub-questions
- **Delegation** — each agent DMs the next, routed by the SynapBus reactor
- **Recursive update** — the writer↔critic loop runs until convergence or
`max_trigger_depth` fires (default 6, giving ~3 full refinement rounds)
Every hop is a subprocess reactive run, subject to the same depth /
budget / cooldown guards as a K8s reactive run. Each hop writes a
`harness_runs` row with usage, cost, duration, and trace id.
## Prereqs
- `gemini` CLI installed and authenticated (`gemini auth login` done once)
- Go 1.25+
- `jq`, `curl`, `sqlite3` available on PATH
- An unused TCP port (default 18088)
## Run it
```bash
./start.sh
./run_task.sh "how does the SynapBus reactor's pending_work flag coalesce bursts of DMs?"
./stop.sh
```
## What happens
- `start.sh` builds `synapbus` from the current checkout, launches a
separate instance on port **18088** with a local `./data` directory,
creates user `algis` (password `algis`), creates three AI agents, and
configures each agent's `harness_config_json` with GEMINI.md, MCP
pointer, role env, and the wrapper script invocation.
- `run_task.sh` kicks off the chain by sending an initial DM from
`algis` to `decomposer-pro` via the admin socket, then polls for a
DM **to** `algis` whose body starts with `FINAL:`. Prints the body
when it arrives (or gives up after 4 min).
- `stop.sh` signals the synapbus PID and waits for it to exit
cleanly.
## View during the run
- **Web UI**: <http://localhost:18088> — log in as `algis` / `algis-demo-pw`
- **Agent detail** (see Harness panel + traces):
- <http://localhost:18088/agents/decomposer-pro>
- <http://localhost:18088/agents/writer-flash>
- <http://localhost:18088/agents/critic-lite>
- **Live slog JSON**: `tail -f synapbus.log | jq -c 'select(.component=="reactor" or .harness)'`
- **All DMs in order**: `./bin/synapbus --socket ./data/synapbus.sock messages list --limit 50`
- **Harness runs**: `sqlite3 ./data/synapbus.db 'SELECT run_id, agent_name, backend, status, duration_ms, tokens_in, tokens_out, cost_usd FROM harness_runs ORDER BY id'`
### OpenTelemetry
Off by default. To ship spans to a collector while you run the task:
```bash
SYNAPBUS_OTEL_ENABLED=1 SYNAPBUS_OTEL_ENDPOINT=otel-collector.synapbus.svc.cluster.local:4318 ./start.sh
```
Or stand up a local collector first using `deploy/kubic/otel-collector.yaml`.
Without a collector, the same information is available in `synapbus.log`
as slog JSON and in the `harness_runs` table.
## Cost
Rough cost per successful run, assuming 2 writer-critic iterations:
| Hops | Model | Cost |
|------|---------------|------|
| 1 | gemini-2.5-pro | ~$0.01 |
| 2 | gemini-2.5-flash | ~$0.01 |
| 3 | gemini-2.5-flash-lite | ~$0.002 |
| **Total** | | ~$0.02 |
The daily trigger budget per agent is capped at 20 (see `start.sh`) so
this example cannot accidentally spend more than pennies per day even
if the reactor loops on a bug.
## Files
- `start.sh` — launch separate synapbus + configure agents
- `run_task.sh` — kickoff DM + poll for final
- `stop.sh` — graceful shutdown
- `wrapper.sh` — shell wrapper used as the agents' `local_command`;
reads `message.json`, calls `gemini`, routes the result back via the
admin socket
- `configs/*.json` — per-agent `harness_config_json` blobs