Borrow what works from GoogleCloudPlatform/scion and paperclipai/paperclip, skip what doesn't, and sketch a minimal harness + OTel integration that fits SynapBus's Go / MCP / SQLite spine.
Both repos converge on the same core idea: a narrow Harness / Adapter interface that abstracts "some external AI CLI" behind a single execute(ctx)→result contract, then registers concrete implementations for Claude Code, Gemini CLI, Codex, OpenCode, etc.
Scion's design is the better template for SynapBus: it's Go, it ships OTel via env-var injection into child processes, and its Harness interface cleanly separates provisioning from invocation — exactly the seam we're missing.
Paperclip contributes two ideas we should adopt: (a) an adapter registry with capability flags so a router can pick the best backend at dispatch time, and (b) a session codec per adapter so long-running agents can be resumed.
SynapBus today has no subprocess executor, no unified runner interface, and no OTel spans — only a K8s-Job path and an HTTP-webhook path living as two disjoint code paths. A small internal/harness/ package would unify both and unlock local-subprocess execution.
Despite the name collision with the SCION internet-architecture project, GoogleCloudPlatform/scion is a multi-agent orchestration harness for evaluating and running "deep agents" (Claude Code, Gemini CLI, Codex, OpenCode) inside isolated containers. It is explicitly not a planner and not a verifier — it is the control plane and observability spine around arbitrary agent CLIs.
pkg/api/harness.go:22–68
type Harness interface { Name() string AdvancedCapabilities() HarnessAdvancedCapabilities GetEnv(agentName, agentHome, unixUsername string) map[string]string GetCommand(task string, resume bool, baseArgs []string) []string DefaultConfigDir() string SkillsDir() string HasSystemPrompt(agentHome string) bool Provision(ctx context.Context, agentName, agentDir, agentHome, agentWorkspace string) error GetEmbedDir() string GetInterruptKey() string GetHarnessEmbedsFS() (embed.FS, string) InjectAgentInstructions(agentHome string, content []byte) error InjectSystemPrompt(agentHome string, content []byte) error // the key OTel seam — returns env vars that the container runtime // will merge into the child process env before exec GetTelemetryEnv() map[string]string ResolveAuth(auth AuthConfig) (*ResolvedAuth, error) }
Three things to notice:
Provision is separate from GetCommand: one-shot setup (write .claude.json, pre-approve tool fingerprints, materialise skill files) versus per-invocation command building.GetEnv / GetTelemetryEnv / ResolveAuth all return maps of env vars. The container runtime layer merges them. This means every harness is credential-injection-agnostic and telemetry-injection-agnostic — you can point a whole pod at a different OTel collector by changing one map.AdvancedCapabilities() lets a dispatcher ask "does this harness support system prompts?" and degrade gracefully (fall back to InjectAgentInstructions) when it doesn't.pkg/harness/harness.go:37–57
func New(name string) Harness { switch name { case "claude": return &ClaudeCode{} case "gemini": return &GeminiCLI{} case "opencode": return &OpenCode{} case "codex": return &Codex{} } if h := pluginMgr.Lookup(name); h != nil { return h } return &Generic{} // universal fallback }
pkg/harness/claude_code.go:311–320
func (c *ClaudeCode) GetTelemetryEnv() map[string]string { return map[string]string{ "CLAUDE_CODE_ENABLE_TELEMETRY": "1", "OTEL_METRICS_EXPORTER": "otlp", "OTEL_LOGS_EXPORTER": "otlp", "OTEL_EXPORTER_OTLP_PROTOCOL": "grpc", "OTEL_EXPORTER_OTLP_ENDPOINT": "http://localhost:4317", "OTEL_METRIC_EXPORT_INTERVAL": "30000", } }
Scion's own Go code emits OTel logs via the OTLP log exporter (pkg/util/logging/otel_provider.go:26–61) and bridges slog into it (pkg/util/logging/otel.go:85–119). W3C traceparent headers are extracted at HTTP ingress (pkg/util/logging/trace.go) so trace context can flow across the dispatcher → runtime → container boundary.
Scion does not decompose tasks. A single task string goes to the agent and the agent's own model decides how to break it up. Coordination between agents happens via a structured StructuredMessage envelope (pkg/messages/types.go:46–61) with fields {sender, recipient, msg, type, urgent, broadcasted, attachments} — an on-disk analogue of a SynapBus channel post.
Harness interface shape, the env-var-injection model for both auth & telemetry, the capability-flags degradation pattern, and the Provision/GetCommand split. Ignore the container runtime abstraction — SynapBus already has K8s-Job + webhook paths and doesn't need a second one.
Paperclip is a Node/Express control plane for running 10–20 agent "companies" with org charts, budgets, and approval gates. Wildly different product — but it has a clean adapter interface worth borrowing.
packages/adapter-utils/src/types.ts:292–331
export interface ServerAdapterModule { type: string; execute(ctx: AdapterExecutionContext): Promise<AdapterExecutionResult>; testEnvironment(ctx: AdapterEnvironmentTestContext): Promise<AdapterEnvironmentTestResult>; listSkills?: (ctx) => Promise<AdapterSkillSnapshot>; syncSkills?: (ctx, desired: string[]) => Promise<AdapterSkillSnapshot>; sessionCodec?: AdapterSessionCodec; // resume / serialize sessions models?: AdapterModel[]; listModels?: () => Promise<AdapterModel[]>; agentConfigurationDoc?: string; onHireApproved?: (payload, cfg) => Promise<HireApprovedHookResult>; getQuotaWindows?: () => Promise<ProviderQuotaResult>; }
AdapterExecutionResult — types.ts:64–95
{ exitCode, signal, timedOut, errorMessage, errorCode,
usage: { inputTokens, outputTokens, cachedInputTokens },
resultJson, costUsd,
question?: { prompt, choices } // can pause for human approval
}
Ten adapters are registered via a mutable map in server/src/adapters/registry.ts:89–222: claude-local, codex-local, cursor, gemini, opencode, pi, openclaw, hermes, http, process. External adapters are loaded from plugins asynchronously (lines 244–270).
server/src/services/heartbeat.ts
No DAG, no queue, no planner. Agents wake on a heartbeat (schedule or event), atomically claim assigned issues via a per-agent start lock (withAgentStartLock(), lines 331–346), run once, and go back to sleep. Concurrency is per-agent (default 1, configurable to 10). Task decomposition is entirely delegated to the agent's own model.
None that's interesting. Exit code 0 = success; timeouts and process-loss retries are tracked; there is no LLM judge, no schema validation, no test runner. Verification is whatever the running agent chooses to self-report in resultJson.
Pino structured logging (server/src/middleware/logger.ts:29–45) + a custom telemetry client (server/src/telemetry.ts:12–26) that batch-flushes events every 60s. No OpenTelemetry. This is the weakest part relative to scion.
sessionCodec idea — each harness knows how to serialise/resume its own session, so SynapBus can carry conversation state across reactive runs; (2) testEnvironment() as a preflight — "is the CLI installed, is auth valid, can it reach the model?"; (3) getQuotaWindows() / cost tracking in the result envelope.
| Aspect | scion (Go) | paperclip (Node) | synapbus today |
|---|---|---|---|
| Core interface | api.Harness — 15 methods, env-var-centric |
ServerAdapterModule — execute() + optional hooks |
k8s.JobRunner (K8s only) + webhooks.EventDispatcher — no unification |
| Backends shipped | claude, gemini, codex, opencode, generic fallback | claude, codex, cursor, gemini, opencode, pi, openclaw, hermes, http, process | K8s Job (one) + outbound HTTP webhook |
| Credential injection | env vars from GetEnv()+ResolveAuth(); HostPath for ~/.claude |
per-adapter config objects; provider SDK auth | K8s env vars from agent's k8s_env_json; HostPath ~/.claude (reactor.go:281–286) |
| Task decomposition | None — passes whole task string to agent | None — agents pull from issue queue themselves | None — reactive trigger wraps one inbound message |
| Verification | Workspace sync + agent logs; no judge | Exit code, token usage, timeout; no judge | K8s Job success/fail + pod logs stored in ReactiveRun |
| Observability | OTel logs via OTLP gRPC, W3C trace-context propagation, slog bridge |
Pino structured logs + custom telemetry client | slog JSON only; Prometheus metrics for reactor; OTel deps present but unused in Go code |
| Coordination | Containers per agent; inter-agent messages via typed envelope | Heartbeat + atomic per-agent lock; org-chart hierarchy | MCP channels & DMs; reactive triggers fire on inbound |
| Capability flags | AdvancedCapabilities() for graceful degradation |
Optional methods on the interface | None — hardcoded paths |
| Session resume | Yes — GetCommand(task, resume bool, ...) |
Yes — per-adapter sessionCodec |
None — each reactive run is fresh |
TriggerMode=reactive reactor.go:51batchv1.Job with env vars SYNAPBUS_MESSAGE_ID/_BODY/_FROM_AGENT/_EVENT/_CHANNEL k8s/runner.go:96–183ReactiveRun reactor/poller.goX-SynapBus-Signature HMAC, X-SynapBus-Depth delivery.go:290–302These paths are two disjoint islands. There is:
Runner/Harness interface — the reactor switches on K8s availability with a NoopRunner fallbackgo.mod but are unimported~/.claude + env vars; webhook path has nonebenchmark/sdk_backend.py returns it but core Go reactor does notThe recent benchmark/sdk_backend.py (commit 0e25fbc) is a Python two-backend fallback (anthropic SDK → claude-agent-sdk) that foreshadows exactly the abstraction we need — but in the benchmark tree, not in core.
internal/harness/package harness type Capabilities struct { SystemPrompt bool SessionResume bool Skills bool OTelNative bool // child process honours OTEL_* env vars MaxConcurrency int } type ExecRequest struct { AgentName string Message *messaging.Message // triggering message Context []*messaging.Message // optional conversation window Budget Budget // tokens, cost, wallclock Env map[string]string // caller-provided overrides Skills []string } type ExecResult struct { ExitCode int Logs string ResultJSON json.RawMessage Usage Usage // { in, out, cached tokens, cost } TraceID string // W3C, for correlation Err error } type Harness interface { Name() string Capabilities() Capabilities TestEnvironment(ctx context.Context) error // preflight Provision(ctx context.Context, agent *agents.Agent) error // one-shot setup Execute(ctx context.Context, req *ExecRequest) (*ExecResult, error) Cancel(ctx context.Context, runID string) error } type Registry struct { /* map[string]Harness + mutex */ } func (r *Registry) Register(h Harness) func (r *Registry) Resolve(agent *agents.Agent) (Harness, error) func (r *Registry) Execute(ctx context.Context, req *ExecRequest) (*ExecResult, error)
| Package | Wraps | Status |
|---|---|---|
internal/harness/k8sjob | existing internal/k8s path | refactor into Harness |
internal/harness/subprocess | os/exec with env-map + workdir + timeout | new |
internal/harness/webhook | existing internal/webhooks/delivery.go | wrap as Harness, async result via DB poll |
internal/harness/stub | in-process fake for tests | new, test-only |
Registry.Resolve(agent) picks a backend based on:
agent.HarnessName field (new column, nullable)K8sImage and k8s.JobRunner.IsAvailable() → k8sjobWebhooks registered → webhookLocalCommand configured → subprocessErrNoBackend016_harness.sql: add agents.harness_name, agents.local_command, agents.harness_config_jsonharness_runs: mirror of ReactiveRun but backend-agnostic, with backend, trace_id, span_id, usage_in, usage_out, cost_usd, result_jsonReactiveRun into harness_runs in a follow-up migrationScion's pattern is the template: (a) initialize an OTel tracer provider in the main process, (b) start a span per harness invocation, (c) inject the trace context into the child as env vars, (d) ship spans via OTLP gRPC to whatever collector is configured.
New file internal/observability/otel.go:
func Init(ctx context.Context, cfg Config) (shutdown func(context.Context) error, err error) { res, _ := resource.New(ctx, resource.WithAttributes(semconv.ServiceName("synapbus")), ) exp, _ := otlptracegrpc.New(ctx, otlptracegrpc.WithEndpoint(cfg.Endpoint), otlptracegrpc.WithInsecure(), ) tp := sdktrace.NewTracerProvider( sdktrace.WithBatcher(exp), sdktrace.WithResource(res), ) otel.SetTracerProvider(tp) otel.SetTextMapPropagator(propagation.TraceContext{}) return tp.Shutdown, nil }
| Span name | Where | Key attributes |
|---|---|---|
mcp.tool.execute | MCP handler entry | mcp.tool, agent.name, message.id |
reactor.dispatch | reactor.Dispatch() | agent.name, trigger.depth, budget.remaining |
harness.resolve | Registry.Resolve | harness.name, fallback.chain |
harness.provision | Harness.Provision | harness.name, agent.home |
harness.execute | Harness.Execute | harness.name, run.id, usage.*, cost.usd, exit.code |
harness.k8s.job.create | k8sjob backend | k8s.job.name, k8s.namespace, k8s.image |
harness.subprocess.exec | subprocess backend | proc.argv[0], proc.pid, proc.workdir |
harness.webhook.deliver | webhook backend | http.url, http.status_code, retry.count |
For each backend, the current span's traceparent is serialised via propagation.TraceContext{}.Inject into an env-var map and merged with Harness.GetTelemetryEnv():
func injectTraceEnv(ctx context.Context, dst map[string]string) { carrier := propagation.MapCarrier{} otel.GetTextMapPropagator().Inject(ctx, carrier) for k, v := range carrier { // OTel convention: TRACEPARENT / TRACESTATE env names dst[strings.ToUpper(k)] = v } dst["OTEL_EXPORTER_OTLP_ENDPOINT"] = cfg.ChildEndpoint // same collector dst["OTEL_SERVICE_NAME"] = "synapbus-agent-" + agentName dst["OTEL_RESOURCE_ATTRIBUTES"] = "synapbus.run_id=" + runID }
For K8s: merged into corev1.EnvVar slice at k8s/runner.go:105–119. For subprocess: merged into cmd.Env. For webhook: added as HTTP headers (traceparent, tracestate) alongside the existing X-SynapBus-* headers.
Keep the existing Prometheus registry (internal/metrics/metrics.go) — it's already wired — but also emit a minimal set via OTel meter, so a single OTLP collector sees both spans and metrics:
synapbus.harness.runs (counter, labels: harness, status)synapbus.harness.duration_ms (histogram)synapbus.harness.tokens_in / tokens_out (counters)synapbus.harness.cost_usd (counter)Three new env vars (matching scion naming, with SYNAPBUS_ prefix for ours):
SYNAPBUS_OTEL_ENDPOINT — e.g. http://otel-collector:4317SYNAPBUS_OTEL_INSECURE — bool, default true for LANSYNAPBUS_OTEL_ENABLED — bool, default false (opt-in)Until a real collector exists on kubic, a file exporter (stdouttrace) or the existing trace.Tracer (SQLite trace table) can back the same interface via an adapter.
GetInterruptKey) — e.g. double-Escape for Claude Code — useful for cancel semantics..claude.json customApiKeyResponses — removes the "did you really want to use this key?" prompt.StructuredMessage envelope — SynapBus messages already have most of this; add a type enum (instruction/input-needed/state-change).testEnvironment() preflight — a health check per harness, runnable from the admin CLI ("can this agent actually dispatch?").sessionCodec — serialise/resume an agent conversation across reactive runs. Gives SynapBus a real "sticky" agent without re-prompting.benchmark/sdk_backend.py, worth lifting into the core result type.docs/harness-otel-design.md (written alongside this report)internal/harness/ with the interface, registry, and a stub backend. Pure Go, no external deps added.Harness interface without changing behaviour. Existing tests stay green.subprocess backend + per-agent local_command config + migration 016.harness_runs. testEnvironment preflight exposed via admin CLI.kubic.home.arpa already, or do we deploy one first (Tempo? Jaeger? stdout only for now)?go-plugin), or is a compile-time registry enough?