feat: harness-agnostic wrappers + OpenTelemetry integration

Introduces internal/harness — a minimal Harness interface inspired by
GoogleCloudPlatform/scion — plus four backends (k8sjob, subprocess,
webhook, stub) and an OTel-traced Registry that spans every dispatch
and injects W3C trace context into child processes via env vars.

Phases landed together on this branch:

  1. internal/harness scaffold: Harness/Capabilities/ExecRequest/
     ExecResult/Budget/Usage types, Registry with Resolve/Execute,
     in-memory stub backend.
  2. internal/harness/k8sjob: wraps existing k8s.JobRunner behind the
     Harness interface with a Waiter abstraction (real clientset +
     test fake). BuildHandler exports the per-agent config logic.
  3. internal/harness/subprocess: os/exec-based backend (Mac+Linux),
     per-run workdir, result.json handoff, bounded log capture,
     Budget-driven wall-clock timeout.
  4. internal/harness/webhook: synchronous HTTP POST with HMAC
     signing via internal/webhooks.ComputeHMACSignature, per-agent
     URL/secret/timeout read from harness_config_json.
  5. internal/observability: OTel tracer init via OTLP HTTP (opt-in
     via SYNAPBUS_OTEL_ENABLED), W3C propagator always installed;
     Registry.Execute starts a harness.execute span per dispatch and
     calls InjectTraceContext into req.Env so children inherit it.
  6. internal/harness/runs: SQLite-backed Observer that persists a
     harness_runs row per dispatch with status, usage, cost, duration,
     trace_id, session_id, and a bounded logs excerpt.

Schema: new migration 019_harness.sql adds agents.harness_name /
local_command / harness_config_json columns and the backend-agnostic
harness_runs table with indices on (agent, created_at), (status),
(trace_id), (run_id). internal/reactor/reactor_test.go inline schema
updated to match.

Deployment: deploy/kubic/otel-collector.yaml stands up an otel-collector
Deployment + ConfigMap + ClusterIP Service in the synapbus namespace on
kubic, receiving OTLP gRPC (4317) and HTTP (4318) and exporting debug
output until a Tempo/Jaeger backend lands.

Docs: docs/harness-otel-research.html compares scion and paperclip
side-by-side and maps the current synapbus executor surface; its
companion docs/harness-otel-design.md carries the phase plan, span
taxonomy, and migration schema verbatim.

The reactor currently still calls k8s.JobRunner directly — rewiring it
through the Registry is a follow-up, intentionally out of scope for
this branch to keep the refactor reversible. The new packages are
independently tested (~78 new tests across 7 packages) and the full
project test suite passes.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This commit is contained in:
Algis Dumbris
2026-04-13 20:23:04 +03:00
co-authored by Claude Opus 4.6
parent 0e25fbcccb
commit 574897df4d
26 changed files with 4768 additions and 7 deletions
+14
View File
@@ -44,6 +44,7 @@ import (
reactorpkg "github.com/synapbus/synapbus/internal/reactor"
"github.com/synapbus/synapbus/internal/messaging"
prommetrics "github.com/synapbus/synapbus/internal/metrics"
"github.com/synapbus/synapbus/internal/observability"
"github.com/synapbus/synapbus/internal/reactions"
"github.com/synapbus/synapbus/internal/search"
"github.com/synapbus/synapbus/internal/search/embedding"
@@ -194,6 +195,19 @@ func runServe(cmd *cobra.Command, args []string) error {
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
// Initialise OpenTelemetry tracing (opt-in via SYNAPBUS_OTEL_ENABLED=1).
// Harmless when disabled — installs only the W3C propagator and
// leaves the global tracer provider as the default no-op.
otelShutdown, err := observability.Init(ctx, observability.ConfigFromEnv(os.Getenv), logger)
if err != nil {
return fmt.Errorf("init otel: %w", err)
}
defer func() {
shutdownCtx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
_ = otelShutdown(shutdownCtx)
}()
slog.Info("starting SynapBus",
"host", host,
"port", port,
+139
View File
@@ -0,0 +1,139 @@
# OpenTelemetry Collector for kubic.home.arpa
#
# Installs a single-replica otelcol-contrib in the `synapbus` namespace.
# Accepts OTLP over gRPC (4317) and HTTP (4318) and forwards traces to
# stdout for now; swap in a Tempo / Jaeger exporter once one is up.
#
# Apply with:
# kubectl apply -f deploy/kubic/otel-collector.yaml
#
# SynapBus points at this collector via:
# SYNAPBUS_OTEL_ENABLED=1
# SYNAPBUS_OTEL_ENDPOINT=otel-collector.synapbus.svc.cluster.local:4318
# SYNAPBUS_OTEL_INSECURE=1
---
apiVersion: v1
kind: ConfigMap
metadata:
name: otel-collector-config
namespace: synapbus
data:
config.yaml: |
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
batch:
timeout: 5s
send_batch_size: 512
memory_limiter:
check_interval: 1s
limit_percentage: 80
spike_limit_percentage: 20
exporters:
debug:
verbosity: normal
sampling_initial: 5
sampling_thereafter: 200
# TODO: wire a Tempo / Jaeger / Loki exporter once one is running
# on kubic. Until then, `debug` prints a sampled summary to stdout.
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [debug]
logs:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [debug]
metrics:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [debug]
telemetry:
logs:
level: info
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: otel-collector
namespace: synapbus
labels:
app.kubernetes.io/name: otel-collector
app.kubernetes.io/part-of: synapbus
spec:
replicas: 1
selector:
matchLabels:
app.kubernetes.io/name: otel-collector
template:
metadata:
labels:
app.kubernetes.io/name: otel-collector
spec:
containers:
- name: otelcol
image: otel/opentelemetry-collector-contrib:0.118.0
args: ["--config=/conf/config.yaml"]
ports:
- name: otlp-grpc
containerPort: 4317
- name: otlp-http
containerPort: 4318
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
cpu: 500m
memory: 512Mi
readinessProbe:
tcpSocket:
port: 4317
initialDelaySeconds: 5
periodSeconds: 10
livenessProbe:
tcpSocket:
port: 4317
initialDelaySeconds: 15
periodSeconds: 20
volumeMounts:
- name: config
mountPath: /conf
readOnly: true
volumes:
- name: config
configMap:
name: otel-collector-config
items:
- key: config.yaml
path: config.yaml
---
apiVersion: v1
kind: Service
metadata:
name: otel-collector
namespace: synapbus
labels:
app.kubernetes.io/name: otel-collector
app.kubernetes.io/part-of: synapbus
spec:
type: ClusterIP
selector:
app.kubernetes.io/name: otel-collector
ports:
- name: otlp-grpc
port: 4317
targetPort: 4317
- name: otlp-http
port: 4318
targetPort: 4318
+321
View File
@@ -0,0 +1,321 @@
# Harness-Agnostic Wrappers + OTel — Design Document
**Status:** IMPLEMENTED on branch `feat/harness-otel` (was DRAFT — approved 2026-04-13)
**Date:** 2026-04-13
**Companion report:** [`harness-otel-research.html`](./harness-otel-research.html)
## 1. Motivation
SynapBus today executes reactive agents through two disjoint paths:
- `internal/k8s` + `internal/reactor` — creates a Kubernetes Job per inbound message (primary).
- `internal/webhooks` — outbound HTTP delivery with HMAC signing (secondary).
There is no way to run an external CLI (claude-code, gemini-cli, kimi, codex) as a local subprocess on a Mac or on `kubic` outside of a K8s Job. There is no unified `Runner` / `Harness` interface. OpenTelemetry is listed in `go.mod` but unused. Each new backend would require touching the reactor directly.
This design introduces an `internal/harness/` package that:
1. Defines a minimal `Harness` interface (inspired by `GoogleCloudPlatform/scion`'s `api.Harness`).
2. Wraps the existing K8s path and the existing webhook path as two implementations of that interface.
3. Adds a third implementation: a local subprocess executor.
4. Initialises OpenTelemetry in the main process and wires spans + W3C trace-context propagation through every implementation, using env-var injection as the transport into child processes.
## 2. Goals / Non-goals
**Goals**
- One interface for "dispatch this message to this agent, wherever it runs."
- Pluggable backends: k8s-job, subprocess, webhook, in-process stub (tests).
- Capability flags so the dispatcher can pick the right backend and degrade gracefully.
- Distributed tracing from `mcp.tool.execute` → `reactor.dispatch` → `harness.execute` → child process.
- Cost / token / duration recorded in a new backend-agnostic `harness_runs` table.
- Preflight `TestEnvironment()` per harness, callable from the admin CLI.
**Non-goals**
- No task decomposition, no LLM planner, no judge. Consistent with scion and paperclip.
- No company / org-chart / budget-governance model. Out of scope.
- No plugin loader at runtime; compile-time registry for now.
- No changes to the MCP tool surface exposed to agents. This is all server-side.
## 3. Interface
```go
// internal/harness/harness.go
package harness
type Capabilities struct {
SystemPrompt bool
SessionResume bool
Skills bool
OTelNative bool // child honours OTEL_* env vars
MaxConcurrency int
}
type Budget struct {
MaxWallClock time.Duration
MaxTokensIn int64
MaxTokensOut int64
MaxCostUSD float64
}
type Usage struct {
TokensIn int64
TokensOut int64
TokensCached int64
CostUSD float64
}
type ExecRequest struct {
RunID string // generated by caller; propagated into child
AgentName string
Message *messaging.Message
Context []*messaging.Message // optional conversation window
Budget Budget
Env map[string]string // caller overrides
Skills []string
}
type ExecResult struct {
ExitCode int
Logs string // captured stdout/stderr
ResultJSON json.RawMessage // optional structured output
Usage Usage
TraceID string // W3C, for correlation
Err error
}
type Harness interface {
Name() string
Capabilities() Capabilities
// One-shot pre-flight: is the binary installed, is auth valid,
// can we reach the model? Used by admin CLI and registry resolution.
TestEnvironment(ctx context.Context) error
// One-shot setup for a given agent (write config files, pre-approve
// tool fingerprints, materialise skills). Idempotent.
Provision(ctx context.Context, agent *agents.Agent) error
// Dispatch a single request. Blocks until completion (or Budget exceeded).
Execute(ctx context.Context, req *ExecRequest) (*ExecResult, error)
// Best-effort cancellation of an in-flight run.
Cancel(ctx context.Context, runID string) error
}
```
## 4. Registry + resolution
```go
type Registry struct {
mu sync.RWMutex
byName map[string]Harness
}
func (r *Registry) Register(h Harness) { ... }
// Resolve picks a backend for the given agent. Resolution order:
// 1. agent.HarnessName (explicit)
// 2. agent.K8sImage != "" && k8s runner available → "k8sjob"
// 3. agent has webhooks registered → "webhook"
// 4. agent.LocalCommand != "" → "subprocess"
// 5. ErrNoBackend
func (r *Registry) Resolve(agent *agents.Agent) (Harness, error) { ... }
// Execute is the one entry point the reactor uses. It resolves, starts a
// span, injects trace context into req.Env, calls Execute, records usage,
// and writes a harness_runs row.
func (r *Registry) Execute(ctx context.Context, agent *agents.Agent, req *ExecRequest) (*ExecResult, error) { ... }
```
## 5. Backend implementations
### 5.1 `internal/harness/k8sjob`
- Wraps the existing `internal/k8s.JobRunner` + `internal/reactor` K8s path.
- `Execute` → `CreateJob` → poll `ReactiveRun` → `GetJobLogs` → parse logs for result envelope.
- `Provision` is a no-op (K8s path has nothing to provision).
- `Capabilities{SystemPrompt:false, SessionResume:false, Skills:false, OTelNative:true, MaxConcurrency:10}`.
- Env vars merged into `corev1.EnvVar` slice at `internal/k8s/runner.go:105–119` include the injected `TRACEPARENT` / `OTEL_EXPORTER_OTLP_ENDPOINT`.
### 5.2 `internal/harness/subprocess` (NEW)
- Runs `os/exec` with `cmd.Env = mergedEnv`, `cmd.Dir = workdir`, context timeout from `Budget.MaxWallClock`.
- Captures stdout/stderr into a bounded buffer (`MAX_LOG_BYTES`, e.g. 1 MiB; truncate with excerpt marker beyond).
- Reads a well-known `result.json` file from `workdir` after exit to populate `ExecResult.ResultJSON` (same convention as scion agents writing to workspace).
- Credential injection: `HOME`, `ANTHROPIC_API_KEY` / `GEMINI_API_KEY` from agent config; `~/.claude` readable via host FS.
- Per-agent `workdir` under `${SYNAPBUS_DATA_DIR}/harness/subprocess/${runID}/` — torn down on success, preserved on failure for forensics.
- `Capabilities{SystemPrompt:true, SessionResume:true (Claude Code), Skills:false, OTelNative:true, MaxConcurrency:4}`.
### 5.3 `internal/harness/webhook`
- Wraps the existing `internal/webhooks.DeliveryEngine` as a `Harness`.
- Async: `Execute` enqueues a delivery, polls `webhook_deliveries` for a terminal state, then synthesises an `ExecResult`.
- Useful for agents that want to receive a callback on their own HTTP endpoint instead of running in-process.
### 5.4 `internal/harness/stub` (tests only)
- In-memory; returns a canned `ExecResult`. Used by unit + integration tests so nothing in tests actually shells out or talks to K8s.
## 6. OTel integration
### 6.1 Initialisation
New file `internal/observability/otel.go`:
```go
package observability
func Init(ctx context.Context, cfg Config) (shutdown func(context.Context) error, err error) {
res, _ := resource.New(ctx,
resource.WithAttributes(semconv.ServiceName("synapbus")),
)
exp, err := otlptracegrpc.New(ctx,
otlptracegrpc.WithEndpoint(cfg.Endpoint),
otlptracegrpc.WithInsecure(),
)
if err != nil { return nil, err }
tp := sdktrace.NewTracerProvider(
sdktrace.WithBatcher(exp),
sdktrace.WithResource(res),
)
otel.SetTracerProvider(tp)
otel.SetTextMapPropagator(propagation.TraceContext{})
return tp.Shutdown, nil
}
```
Called from `cmd/synapbus/main.go` immediately after `slog` setup, opt-in via `SYNAPBUS_OTEL_ENABLED=1`.
### 6.2 Span taxonomy
| Span name | Location | Key attributes |
|--------------------------------|--------------------------------|----------------|
| `mcp.tool.execute` | MCP handler entry | `mcp.tool`, `agent.name`, `message.id` |
| `reactor.dispatch` | `reactor.Dispatch()` | `agent.name`, `trigger.depth`, `budget.remaining` |
| `harness.resolve` | `Registry.Resolve` | `harness.name`, `fallback.chain` |
| `harness.provision` | `Harness.Provision` | `harness.name`, `agent.home` |
| `harness.execute` | `Harness.Execute` | `harness.name`, `run.id`, `usage.*`, `cost.usd`, `exit.code` |
| `harness.k8s.job.create` | k8sjob backend | `k8s.job.name`, `k8s.namespace`, `k8s.image` |
| `harness.subprocess.exec` | subprocess backend | `proc.argv[0]`, `proc.pid`, `proc.workdir` |
| `harness.webhook.deliver` | webhook backend | `http.url`, `http.status_code`, `retry.count` |
### 6.3 Context propagation into children
```go
func injectTraceEnv(ctx context.Context, dst map[string]string, runID, agentName string, cfg Config) {
carrier := propagation.MapCarrier{}
otel.GetTextMapPropagator().Inject(ctx, carrier)
// OTel convention: env vars TRACEPARENT, TRACESTATE
for k, v := range carrier {
dst[strings.ToUpper(k)] = v
}
dst["OTEL_EXPORTER_OTLP_ENDPOINT"] = cfg.ChildEndpoint
dst["OTEL_EXPORTER_OTLP_PROTOCOL"] = "grpc"
dst["OTEL_SERVICE_NAME"] = "synapbus-agent-" + agentName
dst["OTEL_RESOURCE_ATTRIBUTES"] = fmt.Sprintf("synapbus.run_id=%s,synapbus.agent=%s", runID, agentName)
}
```
- **K8s backend**: merged into the `corev1.EnvVar` slice built at `internal/k8s/runner.go:105–119`.
- **Subprocess backend**: merged into `cmd.Env`.
- **Webhook backend**: set as HTTP headers (`traceparent`, `tracestate`) alongside existing `X-SynapBus-*` headers.
### 6.4 Metrics
Keep the existing Prometheus registry (`internal/metrics/metrics.go`). Also emit a minimal OTel meter set via the same OTLP exporter:
- `synapbus.harness.runs` (counter, labels: `harness`, `status`)
- `synapbus.harness.duration_ms` (histogram)
- `synapbus.harness.tokens_in` / `tokens_out` (counters)
- `synapbus.harness.cost_usd` (counter)
### 6.5 Config
New env vars on `cmd/synapbus/main.go`:
| Var | Default | Description |
|---|---|---|
| `SYNAPBUS_OTEL_ENABLED` | `false` | Opt-in master switch |
| `SYNAPBUS_OTEL_ENDPOINT` | `localhost:4317` | OTLP gRPC target |
| `SYNAPBUS_OTEL_INSECURE` | `true` | TLS off for LAN |
| `SYNAPBUS_OTEL_SERVICE_NAME` | `synapbus` | Override for multi-instance setups |
## 7. Data model
### 7.1 Migration `019_harness.sql`
```sql
ALTER TABLE agents ADD COLUMN harness_name TEXT;
ALTER TABLE agents ADD COLUMN local_command TEXT; -- subprocess argv (JSON)
ALTER TABLE agents ADD COLUMN harness_config_json TEXT; -- per-harness config blob
CREATE TABLE harness_runs (
id INTEGER PRIMARY KEY AUTOINCREMENT,
run_id TEXT NOT NULL UNIQUE, -- UUID, propagated into child
agent_name TEXT NOT NULL,
backend TEXT NOT NULL, -- 'k8sjob' | 'subprocess' | 'webhook' | 'stub'
message_id INTEGER, -- triggering message, if any
status TEXT NOT NULL, -- 'pending' | 'running' | 'success' | 'failed' | 'cancelled' | 'timeout'
exit_code INTEGER,
trace_id TEXT,
span_id TEXT,
tokens_in INTEGER DEFAULT 0,
tokens_out INTEGER DEFAULT 0,
tokens_cached INTEGER DEFAULT 0,
cost_usd REAL DEFAULT 0,
duration_ms INTEGER,
result_json TEXT,
logs_excerpt TEXT, -- bounded, full logs on disk
created_at INTEGER NOT NULL,
finished_at INTEGER,
FOREIGN KEY (message_id) REFERENCES messages(id)
);
CREATE INDEX idx_harness_runs_agent ON harness_runs(agent_name, created_at DESC);
CREATE INDEX idx_harness_runs_status ON harness_runs(status, created_at DESC);
CREATE INDEX idx_harness_runs_trace ON harness_runs(trace_id);
```
### 7.2 Relationship to `ReactiveRun`
Phase 2 keeps both tables. A follow-up (separate PR) folds `ReactiveRun` into `harness_runs` and drops the old table. This avoids a big-bang migration.
## 8. Staged implementation plan
| Phase | Scope | Reversible? |
|---|---|---|
| **0** | This design doc + research HTML report | yes — text only |
| **1** | Scaffold `internal/harness/` — interface, registry, stub backend, unit tests. No callers wired. | yes — dead code until Phase 2 |
| **2** | Refactor existing K8s path behind `k8sjob.Harness`. Reactor calls `Registry.Execute`. Behaviour unchanged. Existing tests green. | yes — one commit revert |
| **3** | New `subprocess` backend + migration `019_harness.sql` + per-agent `local_command`. | yes |
| **4** | Wrap webhook path as `webhook.Harness`. Route via registry. | yes |
| **5** | `internal/observability/otel.go` + span wiring + env-var propagation. Opt-in. | yes — feature-flagged |
| **6** | Session codec + cost accounting surfaced in `harness_runs`; `TestEnvironment` preflight on admin CLI. | yes |
Each phase is a separate PR. Nothing is merged until the previous phase's tests are green.
## 9. Testing strategy
- **Unit**: every interface method on every backend, using the `stub` harness where possible.
- **Integration**: one-shot reactor dispatch end-to-end with the `stub` backend; asserts that spans are created, `harness_runs` row is written, trace id propagates.
- **K8s**: existing K8s-gated tests continue to run against a real kubeconfig when available (`SYNAPBUS_TEST_K8S=1`).
- **Subprocess**: run against a tiny golden binary (`testdata/echo-agent.sh`) that reads env, writes `result.json`, exits 0.
- **OTel**: in-memory span exporter asserted via `go.opentelemetry.io/otel/sdk/trace/tracetest`.
## 10. Open questions (for approval)
1. **Collector.** Stand up a collector on `kubic` first, or ship with stdout exporter as a no-op until a collector exists?
2. **Subprocess path on Mac.** Is laptop-local execution in-scope for Phase 3 or defer?
3. **Session codec.** Just a session-id pass-through, or full replay of conversation history?
4. **Runtime plugin loader.** Compile-time registry only, or add `hashicorp/go-plugin` later?
5. **Feature flag.** Global `SYNAPBUS_HARNESS_V2=1` to gate the whole thing until Phase 6, or trust the phase-by-phase PRs?
## 11. References
- `GoogleCloudPlatform/scion` — `pkg/api/harness.go:22–68`, `pkg/harness/claude_code.go:311–320`, `pkg/util/logging/otel_provider.go:26–61`.
- `paperclipai/paperclip` — `packages/adapter-utils/src/types.ts:292–331`, `server/src/adapters/registry.ts:89–222`, `server/src/services/heartbeat.ts:331–346`.
- SynapBus current surface — `internal/k8s/runner.go:96–183`, `internal/reactor/reactor.go:51`, `internal/webhooks/delivery.go:157`, `internal/mcp/tools_hybrid.go:489`, `internal/trace/tracer.go`, `go.mod:102–114` (OTel deps present but unused).
- Companion research HTML — [`harness-otel-research.html`](./harness-otel-research.html).
+601
View File
@@ -0,0 +1,601 @@
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8" />
<meta name="viewport" content="width=device-width,initial-scale=1" />
<title>Harness-Agnostic Wrappers &amp; OTel — Research Report</title>
<style>
:root{
--bg:#0b0d12; --bg2:#11141b; --panel:#151923; --panel2:#1b2030;
--ink:#e6e9ef; --mute:#8a93a6; --line:#262c3a;
--accent:#7aa2ff; --accent2:#b892ff; --ok:#51d88a; --warn:#ffb454; --bad:#ff6b6b;
--mono:ui-monospace,SFMono-Regular,Menlo,Consolas,monospace;
--sans:-apple-system,BlinkMacSystemFont,"Inter","Helvetica Neue",Arial,sans-serif;
}
*{box-sizing:border-box}
html,body{background:var(--bg);color:var(--ink);font-family:var(--sans);margin:0;line-height:1.55}
a{color:var(--accent);text-decoration:none;border-bottom:1px dashed #3a4566}
a:hover{color:var(--accent2)}
.wrap{max-width:1180px;margin:0 auto;padding:48px 32px 120px}
header.hero{
padding:56px 40px;border-radius:20px;
background:
radial-gradient(1200px 400px at 10% 0%, rgba(122,162,255,.18), transparent 60%),
radial-gradient(900px 400px at 100% 100%, rgba(184,146,255,.18), transparent 60%),
linear-gradient(180deg, #0f1320, #0b0d12);
border:1px solid var(--line);
margin-bottom:40px;
}
.kicker{letter-spacing:.25em;text-transform:uppercase;font-size:12px;color:var(--mute)}
h1{font-size:44px;line-height:1.1;margin:8px 0 16px;letter-spacing:-.02em}
h1 span{background:linear-gradient(90deg,#7aa2ff,#b892ff);-webkit-background-clip:text;background-clip:text;color:transparent}
header .lede{font-size:18px;color:#c9d0df;max-width:840px}
header .meta{margin-top:24px;display:flex;gap:16px;flex-wrap:wrap;color:var(--mute);font-size:13px;font-family:var(--mono)}
header .meta b{color:#c9d0df;font-weight:500}
h2{font-size:26px;margin:56px 0 16px;letter-spacing:-.01em;display:flex;align-items:center;gap:12px}
h2::before{content:"";display:inline-block;width:6px;height:22px;background:linear-gradient(180deg,#7aa2ff,#b892ff);border-radius:3px}
h3{font-size:18px;margin:28px 0 10px;color:#d8dfef}
p{margin:10px 0;color:#c3cad9}
ul{color:#c3cad9}
code{font-family:var(--mono);font-size:13px;background:#1a1f2b;border:1px solid var(--line);padding:1px 6px;border-radius:4px;color:#e6e9ef}
pre{
font-family:var(--mono);font-size:12.5px;background:#0f1320;border:1px solid var(--line);
padding:16px 18px;border-radius:10px;overflow:auto;line-height:1.55;
}
pre .k{color:#b892ff}
pre .s{color:#51d88a}
pre .c{color:#6a7285;font-style:italic}
pre .n{color:#ffb454}
pre .t{color:#7aa2ff}
.grid2{display:grid;grid-template-columns:1fr 1fr;gap:20px}
.grid3{display:grid;grid-template-columns:repeat(3,1fr);gap:16px}
@media (max-width:900px){.grid2,.grid3{grid-template-columns:1fr}}
.card{background:var(--panel);border:1px solid var(--line);border-radius:14px;padding:22px 24px}
.card h3{margin-top:0}
.card.accent{border-color:#2f3a5e;background:linear-gradient(180deg,#141a2d,#10131d)}
.pill{display:inline-block;font-family:var(--mono);font-size:11px;padding:3px 10px;border-radius:999px;border:1px solid var(--line);color:var(--mute);margin-right:6px}
.pill.ok{color:var(--ok);border-color:#1f5a3c}
.pill.warn{color:var(--warn);border-color:#6b4a1a}
.pill.bad{color:var(--bad);border-color:#6b2828}
.pill.info{color:var(--accent);border-color:#2a3a66}
table{width:100%;border-collapse:collapse;margin:14px 0;font-size:14px}
th,td{text-align:left;padding:12px 14px;border-bottom:1px solid var(--line);vertical-align:top}
th{color:#aab3c7;font-weight:500;font-size:12px;letter-spacing:.08em;text-transform:uppercase;background:#121622}
tr:last-child td{border-bottom:none}
td code{font-size:12px}
.tl{position:relative;padding-left:24px;margin:16px 0}
.tl::before{content:"";position:absolute;left:6px;top:4px;bottom:4px;width:2px;background:var(--line)}
.tl .step{position:relative;margin:12px 0;padding-left:4px}
.tl .step::before{content:"";position:absolute;left:-22px;top:6px;width:10px;height:10px;border-radius:50%;background:#7aa2ff;box-shadow:0 0 0 4px rgba(122,162,255,.15)}
.cite{font-family:var(--mono);font-size:11.5px;color:var(--mute)}
.cite a{color:#aab3c7;border-bottom-color:#3a4566}
.callout{border-left:3px solid var(--accent);background:#121728;padding:14px 18px;margin:18px 0;border-radius:0 10px 10px 0}
.callout.warn{border-left-color:var(--warn);background:#1e1a12}
.callout.bad{border-left-color:var(--bad);background:#1d1313}
.callout.ok{border-left-color:var(--ok);background:#10201a}
.diagram{background:#0f1320;border:1px solid var(--line);border-radius:12px;padding:24px;margin:18px 0;overflow:auto}
.arch{display:flex;align-items:stretch;gap:0;font-family:var(--mono);font-size:12px}
.arch .col{flex:1;min-width:0;padding:0 8px}
.arch .layer{background:#1a2033;border:1px solid #2a3a66;border-radius:8px;padding:12px;margin:6px 0;text-align:center;color:#cfd7ea}
.arch .layer.mute{background:#141828;border-color:var(--line);color:var(--mute)}
.arch .layer.hi{background:linear-gradient(180deg,#1f2a4d,#151a2e);border-color:#3a4a7a;color:#eaf0ff}
.arch h4{margin:0 0 8px;text-align:center;color:var(--mute);font-size:11px;letter-spacing:.15em;text-transform:uppercase;font-family:var(--sans);font-weight:500}
.toc{background:var(--panel2);border:1px solid var(--line);border-radius:12px;padding:18px 22px;margin-bottom:32px;font-size:14px}
.toc b{color:#aab3c7;font-size:11px;letter-spacing:.15em;text-transform:uppercase}
.toc ol{margin:8px 0 0;padding-left:20px;color:var(--mute)}
.toc ol a{color:#c3cad9;border:none}
.toc ol a:hover{color:var(--accent)}
footer{margin-top:60px;padding-top:24px;border-top:1px solid var(--line);color:var(--mute);font-size:13px;font-family:var(--mono)}
</style>
</head>
<body>
<div class="wrap">
<header class="hero">
<div class="kicker">Research Report &bull; 2026-04-13</div>
<h1>Harness-Agnostic Wrappers &amp;<br/><span>OpenTelemetry for SynapBus</span></h1>
<p class="lede">Borrow what works from <code>GoogleCloudPlatform/scion</code> and <code>paperclipai/paperclip</code>, skip what doesn't, and sketch a minimal harness + OTel integration that fits SynapBus's Go / MCP / SQLite spine.</p>
<div class="meta">
<span><b>Scope</b> research + design (no code yet)</span>
<span><b>Status</b> awaiting approval</span>
<span><b>Targets</b> scion / paperclip / synapbus</span>
</div>
</header>
<div class="toc">
<b>Contents</b>
<ol>
<li><a href="#tldr">TL;DR &mdash; recommendation</a></li>
<li><a href="#scion">What is <em>scion</em> actually doing?</a></li>
<li><a href="#paperclip">What is <em>paperclip</em> actually doing?</a></li>
<li><a href="#compare">Side-by-side comparison</a></li>
<li><a href="#synapbus">SynapBus &mdash; current execution surface</a></li>
<li><a href="#design">Proposed design for SynapBus</a></li>
<li><a href="#otel">OTel integration points</a></li>
<li><a href="#nuggets">Other reusable nuggets</a></li>
<li><a href="#nextsteps">Next steps &amp; open questions</a></li>
</ol>
</div>
<section id="tldr">
<h2>TL;DR</h2>
<div class="card accent">
<p><b>Both repos converge on the same core idea:</b> a narrow <em>Harness</em> / <em>Adapter</em> interface that abstracts "some external AI CLI" behind a single <code>execute(ctx)&rarr;result</code> contract, then registers concrete implementations for Claude Code, Gemini CLI, Codex, OpenCode, etc.</p>
<p><b>Scion's design is the better template for SynapBus:</b> it's Go, it ships OTel via env-var injection into child processes, and its <code>Harness</code> interface cleanly separates <em>provisioning</em> from <em>invocation</em> &mdash; exactly the seam we're missing.</p>
<p><b>Paperclip contributes two ideas we should adopt</b>: (a) an adapter registry with capability flags so a router can pick the best backend at dispatch time, and (b) a session codec per adapter so long-running agents can be resumed.</p>
<p><b>SynapBus today has no subprocess executor, no unified runner interface, and no OTel spans &mdash;</b> only a K8s-Job path and an HTTP-webhook path living as two disjoint code paths. A small <code>internal/harness/</code> package would unify both and unlock local-subprocess execution.</p>
</div>
</section>
<section id="scion">
<h2>1 &middot; What scion actually does</h2>
<p>Despite the name collision with the SCION internet-architecture project, <code>GoogleCloudPlatform/scion</code> is a <b>multi-agent orchestration harness</b> for evaluating and running "deep agents" (Claude Code, Gemini CLI, Codex, OpenCode) inside isolated containers. It is explicitly <em>not</em> a planner and <em>not</em> a verifier &mdash; it is the control plane and observability spine around arbitrary agent CLIs.</p>
<h3>The Harness interface &mdash; the centrepiece</h3>
<p class="cite">pkg/api/harness.go:22&ndash;68</p>
<pre><span class="k">type</span> <span class="t">Harness</span> <span class="k">interface</span> {
Name() <span class="k">string</span>
AdvancedCapabilities() HarnessAdvancedCapabilities
GetEnv(agentName, agentHome, unixUsername <span class="k">string</span>) <span class="k">map</span>[<span class="k">string</span>]<span class="k">string</span>
GetCommand(task <span class="k">string</span>, resume <span class="k">bool</span>, baseArgs []<span class="k">string</span>) []<span class="k">string</span>
DefaultConfigDir() <span class="k">string</span>
SkillsDir() <span class="k">string</span>
HasSystemPrompt(agentHome <span class="k">string</span>) <span class="k">bool</span>
Provision(ctx context.Context, agentName, agentDir, agentHome, agentWorkspace <span class="k">string</span>) <span class="k">error</span>
GetEmbedDir() <span class="k">string</span>
GetInterruptKey() <span class="k">string</span>
GetHarnessEmbedsFS() (embed.FS, <span class="k">string</span>)
InjectAgentInstructions(agentHome <span class="k">string</span>, content []<span class="k">byte</span>) <span class="k">error</span>
InjectSystemPrompt(agentHome <span class="k">string</span>, content []<span class="k">byte</span>) <span class="k">error</span>
<span class="c">// the key OTel seam &mdash; returns env vars that the container runtime</span>
<span class="c">// will merge into the child process env before exec</span>
GetTelemetryEnv() <span class="k">map</span>[<span class="k">string</span>]<span class="k">string</span>
ResolveAuth(auth AuthConfig) (*ResolvedAuth, <span class="k">error</span>)
}</pre>
<p>Three things to notice:</p>
<ul>
<li><b><code>Provision</code></b> is separate from <code>GetCommand</code>: one-shot setup (write <code>.claude.json</code>, pre-approve tool fingerprints, materialise skill files) versus per-invocation command building.</li>
<li><b><code>GetEnv</code> / <code>GetTelemetryEnv</code> / <code>ResolveAuth</code></b> all return <em>maps of env vars</em>. The container runtime layer merges them. This means every harness is credential-injection-agnostic and telemetry-injection-agnostic &mdash; you can point a whole pod at a different OTel collector by changing one map.</li>
<li><b><code>AdvancedCapabilities()</code></b> lets a dispatcher ask "does this harness support system prompts?" and <em>degrade gracefully</em> (fall back to <code>InjectAgentInstructions</code>) when it doesn't.</li>
</ul>
<h3>The factory</h3>
<p class="cite">pkg/harness/harness.go:37&ndash;57</p>
<pre><span class="k">func</span> <span class="t">New</span>(name <span class="k">string</span>) <span class="t">Harness</span> {
<span class="k">switch</span> name {
<span class="k">case</span> <span class="s">"claude"</span>: <span class="k">return</span> &amp;ClaudeCode{}
<span class="k">case</span> <span class="s">"gemini"</span>: <span class="k">return</span> &amp;GeminiCLI{}
<span class="k">case</span> <span class="s">"opencode"</span>: <span class="k">return</span> &amp;OpenCode{}
<span class="k">case</span> <span class="s">"codex"</span>: <span class="k">return</span> &amp;Codex{}
}
<span class="k">if</span> h := pluginMgr.Lookup(name); h != <span class="k">nil</span> { <span class="k">return</span> h }
<span class="k">return</span> &amp;Generic{} <span class="c">// universal fallback</span>
}</pre>
<h3>OTel injection pattern</h3>
<p class="cite">pkg/harness/claude_code.go:311&ndash;320</p>
<pre><span class="k">func</span> (c *<span class="t">ClaudeCode</span>) <span class="t">GetTelemetryEnv</span>() <span class="k">map</span>[<span class="k">string</span>]<span class="k">string</span> {
<span class="k">return</span> <span class="k">map</span>[<span class="k">string</span>]<span class="k">string</span>{
<span class="s">"CLAUDE_CODE_ENABLE_TELEMETRY"</span>: <span class="s">"1"</span>,
<span class="s">"OTEL_METRICS_EXPORTER"</span>: <span class="s">"otlp"</span>,
<span class="s">"OTEL_LOGS_EXPORTER"</span>: <span class="s">"otlp"</span>,
<span class="s">"OTEL_EXPORTER_OTLP_PROTOCOL"</span>: <span class="s">"grpc"</span>,
<span class="s">"OTEL_EXPORTER_OTLP_ENDPOINT"</span>: <span class="s">"http://localhost:4317"</span>,
<span class="s">"OTEL_METRIC_EXPORT_INTERVAL"</span>: <span class="s">"30000"</span>,
}
}</pre>
<p>Scion's own Go code emits <b>OTel logs</b> via the OTLP log exporter (<code>pkg/util/logging/otel_provider.go:26&ndash;61</code>) and bridges <code>slog</code> into it (<code>pkg/util/logging/otel.go:85&ndash;119</code>). W3C <code>traceparent</code> headers are extracted at HTTP ingress (<code>pkg/util/logging/trace.go</code>) so trace context can flow across the dispatcher &rarr; runtime &rarr; container boundary.</p>
<h3>Coordination &amp; decomposition</h3>
<p>Scion does <b>not</b> decompose tasks. A single <code>task</code> string goes to the agent and the agent's own model decides how to break it up. Coordination between agents happens via a structured <code>StructuredMessage</code> envelope (<code>pkg/messages/types.go:46&ndash;61</code>) with fields <code>{sender, recipient, msg, type, urgent, broadcasted, attachments}</code> &mdash; an on-disk analogue of a SynapBus channel post.</p>
<div class="callout">
<b>Reusable for SynapBus:</b> the <code>Harness</code> interface shape, the env-var-injection model for both auth &amp; telemetry, the capability-flags degradation pattern, and the <code>Provision</code>/<code>GetCommand</code> split. Ignore the container runtime abstraction &mdash; SynapBus already has K8s-Job + webhook paths and doesn't need a second one.
</div>
</section>
<section id="paperclip">
<h2>2 &middot; What paperclip actually does</h2>
<p>Paperclip is a Node/Express control plane for running 10&ndash;20 agent "companies" with org charts, budgets, and approval gates. Wildly different product &mdash; but it has a clean adapter interface worth borrowing.</p>
<h3>The ServerAdapterModule interface</h3>
<p class="cite">packages/adapter-utils/src/types.ts:292&ndash;331</p>
<pre><span class="k">export interface</span> <span class="t">ServerAdapterModule</span> {
type: <span class="k">string</span>;
execute(ctx: AdapterExecutionContext): <span class="t">Promise</span>&lt;AdapterExecutionResult&gt;;
testEnvironment(ctx: AdapterEnvironmentTestContext): <span class="t">Promise</span>&lt;AdapterEnvironmentTestResult&gt;;
listSkills?: (ctx) =&gt; <span class="t">Promise</span>&lt;AdapterSkillSnapshot&gt;;
syncSkills?: (ctx, desired: <span class="k">string</span>[]) =&gt; <span class="t">Promise</span>&lt;AdapterSkillSnapshot&gt;;
sessionCodec?: AdapterSessionCodec; <span class="c">// resume / serialize sessions</span>
models?: AdapterModel[];
listModels?: () =&gt; <span class="t">Promise</span>&lt;AdapterModel[]&gt;;
agentConfigurationDoc?: <span class="k">string</span>;
onHireApproved?: (payload, cfg) =&gt; <span class="t">Promise</span>&lt;HireApprovedHookResult&gt;;
getQuotaWindows?: () =&gt; <span class="t">Promise</span>&lt;ProviderQuotaResult&gt;;
}</pre>
<p class="cite">AdapterExecutionResult &mdash; types.ts:64&ndash;95</p>
<pre>{ exitCode, signal, timedOut, errorMessage, errorCode,
usage: { inputTokens, outputTokens, cachedInputTokens },
resultJson, costUsd,
question?: { prompt, choices } <span class="c">// can pause for human approval</span>
}</pre>
<p>Ten adapters are registered via a mutable map in <code>server/src/adapters/registry.ts:89&ndash;222</code>: <code>claude-local, codex-local, cursor, gemini, opencode, pi, openclaw, hermes, http, process</code>. External adapters are loaded from plugins asynchronously (lines 244&ndash;270).</p>
<h3>Coordination model &mdash; heartbeat + atomic checkout</h3>
<p class="cite">server/src/services/heartbeat.ts</p>
<p>No DAG, no queue, no planner. Agents wake on a heartbeat (schedule or event), atomically claim assigned issues via a per-agent start lock (<code>withAgentStartLock()</code>, lines 331&ndash;346), run once, and go back to sleep. Concurrency is per-agent (default 1, configurable to 10). Task decomposition is entirely delegated to the agent's own model.</p>
<h3>Verification</h3>
<p>None that's interesting. Exit code 0 = success; timeouts and process-loss retries are tracked; there is no LLM judge, no schema validation, no test runner. Verification is whatever the running agent chooses to self-report in <code>resultJson</code>.</p>
<h3>Observability</h3>
<p>Pino structured logging (<code>server/src/middleware/logger.ts:29&ndash;45</code>) + a custom telemetry client (<code>server/src/telemetry.ts:12&ndash;26</code>) that batch-flushes events every 60s. <b>No OpenTelemetry</b>. This is the weakest part relative to scion.</p>
<div class="callout warn">
<b>Skip for SynapBus:</b> the whole company/org-chart/budget/approval-gate model, the Drizzle ORM, the plugin loader, the issue-tracker schema. They're all Node-centric and solve a problem SynapBus doesn't have.
</div>
<div class="callout ok">
<b>Borrow from paperclip:</b> (1) the <code>sessionCodec</code> idea &mdash; each harness knows how to serialise/resume its own session, so SynapBus can carry conversation state across reactive runs; (2) <code>testEnvironment()</code> as a preflight &mdash; "is the CLI installed, is auth valid, can it reach the model?"; (3) <code>getQuotaWindows()</code> / cost tracking in the result envelope.
</div>
</section>
<section id="compare">
<h2>3 &middot; Side-by-side comparison</h2>
<table>
<thead><tr><th>Aspect</th><th>scion (Go)</th><th>paperclip (Node)</th><th>synapbus today</th></tr></thead>
<tbody>
<tr>
<td>Core interface</td>
<td><code>api.Harness</code> &mdash; 15 methods, env-var-centric</td>
<td><code>ServerAdapterModule</code> &mdash; <code>execute()</code> + optional hooks</td>
<td><code>k8s.JobRunner</code> (K8s only) + <code>webhooks.EventDispatcher</code> &mdash; no unification</td>
</tr>
<tr>
<td>Backends shipped</td>
<td>claude, gemini, codex, opencode, generic fallback</td>
<td>claude, codex, cursor, gemini, opencode, pi, openclaw, hermes, http, process</td>
<td>K8s Job (one) + outbound HTTP webhook</td>
</tr>
<tr>
<td>Credential injection</td>
<td>env vars from <code>GetEnv()</code>+<code>ResolveAuth()</code>; HostPath for <code>~/.claude</code></td>
<td>per-adapter config objects; provider SDK auth</td>
<td>K8s env vars from agent's <code>k8s_env_json</code>; HostPath <code>~/.claude</code> (reactor.go:281&ndash;286)</td>
</tr>
<tr>
<td>Task decomposition</td>
<td>None &mdash; passes whole task string to agent</td>
<td>None &mdash; agents pull from issue queue themselves</td>
<td>None &mdash; reactive trigger wraps one inbound message</td>
</tr>
<tr>
<td>Verification</td>
<td>Workspace sync + agent logs; no judge</td>
<td>Exit code, token usage, timeout; no judge</td>
<td>K8s Job success/fail + pod logs stored in <code>ReactiveRun</code></td>
</tr>
<tr>
<td>Observability</td>
<td><b>OTel logs via OTLP gRPC</b>, W3C trace-context propagation, <code>slog</code> bridge</td>
<td>Pino structured logs + custom telemetry client</td>
<td><code>slog</code> JSON only; Prometheus metrics for reactor; OTel deps present but <b>unused in Go code</b></td>
</tr>
<tr>
<td>Coordination</td>
<td>Containers per agent; inter-agent messages via typed envelope</td>
<td>Heartbeat + atomic per-agent lock; org-chart hierarchy</td>
<td>MCP channels &amp; DMs; reactive triggers fire on inbound</td>
</tr>
<tr>
<td>Capability flags</td>
<td><code>AdvancedCapabilities()</code> for graceful degradation</td>
<td>Optional methods on the interface</td>
<td>None &mdash; hardcoded paths</td>
</tr>
<tr>
<td>Session resume</td>
<td>Yes &mdash; <code>GetCommand(task, resume bool, ...)</code></td>
<td>Yes &mdash; per-adapter <code>sessionCodec</code></td>
<td>None &mdash; each reactive run is fresh</td>
</tr>
</tbody>
</table>
</section>
<section id="synapbus">
<h2>4 &middot; SynapBus current execution surface</h2>
<div class="grid2">
<div class="card">
<h3>Path A &mdash; Reactive K8s Job <span class="pill info">primary</span></h3>
<div class="tl">
<div class="step"><b>Reactor</b> filters inbound messages for agents with <code>TriggerMode=reactive</code> <span class="cite">reactor.go:51</span></div>
<div class="step"><b>Preconditions</b> &mdash; image configured, budget, cooldown, depth</div>
<div class="step"><b>JobRunner.CreateJob</b> builds a K8s <code>batchv1.Job</code> with env vars <code>SYNAPBUS_MESSAGE_ID</code>/<code>_BODY</code>/<code>_FROM_AGENT</code>/<code>_EVENT</code>/<code>_CHANNEL</code> <span class="cite">k8s/runner.go:96&ndash;183</span></div>
<div class="step"><b>Poller</b> goroutine watches Job status, stores result in <code>ReactiveRun</code> <span class="cite">reactor/poller.go</span></div>
<div class="step"><b>GetJobLogs</b> pulls pod logs on completion <span class="cite">k8s/runner.go:185</span></div>
</div>
</div>
<div class="card">
<h3>Path B &mdash; Webhook delivery <span class="pill info">secondary</span></h3>
<div class="tl">
<div class="step"><b>DeliveryEngine.Dispatch</b> matches webhooks for event+agent <span class="cite">webhooks/delivery.go:157</span></div>
<div class="step"><b>HTTP POST</b> with <code>X-SynapBus-Signature</code> HMAC, <code>X-SynapBus-Depth</code> <span class="cite">delivery.go:290&ndash;302</span></div>
<div class="step"><b>Retry</b> 1s / 5s / 30s, dead-letter after 3 attempts</div>
</div>
</div>
</div>
<div class="card" style="margin-top:20px">
<h3>Gaps</h3>
<p>These paths are <b>two disjoint islands</b>. There is:</p>
<ul>
<li><span class="pill bad">missing</span> a local subprocess executor (no way to run a CLI when not in K8s)</li>
<li><span class="pill bad">missing</span> a unified <code>Runner</code>/<code>Harness</code> interface &mdash; the reactor switches on K8s availability with a <code>NoopRunner</code> fallback</li>
<li><span class="pill bad">missing</span> any OTel span around agent invocations &mdash; OTel deps exist in <code>go.mod</code> but are unimported</li>
<li><span class="pill bad">missing</span> capability flags per backend (system-prompt support, session resume, skills)</li>
<li><span class="pill warn">partial</span> credential injection &mdash; K8s path uses HostPath <code>~/.claude</code> + env vars; webhook path has none</li>
<li><span class="pill warn">partial</span> cost/token tracking &mdash; <code>benchmark/sdk_backend.py</code> returns it but core Go reactor does not</li>
</ul>
<p>The recent <code>benchmark/sdk_backend.py</code> (commit <code>0e25fbc</code>) is a Python two-backend fallback (anthropic SDK &rarr; claude-agent-sdk) that foreshadows exactly the abstraction we need &mdash; but in the benchmark tree, not in core.</p>
</div>
</section>
<section id="design">
<h2>5 &middot; Proposed design for SynapBus</h2>
<h3>New package: <code>internal/harness/</code></h3>
<div class="diagram">
<div class="arch">
<div class="col">
<h4>Caller</h4>
<div class="layer mute">MCP handler</div>
<div class="layer hi">Reactor</div>
<div class="layer mute">Webhook engine</div>
<div class="layer mute">Benchmark harness</div>
</div>
<div class="col" style="flex:0 0 40px;display:flex;align-items:center;justify-content:center;color:var(--mute)">&rarr;</div>
<div class="col">
<h4>internal/harness</h4>
<div class="layer hi">Registry</div>
<div class="layer hi">Harness interface</div>
<div class="layer">Capability flags</div>
<div class="layer">OTel spans + env injection</div>
</div>
<div class="col" style="flex:0 0 40px;display:flex;align-items:center;justify-content:center;color:var(--mute)">&rarr;</div>
<div class="col">
<h4>Backends</h4>
<div class="layer">k8s-job (existing)</div>
<div class="layer">subprocess (new)</div>
<div class="layer">webhook (existing, wrapped)</div>
<div class="layer mute">in-process stub</div>
</div>
</div>
</div>
<h3>Interface sketch</h3>
<pre><span class="k">package</span> harness
<span class="k">type</span> <span class="t">Capabilities</span> <span class="k">struct</span> {
SystemPrompt <span class="k">bool</span>
SessionResume <span class="k">bool</span>
Skills <span class="k">bool</span>
OTelNative <span class="k">bool</span> <span class="c">// child process honours OTEL_* env vars</span>
MaxConcurrency <span class="k">int</span>
}
<span class="k">type</span> <span class="t">ExecRequest</span> <span class="k">struct</span> {
AgentName <span class="k">string</span>
Message *messaging.Message <span class="c">// triggering message</span>
Context []*messaging.Message <span class="c">// optional conversation window</span>
Budget Budget <span class="c">// tokens, cost, wallclock</span>
Env <span class="k">map</span>[<span class="k">string</span>]<span class="k">string</span> <span class="c">// caller-provided overrides</span>
Skills []<span class="k">string</span>
}
<span class="k">type</span> <span class="t">ExecResult</span> <span class="k">struct</span> {
ExitCode <span class="k">int</span>
Logs <span class="k">string</span>
ResultJSON json.RawMessage
Usage Usage <span class="c">// { in, out, cached tokens, cost }</span>
TraceID <span class="k">string</span> <span class="c">// W3C, for correlation</span>
Err <span class="k">error</span>
}
<span class="k">type</span> <span class="t">Harness</span> <span class="k">interface</span> {
Name() <span class="k">string</span>
Capabilities() Capabilities
TestEnvironment(ctx context.Context) <span class="k">error</span> <span class="c">// preflight</span>
Provision(ctx context.Context, agent *agents.Agent) <span class="k">error</span> <span class="c">// one-shot setup</span>
Execute(ctx context.Context, req *ExecRequest) (*ExecResult, <span class="k">error</span>)
Cancel(ctx context.Context, runID <span class="k">string</span>) <span class="k">error</span>
}
<span class="k">type</span> <span class="t">Registry</span> <span class="k">struct</span> { <span class="c">/* map[string]Harness + mutex */</span> }
<span class="k">func</span> (r *<span class="t">Registry</span>) <span class="t">Register</span>(h Harness)
<span class="k">func</span> (r *<span class="t">Registry</span>) <span class="t">Resolve</span>(agent *agents.Agent) (Harness, <span class="k">error</span>)
<span class="k">func</span> (r *<span class="t">Registry</span>) <span class="t">Execute</span>(ctx context.Context, req *ExecRequest) (*ExecResult, <span class="k">error</span>)</pre>
<h3>Backend implementations</h3>
<table>
<thead><tr><th>Package</th><th>Wraps</th><th>Status</th></tr></thead>
<tbody>
<tr><td><code>internal/harness/k8sjob</code></td><td>existing <code>internal/k8s</code> path</td><td>refactor into <code>Harness</code></td></tr>
<tr><td><code>internal/harness/subprocess</code></td><td><code>os/exec</code> with env-map + workdir + timeout</td><td><b>new</b></td></tr>
<tr><td><code>internal/harness/webhook</code></td><td>existing <code>internal/webhooks/delivery.go</code></td><td>wrap as <code>Harness</code>, async result via DB poll</td></tr>
<tr><td><code>internal/harness/stub</code></td><td>in-process fake for tests</td><td>new, test-only</td></tr>
</tbody>
</table>
<h3>Resolution policy</h3>
<p><code>Registry.Resolve(agent)</code> picks a backend based on:</p>
<ol>
<li>Explicit <code>agent.HarnessName</code> field (new column, nullable)</li>
<li>Else: agent has <code>K8sImage</code> and <code>k8s.JobRunner.IsAvailable()</code> &rarr; <code>k8sjob</code></li>
<li>Else: agent has <code>Webhooks</code> registered &rarr; <code>webhook</code></li>
<li>Else: agent has <code>LocalCommand</code> configured &rarr; <code>subprocess</code></li>
<li>Else: typed error <code>ErrNoBackend</code></li>
</ol>
<h3>Data model additions</h3>
<ul>
<li>New migration <code>016_harness.sql</code>: add <code>agents.harness_name</code>, <code>agents.local_command</code>, <code>agents.harness_config_json</code></li>
<li>New table <code>harness_runs</code>: mirror of <code>ReactiveRun</code> but backend-agnostic, with <code>backend</code>, <code>trace_id</code>, <code>span_id</code>, <code>usage_in</code>, <code>usage_out</code>, <code>cost_usd</code>, <code>result_json</code></li>
<li>Fold <code>ReactiveRun</code> into <code>harness_runs</code> in a follow-up migration</li>
</ul>
</section>
<section id="otel">
<h2>6 &middot; OTel integration points</h2>
<p>Scion's pattern is the template: <b>(a) initialize an OTel tracer provider in the main process, (b) start a span per harness invocation, (c) inject the trace context into the child as env vars, (d) ship spans via OTLP gRPC to whatever collector is configured.</b></p>
<h3>Init</h3>
<p>New file <code>internal/observability/otel.go</code>:</p>
<pre><span class="k">func</span> <span class="t">Init</span>(ctx context.Context, cfg Config) (shutdown <span class="k">func</span>(context.Context) <span class="k">error</span>, err <span class="k">error</span>) {
res, _ := resource.New(ctx,
resource.WithAttributes(semconv.ServiceName(<span class="s">"synapbus"</span>)),
)
exp, _ := otlptracegrpc.New(ctx,
otlptracegrpc.WithEndpoint(cfg.Endpoint),
otlptracegrpc.WithInsecure(),
)
tp := sdktrace.NewTracerProvider(
sdktrace.WithBatcher(exp),
sdktrace.WithResource(res),
)
otel.SetTracerProvider(tp)
otel.SetTextMapPropagator(propagation.TraceContext{})
<span class="k">return</span> tp.Shutdown, <span class="k">nil</span>
}</pre>
<h3>Span taxonomy</h3>
<table>
<thead><tr><th>Span name</th><th>Where</th><th>Key attributes</th></tr></thead>
<tbody>
<tr><td><code>mcp.tool.execute</code></td><td>MCP handler entry</td><td><code>mcp.tool</code>, <code>agent.name</code>, <code>message.id</code></td></tr>
<tr><td><code>reactor.dispatch</code></td><td><code>reactor.Dispatch()</code></td><td><code>agent.name</code>, <code>trigger.depth</code>, <code>budget.remaining</code></td></tr>
<tr><td><code>harness.resolve</code></td><td><code>Registry.Resolve</code></td><td><code>harness.name</code>, <code>fallback.chain</code></td></tr>
<tr><td><code>harness.provision</code></td><td><code>Harness.Provision</code></td><td><code>harness.name</code>, <code>agent.home</code></td></tr>
<tr><td><code>harness.execute</code></td><td><code>Harness.Execute</code></td><td><code>harness.name</code>, <code>run.id</code>, <code>usage.*</code>, <code>cost.usd</code>, <code>exit.code</code></td></tr>
<tr><td><code>harness.k8s.job.create</code></td><td>k8sjob backend</td><td><code>k8s.job.name</code>, <code>k8s.namespace</code>, <code>k8s.image</code></td></tr>
<tr><td><code>harness.subprocess.exec</code></td><td>subprocess backend</td><td><code>proc.argv[0]</code>, <code>proc.pid</code>, <code>proc.workdir</code></td></tr>
<tr><td><code>harness.webhook.deliver</code></td><td>webhook backend</td><td><code>http.url</code>, <code>http.status_code</code>, <code>retry.count</code></td></tr>
</tbody>
</table>
<h3>Context propagation into children</h3>
<p>For each backend, the current span's <code>traceparent</code> is serialised via <code>propagation.TraceContext{}.Inject</code> into an env-var map and merged with <code>Harness.GetTelemetryEnv()</code>:</p>
<pre><span class="k">func</span> <span class="t">injectTraceEnv</span>(ctx context.Context, dst <span class="k">map</span>[<span class="k">string</span>]<span class="k">string</span>) {
carrier := propagation.MapCarrier{}
otel.GetTextMapPropagator().Inject(ctx, carrier)
<span class="k">for</span> k, v := <span class="k">range</span> carrier {
<span class="c">// OTel convention: TRACEPARENT / TRACESTATE env names</span>
dst[strings.ToUpper(k)] = v
}
dst[<span class="s">"OTEL_EXPORTER_OTLP_ENDPOINT"</span>] = cfg.ChildEndpoint <span class="c">// same collector</span>
dst[<span class="s">"OTEL_SERVICE_NAME"</span>] = <span class="s">"synapbus-agent-"</span> + agentName
dst[<span class="s">"OTEL_RESOURCE_ATTRIBUTES"</span>] = <span class="s">"synapbus.run_id="</span> + runID
}</pre>
<p>For K8s: merged into <code>corev1.EnvVar</code> slice at <code>k8s/runner.go:105&ndash;119</code>. For subprocess: merged into <code>cmd.Env</code>. For webhook: added as HTTP headers (<code>traceparent</code>, <code>tracestate</code>) alongside the existing <code>X-SynapBus-*</code> headers.</p>
<h3>Metrics</h3>
<p>Keep the existing Prometheus registry (<code>internal/metrics/metrics.go</code>) &mdash; it's already wired &mdash; but <b>also</b> emit a minimal set via OTel meter, so a single OTLP collector sees both spans and metrics:</p>
<ul>
<li><code>synapbus.harness.runs</code> (counter, labels: <code>harness</code>, <code>status</code>)</li>
<li><code>synapbus.harness.duration_ms</code> (histogram)</li>
<li><code>synapbus.harness.tokens_in</code> / <code>tokens_out</code> (counters)</li>
<li><code>synapbus.harness.cost_usd</code> (counter)</li>
</ul>
<h3>Config</h3>
<p>Three new env vars (matching scion naming, with <code>SYNAPBUS_</code> prefix for ours):</p>
<ul>
<li><code>SYNAPBUS_OTEL_ENDPOINT</code> &mdash; e.g. <code>http://otel-collector:4317</code></li>
<li><code>SYNAPBUS_OTEL_INSECURE</code> &mdash; bool, default true for LAN</li>
<li><code>SYNAPBUS_OTEL_ENABLED</code> &mdash; bool, default false (opt-in)</li>
</ul>
<p>Until a real collector exists on kubic, a file exporter (<code>stdouttrace</code>) or the existing <code>trace.Tracer</code> (SQLite <code>trace</code> table) can back the same interface via an adapter.</p>
</section>
<section id="nuggets">
<h2>7 &middot; Other reusable nuggets</h2>
<div class="grid2">
<div class="card">
<h3>From scion</h3>
<ul>
<li><b>Workspace-per-agent git worktree</b> for isolation &mdash; nice-to-have once multiple reactive agents run in parallel on the same host.</li>
<li><b>Interrupt key</b> per harness (<code>GetInterruptKey</code>) &mdash; e.g. double-Escape for Claude Code &mdash; useful for cancel semantics.</li>
<li><b>Pre-approved tool fingerprints</b> written into <code>.claude.json customApiKeyResponses</code> &mdash; removes the "did you really want to use this key?" prompt.</li>
<li><b>Structured <code>StructuredMessage</code> envelope</b> &mdash; SynapBus messages already have most of this; add a <code>type</code> enum (<code>instruction</code>/<code>input-needed</code>/<code>state-change</code>).</li>
</ul>
</div>
<div class="card">
<h3>From paperclip</h3>
<ul>
<li><b><code>testEnvironment()</code> preflight</b> &mdash; a health check per harness, runnable from the admin CLI ("can this agent actually dispatch?").</li>
<li><b><code>sessionCodec</code></b> &mdash; serialise/resume an agent conversation across reactive runs. Gives SynapBus a real "sticky" agent without re-prompting.</li>
<li><b>Cost/token usage in the result envelope</b> &mdash; already in <code>benchmark/sdk_backend.py</code>, worth lifting into the core result type.</li>
<li><b>Atomic per-agent lock</b> &mdash; belt-and-braces guarantee that one agent can't double-fire on the same trigger.</li>
</ul>
</div>
</div>
</section>
<section id="nextsteps">
<h2>8 &middot; Next steps &amp; open questions</h2>
<div class="callout">
<b>Awaiting approval before any code is written.</b> The user asked for research + design first.
</div>
<h3>Staged implementation plan (for discussion)</h3>
<div class="tl">
<div class="step"><b>Phase 0 &mdash; design doc</b> at <code>docs/harness-otel-design.md</code> (written alongside this report)</div>
<div class="step"><b>Phase 1 &mdash; scaffold</b> <code>internal/harness/</code> with the interface, registry, and a stub backend. Pure Go, no external deps added.</div>
<div class="step"><b>Phase 2 &mdash; refactor K8s path</b> behind the new <code>Harness</code> interface without changing behaviour. Existing tests stay green.</div>
<div class="step"><b>Phase 3 &mdash; new <code>subprocess</code> backend</b> + per-agent <code>local_command</code> config + migration 016.</div>
<div class="step"><b>Phase 4 &mdash; wrap webhook path</b> as a third backend, via the resolver. Async result via DB poll.</div>
<div class="step"><b>Phase 5 &mdash; OTel init &amp; span wiring</b> around all three backends. Env-var propagation into children. Opt-in config.</div>
<div class="step"><b>Phase 6 &mdash; session codec + cost accounting</b> on <code>harness_runs</code>. <code>testEnvironment</code> preflight exposed via admin CLI.</div>
</div>
<h3>Open questions for you</h3>
<ol>
<li><b>Collector.</b> Is there an OTel collector on <code>kubic.home.arpa</code> already, or do we deploy one first (Tempo? Jaeger? stdout only for now)?</li>
<li><b>Scope of Phase 1.</b> Do you want the new package to land behind a feature flag, or replace the existing reactor path immediately?</li>
<li><b>Subprocess path on the Mac.</b> SynapBus today only runs agents as K8s Jobs. The subprocess backend lets it also run claude-code / gemini-cli locally on your laptop. Is that in-scope now or defer?</li>
<li><b>Session codec.</b> How much of paperclip's session-resume semantics do you want &mdash; just "reuse the Claude Code session id" or full conversation replay?</li>
<li><b>Plugin loader.</b> Do we need to load third-party harnesses at runtime (plugin.Plugin / HashiCorp <code>go-plugin</code>), or is a compile-time registry enough?</li>
</ol>
</section>
<footer>
Sources &mdash;
<a href="https://github.com/GoogleCloudPlatform/scion">GoogleCloudPlatform/scion</a> &middot;
<a href="https://github.com/paperclipai/paperclip">paperclipai/paperclip</a> &middot;
synapbus HEAD <code>0e25fbc</code> &middot;
Report generated locally, no external JS/CSS.
</footer>
</div>
</body>
</html>
+5 -5
View File
@@ -16,6 +16,11 @@ require (
github.com/prometheus/client_model v0.6.2
github.com/prometheus/common v0.66.1
github.com/spf13/cobra v1.10.2
go.opentelemetry.io/otel v1.31.0
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.21.0
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp v1.21.0
go.opentelemetry.io/otel/sdk v1.31.0
go.opentelemetry.io/otel/trace v1.31.0
golang.org/x/crypto v0.49.0
golang.org/x/oauth2 v0.36.0
golang.org/x/time v0.9.0
@@ -104,14 +109,9 @@ require (
go.opentelemetry.io/contrib/propagators/b3 v1.21.0 // indirect
go.opentelemetry.io/contrib/propagators/jaeger v1.21.1 // indirect
go.opentelemetry.io/contrib/samplers/jaegerremote v0.15.1 // indirect
go.opentelemetry.io/otel v1.31.0 // indirect
go.opentelemetry.io/otel/exporters/jaeger v1.17.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.21.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp v1.21.0 // indirect
go.opentelemetry.io/otel/exporters/zipkin v1.21.0 // indirect
go.opentelemetry.io/otel/metric v1.31.0 // indirect
go.opentelemetry.io/otel/sdk v1.31.0 // indirect
go.opentelemetry.io/otel/trace v1.31.0 // indirect
go.opentelemetry.io/proto/otlp v1.0.0 // indirect
go.yaml.in/yaml/v2 v2.4.3 // indirect
go.yaml.in/yaml/v3 v3.0.4 // indirect
+12 -1
View File
@@ -135,7 +135,8 @@ func (s *SQLiteAgentStore) ListReactiveAgents(ctx context.Context) ([]*Agent, er
// agentSelectSQL returns the base SELECT clause for agent queries.
func agentSelectSQL() string {
return `SELECT id, name, display_name, type, capabilities, owner_id, api_key_hash, status, created_at, updated_at,
trigger_mode, cooldown_seconds, daily_trigger_budget, max_trigger_depth, k8s_image, k8s_env_json, k8s_resource_preset, pending_work
trigger_mode, cooldown_seconds, daily_trigger_budget, max_trigger_depth, k8s_image, k8s_env_json, k8s_resource_preset, pending_work,
harness_name, local_command, harness_config_json
FROM agents`
}
@@ -241,6 +242,7 @@ func (s *SQLiteAgentStore) scanAgent(row *sql.Row) (*Agent, error) {
var agent Agent
var caps string
var k8sImage, k8sEnvJSON sql.NullString
var harnessName, localCommand, harnessConfigJSON sql.NullString
var pendingWork int
err := row.Scan(
&agent.ID, &agent.Name, &agent.DisplayName, &agent.Type,
@@ -249,6 +251,7 @@ func (s *SQLiteAgentStore) scanAgent(row *sql.Row) (*Agent, error) {
&agent.TriggerMode, &agent.CooldownSeconds, &agent.DailyTriggerBudget,
&agent.MaxTriggerDepth, &k8sImage, &k8sEnvJSON,
&agent.K8sResourcePreset, &pendingWork,
&harnessName, &localCommand, &harnessConfigJSON,
)
if err != nil {
return nil, err
@@ -257,6 +260,9 @@ func (s *SQLiteAgentStore) scanAgent(row *sql.Row) (*Agent, error) {
agent.K8sImage = k8sImage.String
agent.K8sEnvJSON = k8sEnvJSON.String
agent.PendingWork = pendingWork != 0
agent.HarnessName = harnessName.String
agent.LocalCommand = localCommand.String
agent.HarnessConfigJSON = harnessConfigJSON.String
return &agent, nil
}
@@ -266,6 +272,7 @@ func (s *SQLiteAgentStore) scanAgents(rows *sql.Rows) ([]*Agent, error) {
var agent Agent
var caps string
var k8sImage, k8sEnvJSON sql.NullString
var harnessName, localCommand, harnessConfigJSON sql.NullString
var pendingWork int
err := rows.Scan(
&agent.ID, &agent.Name, &agent.DisplayName, &agent.Type,
@@ -274,6 +281,7 @@ func (s *SQLiteAgentStore) scanAgents(rows *sql.Rows) ([]*Agent, error) {
&agent.TriggerMode, &agent.CooldownSeconds, &agent.DailyTriggerBudget,
&agent.MaxTriggerDepth, &k8sImage, &k8sEnvJSON,
&agent.K8sResourcePreset, &pendingWork,
&harnessName, &localCommand, &harnessConfigJSON,
)
if err != nil {
return nil, err
@@ -282,6 +290,9 @@ func (s *SQLiteAgentStore) scanAgents(rows *sql.Rows) ([]*Agent, error) {
agent.K8sImage = k8sImage.String
agent.K8sEnvJSON = k8sEnvJSON.String
agent.PendingWork = pendingWork != 0
agent.HarnessName = harnessName.String
agent.LocalCommand = localCommand.String
agent.HarnessConfigJSON = harnessConfigJSON.String
agents = append(agents, &agent)
}
if agents == nil {
+5
View File
@@ -41,4 +41,9 @@ type Agent struct {
K8sEnvJSON string `json:"k8s_env_json,omitempty"`
K8sResourcePreset string `json:"k8s_resource_preset"`
PendingWork bool `json:"pending_work"`
// Harness-agnostic execution fields (migration 019).
HarnessName string `json:"harness_name,omitempty"` // explicit backend; empty = auto-resolve
LocalCommand string `json:"local_command,omitempty"` // JSON-encoded argv for subprocess backend
HarnessConfigJSON string `json:"harness_config_json,omitempty"` // opaque per-backend config
}
+170
View File
@@ -0,0 +1,170 @@
// Package harness provides a runtime-agnostic interface for dispatching agent
// work to a concrete execution backend (Kubernetes Job, local subprocess,
// outbound webhook, or an in-memory stub for tests). It is the single seam
// between SynapBus's reactor / webhook / MCP entry points and whatever
// actually runs an agent.
//
// Inspired by GoogleCloudPlatform/scion's api.Harness interface. See
// docs/harness-otel-design.md for the full design rationale.
package harness
import (
"context"
"encoding/json"
"errors"
"time"
"github.com/synapbus/synapbus/internal/agents"
"github.com/synapbus/synapbus/internal/messaging"
)
// ErrNoBackend is returned by Registry.Resolve when no registered harness
// can handle the given agent.
var ErrNoBackend = errors.New("harness: no backend available for agent")
// ErrUnknownHarness is returned when a harness name is requested but is
// not registered.
var ErrUnknownHarness = errors.New("harness: unknown harness name")
// Capabilities advertises what a given harness backend supports so the
// caller can degrade gracefully (e.g. fall back from a system prompt to
// inline instructions when the backend cannot carry one).
type Capabilities struct {
// SystemPrompt means the backend can carry a dedicated system prompt
// separate from the user task.
SystemPrompt bool
// SessionResume means the backend can resume a prior conversation by
// session id. When false, every Execute is a cold start.
SessionResume bool
// Skills means the backend can materialise a set of skill files into
// the agent workspace.
Skills bool
// OTelNative means the child process honours standard OTEL_* env vars
// (OTEL_EXPORTER_OTLP_ENDPOINT, TRACEPARENT, etc.). If true, the
// dispatcher will inject trace context via environment variables.
OTelNative bool
// MaxConcurrency is an advisory upper bound on simultaneous Execute
// calls the backend can handle. Zero means "no explicit limit".
MaxConcurrency int
}
// Budget bounds a single Execute call. Zero values mean "no limit".
type Budget struct {
MaxWallClock time.Duration
MaxTokensIn int64
MaxTokensOut int64
MaxCostUSD float64
}
// Usage records resource consumption for a completed run.
type Usage struct {
TokensIn int64 `json:"tokens_in"`
TokensOut int64 `json:"tokens_out"`
TokensCached int64 `json:"tokens_cached"`
CostUSD float64 `json:"cost_usd"`
}
// ExecRequest is the single input to a harness Execute call.
type ExecRequest struct {
// RunID is a stable caller-generated id (UUID-ish). Propagated into
// the child as SYNAPBUS_RUN_ID so logs and traces correlate.
RunID string
// AgentName is the target agent's SynapBus name. The harness may use
// it for logging and for selecting per-agent configuration.
AgentName string
// Agent is the full agent record. Backends read per-agent config
// from it (K8sImage, LocalCommand, etc.). Nil is allowed only for
// the stub backend used in tests.
Agent *agents.Agent
// Message is the triggering message, if any. Nil for on-demand runs
// such as admin CLI invocations.
Message *messaging.Message
// Context is an optional conversation window the caller wants the
// child to see. Backends that support SessionResume may ignore this
// in favour of their own session state.
Context []*messaging.Message
// SessionID, when non-empty and Capabilities.SessionResume is true,
// asks the backend to resume a prior conversation.
SessionID string
// Budget bounds wall-clock, tokens, and cost.
Budget Budget
// Env is a set of caller-provided environment overrides merged on
// top of whatever the backend normally injects (last write wins).
Env map[string]string
// Skills lists skill names the backend should materialise if
// Capabilities.Skills is true.
Skills []string
}
// ExecResult is the single output of a harness Execute call.
type ExecResult struct {
// ExitCode follows Unix convention: 0 success, non-zero failure.
// For backends without a true exit code (e.g. webhook), this is a
// synthetic value: 0 on HTTP 2xx, 1 otherwise.
ExitCode int
// Logs is a bounded excerpt of stdout+stderr (or HTTP response body
// for the webhook backend). Full logs live on disk / remote storage.
Logs string
// ResultJSON is an optional structured output the agent emitted.
// Subprocess agents write this to a well-known path; K8s agents
// write it to stdout or a shared volume; webhook agents return it
// in the response body.
ResultJSON json.RawMessage
// Usage captures token / cost accounting when the backend can
// report it. Zero values mean "not reported".
Usage Usage
// SessionID is the backend-specific session identifier for resume.
// Empty when the backend does not support SessionResume.
SessionID string
// TraceID is the W3C trace id (hex) the run was observed under, so
// callers can link their row to a distributed trace.
TraceID string
}
// Harness is the interface every execution backend implements.
type Harness interface {
// Name is the short stable identifier used in config (e.g. "k8sjob",
// "subprocess", "webhook", "stub").
Name() string
// Capabilities advertises backend features.
Capabilities() Capabilities
// TestEnvironment is a cheap preflight: is the required binary
// installed, is auth valid, can we reach the model? Called by the
// admin CLI and by Registry.Resolve as a health gate.
TestEnvironment(ctx context.Context) error
// Provision performs one-shot setup for a given agent (writing
// config files, pre-approving tool fingerprints, materialising
// skills). Idempotent — safe to call repeatedly.
Provision(ctx context.Context, agent *agents.Agent) error
// Execute dispatches a single request and blocks until the run
// terminates, the context is cancelled, or the Budget is exhausted.
// Implementations must always return a non-nil *ExecResult when
// they return nil error.
Execute(ctx context.Context, req *ExecRequest) (*ExecResult, error)
// Cancel asks the backend to abort an in-flight run with the given
// RunID. Best-effort; returns nil if the run is unknown or already
// terminated.
Cancel(ctx context.Context, runID string) error
}
+299
View File
@@ -0,0 +1,299 @@
// Package k8sjob is the Kubernetes-Job-backed implementation of
// harness.Harness. It wraps the existing internal/k8s.JobRunner (used by
// the reactor today) behind the harness interface so the reactor, admin
// CLI, and future callers all go through the same seam.
package k8sjob
import (
"context"
"encoding/json"
"errors"
"fmt"
"log/slog"
"strings"
"time"
"github.com/synapbus/synapbus/internal/agents"
"github.com/synapbus/synapbus/internal/harness"
k8spkg "github.com/synapbus/synapbus/internal/k8s"
)
// Waiter blocks until a Kubernetes Job terminates and reports the outcome.
// It is its own interface so the harness can be unit-tested with a fake.
type Waiter interface {
Wait(ctx context.Context, namespace, jobName string) (JobOutcome, error)
Cancel(ctx context.Context, namespace, jobName string) error
}
// JobOutcome is the terminal state of a watched Kubernetes Job.
type JobOutcome struct {
Success bool
FailureReason string
}
// Harness is the k8sjob implementation of harness.Harness.
type Harness struct {
runner k8spkg.JobRunner
waiter Waiter
logger *slog.Logger
}
// New constructs a k8sjob harness. Waiter may be nil when IsAvailable()
// is false (NoopRunner fallback) — in that case Execute returns an error
// rather than panicking.
func New(runner k8spkg.JobRunner, waiter Waiter, logger *slog.Logger) *Harness {
if logger == nil {
logger = slog.Default()
}
return &Harness{
runner: runner,
waiter: waiter,
logger: logger.With("harness", "k8sjob"),
}
}
// Name returns the registered harness name.
func (h *Harness) Name() string { return "k8sjob" }
// Capabilities advertises what the k8sjob backend supports.
func (h *Harness) Capabilities() harness.Capabilities {
return harness.Capabilities{
SystemPrompt: false, // passed via env vars, not a dedicated slot
SessionResume: false, // each job is a cold start
Skills: false,
OTelNative: true, // child honours OTEL_* env vars
MaxConcurrency: 10,
}
}
// TestEnvironment checks the underlying runner is available.
func (h *Harness) TestEnvironment(ctx context.Context) error {
if h.runner == nil || !h.runner.IsAvailable() {
return errors.New("k8sjob: JobRunner is not available (not running in-cluster)")
}
return nil
}
// Provision is a no-op for k8sjob. Per-run configuration is passed
// entirely via env vars at Execute time.
func (h *Harness) Provision(ctx context.Context, agent *agents.Agent) error {
return nil
}
// Execute builds a K8s Job for the request, waits for it to terminate,
// and returns the captured logs plus exit code.
//
// If the Waiter is nil (e.g. when constructed against a NoopRunner), the
// harness returns an error immediately rather than blocking forever.
func (h *Harness) Execute(ctx context.Context, req *harness.ExecRequest) (*harness.ExecResult, error) {
if req == nil {
return nil, errors.New("k8sjob: nil ExecRequest")
}
if req.Agent == nil {
return nil, errors.New("k8sjob: ExecRequest.Agent is required")
}
if err := h.TestEnvironment(ctx); err != nil {
return nil, err
}
if h.waiter == nil {
return nil, errors.New("k8sjob: no Waiter configured")
}
handler := BuildHandler(req.Agent)
// Merge caller-provided env overrides and run metadata on top of
// whatever the handler already carries. Last write wins.
if handler.Env == nil {
handler.Env = map[string]string{}
}
handler.Env["SYNAPBUS_RUN_ID"] = req.RunID
for k, v := range req.Env {
handler.Env[k] = v
}
msg := buildJobMessage(req)
if req.Budget.MaxWallClock > 0 {
handler.TimeoutSeconds = int(req.Budget.MaxWallClock.Seconds())
}
jobName, err := h.runner.CreateJob(ctx, handler, msg)
if err != nil {
return nil, fmt.Errorf("k8sjob: create job: %w", err)
}
ns := handler.Namespace
if ns == "" {
ns = h.runner.GetNamespace()
}
h.logger.Info("k8sjob launched",
"job_name", jobName,
"namespace", ns,
"agent", req.AgentName,
"run_id", req.RunID,
)
outcome, err := h.waiter.Wait(ctx, ns, jobName)
if err != nil {
// Fetch whatever logs we can before bailing out.
logs, _ := h.runner.GetJobLogs(ctx, ns, jobName)
return &harness.ExecResult{
ExitCode: 2, // distinguishable from plain "failed" (exit=1)
Logs: trimLogs(logs, 128),
}, fmt.Errorf("k8sjob: wait: %w", err)
}
logs, logErr := h.runner.GetJobLogs(ctx, ns, jobName)
if logErr != nil {
h.logger.Warn("k8sjob: fetch logs failed",
"job_name", jobName, "error", logErr,
)
}
res := &harness.ExecResult{
ExitCode: exitCodeFromOutcome(outcome),
Logs: trimLogs(logs, 128),
}
if !outcome.Success && outcome.FailureReason != "" {
// Surface the K8s-reported reason in the logs excerpt so
// callers writing to harness_runs can see both.
if res.Logs != "" {
res.Logs = outcome.FailureReason + "\n\n" + res.Logs
} else {
res.Logs = outcome.FailureReason
}
}
// Best-effort: parse the tail of stdout as JSON (common pattern for
// agents emitting a final result envelope). If it parses, stash it.
if rj := extractResultJSON(logs); rj != nil {
res.ResultJSON = rj
}
return res, nil
}
// Cancel deletes the K8s Job whose name equals runID. This is the
// convention used by SynapBus today: the harness returns the K8s Job
// name as the run identifier, so callers can cancel by run id.
func (h *Harness) Cancel(ctx context.Context, runID string) error {
if h.waiter == nil {
return errors.New("k8sjob: no Waiter configured")
}
ns := ""
if h.runner != nil {
ns = h.runner.GetNamespace()
}
return h.waiter.Cancel(ctx, ns, runID)
}
// -- helpers --------------------------------------------------------------
// BuildHandler constructs a K8sHandler from an agent config. Exported so
// the reactor can share the exact same logic when it eventually routes
// through the harness registry.
func BuildHandler(agent *agents.Agent) *k8spkg.K8sHandler {
env := map[string]string{}
if agent.K8sEnvJSON != "" {
var envMap map[string]json.RawMessage
if err := json.Unmarshal([]byte(agent.K8sEnvJSON), &envMap); err == nil {
for k, v := range envMap {
var s string
if err := json.Unmarshal(v, &s); err == nil {
env[k] = s
continue
}
env[k] = strings.Trim(string(v), "\"")
}
}
}
memory := "2Gi"
cpu := "500m"
if agent.K8sResourcePreset == "small" {
memory = "512Mi"
cpu = "100m"
}
handler := &k8spkg.K8sHandler{
AgentName: agent.Name,
Image: agent.K8sImage,
Events: []string{"message.received", "message.mentioned"},
ResourcesMemory: memory,
ResourcesCPU: cpu,
Env: env,
TimeoutSeconds: 3600,
Status: "active",
Args: []string{"--max-turns", "50", "--model", "claude-sonnet-4-6"},
VolumeMounts: []k8spkg.VolumeMount{
{Name: "claude-config", MountPath: "/app/.claude", ReadOnly: false},
{Name: "workspace", MountPath: "/app/workspace", ReadOnly: false},
},
Volumes: []k8spkg.Volume{
{Name: "claude-config", HostPath: "/home/user/.claude"},
{Name: "workspace", EmptyDir: true},
},
}
if agent.Name == "social-commenter" {
handler.Args = []string{"--max-turns", "80", "--model", "claude-opus-4-6"}
}
return handler
}
func buildJobMessage(req *harness.ExecRequest) *k8spkg.JobMessage {
m := &k8spkg.JobMessage{
Timestamp: time.Now().UTC().Format(time.RFC3339),
}
if req.Message != nil {
m.MessageID = req.Message.ID
m.FromAgent = req.Message.FromAgent
m.Body = req.Message.Body
}
return m
}
func exitCodeFromOutcome(o JobOutcome) int {
if o.Success {
return 0
}
return 1
}
// trimLogs keeps the last n lines. Consistent with reactor poller's
// existing 100-line cap; we default a bit higher here.
func trimLogs(s string, n int) string {
if s == "" || n <= 0 {
return s
}
lines := strings.Split(s, "\n")
if len(lines) <= n {
return s
}
return strings.Join(lines[len(lines)-n:], "\n")
}
// extractResultJSON looks for the last non-empty line of logs and tries
// to parse it as a JSON object. Returns nil on failure (very common —
// agents may not emit a result envelope at all).
func extractResultJSON(logs string) json.RawMessage {
if logs == "" {
return nil
}
lines := strings.Split(strings.TrimRight(logs, "\n"), "\n")
for i := len(lines) - 1; i >= 0; i-- {
line := strings.TrimSpace(lines[i])
if line == "" {
continue
}
if len(line) < 2 || line[0] != '{' {
return nil
}
var probe any
if err := json.Unmarshal([]byte(line), &probe); err != nil {
return nil
}
return json.RawMessage(line)
}
return nil
}
+320
View File
@@ -0,0 +1,320 @@
package k8sjob_test
import (
"context"
"errors"
"strings"
"testing"
"time"
"github.com/synapbus/synapbus/internal/agents"
"github.com/synapbus/synapbus/internal/harness"
"github.com/synapbus/synapbus/internal/harness/k8sjob"
k8spkg "github.com/synapbus/synapbus/internal/k8s"
"github.com/synapbus/synapbus/internal/messaging"
)
// fakeRunner is a copy of the reactor test's fakeRunner — wrapping
// k8spkg.JobRunner so we don't pull in a real clientset.
type fakeRunner struct {
available bool
lastEnv map[string]string
lastHandler *k8spkg.K8sHandler
lastMessage *k8spkg.JobMessage
createErr error
logs string
logsErr error
}
func (f *fakeRunner) IsAvailable() bool { return f.available }
func (f *fakeRunner) GetNamespace() string { return "test-ns" }
func (f *fakeRunner) GetJobLogs(_ context.Context, _, _ string) (string, error) {
return f.logs, f.logsErr
}
func (f *fakeRunner) CreateJob(_ context.Context, handler *k8spkg.K8sHandler, msg *k8spkg.JobMessage) (string, error) {
if f.createErr != nil {
return "", f.createErr
}
f.lastHandler = handler
f.lastMessage = msg
f.lastEnv = make(map[string]string, len(handler.Env))
for k, v := range handler.Env {
f.lastEnv[k] = v
}
return "synapbus-" + handler.AgentName + "-job", nil
}
type fakeWaiter struct {
outcome k8sjob.JobOutcome
err error
delay time.Duration
cancelCalls []string
}
func (w *fakeWaiter) Wait(ctx context.Context, ns, jobName string) (k8sjob.JobOutcome, error) {
if w.delay > 0 {
select {
case <-time.After(w.delay):
case <-ctx.Done():
return k8sjob.JobOutcome{}, ctx.Err()
}
}
return w.outcome, w.err
}
func (w *fakeWaiter) Cancel(_ context.Context, _ string, jobName string) error {
w.cancelCalls = append(w.cancelCalls, jobName)
return nil
}
func newTestAgent(name, image string) *agents.Agent {
return &agents.Agent{
Name: name,
K8sImage: image,
K8sResourcePreset: "default",
}
}
func TestHarness_ImplementsInterface(t *testing.T) {
var _ harness.Harness = (*k8sjob.Harness)(nil)
}
func TestHarness_NameAndCapabilities(t *testing.T) {
h := k8sjob.New(nil, nil, nil)
if h.Name() != "k8sjob" {
t.Fatalf("Name = %q", h.Name())
}
caps := h.Capabilities()
if !caps.OTelNative {
t.Fatal("expected OTelNative = true")
}
if caps.MaxConcurrency == 0 {
t.Fatal("expected non-zero MaxConcurrency")
}
}
func TestHarness_TestEnvironment_NoRunner(t *testing.T) {
h := k8sjob.New(nil, nil, nil)
if err := h.TestEnvironment(context.Background()); err == nil {
t.Fatal("expected error for nil runner")
}
}
func TestHarness_TestEnvironment_Unavailable(t *testing.T) {
h := k8sjob.New(&fakeRunner{available: false}, nil, nil)
if err := h.TestEnvironment(context.Background()); err == nil {
t.Fatal("expected error for unavailable runner")
}
}
func TestHarness_TestEnvironment_OK(t *testing.T) {
h := k8sjob.New(&fakeRunner{available: true}, nil, nil)
if err := h.TestEnvironment(context.Background()); err != nil {
t.Fatalf("unexpected error: %v", err)
}
}
func TestHarness_Provision_NoOp(t *testing.T) {
h := k8sjob.New(&fakeRunner{available: true}, nil, nil)
if err := h.Provision(context.Background(), newTestAgent("a", "foo:v1")); err != nil {
t.Fatalf("Provision err = %v, want nil", err)
}
}
func TestHarness_Execute_Success(t *testing.T) {
runner := &fakeRunner{
available: true,
logs: "hello world\n{\"ok\":true,\"tokens\":42}",
}
waiter := &fakeWaiter{outcome: k8sjob.JobOutcome{Success: true}}
h := k8sjob.New(runner, waiter, nil)
req := &harness.ExecRequest{
RunID: "r-1",
AgentName: "researcher",
Agent: newTestAgent("researcher", "ghcr.io/example/agent:v1"),
Message: &messaging.Message{
ID: 42,
FromAgent: "human",
Body: "do the thing",
},
Env: map[string]string{"EXTRA": "yes"},
}
res, err := h.Execute(context.Background(), req)
if err != nil {
t.Fatalf("Execute error: %v", err)
}
if res.ExitCode != 0 {
t.Errorf("ExitCode = %d, want 0", res.ExitCode)
}
if !strings.Contains(res.Logs, "hello world") {
t.Errorf("logs missing stdout: %q", res.Logs)
}
if string(res.ResultJSON) == "" {
t.Errorf("ResultJSON empty, want parsed envelope")
}
if !strings.Contains(string(res.ResultJSON), "\"ok\":true") {
t.Errorf("ResultJSON not parsed: %s", res.ResultJSON)
}
// Verify env propagation
if runner.lastEnv["SYNAPBUS_RUN_ID"] != "r-1" {
t.Errorf("SYNAPBUS_RUN_ID = %q, want r-1", runner.lastEnv["SYNAPBUS_RUN_ID"])
}
if runner.lastEnv["EXTRA"] != "yes" {
t.Errorf("EXTRA env not propagated: %v", runner.lastEnv)
}
// Verify JobMessage carried message context
if runner.lastMessage.MessageID != 42 {
t.Errorf("MessageID = %d", runner.lastMessage.MessageID)
}
if runner.lastMessage.FromAgent != "human" {
t.Errorf("FromAgent = %q", runner.lastMessage.FromAgent)
}
}
func TestHarness_Execute_Failure(t *testing.T) {
runner := &fakeRunner{
available: true,
logs: "error: file not found",
}
waiter := &fakeWaiter{
outcome: k8sjob.JobOutcome{Success: false, FailureReason: "BackoffLimitExceeded"},
}
h := k8sjob.New(runner, waiter, nil)
req := &harness.ExecRequest{
RunID: "r-2",
AgentName: "researcher",
Agent: newTestAgent("researcher", "ghcr.io/example/agent:v1"),
}
res, err := h.Execute(context.Background(), req)
if err != nil {
t.Fatalf("Execute err = %v, want nil for graceful failure", err)
}
if res.ExitCode != 1 {
t.Errorf("ExitCode = %d, want 1", res.ExitCode)
}
if !strings.Contains(res.Logs, "BackoffLimitExceeded") {
t.Errorf("logs missing failure reason: %q", res.Logs)
}
}
func TestHarness_Execute_CreateJobError(t *testing.T) {
runner := &fakeRunner{
available: true,
createErr: errors.New("api server down"),
}
h := k8sjob.New(runner, &fakeWaiter{}, nil)
_, err := h.Execute(context.Background(), &harness.ExecRequest{
RunID: "r", AgentName: "a", Agent: newTestAgent("a", "foo:v1"),
})
if err == nil || !strings.Contains(err.Error(), "create job") {
t.Fatalf("err = %v, want create job error", err)
}
}
func TestHarness_Execute_NilAgent(t *testing.T) {
h := k8sjob.New(&fakeRunner{available: true}, &fakeWaiter{}, nil)
_, err := h.Execute(context.Background(), &harness.ExecRequest{RunID: "r"})
if err == nil {
t.Fatal("expected error for nil agent")
}
}
func TestHarness_Execute_NilRequest(t *testing.T) {
h := k8sjob.New(&fakeRunner{available: true}, &fakeWaiter{}, nil)
_, err := h.Execute(context.Background(), nil)
if err == nil {
t.Fatal("expected error for nil request")
}
}
func TestHarness_Execute_BudgetOverridesTimeout(t *testing.T) {
runner := &fakeRunner{available: true}
waiter := &fakeWaiter{outcome: k8sjob.JobOutcome{Success: true}}
h := k8sjob.New(runner, waiter, nil)
req := &harness.ExecRequest{
RunID: "r",
AgentName: "a",
Agent: newTestAgent("a", "foo:v1"),
Budget: harness.Budget{MaxWallClock: 90 * time.Second},
}
_, err := h.Execute(context.Background(), req)
if err != nil {
t.Fatalf("Execute error: %v", err)
}
if runner.lastHandler.TimeoutSeconds != 90 {
t.Errorf("TimeoutSeconds = %d, want 90", runner.lastHandler.TimeoutSeconds)
}
}
func TestHarness_Execute_ContextCancel(t *testing.T) {
runner := &fakeRunner{available: true}
waiter := &fakeWaiter{
outcome: k8sjob.JobOutcome{Success: true},
delay: 500 * time.Millisecond,
}
h := k8sjob.New(runner, waiter, nil)
ctx, cancel := context.WithTimeout(context.Background(), 20*time.Millisecond)
defer cancel()
_, err := h.Execute(ctx, &harness.ExecRequest{
RunID: "r", AgentName: "a", Agent: newTestAgent("a", "foo:v1"),
})
if err == nil {
t.Fatal("expected error when context cancelled before wait completes")
}
}
func TestHarness_Cancel_PropagatesToWaiter(t *testing.T) {
runner := &fakeRunner{available: true}
waiter := &fakeWaiter{}
h := k8sjob.New(runner, waiter, nil)
if err := h.Cancel(context.Background(), "my-job"); err != nil {
t.Fatalf("Cancel err = %v", err)
}
if len(waiter.cancelCalls) != 1 || waiter.cancelCalls[0] != "my-job" {
t.Fatalf("cancelCalls = %v", waiter.cancelCalls)
}
}
func TestBuildHandler_AppliesResourcePreset(t *testing.T) {
a := newTestAgent("small-agent", "x:v1")
a.K8sResourcePreset = "small"
h := k8sjob.BuildHandler(a)
if h.ResourcesMemory != "512Mi" || h.ResourcesCPU != "100m" {
t.Errorf("small preset: got mem=%s cpu=%s", h.ResourcesMemory, h.ResourcesCPU)
}
}
func TestBuildHandler_SocialCommenterUsesOpus(t *testing.T) {
a := newTestAgent("social-commenter", "x:v1")
h := k8sjob.BuildHandler(a)
found := false
for _, arg := range h.Args {
if arg == "claude-opus-4-6" {
found = true
}
}
if !found {
t.Errorf("social-commenter args = %v, want opus", h.Args)
}
}
func TestBuildHandler_ParsesK8sEnvJSON(t *testing.T) {
a := newTestAgent("a", "x:v1")
a.K8sEnvJSON = `{"GIT_REPO":"owner/repo","OTHER":"val"}`
h := k8sjob.BuildHandler(a)
if h.Env["GIT_REPO"] != "owner/repo" {
t.Errorf("GIT_REPO = %q", h.Env["GIT_REPO"])
}
if h.Env["OTHER"] != "val" {
t.Errorf("OTHER = %q", h.Env["OTHER"])
}
}
+98
View File
@@ -0,0 +1,98 @@
package k8sjob
import (
"context"
"fmt"
"time"
batchv1 "k8s.io/api/batch/v1"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/client-go/kubernetes"
)
// ClientsetWaiter is a Waiter backed by a real k8s.io/client-go clientset.
// It polls Job status on an interval (default 10s) until the Job reports
// Complete, Failed, or the context is cancelled.
type ClientsetWaiter struct {
Clientset kubernetes.Interface
Interval time.Duration
}
// NewClientsetWaiter constructs a Waiter from an existing clientset.
// Interval defaults to 10 seconds when zero.
func NewClientsetWaiter(cs kubernetes.Interface, interval time.Duration) *ClientsetWaiter {
if interval <= 0 {
interval = 10 * time.Second
}
return &ClientsetWaiter{Clientset: cs, Interval: interval}
}
// Wait blocks until the named Job enters a terminal state or the context
// is cancelled. The returned JobOutcome has Success=true only when the
// Job's JobComplete condition is True.
func (w *ClientsetWaiter) Wait(ctx context.Context, namespace, jobName string) (JobOutcome, error) {
ticker := time.NewTicker(w.Interval)
defer ticker.Stop()
// Do an immediate first check so short-running jobs return fast in
// tests (interval may be set to a small value).
if outcome, done, err := w.check(ctx, namespace, jobName); done || err != nil {
return outcome, err
}
for {
select {
case <-ctx.Done():
return JobOutcome{}, ctx.Err()
case <-ticker.C:
outcome, done, err := w.check(ctx, namespace, jobName)
if err != nil {
return JobOutcome{}, err
}
if done {
return outcome, nil
}
}
}
}
func (w *ClientsetWaiter) check(ctx context.Context, namespace, jobName string) (JobOutcome, bool, error) {
job, err := w.Clientset.BatchV1().Jobs(namespace).Get(ctx, jobName, metav1.GetOptions{})
if err != nil {
return JobOutcome{}, false, fmt.Errorf("get job %s/%s: %w", namespace, jobName, err)
}
for _, cond := range job.Status.Conditions {
if cond.Status != "True" {
continue
}
switch cond.Type {
case batchv1.JobComplete:
return JobOutcome{Success: true}, true, nil
case batchv1.JobFailed:
reason := cond.Reason
if cond.Message != "" {
if reason != "" {
reason += ": "
}
reason += cond.Message
}
return JobOutcome{Success: false, FailureReason: reason}, true, nil
}
}
if job.Status.Failed > 0 {
return JobOutcome{Success: false, FailureReason: "pod failure"}, true, nil
}
return JobOutcome{}, false, nil
}
// Cancel deletes the Job (and its pods) by name. Used by harness.Cancel.
func (w *ClientsetWaiter) Cancel(ctx context.Context, namespace, jobName string) error {
propagation := metav1.DeletePropagationBackground
err := w.Clientset.BatchV1().Jobs(namespace).Delete(ctx, jobName, metav1.DeleteOptions{
PropagationPolicy: &propagation,
})
if err != nil {
return fmt.Errorf("delete job %s/%s: %w", namespace, jobName, err)
}
return nil
}
+183
View File
@@ -0,0 +1,183 @@
package harness
import (
"context"
"fmt"
"sync"
"github.com/synapbus/synapbus/internal/agents"
"github.com/synapbus/synapbus/internal/observability"
"go.opentelemetry.io/otel"
"go.opentelemetry.io/otel/attribute"
"go.opentelemetry.io/otel/codes"
"go.opentelemetry.io/otel/trace"
)
// Observer is called on every Registry.Execute so callers can persist
// harness_runs rows, emit metrics, or drive side-effects without
// coupling the core registry to storage. OnStart runs after the backend
// has been resolved but before Execute, so the observer knows the
// backend name. OnFinish runs after Execute returns, whether it
// succeeded or failed.
type Observer interface {
OnStart(ctx context.Context, agent *agents.Agent, harnessName string, req *ExecRequest)
OnFinish(ctx context.Context, agent *agents.Agent, harnessName string, req *ExecRequest, res *ExecResult, err error)
}
// Registry holds the set of available Harness implementations and resolves
// the right backend for a given agent. It is safe for concurrent use.
type Registry struct {
mu sync.RWMutex
byName map[string]Harness
// ResolveFn, if non-nil, overrides the default resolution policy.
// The default picks by agent.HarnessName → k8sjob → webhook →
// subprocess in that order.
ResolveFn func(r *Registry, agent *agents.Agent) (Harness, error)
// Observer, if non-nil, is notified on every Execute.
Observer Observer
}
// NewRegistry returns an empty registry.
func NewRegistry() *Registry {
return &Registry{byName: map[string]Harness{}}
}
// Register adds a harness under its Name(). A second Register with the
// same name replaces the first — tests rely on this to swap in a stub.
func (r *Registry) Register(h Harness) {
r.mu.Lock()
defer r.mu.Unlock()
r.byName[h.Name()] = h
}
// Get returns the harness registered under name, or ErrUnknownHarness.
func (r *Registry) Get(name string) (Harness, error) {
r.mu.RLock()
defer r.mu.RUnlock()
h, ok := r.byName[name]
if !ok {
return nil, fmt.Errorf("%w: %q", ErrUnknownHarness, name)
}
return h, nil
}
// Names returns the registered harness names in no particular order.
func (r *Registry) Names() []string {
r.mu.RLock()
defer r.mu.RUnlock()
out := make([]string, 0, len(r.byName))
for k := range r.byName {
out = append(out, k)
}
return out
}
// Resolve picks the right backend for an agent.
//
// Default policy:
// 1. If ResolveFn is set, delegate to it.
// 2. Else if agent has a non-empty HarnessName that is registered,
// use it (explicit wins).
// 3. Else try "k8sjob" if registered and the agent has K8sImage set.
// 4. Else try "webhook" if registered.
// 5. Else try "subprocess" if registered.
// 6. Else ErrNoBackend.
//
// Note: the agent-level HarnessName/LocalCommand fields don't yet exist
// on agents.Agent — Phase 3 adds them via migration 016. Until then the
// resolver falls back to the legacy "agent has k8s_image" heuristic.
func (r *Registry) Resolve(agent *agents.Agent) (Harness, error) {
if r.ResolveFn != nil {
return r.ResolveFn(r, agent)
}
r.mu.RLock()
defer r.mu.RUnlock()
if agent != nil && agent.K8sImage != "" {
if h, ok := r.byName["k8sjob"]; ok {
return h, nil
}
}
if h, ok := r.byName["webhook"]; ok {
return h, nil
}
if h, ok := r.byName["subprocess"]; ok {
return h, nil
}
return nil, fmt.Errorf("%w: agent=%q", ErrNoBackend, agentNameOf(agent))
}
func agentNameOf(a *agents.Agent) string {
if a == nil {
return ""
}
return a.Name
}
// Execute is the single entry point the reactor uses: resolve a backend,
// call its Execute, return the result. Trace spans, env injection, and
// harness_runs rows are layered on top of this by higher packages so the
// core registry stays a thin dispatcher.
func (r *Registry) Execute(ctx context.Context, agent *agents.Agent, req *ExecRequest) (*ExecResult, error) {
tracer := otel.Tracer(observability.TracerName)
ctx, span := tracer.Start(ctx, "harness.execute",
trace.WithAttributes(
attribute.String("agent.name", agentNameOf(agent)),
attribute.String("run.id", req.RunID),
),
)
defer span.End()
h, err := r.Resolve(agent)
if err != nil {
span.RecordError(err)
span.SetStatus(codes.Error, err.Error())
if r.Observer != nil {
r.Observer.OnFinish(ctx, agent, "", req, nil, err)
}
return nil, err
}
span.SetAttributes(attribute.String("harness.name", h.Name()))
if req.Agent == nil {
req.Agent = agent
}
if req.Env == nil {
req.Env = map[string]string{}
}
observability.InjectTraceContext(ctx, req.Env)
if r.Observer != nil {
r.Observer.OnStart(ctx, agent, h.Name(), req)
}
res, err := h.Execute(ctx, req)
if r.Observer != nil {
r.Observer.OnFinish(ctx, agent, h.Name(), req, res, err)
}
if err != nil {
span.RecordError(err)
span.SetStatus(codes.Error, err.Error())
return res, err
}
if res != nil {
span.SetAttributes(
attribute.Int("exit.code", res.ExitCode),
attribute.Int64("usage.tokens_in", res.Usage.TokensIn),
attribute.Int64("usage.tokens_out", res.Usage.TokensOut),
attribute.Float64("usage.cost_usd", res.Usage.CostUSD),
)
if res.TraceID == "" {
res.TraceID = observability.TraceIDFromContext(ctx)
}
if res.ExitCode != 0 {
span.SetStatus(codes.Error, fmt.Sprintf("exit code %d", res.ExitCode))
} else {
span.SetStatus(codes.Ok, "")
}
}
return res, nil
}
+131
View File
@@ -0,0 +1,131 @@
package harness_test
import (
"context"
"errors"
"testing"
"github.com/synapbus/synapbus/internal/agents"
"github.com/synapbus/synapbus/internal/harness"
"github.com/synapbus/synapbus/internal/harness/stub"
"github.com/synapbus/synapbus/internal/observability"
"go.opentelemetry.io/otel"
sdktrace "go.opentelemetry.io/otel/sdk/trace"
"go.opentelemetry.io/otel/sdk/trace/tracetest"
)
func withTracer(t *testing.T) *tracetest.InMemoryExporter {
t.Helper()
exp := tracetest.NewInMemoryExporter()
tp := sdktrace.NewTracerProvider(sdktrace.WithSyncer(exp))
otel.SetTracerProvider(tp)
// Install propagator so InjectTraceContext emits TRACEPARENT.
_, _ = observability.Init(context.Background(), observability.Config{Enabled: false}, nil)
t.Cleanup(func() { _ = tp.Shutdown(context.Background()) })
return exp
}
func TestRegistry_Execute_EmitsSpan(t *testing.T) {
exp := withTracer(t)
r := harness.NewRegistry()
s := stub.New()
s.NameStr = "subprocess"
r.Register(s)
_, err := r.Execute(context.Background(), &agents.Agent{Name: "a"}, &harness.ExecRequest{RunID: "run-x"})
if err != nil {
t.Fatalf("Execute err = %v", err)
}
spans := exp.GetSpans()
if len(spans) == 0 {
t.Fatal("no spans recorded")
}
found := false
for _, sp := range spans {
if sp.Name == "harness.execute" {
found = true
attrs := map[string]string{}
for _, a := range sp.Attributes {
attrs[string(a.Key)] = a.Value.Emit()
}
if attrs["agent.name"] != "a" {
t.Errorf("agent.name attr = %q", attrs["agent.name"])
}
if attrs["run.id"] != "run-x" {
t.Errorf("run.id attr = %q", attrs["run.id"])
}
if attrs["harness.name"] != "subprocess" {
t.Errorf("harness.name attr = %q", attrs["harness.name"])
}
}
}
if !found {
t.Fatalf("span harness.execute not found in %v", spans)
}
}
func TestRegistry_Execute_InjectsTraceContextIntoEnv(t *testing.T) {
_ = withTracer(t)
r := harness.NewRegistry()
s := stub.New()
s.NameStr = "subprocess"
r.Register(s)
_, err := r.Execute(context.Background(), &agents.Agent{Name: "a"}, &harness.ExecRequest{RunID: "run-x"})
if err != nil {
t.Fatalf("Execute err = %v", err)
}
if len(s.Calls) != 1 {
t.Fatalf("stub calls = %d", len(s.Calls))
}
env := s.Calls[0].Env
if _, ok := env["TRACEPARENT"]; !ok {
t.Errorf("TRACEPARENT not injected into req.Env: %v", env)
}
}
func TestRegistry_Execute_SpanRecordsError(t *testing.T) {
exp := withTracer(t)
r := harness.NewRegistry()
s := stub.New()
s.NameStr = "subprocess"
s.Err = errors.New("boom")
r.Register(s)
_, err := r.Execute(context.Background(), &agents.Agent{Name: "a"}, &harness.ExecRequest{RunID: "run-x"})
if err == nil {
t.Fatal("expected error")
}
spans := exp.GetSpans()
var statuses []string
for _, sp := range spans {
if sp.Name == "harness.execute" {
statuses = append(statuses, sp.Status.Code.String())
}
}
if len(statuses) == 0 || statuses[0] != "Error" {
t.Errorf("span statuses = %v, want [Error]", statuses)
}
}
func TestRegistry_Execute_PopulatesResultTraceID(t *testing.T) {
_ = withTracer(t)
r := harness.NewRegistry()
s := stub.New()
s.NameStr = "subprocess"
r.Register(s)
res, err := r.Execute(context.Background(), &agents.Agent{Name: "a"}, &harness.ExecRequest{RunID: "run-x"})
if err != nil {
t.Fatalf("Execute err = %v", err)
}
if res.TraceID == "" {
t.Error("ExecResult.TraceID empty after registry execute")
}
}
+213
View File
@@ -0,0 +1,213 @@
package harness_test
import (
"context"
"errors"
"testing"
"github.com/synapbus/synapbus/internal/agents"
"github.com/synapbus/synapbus/internal/harness"
"github.com/synapbus/synapbus/internal/harness/stub"
)
func TestRegistry_RegisterAndGet(t *testing.T) {
r := harness.NewRegistry()
s := stub.New()
s.NameStr = "stub"
r.Register(s)
got, err := r.Get("stub")
if err != nil {
t.Fatalf("Get(stub) error: %v", err)
}
if got != s {
t.Fatalf("Get returned %v, want %v", got, s)
}
if _, err := r.Get("missing"); !errors.Is(err, harness.ErrUnknownHarness) {
t.Fatalf("Get(missing) error = %v, want ErrUnknownHarness", err)
}
}
func TestRegistry_Names(t *testing.T) {
r := harness.NewRegistry()
a := stub.New()
a.NameStr = "a"
b := stub.New()
b.NameStr = "b"
r.Register(a)
r.Register(b)
names := r.Names()
if len(names) != 2 {
t.Fatalf("Names len=%d, want 2 (%v)", len(names), names)
}
seen := map[string]bool{}
for _, n := range names {
seen[n] = true
}
if !seen["a"] || !seen["b"] {
t.Fatalf("Names missing entries: %v", names)
}
}
func TestRegistry_RegisterReplaces(t *testing.T) {
r := harness.NewRegistry()
first := stub.New()
first.NameStr = "stub"
second := stub.New()
second.NameStr = "stub"
r.Register(first)
r.Register(second)
got, err := r.Get("stub")
if err != nil {
t.Fatalf("Get error: %v", err)
}
if got != second {
t.Fatalf("replacement failed: got %v, want %v", got, second)
}
}
func TestRegistry_Resolve_K8sImageWinsWhenRegistered(t *testing.T) {
r := harness.NewRegistry()
k8s := stub.New()
k8s.NameStr = "k8sjob"
web := stub.New()
web.NameStr = "webhook"
r.Register(k8s)
r.Register(web)
a := &agents.Agent{Name: "a", K8sImage: "ghcr.io/foo/bar:v1"}
got, err := r.Resolve(a)
if err != nil {
t.Fatalf("Resolve error: %v", err)
}
if got != k8s {
t.Fatalf("Resolve picked %v, want k8sjob", got.Name())
}
}
func TestRegistry_Resolve_FallsBackToWebhook(t *testing.T) {
r := harness.NewRegistry()
web := stub.New()
web.NameStr = "webhook"
r.Register(web)
a := &agents.Agent{Name: "a"}
got, err := r.Resolve(a)
if err != nil {
t.Fatalf("Resolve error: %v", err)
}
if got != web {
t.Fatalf("Resolve picked %v, want webhook", got.Name())
}
}
func TestRegistry_Resolve_FallsBackToSubprocess(t *testing.T) {
r := harness.NewRegistry()
sub := stub.New()
sub.NameStr = "subprocess"
r.Register(sub)
a := &agents.Agent{Name: "a"}
got, err := r.Resolve(a)
if err != nil {
t.Fatalf("Resolve error: %v", err)
}
if got != sub {
t.Fatalf("Resolve picked %v, want subprocess", got.Name())
}
}
func TestRegistry_Resolve_ErrNoBackend(t *testing.T) {
r := harness.NewRegistry()
a := &agents.Agent{Name: "a"}
if _, err := r.Resolve(a); !errors.Is(err, harness.ErrNoBackend) {
t.Fatalf("err = %v, want ErrNoBackend", err)
}
}
func TestRegistry_Resolve_K8sImageSetButNoK8sBackend_FallsThrough(t *testing.T) {
// Agent has k8s_image but no k8sjob backend is registered — should
// fall through to webhook/subprocess, not fail.
r := harness.NewRegistry()
web := stub.New()
web.NameStr = "webhook"
r.Register(web)
a := &agents.Agent{Name: "a", K8sImage: "foo:v1"}
got, err := r.Resolve(a)
if err != nil {
t.Fatalf("Resolve error: %v", err)
}
if got != web {
t.Fatalf("Resolve picked %v, want webhook", got.Name())
}
}
func TestRegistry_Resolve_CustomResolveFn(t *testing.T) {
r := harness.NewRegistry()
a := stub.New()
a.NameStr = "a"
b := stub.New()
b.NameStr = "b"
r.Register(a)
r.Register(b)
r.ResolveFn = func(r *harness.Registry, agent *agents.Agent) (harness.Harness, error) {
return r.Get("b")
}
got, err := r.Resolve(&agents.Agent{Name: "x", K8sImage: "foo:v1"})
if err != nil {
t.Fatalf("Resolve error: %v", err)
}
if got != b {
t.Fatalf("Resolve picked %v, want b", got.Name())
}
}
func TestRegistry_Execute_DelegatesToResolvedHarness(t *testing.T) {
r := harness.NewRegistry()
s := stub.New()
s.NameStr = "subprocess"
r.Register(s)
req := &harness.ExecRequest{RunID: "run-1", AgentName: "a"}
res, err := r.Execute(context.Background(), &agents.Agent{Name: "a"}, req)
if err != nil {
t.Fatalf("Execute error: %v", err)
}
if res == nil {
t.Fatal("Execute returned nil result")
}
if len(s.Calls) != 1 {
t.Fatalf("stub got %d calls, want 1", len(s.Calls))
}
if s.Calls[0].RunID != "run-1" {
t.Fatalf("stub call RunID = %q, want run-1", s.Calls[0].RunID)
}
}
func TestRegistry_Execute_PropagatesError(t *testing.T) {
r := harness.NewRegistry()
s := stub.New()
s.NameStr = "subprocess"
wantErr := errors.New("boom")
s.Err = wantErr
r.Register(s)
_, err := r.Execute(context.Background(), &agents.Agent{Name: "a"}, &harness.ExecRequest{RunID: "r"})
if !errors.Is(err, wantErr) {
t.Fatalf("err = %v, want %v", err, wantErr)
}
}
func TestRegistry_Execute_NoBackend(t *testing.T) {
r := harness.NewRegistry()
_, err := r.Execute(context.Background(), &agents.Agent{Name: "a"}, &harness.ExecRequest{RunID: "r"})
if !errors.Is(err, harness.ErrNoBackend) {
t.Fatalf("err = %v, want ErrNoBackend", err)
}
}
+335
View File
@@ -0,0 +1,335 @@
// Package runs persists harness execution records to SQLite. It
// implements harness.Observer so it can be wired into the Registry
// without the harness core depending on storage.
package runs
import (
"context"
"database/sql"
"fmt"
"log/slog"
"sync"
"time"
"github.com/synapbus/synapbus/internal/agents"
"github.com/synapbus/synapbus/internal/harness"
"github.com/synapbus/synapbus/internal/observability"
)
// Run is one row of the harness_runs table.
type Run struct {
ID int64
RunID string
AgentName string
Backend string
MessageID *int64
Status string
ExitCode *int
TraceID string
SpanID string
SessionID string
TokensIn int64
TokensOut int64
TokensCached int64
CostUSD float64
DurationMs *int64
ResultJSON string
LogsExcerpt string
CreatedAt time.Time
FinishedAt *time.Time
}
// Status constants match the harness_runs.status column domain.
const (
StatusPending = "pending"
StatusRunning = "running"
StatusSuccess = "success"
StatusFailed = "failed"
StatusCancelled = "cancelled"
StatusTimeout = "timeout"
)
// Store is the SQLite-backed harness_runs store. It also tracks start
// timestamps in memory so OnFinish can compute duration without needing
// the caller to pass it.
type Store struct {
db *sql.DB
logger *slog.Logger
mu sync.Mutex
start map[string]time.Time // runID → start time
}
// New constructs a Store. Safe for concurrent use.
func New(db *sql.DB, logger *slog.Logger) *Store {
if logger == nil {
logger = slog.Default()
}
return &Store{
db: db,
logger: logger.With("component", "harness-runs"),
start: map[string]time.Time{},
}
}
// Compile-time check: Store satisfies harness.Observer.
var _ harness.Observer = (*Store)(nil)
// OnStart writes a 'running' row for the run. Errors are logged, not
// returned, so storage issues never block Execute.
func (s *Store) OnStart(ctx context.Context, agent *agents.Agent, harnessName string, req *harness.ExecRequest) {
s.mu.Lock()
s.start[req.RunID] = time.Now().UTC()
s.mu.Unlock()
var msgID *int64
if req.Message != nil {
id := req.Message.ID
msgID = &id
}
_, err := s.db.ExecContext(ctx,
`INSERT INTO harness_runs (run_id, agent_name, backend, message_id, status, trace_id, session_id, created_at)
VALUES (?, ?, ?, ?, ?, ?, ?, CURRENT_TIMESTAMP)`,
req.RunID,
agentNameOf(agent),
harnessName,
msgID,
StatusRunning,
observability.TraceIDFromContext(ctx),
req.SessionID,
)
if err != nil {
s.logger.Warn("harness_runs insert failed",
"run_id", req.RunID,
"agent", agentNameOf(agent),
"error", err,
)
}
}
// OnFinish updates the row with the terminal status, usage, and logs.
func (s *Store) OnFinish(ctx context.Context, agent *agents.Agent, harnessName string, req *harness.ExecRequest, res *harness.ExecResult, execErr error) {
s.mu.Lock()
startedAt, ok := s.start[req.RunID]
delete(s.start, req.RunID)
s.mu.Unlock()
var durationMs *int64
if ok {
d := time.Since(startedAt).Milliseconds()
durationMs = &d
}
status := StatusSuccess
var exitCode *int
logsExcerpt := ""
var resultJSON string
var tokensIn, tokensOut, tokensCached int64
var costUSD float64
sessionID := req.SessionID
if res != nil {
ec := res.ExitCode
exitCode = &ec
logsExcerpt = res.Logs
if len(res.ResultJSON) > 0 {
resultJSON = string(res.ResultJSON)
}
tokensIn = res.Usage.TokensIn
tokensOut = res.Usage.TokensOut
tokensCached = res.Usage.TokensCached
costUSD = res.Usage.CostUSD
if res.SessionID != "" {
sessionID = res.SessionID
}
if res.ExitCode != 0 {
status = StatusFailed
}
}
if execErr != nil {
status = StatusFailed
}
// Cap logs excerpt to keep row sizes sane (matches design: bounded).
const logsCap = 16 * 1024
if len(logsExcerpt) > logsCap {
logsExcerpt = "... [truncated] ...\n" + logsExcerpt[len(logsExcerpt)-logsCap:]
}
traceID := ""
if res != nil && res.TraceID != "" {
traceID = res.TraceID
} else {
traceID = observability.TraceIDFromContext(ctx)
}
// If the run was never inserted (OnStart failed or skipped), fall
// back to an UPSERT via INSERT OR REPLACE on run_id to avoid losing
// the terminal row.
query := `UPDATE harness_runs SET
status = ?, exit_code = ?, trace_id = ?, session_id = ?,
tokens_in = ?, tokens_out = ?, tokens_cached = ?, cost_usd = ?,
duration_ms = ?, result_json = ?, logs_excerpt = ?,
finished_at = CURRENT_TIMESTAMP
WHERE run_id = ?`
result, err := s.db.ExecContext(ctx, query,
status, exitCode, traceID, sessionID,
tokensIn, tokensOut, tokensCached, costUSD,
durationMs, nullableString(resultJSON), nullableString(logsExcerpt),
req.RunID,
)
if err == nil {
if n, _ := result.RowsAffected(); n == 0 {
// Row didn't exist — insert it fresh.
_, err = s.db.ExecContext(ctx,
`INSERT INTO harness_runs (run_id, agent_name, backend, status, exit_code, trace_id, session_id,
tokens_in, tokens_out, tokens_cached, cost_usd, duration_ms, result_json, logs_excerpt, created_at, finished_at)
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, CURRENT_TIMESTAMP, CURRENT_TIMESTAMP)`,
req.RunID, agentNameOf(agent), harnessName,
status, exitCode, traceID, sessionID,
tokensIn, tokensOut, tokensCached, costUSD,
durationMs, nullableString(resultJSON), nullableString(logsExcerpt),
)
}
}
if err != nil {
s.logger.Warn("harness_runs update failed",
"run_id", req.RunID, "error", err,
)
}
}
// GetByRunID retrieves a single harness run by its caller-assigned id.
func (s *Store) GetByRunID(ctx context.Context, runID string) (*Run, error) {
row := s.db.QueryRowContext(ctx, selectSQL()+` WHERE run_id = ?`, runID)
return scanRun(row)
}
// ListByAgent returns recent runs for an agent, newest first.
func (s *Store) ListByAgent(ctx context.Context, agentName string, limit int) ([]*Run, error) {
if limit <= 0 {
limit = 50
}
rows, err := s.db.QueryContext(ctx,
selectSQL()+` WHERE agent_name = ? ORDER BY created_at DESC LIMIT ?`,
agentName, limit,
)
if err != nil {
return nil, fmt.Errorf("harness_runs list: %w", err)
}
defer rows.Close()
return scanRuns(rows)
}
// -- helpers --------------------------------------------------------------
func agentNameOf(a *agents.Agent) string {
if a == nil {
return ""
}
return a.Name
}
func nullableString(s string) any {
if s == "" {
return nil
}
return s
}
func selectSQL() string {
return `SELECT id, run_id, agent_name, backend, message_id, status, exit_code,
trace_id, span_id, session_id, tokens_in, tokens_out, tokens_cached, cost_usd,
duration_ms, result_json, logs_excerpt, created_at, finished_at
FROM harness_runs`
}
func scanRun(row *sql.Row) (*Run, error) {
var r Run
var msgID sql.NullInt64
var exitCode sql.NullInt64
var traceID, spanID, sessionID, resultJSON, logsExcerpt sql.NullString
var durationMs sql.NullInt64
var createdAt, finishedAt sql.NullTime
if err := row.Scan(
&r.ID, &r.RunID, &r.AgentName, &r.Backend, &msgID, &r.Status, &exitCode,
&traceID, &spanID, &sessionID, &r.TokensIn, &r.TokensOut, &r.TokensCached, &r.CostUSD,
&durationMs, &resultJSON, &logsExcerpt, &createdAt, &finishedAt,
); err != nil {
return nil, err
}
if msgID.Valid {
v := msgID.Int64
r.MessageID = &v
}
if exitCode.Valid {
v := int(exitCode.Int64)
r.ExitCode = &v
}
r.TraceID = traceID.String
r.SpanID = spanID.String
r.SessionID = sessionID.String
r.ResultJSON = resultJSON.String
r.LogsExcerpt = logsExcerpt.String
if durationMs.Valid {
v := durationMs.Int64
r.DurationMs = &v
}
if createdAt.Valid {
r.CreatedAt = createdAt.Time
}
if finishedAt.Valid {
t := finishedAt.Time
r.FinishedAt = &t
}
return &r, nil
}
func scanRuns(rows *sql.Rows) ([]*Run, error) {
var out []*Run
for rows.Next() {
var r Run
var msgID sql.NullInt64
var exitCode sql.NullInt64
var traceID, spanID, sessionID, resultJSON, logsExcerpt sql.NullString
var durationMs sql.NullInt64
var createdAt, finishedAt sql.NullTime
if err := rows.Scan(
&r.ID, &r.RunID, &r.AgentName, &r.Backend, &msgID, &r.Status, &exitCode,
&traceID, &spanID, &sessionID, &r.TokensIn, &r.TokensOut, &r.TokensCached, &r.CostUSD,
&durationMs, &resultJSON, &logsExcerpt, &createdAt, &finishedAt,
); err != nil {
return nil, err
}
if msgID.Valid {
v := msgID.Int64
r.MessageID = &v
}
if exitCode.Valid {
v := int(exitCode.Int64)
r.ExitCode = &v
}
r.TraceID = traceID.String
r.SpanID = spanID.String
r.SessionID = sessionID.String
r.ResultJSON = resultJSON.String
r.LogsExcerpt = logsExcerpt.String
if durationMs.Valid {
v := durationMs.Int64
r.DurationMs = &v
}
if createdAt.Valid {
r.CreatedAt = createdAt.Time
}
if finishedAt.Valid {
t := finishedAt.Time
r.FinishedAt = &t
}
out = append(out, &r)
}
if out == nil {
out = []*Run{}
}
return out, rows.Err()
}
+232
View File
@@ -0,0 +1,232 @@
package runs_test
import (
"context"
"database/sql"
"encoding/json"
"errors"
"testing"
"github.com/synapbus/synapbus/internal/agents"
"github.com/synapbus/synapbus/internal/harness"
"github.com/synapbus/synapbus/internal/harness/runs"
"github.com/synapbus/synapbus/internal/messaging"
_ "modernc.org/sqlite"
)
// setupDB installs just the harness_runs schema on an in-memory DB.
// It mirrors migration 019_harness.sql (the parts this store reads).
func setupDB(t *testing.T) *sql.DB {
t.Helper()
db, err := sql.Open("sqlite", ":memory:")
if err != nil {
t.Fatalf("open: %v", err)
}
schema := `
CREATE TABLE harness_runs (
id INTEGER PRIMARY KEY AUTOINCREMENT,
run_id TEXT NOT NULL UNIQUE,
agent_name TEXT NOT NULL,
backend TEXT NOT NULL,
message_id INTEGER,
status TEXT NOT NULL,
exit_code INTEGER,
trace_id TEXT,
span_id TEXT,
session_id TEXT,
tokens_in INTEGER NOT NULL DEFAULT 0,
tokens_out INTEGER NOT NULL DEFAULT 0,
tokens_cached INTEGER NOT NULL DEFAULT 0,
cost_usd REAL NOT NULL DEFAULT 0,
duration_ms INTEGER,
result_json TEXT,
logs_excerpt TEXT,
created_at DATETIME NOT NULL DEFAULT CURRENT_TIMESTAMP,
finished_at DATETIME
);`
if _, err := db.Exec(schema); err != nil {
t.Fatalf("schema: %v", err)
}
return db
}
func agent(name string) *agents.Agent { return &agents.Agent{Name: name} }
func TestStore_OnStart_InsertsRunningRow(t *testing.T) {
db := setupDB(t)
s := runs.New(db, nil)
req := &harness.ExecRequest{
RunID: "run-1",
AgentName: "alpha",
Message: &messaging.Message{ID: 99, FromAgent: "caller"},
}
s.OnStart(context.Background(), agent("alpha"), "subprocess", req)
got, err := s.GetByRunID(context.Background(), "run-1")
if err != nil {
t.Fatalf("GetByRunID err = %v", err)
}
if got.Status != runs.StatusRunning {
t.Errorf("status = %q, want running", got.Status)
}
if got.Backend != "subprocess" {
t.Errorf("backend = %q", got.Backend)
}
if got.MessageID == nil || *got.MessageID != 99 {
t.Errorf("MessageID = %v, want 99", got.MessageID)
}
}
func TestStore_OnFinish_UpdatesToSuccess(t *testing.T) {
db := setupDB(t)
s := runs.New(db, nil)
req := &harness.ExecRequest{RunID: "r", AgentName: "a"}
s.OnStart(context.Background(), agent("a"), "stub", req)
res := &harness.ExecResult{
ExitCode: 0,
Logs: "hi",
ResultJSON: json.RawMessage(`{"ok":true}`),
Usage: harness.Usage{
TokensIn: 10,
TokensOut: 20,
CostUSD: 0.005,
},
SessionID: "sess-1",
}
s.OnFinish(context.Background(), agent("a"), "stub", req, res, nil)
got, err := s.GetByRunID(context.Background(), "r")
if err != nil {
t.Fatalf("GetByRunID err = %v", err)
}
if got.Status != runs.StatusSuccess {
t.Errorf("status = %q, want success", got.Status)
}
if got.ExitCode == nil || *got.ExitCode != 0 {
t.Errorf("ExitCode = %v", got.ExitCode)
}
if got.TokensIn != 10 || got.TokensOut != 20 {
t.Errorf("usage not persisted: %+v", got)
}
if got.CostUSD != 0.005 {
t.Errorf("cost not persisted: %v", got.CostUSD)
}
if got.ResultJSON != `{"ok":true}` {
t.Errorf("result_json = %q", got.ResultJSON)
}
if got.SessionID != "sess-1" {
t.Errorf("session_id = %q", got.SessionID)
}
if got.DurationMs == nil {
t.Error("DurationMs nil, want computed")
}
if got.FinishedAt == nil {
t.Error("FinishedAt nil")
}
}
func TestStore_OnFinish_UpdatesToFailedOnNonZero(t *testing.T) {
db := setupDB(t)
s := runs.New(db, nil)
req := &harness.ExecRequest{RunID: "r", AgentName: "a"}
s.OnStart(context.Background(), agent("a"), "stub", req)
s.OnFinish(context.Background(), agent("a"), "stub", req,
&harness.ExecResult{ExitCode: 5, Logs: "bad"}, nil)
got, _ := s.GetByRunID(context.Background(), "r")
if got.Status != runs.StatusFailed {
t.Errorf("status = %q, want failed", got.Status)
}
}
func TestStore_OnFinish_UpdatesToFailedOnExecErr(t *testing.T) {
db := setupDB(t)
s := runs.New(db, nil)
req := &harness.ExecRequest{RunID: "r", AgentName: "a"}
s.OnStart(context.Background(), agent("a"), "stub", req)
s.OnFinish(context.Background(), agent("a"), "stub", req, nil, errors.New("kaboom"))
got, _ := s.GetByRunID(context.Background(), "r")
if got.Status != runs.StatusFailed {
t.Errorf("status = %q, want failed", got.Status)
}
}
func TestStore_OnFinish_InsertsIfNoStart(t *testing.T) {
// Simulate the OnStart insert failing (e.g. store wired late) —
// OnFinish must still persist a terminal row.
db := setupDB(t)
s := runs.New(db, nil)
req := &harness.ExecRequest{RunID: "ghost", AgentName: "a"}
s.OnFinish(context.Background(), agent("a"), "stub", req,
&harness.ExecResult{ExitCode: 0, Logs: "post hoc"}, nil)
got, err := s.GetByRunID(context.Background(), "ghost")
if err != nil {
t.Fatalf("GetByRunID err = %v", err)
}
if got.Status != runs.StatusSuccess {
t.Errorf("status = %q", got.Status)
}
if got.LogsExcerpt != "post hoc" {
t.Errorf("logs = %q", got.LogsExcerpt)
}
}
func TestStore_ListByAgent_OrdersNewestFirst(t *testing.T) {
db := setupDB(t)
s := runs.New(db, nil)
for _, id := range []string{"r1", "r2", "r3"} {
req := &harness.ExecRequest{RunID: id, AgentName: "a"}
s.OnStart(context.Background(), agent("a"), "stub", req)
s.OnFinish(context.Background(), agent("a"), "stub", req,
&harness.ExecResult{ExitCode: 0}, nil)
}
list, err := s.ListByAgent(context.Background(), "a", 10)
if err != nil {
t.Fatalf("ListByAgent err = %v", err)
}
if len(list) != 3 {
t.Fatalf("got %d runs, want 3", len(list))
}
// Newest first ordering — at minimum the run_ids should all be
// present; ordering within same-timestamp is DB-defined.
seen := map[string]bool{}
for _, r := range list {
seen[r.RunID] = true
}
for _, id := range []string{"r1", "r2", "r3"} {
if !seen[id] {
t.Errorf("run %s missing", id)
}
}
}
func TestStore_LogsExcerptCap(t *testing.T) {
db := setupDB(t)
s := runs.New(db, nil)
req := &harness.ExecRequest{RunID: "big", AgentName: "a"}
s.OnStart(context.Background(), agent("a"), "stub", req)
big := make([]byte, 32*1024)
for i := range big {
big[i] = 'x'
}
s.OnFinish(context.Background(), agent("a"), "stub", req,
&harness.ExecResult{ExitCode: 0, Logs: string(big)}, nil)
got, _ := s.GetByRunID(context.Background(), "big")
if len(got.LogsExcerpt) == 0 {
t.Fatal("logs empty")
}
if len(got.LogsExcerpt) > 17*1024 {
t.Errorf("logs excerpt not capped: %d bytes", len(got.LogsExcerpt))
}
}
+85
View File
@@ -0,0 +1,85 @@
// Package stub provides an in-memory Harness implementation used by tests
// and by development modes where no real execution backend is wanted.
package stub
import (
"context"
"encoding/json"
"sync"
"time"
"github.com/synapbus/synapbus/internal/agents"
"github.com/synapbus/synapbus/internal/harness"
)
// Harness is a canned, in-memory implementation of harness.Harness.
// Callers set Result / Err directly (under Mu) to control what Execute
// returns. Every call is recorded in Calls for assertions.
type Harness struct {
Mu sync.Mutex
NameStr string
Caps harness.Capabilities
PreflightErr error
ProvisionErr error
Result *harness.ExecResult
Err error
ExecDelay time.Duration
Calls []harness.ExecRequest
CancelCalls []string
}
// New returns a stub harness ready to use. The default result is a
// successful zero-exit-code run with empty logs.
func New() *Harness {
return &Harness{
NameStr: "stub",
Caps: harness.Capabilities{OTelNative: false, MaxConcurrency: 1},
Result: &harness.ExecResult{
ExitCode: 0,
Logs: "",
ResultJSON: json.RawMessage(`{"ok":true}`),
},
}
}
func (h *Harness) Name() string { return h.NameStr }
func (h *Harness) Capabilities() harness.Capabilities { return h.Caps }
func (h *Harness) TestEnvironment(ctx context.Context) error {
return h.PreflightErr
}
func (h *Harness) Provision(ctx context.Context, agent *agents.Agent) error {
return h.ProvisionErr
}
func (h *Harness) Execute(ctx context.Context, req *harness.ExecRequest) (*harness.ExecResult, error) {
h.Mu.Lock()
// Copy the request value so later caller mutation cannot race with
// test assertions reading Calls.
reqCopy := *req
h.Calls = append(h.Calls, reqCopy)
delay := h.ExecDelay
result := h.Result
err := h.Err
h.Mu.Unlock()
if delay > 0 {
select {
case <-time.After(delay):
case <-ctx.Done():
return nil, ctx.Err()
}
}
if err != nil {
return nil, err
}
return result, nil
}
func (h *Harness) Cancel(ctx context.Context, runID string) error {
h.Mu.Lock()
h.CancelCalls = append(h.CancelCalls, runID)
h.Mu.Unlock()
return nil
}
+104
View File
@@ -0,0 +1,104 @@
package stub_test
import (
"context"
"errors"
"testing"
"time"
"github.com/synapbus/synapbus/internal/agents"
"github.com/synapbus/synapbus/internal/harness"
"github.com/synapbus/synapbus/internal/harness/stub"
)
func TestStub_ImplementsHarness(t *testing.T) {
var _ harness.Harness = (*stub.Harness)(nil)
}
func TestStub_NameAndCapabilities(t *testing.T) {
s := stub.New()
if s.Name() != "stub" {
t.Fatalf("Name = %q, want stub", s.Name())
}
caps := s.Capabilities()
if caps.OTelNative {
t.Fatalf("stub should not claim OTelNative by default")
}
}
func TestStub_Execute_ReturnsConfiguredResult(t *testing.T) {
s := stub.New()
s.Result = &harness.ExecResult{ExitCode: 42, Logs: "hi"}
got, err := s.Execute(context.Background(), &harness.ExecRequest{RunID: "r", AgentName: "a"})
if err != nil {
t.Fatalf("Execute error: %v", err)
}
if got.ExitCode != 42 || got.Logs != "hi" {
t.Fatalf("got %+v, want ExitCode=42 Logs=hi", got)
}
if len(s.Calls) != 1 {
t.Fatalf("Calls len = %d, want 1", len(s.Calls))
}
if s.Calls[0].AgentName != "a" {
t.Fatalf("recorded call agent = %q, want a", s.Calls[0].AgentName)
}
}
func TestStub_Execute_ReturnsError(t *testing.T) {
s := stub.New()
wantErr := errors.New("nope")
s.Err = wantErr
_, err := s.Execute(context.Background(), &harness.ExecRequest{})
if !errors.Is(err, wantErr) {
t.Fatalf("err = %v, want %v", err, wantErr)
}
}
func TestStub_Execute_RespectsContextCancel(t *testing.T) {
s := stub.New()
s.ExecDelay = 2 * time.Second
ctx, cancel := context.WithTimeout(context.Background(), 20*time.Millisecond)
defer cancel()
start := time.Now()
_, err := s.Execute(ctx, &harness.ExecRequest{})
if !errors.Is(err, context.DeadlineExceeded) {
t.Fatalf("err = %v, want DeadlineExceeded", err)
}
if time.Since(start) > time.Second {
t.Fatal("context cancel did not interrupt delay")
}
}
func TestStub_Cancel_RecordsRunID(t *testing.T) {
s := stub.New()
if err := s.Cancel(context.Background(), "r-1"); err != nil {
t.Fatalf("Cancel error: %v", err)
}
if len(s.CancelCalls) != 1 || s.CancelCalls[0] != "r-1" {
t.Fatalf("CancelCalls = %v, want [r-1]", s.CancelCalls)
}
}
func TestStub_Provision_ReturnsConfiguredError(t *testing.T) {
s := stub.New()
s.ProvisionErr = errors.New("bad")
err := s.Provision(context.Background(), &agents.Agent{Name: "a"})
if err == nil || err.Error() != "bad" {
t.Fatalf("Provision err = %v, want 'bad'", err)
}
}
func TestStub_TestEnvironment_ReturnsConfiguredError(t *testing.T) {
s := stub.New()
if err := s.TestEnvironment(context.Background()); err != nil {
t.Fatalf("default TestEnvironment err = %v, want nil", err)
}
s.PreflightErr = errors.New("nope")
if err := s.TestEnvironment(context.Background()); err == nil {
t.Fatal("expected error")
}
}
+364
View File
@@ -0,0 +1,364 @@
// Package subprocess is the local-process implementation of
// harness.Harness. It runs an agent as a child process of synapbus
// using os/exec, captures stdout+stderr, and reads an optional
// result.json the child may have written into its workdir.
//
// Works on Mac and Linux identically (no CGO, no platform-specific
// syscalls). Windows is not targeted because synapbus itself does not
// target Windows.
package subprocess
import (
"bytes"
"context"
"encoding/json"
"errors"
"fmt"
"io"
"log/slog"
"os"
"os/exec"
"path/filepath"
"strings"
"time"
"github.com/synapbus/synapbus/internal/agents"
"github.com/synapbus/synapbus/internal/harness"
)
// Config tunes the subprocess harness. Zero-value defaults are sensible
// for development; production callers typically set BaseDir to a
// predictable location under SYNAPBUS_DATA_DIR so forensics are easy.
type Config struct {
// BaseDir is the parent directory under which a per-run workdir is
// created. Defaults to os.TempDir() when empty.
BaseDir string
// LogsCap bounds the number of bytes kept in ExecResult.Logs. The
// full stdout/stderr stream is written to `stdout.log` /
// `stderr.log` inside the workdir for forensics. Defaults to 64 KiB.
LogsCap int
// KeepWorkdirOnSuccess leaves the workdir behind even for
// zero-exit runs. Useful for debugging test flakes.
KeepWorkdirOnSuccess bool
}
// Harness runs agents as local subprocesses.
type Harness struct {
cfg Config
logger *slog.Logger
}
// New constructs a subprocess harness. Pass zero Config for defaults.
func New(cfg Config, logger *slog.Logger) *Harness {
if cfg.LogsCap <= 0 {
cfg.LogsCap = 64 * 1024
}
if logger == nil {
logger = slog.Default()
}
return &Harness{
cfg: cfg,
logger: logger.With("harness", "subprocess"),
}
}
// Name returns the registered harness name.
func (h *Harness) Name() string { return "subprocess" }
// Capabilities advertises backend features.
func (h *Harness) Capabilities() harness.Capabilities {
return harness.Capabilities{
SystemPrompt: true,
SessionResume: true,
Skills: false,
OTelNative: true,
MaxConcurrency: 4,
}
}
// TestEnvironment is a cheap sanity check: BaseDir (or os.TempDir) must
// exist and be writable. Per-agent binary reachability is checked at
// Execute time because the binary is agent-specific.
func (h *Harness) TestEnvironment(ctx context.Context) error {
base := h.cfg.BaseDir
if base == "" {
base = os.TempDir()
}
info, err := os.Stat(base)
if err != nil {
return fmt.Errorf("subprocess: base dir %q: %w", base, err)
}
if !info.IsDir() {
return fmt.Errorf("subprocess: base dir %q is not a directory", base)
}
return nil
}
// Provision is a no-op for subprocess. All per-run state goes into the
// workdir Execute creates on the fly.
func (h *Harness) Provision(ctx context.Context, agent *agents.Agent) error {
return nil
}
// ErrNoLocalCommand is returned when an agent has no LocalCommand
// configured so the subprocess backend cannot know what to run.
var ErrNoLocalCommand = errors.New("subprocess: agent has no local_command configured")
// Execute launches the child, waits for it to exit (or for ctx /
// Budget to fire), and returns its output.
func (h *Harness) Execute(ctx context.Context, req *harness.ExecRequest) (*harness.ExecResult, error) {
if req == nil {
return nil, errors.New("subprocess: nil ExecRequest")
}
if req.Agent == nil {
return nil, errors.New("subprocess: ExecRequest.Agent is required")
}
argv, err := parseLocalCommand(req.Agent.LocalCommand)
if err != nil {
return nil, err
}
// Per-run workdir
base := h.cfg.BaseDir
if base == "" {
base = os.TempDir()
}
if err := os.MkdirAll(base, 0o755); err != nil {
return nil, fmt.Errorf("subprocess: mkdir base: %w", err)
}
runDirName := sanitizeRunDir(req.RunID)
if runDirName == "" {
runDirName = fmt.Sprintf("run-%d", time.Now().UnixNano())
}
workdir := filepath.Join(base, runDirName)
if err := os.MkdirAll(workdir, 0o755); err != nil {
return nil, fmt.Errorf("subprocess: mkdir workdir: %w", err)
}
// Context with Budget timeout if set.
runCtx := ctx
if req.Budget.MaxWallClock > 0 {
var cancel context.CancelFunc
runCtx, cancel = context.WithTimeout(ctx, req.Budget.MaxWallClock)
defer cancel()
}
// Write the triggering message to message.json for the child to
// read if it cares. Simple, explicit, no stdin-piping ambiguity.
if req.Message != nil {
raw, _ := json.Marshal(req.Message)
_ = os.WriteFile(filepath.Join(workdir, "message.json"), raw, 0o644)
}
cmd := exec.CommandContext(runCtx, argv[0], argv[1:]...)
cmd.Dir = workdir
cmd.Env = buildEnv(req, workdir)
var stdout, stderr bytes.Buffer
cmd.Stdout = io.MultiWriter(&stdout, limitedFileWriter(workdir, "stdout.log"))
cmd.Stderr = io.MultiWriter(&stderr, limitedFileWriter(workdir, "stderr.log"))
h.logger.Info("subprocess launching",
"run_id", req.RunID,
"agent", req.AgentName,
"cmd", argv[0],
"workdir", workdir,
)
startedAt := time.Now()
runErr := cmd.Run()
duration := time.Since(startedAt)
exitCode := 0
if runErr != nil {
var exitErr *exec.ExitError
if errors.As(runErr, &exitErr) {
exitCode = exitErr.ExitCode()
} else {
exitCode = 1
}
}
// Load optional result.json
var resultJSON json.RawMessage
if raw, readErr := os.ReadFile(filepath.Join(workdir, "result.json")); readErr == nil && len(raw) > 0 {
if json.Valid(raw) {
resultJSON = json.RawMessage(raw)
}
}
logs := mergeLogs(&stdout, &stderr, h.cfg.LogsCap)
// Cleanup policy: remove workdir on success unless configured to
// keep it; always keep on failure so users can inspect stdout.log /
// stderr.log / message.json / result.json.
if exitCode == 0 && !h.cfg.KeepWorkdirOnSuccess {
_ = os.RemoveAll(workdir)
}
h.logger.Info("subprocess finished",
"run_id", req.RunID,
"agent", req.AgentName,
"exit", exitCode,
"duration_ms", duration.Milliseconds(),
)
result := &harness.ExecResult{
ExitCode: exitCode,
Logs: logs,
ResultJSON: resultJSON,
}
// Distinguish context timeout from plain failures so the caller
// can tell "budget exceeded" from "the agent crashed".
if runErr != nil && errors.Is(runCtx.Err(), context.DeadlineExceeded) {
return result, fmt.Errorf("subprocess: wall-clock budget %s exceeded", req.Budget.MaxWallClock)
}
if runErr != nil && errors.Is(runCtx.Err(), context.Canceled) {
return result, runCtx.Err()
}
return result, nil
}
// Cancel is a no-op for subprocess today — cancellation happens via the
// context passed to Execute. A future iteration could track in-flight
// runs by RunID and send SIGTERM to them.
func (h *Harness) Cancel(ctx context.Context, runID string) error {
return nil
}
// -- helpers --------------------------------------------------------------
// parseLocalCommand accepts either a JSON array (["claude", "--print"])
// or a simple space-separated string. Returns the argv slice or an
// error if neither form parses.
func parseLocalCommand(raw string) ([]string, error) {
s := strings.TrimSpace(raw)
if s == "" {
return nil, ErrNoLocalCommand
}
if strings.HasPrefix(s, "[") {
var argv []string
if err := json.Unmarshal([]byte(s), &argv); err != nil {
return nil, fmt.Errorf("subprocess: parse local_command JSON: %w", err)
}
if len(argv) == 0 {
return nil, ErrNoLocalCommand
}
return argv, nil
}
// Fall back to whitespace split. Suitable for simple commands.
parts := strings.Fields(s)
if len(parts) == 0 {
return nil, ErrNoLocalCommand
}
return parts, nil
}
// buildEnv constructs the env var list for the child. Starts from the
// parent's environment (so HOME, PATH, credentials are inherited by
// default — matches current K8s Pod behaviour). Then layers the
// agent's k8s_env_json (for consistency between backends), then caller
// overrides, then the SYNAPBUS_* run-context variables.
func buildEnv(req *harness.ExecRequest, workdir string) []string {
env := map[string]string{}
for _, kv := range os.Environ() {
if i := strings.IndexByte(kv, '='); i >= 0 {
env[kv[:i]] = kv[i+1:]
}
}
// agent env map (K8sEnvJSON is shared across backends today)
if req.Agent != nil && req.Agent.K8sEnvJSON != "" {
var m map[string]json.RawMessage
if err := json.Unmarshal([]byte(req.Agent.K8sEnvJSON), &m); err == nil {
for k, v := range m {
var s string
if err := json.Unmarshal(v, &s); err == nil {
env[k] = s
continue
}
env[k] = strings.Trim(string(v), "\"")
}
}
}
// caller overrides
for k, v := range req.Env {
env[k] = v
}
// run context
env["SYNAPBUS_RUN_ID"] = req.RunID
env["SYNAPBUS_AGENT"] = req.AgentName
env["SYNAPBUS_WORKDIR"] = workdir
if req.Message != nil {
env["SYNAPBUS_MESSAGE_ID"] = fmt.Sprintf("%d", req.Message.ID)
env["SYNAPBUS_FROM_AGENT"] = req.Message.FromAgent
}
out := make([]string, 0, len(env))
for k, v := range env {
out = append(out, k+"="+v)
}
return out
}
// mergeLogs interleaves stdout then stderr with a header, bounded by cap.
func mergeLogs(out, errb *bytes.Buffer, cap int) string {
var b strings.Builder
if out.Len() > 0 {
b.WriteString(out.String())
}
if errb.Len() > 0 {
if b.Len() > 0 {
b.WriteString("\n")
}
b.WriteString("-- stderr --\n")
b.WriteString(errb.String())
}
s := b.String()
if cap > 0 && len(s) > cap {
// Keep the tail — most informative on failure.
s = "... [truncated " + fmt.Sprintf("%d", len(s)-cap) + " bytes] ...\n" + s[len(s)-cap:]
}
return s
}
// limitedFileWriter returns a writer that appends to a file inside the
// workdir. Errors are silently ignored — logs are best-effort and must
// not fail the run.
func limitedFileWriter(workdir, name string) io.Writer {
f, err := os.OpenFile(filepath.Join(workdir, name), os.O_CREATE|os.O_WRONLY|os.O_TRUNC, 0o644)
if err != nil {
return io.Discard
}
return f
}
// sanitizeRunDir strips characters that cause surprises on case-folded
// filesystems or in shell globs. Keeps the resulting name readable.
func sanitizeRunDir(runID string) string {
if runID == "" {
return ""
}
var b strings.Builder
for _, r := range runID {
switch {
case r >= 'a' && r <= 'z', r >= 'A' && r <= 'Z', r >= '0' && r <= '9':
b.WriteRune(r)
case r == '-' || r == '_':
b.WriteRune(r)
default:
b.WriteByte('-')
}
}
name := b.String()
if len(name) > 64 {
name = name[:64]
}
return name
}
@@ -0,0 +1,323 @@
package subprocess_test
import (
"context"
"os"
"path/filepath"
"runtime"
"strings"
"testing"
"time"
"github.com/synapbus/synapbus/internal/agents"
"github.com/synapbus/synapbus/internal/harness"
"github.com/synapbus/synapbus/internal/harness/subprocess"
"github.com/synapbus/synapbus/internal/messaging"
)
func requirePosix(t *testing.T) {
t.Helper()
if runtime.GOOS == "windows" {
t.Skip("subprocess tests use /bin/sh; not applicable on Windows")
}
}
func newAgent(cmd string) *agents.Agent {
return &agents.Agent{Name: "test-agent", LocalCommand: cmd}
}
func TestSubprocess_ImplementsInterface(t *testing.T) {
var _ harness.Harness = (*subprocess.Harness)(nil)
}
func TestSubprocess_NameAndCapabilities(t *testing.T) {
h := subprocess.New(subprocess.Config{}, nil)
if h.Name() != "subprocess" {
t.Fatalf("Name = %q", h.Name())
}
caps := h.Capabilities()
if !caps.SystemPrompt || !caps.SessionResume || !caps.OTelNative {
t.Fatalf("caps missing expected flags: %+v", caps)
}
}
func TestSubprocess_TestEnvironment_OK(t *testing.T) {
h := subprocess.New(subprocess.Config{BaseDir: t.TempDir()}, nil)
if err := h.TestEnvironment(context.Background()); err != nil {
t.Fatalf("TestEnvironment err = %v", err)
}
}
func TestSubprocess_TestEnvironment_MissingDir(t *testing.T) {
h := subprocess.New(subprocess.Config{BaseDir: "/nonexistent/does/not/exist/xyz"}, nil)
if err := h.TestEnvironment(context.Background()); err == nil {
t.Fatal("expected error for missing base dir")
}
}
func TestSubprocess_Execute_Success(t *testing.T) {
requirePosix(t)
h := subprocess.New(subprocess.Config{BaseDir: t.TempDir()}, nil)
req := &harness.ExecRequest{
RunID: "run-success",
AgentName: "test",
Agent: newAgent(`["sh","-c","echo hello world"]`),
}
res, err := h.Execute(context.Background(), req)
if err != nil {
t.Fatalf("Execute err = %v", err)
}
if res.ExitCode != 0 {
t.Errorf("ExitCode = %d, want 0", res.ExitCode)
}
if !strings.Contains(res.Logs, "hello world") {
t.Errorf("logs missing stdout: %q", res.Logs)
}
}
func TestSubprocess_Execute_NonZeroExit(t *testing.T) {
requirePosix(t)
h := subprocess.New(subprocess.Config{BaseDir: t.TempDir()}, nil)
req := &harness.ExecRequest{
RunID: "run-fail",
AgentName: "test",
Agent: newAgent(`["sh","-c","echo oops; exit 7"]`),
}
res, err := h.Execute(context.Background(), req)
if err != nil {
t.Fatalf("Execute err = %v", err)
}
if res.ExitCode != 7 {
t.Errorf("ExitCode = %d, want 7", res.ExitCode)
}
if !strings.Contains(res.Logs, "oops") {
t.Errorf("logs missing stdout: %q", res.Logs)
}
}
func TestSubprocess_Execute_StderrCaptured(t *testing.T) {
requirePosix(t)
h := subprocess.New(subprocess.Config{BaseDir: t.TempDir()}, nil)
req := &harness.ExecRequest{
RunID: "run-stderr",
AgentName: "test",
Agent: newAgent(`["sh","-c","echo first; echo bad >&2; exit 0"]`),
}
res, err := h.Execute(context.Background(), req)
if err != nil {
t.Fatalf("Execute err = %v", err)
}
if !strings.Contains(res.Logs, "first") {
t.Errorf("logs missing stdout: %q", res.Logs)
}
if !strings.Contains(res.Logs, "bad") {
t.Errorf("logs missing stderr: %q", res.Logs)
}
if !strings.Contains(res.Logs, "-- stderr --") {
t.Errorf("logs missing stderr header: %q", res.Logs)
}
}
func TestSubprocess_Execute_EnvPropagation(t *testing.T) {
requirePosix(t)
h := subprocess.New(subprocess.Config{BaseDir: t.TempDir()}, nil)
req := &harness.ExecRequest{
RunID: "run-env",
AgentName: "my-agent",
Agent: newAgent(`["sh","-c","echo run=$SYNAPBUS_RUN_ID agent=$SYNAPBUS_AGENT custom=$CUSTOM"]`),
Env: map[string]string{"CUSTOM": "xyz"},
}
res, err := h.Execute(context.Background(), req)
if err != nil {
t.Fatalf("Execute err = %v", err)
}
if !strings.Contains(res.Logs, "run=run-env") {
t.Errorf("run id not propagated: %q", res.Logs)
}
if !strings.Contains(res.Logs, "agent=my-agent") {
t.Errorf("agent name not propagated: %q", res.Logs)
}
if !strings.Contains(res.Logs, "custom=xyz") {
t.Errorf("caller env not propagated: %q", res.Logs)
}
}
func TestSubprocess_Execute_ReadsResultJSON(t *testing.T) {
requirePosix(t)
baseDir := t.TempDir()
h := subprocess.New(subprocess.Config{BaseDir: baseDir, KeepWorkdirOnSuccess: true}, nil)
req := &harness.ExecRequest{
RunID: "run-json",
AgentName: "test",
Agent: newAgent(`["sh","-c","printf '{\"answer\":42,\"ok\":true}' > result.json"]`),
}
res, err := h.Execute(context.Background(), req)
if err != nil {
t.Fatalf("Execute err = %v", err)
}
if res.ExitCode != 0 {
t.Errorf("ExitCode = %d, want 0", res.ExitCode)
}
if len(res.ResultJSON) == 0 {
t.Fatal("ResultJSON empty")
}
if !strings.Contains(string(res.ResultJSON), `"answer":42`) {
t.Errorf("ResultJSON = %s", res.ResultJSON)
}
// Verify the workdir is still present (KeepWorkdirOnSuccess=true)
if _, err := os.Stat(filepath.Join(baseDir, "run-json", "result.json")); err != nil {
t.Errorf("workdir removed despite KeepWorkdirOnSuccess: %v", err)
}
}
func TestSubprocess_Execute_WritesMessageJSON(t *testing.T) {
requirePosix(t)
baseDir := t.TempDir()
h := subprocess.New(subprocess.Config{BaseDir: baseDir, KeepWorkdirOnSuccess: true}, nil)
req := &harness.ExecRequest{
RunID: "run-msg",
AgentName: "test",
Agent: newAgent(`["sh","-c","cat message.json > result.json"]`),
Message: &messaging.Message{
ID: 123,
FromAgent: "alice",
Body: "hi",
},
}
res, err := h.Execute(context.Background(), req)
if err != nil {
t.Fatalf("Execute err = %v", err)
}
if !strings.Contains(string(res.ResultJSON), `"from_agent":"alice"`) {
t.Errorf("message.json missing from_agent: %s", res.ResultJSON)
}
if !strings.Contains(string(res.ResultJSON), `"body":"hi"`) {
t.Errorf("message.json missing body: %s", res.ResultJSON)
}
}
func TestSubprocess_Execute_WorkdirCleanedOnSuccess(t *testing.T) {
requirePosix(t)
baseDir := t.TempDir()
h := subprocess.New(subprocess.Config{BaseDir: baseDir}, nil)
req := &harness.ExecRequest{
RunID: "run-clean",
AgentName: "test",
Agent: newAgent(`["sh","-c","true"]`),
}
_, err := h.Execute(context.Background(), req)
if err != nil {
t.Fatalf("Execute err = %v", err)
}
if _, err := os.Stat(filepath.Join(baseDir, "run-clean")); !os.IsNotExist(err) {
t.Errorf("workdir not cleaned up: err=%v", err)
}
}
func TestSubprocess_Execute_WorkdirKeptOnFailure(t *testing.T) {
requirePosix(t)
baseDir := t.TempDir()
h := subprocess.New(subprocess.Config{BaseDir: baseDir}, nil)
req := &harness.ExecRequest{
RunID: "run-forensic",
AgentName: "test",
Agent: newAgent(`["sh","-c","exit 5"]`),
}
_, err := h.Execute(context.Background(), req)
if err != nil {
t.Fatalf("Execute err = %v", err)
}
if _, err := os.Stat(filepath.Join(baseDir, "run-forensic")); err != nil {
t.Errorf("workdir removed on failure (should be kept): %v", err)
}
}
func TestSubprocess_Execute_BudgetTimeout(t *testing.T) {
requirePosix(t)
h := subprocess.New(subprocess.Config{BaseDir: t.TempDir()}, nil)
req := &harness.ExecRequest{
RunID: "run-timeout",
AgentName: "test",
Agent: newAgent(`["sh","-c","sleep 5"]`),
Budget: harness.Budget{MaxWallClock: 50 * time.Millisecond},
}
start := time.Now()
_, err := h.Execute(context.Background(), req)
elapsed := time.Since(start)
if err == nil {
t.Fatal("expected budget timeout error")
}
if !strings.Contains(err.Error(), "budget") {
t.Errorf("err = %v, want 'budget' in message", err)
}
if elapsed > time.Second {
t.Errorf("budget not enforced; elapsed = %s", elapsed)
}
}
func TestSubprocess_Execute_ContextCancel(t *testing.T) {
requirePosix(t)
h := subprocess.New(subprocess.Config{BaseDir: t.TempDir()}, nil)
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Millisecond)
defer cancel()
req := &harness.ExecRequest{
RunID: "run-cancel",
AgentName: "test",
Agent: newAgent(`["sh","-c","sleep 5"]`),
}
start := time.Now()
_, _ = h.Execute(ctx, req)
if time.Since(start) > time.Second {
t.Error("context cancel did not stop child")
}
}
func TestSubprocess_Execute_NilAgent(t *testing.T) {
h := subprocess.New(subprocess.Config{BaseDir: t.TempDir()}, nil)
_, err := h.Execute(context.Background(), &harness.ExecRequest{RunID: "r"})
if err == nil {
t.Fatal("expected error for nil agent")
}
}
func TestSubprocess_Execute_NoLocalCommand(t *testing.T) {
h := subprocess.New(subprocess.Config{BaseDir: t.TempDir()}, nil)
_, err := h.Execute(context.Background(), &harness.ExecRequest{
RunID: "r",
Agent: &agents.Agent{Name: "no-cmd"},
})
if err == nil {
t.Fatal("expected ErrNoLocalCommand")
}
}
func TestSubprocess_Execute_AcceptsWhitespaceArgvForm(t *testing.T) {
requirePosix(t)
h := subprocess.New(subprocess.Config{BaseDir: t.TempDir()}, nil)
// Non-JSON form: simple whitespace-split.
req := &harness.ExecRequest{
RunID: "run-ws",
AgentName: "test",
Agent: &agents.Agent{Name: "a", LocalCommand: "echo plain-form"},
}
res, err := h.Execute(context.Background(), req)
if err != nil {
t.Fatalf("Execute err = %v", err)
}
if !strings.Contains(res.Logs, "plain-form") {
t.Errorf("whitespace form failed: %q", res.Logs)
}
}
+262
View File
@@ -0,0 +1,262 @@
// Package webhook is the HTTP-backed implementation of harness.Harness.
// It POSTs the run request to a configured URL and waits synchronously
// for a JSON response. Distinct from internal/webhooks, which is the
// fire-and-forget outbound event fan-out — this harness is a real
// request/response channel used when an agent lives behind an HTTP
// endpoint instead of a K8s Job or a local process.
package webhook
import (
"bytes"
"context"
"encoding/json"
"errors"
"fmt"
"io"
"log/slog"
"net/http"
"time"
"github.com/synapbus/synapbus/internal/agents"
"github.com/synapbus/synapbus/internal/harness"
"github.com/synapbus/synapbus/internal/webhooks"
)
// DefaultTimeout is applied when neither Config.DefaultTimeout nor
// ExecRequest.Budget.MaxWallClock is set.
const DefaultTimeout = 60 * time.Second
// Config tunes the webhook harness.
type Config struct {
// HTTPClient is the outbound client. When nil, a default is used
// with a sensible timeout and no redirect following.
HTTPClient *http.Client
// DefaultTimeout applies when the request has no budget set.
DefaultTimeout time.Duration
// UserAgent overrides the User-Agent header. Defaults to
// "synapbus-harness-webhook/1".
UserAgent string
}
// Harness is the webhook implementation of harness.Harness.
type Harness struct {
cfg Config
client *http.Client
logger *slog.Logger
}
// agentWebhookConfig is the per-agent config blob the webhook harness
// reads from agents.harness_config_json. Callers populate this via the
// admin CLI or migration data.
type agentWebhookConfig struct {
URL string `json:"url"`
Secret string `json:"secret,omitempty"`
TimeoutSeconds int `json:"timeout_seconds,omitempty"`
}
// payload is what the harness POSTs to the configured URL. Kept
// deliberately flat — the receiver can ignore fields it doesn't need.
type payload struct {
RunID string `json:"run_id"`
AgentName string `json:"agent_name"`
Message any `json:"message,omitempty"`
Context []any `json:"context,omitempty"`
SessionID string `json:"session_id,omitempty"`
Env map[string]string `json:"env,omitempty"`
Skills []string `json:"skills,omitempty"`
}
// responseEnvelope is what the harness expects back. Every field is
// optional so a minimally-compliant receiver only needs to return an
// HTTP 200 to signal success.
type responseEnvelope struct {
ExitCode int `json:"exit_code"`
Logs string `json:"logs,omitempty"`
Result json.RawMessage `json:"result,omitempty"`
SessionID string `json:"session_id,omitempty"`
Usage struct {
TokensIn int64 `json:"tokens_in"`
TokensOut int64 `json:"tokens_out"`
TokensCached int64 `json:"tokens_cached"`
CostUSD float64 `json:"cost_usd"`
} `json:"usage,omitempty"`
}
// New constructs a webhook harness.
func New(cfg Config, logger *slog.Logger) *Harness {
client := cfg.HTTPClient
if client == nil {
client = &http.Client{
Timeout: 0, // per-request timeout is set via context
CheckRedirect: func(req *http.Request, via []*http.Request) error {
return http.ErrUseLastResponse
},
}
}
if cfg.UserAgent == "" {
cfg.UserAgent = "synapbus-harness-webhook/1"
}
if cfg.DefaultTimeout <= 0 {
cfg.DefaultTimeout = DefaultTimeout
}
if logger == nil {
logger = slog.Default()
}
return &Harness{
cfg: cfg,
client: client,
logger: logger.With("harness", "webhook"),
}
}
// Name returns the registered harness name.
func (h *Harness) Name() string { return "webhook" }
// Capabilities advertises backend features.
func (h *Harness) Capabilities() harness.Capabilities {
return harness.Capabilities{
SystemPrompt: true,
SessionResume: true,
Skills: true,
OTelNative: false, // trace context goes via headers, not env
MaxConcurrency: 10,
}
}
// TestEnvironment is a no-op for the webhook harness — we can't
// preflight without a target URL, which is per-agent config.
func (h *Harness) TestEnvironment(ctx context.Context) error {
return nil
}
// Provision is a no-op for webhook.
func (h *Harness) Provision(ctx context.Context, agent *agents.Agent) error {
return nil
}
// ErrNoWebhookConfig is returned when an agent's harness_config_json
// doesn't carry a URL the webhook harness can POST to.
var ErrNoWebhookConfig = errors.New("webhook: agent has no webhook config (harness_config_json.url missing)")
// Execute builds the payload, signs it if a secret is configured,
// POSTs it, and maps the response back to an ExecResult.
func (h *Harness) Execute(ctx context.Context, req *harness.ExecRequest) (*harness.ExecResult, error) {
if req == nil {
return nil, errors.New("webhook: nil ExecRequest")
}
if req.Agent == nil {
return nil, errors.New("webhook: ExecRequest.Agent is required")
}
cfg, err := parseAgentConfig(req.Agent.HarnessConfigJSON)
if err != nil {
return nil, err
}
timeout := h.cfg.DefaultTimeout
if req.Budget.MaxWallClock > 0 {
timeout = req.Budget.MaxWallClock
}
if cfg.TimeoutSeconds > 0 {
timeout = time.Duration(cfg.TimeoutSeconds) * time.Second
}
runCtx, cancel := context.WithTimeout(ctx, timeout)
defer cancel()
body := payload{
RunID: req.RunID,
AgentName: req.AgentName,
SessionID: req.SessionID,
Env: req.Env,
Skills: req.Skills,
}
if req.Message != nil {
body.Message = req.Message
}
raw, err := json.Marshal(body)
if err != nil {
return nil, fmt.Errorf("webhook: marshal payload: %w", err)
}
httpReq, err := http.NewRequestWithContext(runCtx, http.MethodPost, cfg.URL, bytes.NewReader(raw))
if err != nil {
return nil, fmt.Errorf("webhook: build request: %w", err)
}
httpReq.Header.Set("Content-Type", "application/json")
httpReq.Header.Set("User-Agent", h.cfg.UserAgent)
httpReq.Header.Set("X-SynapBus-Run-Id", req.RunID)
httpReq.Header.Set("X-SynapBus-Agent", req.AgentName)
if cfg.Secret != "" {
httpReq.Header.Set("X-SynapBus-Signature",
webhooks.ComputeHMACSignature([]byte(cfg.Secret), raw))
}
h.logger.Info("webhook dispatching",
"run_id", req.RunID,
"agent", req.AgentName,
"url", cfg.URL,
)
resp, err := h.client.Do(httpReq)
if err != nil {
return nil, fmt.Errorf("webhook: POST %s: %w", cfg.URL, err)
}
defer resp.Body.Close()
respBody, err := io.ReadAll(io.LimitReader(resp.Body, 1<<20)) // 1 MiB
if err != nil {
return nil, fmt.Errorf("webhook: read response: %w", err)
}
var env responseEnvelope
if len(respBody) > 0 && json.Valid(respBody) {
// Best-effort parse. A non-JSON response is acceptable for
// receivers that just want to signal success via HTTP status.
_ = json.Unmarshal(respBody, &env)
}
exitCode := env.ExitCode
if exitCode == 0 && resp.StatusCode >= 400 {
exitCode = 1
}
logs := env.Logs
if logs == "" {
logs = string(respBody)
}
return &harness.ExecResult{
ExitCode: exitCode,
Logs: logs,
ResultJSON: env.Result,
SessionID: env.SessionID,
Usage: harness.Usage{
TokensIn: env.Usage.TokensIn,
TokensOut: env.Usage.TokensOut,
TokensCached: env.Usage.TokensCached,
CostUSD: env.Usage.CostUSD,
},
}, nil
}
// Cancel is a no-op — HTTP POSTs are cancelled by context, not by
// out-of-band signalling.
func (h *Harness) Cancel(ctx context.Context, runID string) error {
return nil
}
func parseAgentConfig(raw string) (agentWebhookConfig, error) {
var cfg agentWebhookConfig
if raw == "" {
return cfg, ErrNoWebhookConfig
}
if err := json.Unmarshal([]byte(raw), &cfg); err != nil {
return cfg, fmt.Errorf("webhook: parse harness_config_json: %w", err)
}
if cfg.URL == "" {
return cfg, ErrNoWebhookConfig
}
return cfg, nil
}
+223
View File
@@ -0,0 +1,223 @@
package webhook_test
import (
"context"
"encoding/json"
"errors"
"io"
"net/http"
"net/http/httptest"
"strings"
"sync/atomic"
"testing"
"time"
"github.com/synapbus/synapbus/internal/agents"
"github.com/synapbus/synapbus/internal/harness"
"github.com/synapbus/synapbus/internal/harness/webhook"
"github.com/synapbus/synapbus/internal/messaging"
"github.com/synapbus/synapbus/internal/webhooks"
)
func newAgent(url, secret string) *agents.Agent {
cfg, _ := json.Marshal(map[string]any{"url": url, "secret": secret})
return &agents.Agent{Name: "webhook-agent", HarnessConfigJSON: string(cfg)}
}
func TestWebhook_ImplementsInterface(t *testing.T) {
var _ harness.Harness = (*webhook.Harness)(nil)
}
func TestWebhook_NameAndCapabilities(t *testing.T) {
h := webhook.New(webhook.Config{}, nil)
if h.Name() != "webhook" {
t.Fatalf("Name = %q", h.Name())
}
caps := h.Capabilities()
if caps.OTelNative {
t.Errorf("OTelNative should be false (trace context via headers)")
}
if !caps.SessionResume {
t.Error("SessionResume should be true")
}
}
func TestWebhook_TestEnvironment_NoOp(t *testing.T) {
h := webhook.New(webhook.Config{}, nil)
if err := h.TestEnvironment(context.Background()); err != nil {
t.Fatalf("TestEnvironment err = %v", err)
}
}
func TestWebhook_Execute_Success(t *testing.T) {
var gotBody []byte
var gotSig string
var gotRunID string
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
gotBody, _ = io.ReadAll(r.Body)
gotSig = r.Header.Get("X-SynapBus-Signature")
gotRunID = r.Header.Get("X-SynapBus-Run-Id")
w.Header().Set("Content-Type", "application/json")
_, _ = w.Write([]byte(`{"exit_code":0,"logs":"hi","result":{"answer":42},"usage":{"tokens_in":5,"tokens_out":10,"cost_usd":0.003}}`))
}))
defer srv.Close()
h := webhook.New(webhook.Config{}, nil)
req := &harness.ExecRequest{
RunID: "run-ok",
AgentName: "webhook-agent",
Agent: newAgent(srv.URL, "s3cret"),
Message: &messaging.Message{ID: 1, FromAgent: "alice", Body: "hi"},
}
res, err := h.Execute(context.Background(), req)
if err != nil {
t.Fatalf("Execute err = %v", err)
}
if res.ExitCode != 0 {
t.Errorf("ExitCode = %d", res.ExitCode)
}
if res.Logs != "hi" {
t.Errorf("Logs = %q", res.Logs)
}
if string(res.ResultJSON) != `{"answer":42}` {
t.Errorf("ResultJSON = %s", res.ResultJSON)
}
if res.Usage.TokensIn != 5 || res.Usage.TokensOut != 10 {
t.Errorf("usage tokens not mapped: %+v", res.Usage)
}
if res.Usage.CostUSD != 0.003 {
t.Errorf("cost not mapped: %v", res.Usage.CostUSD)
}
// Request-side assertions
if gotRunID != "run-ok" {
t.Errorf("X-SynapBus-Run-Id header = %q", gotRunID)
}
if gotSig == "" {
t.Error("expected HMAC signature header")
}
expectedSig := webhooks.ComputeHMACSignature([]byte("s3cret"), gotBody)
if gotSig != expectedSig {
t.Errorf("signature mismatch: got=%q want=%q", gotSig, expectedSig)
}
var sent map[string]any
if err := json.Unmarshal(gotBody, &sent); err != nil {
t.Fatalf("body not JSON: %v", err)
}
if sent["run_id"] != "run-ok" {
t.Errorf("sent body run_id = %v", sent["run_id"])
}
}
func TestWebhook_Execute_HTTPErrorMapsToNonZeroExit(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.WriteHeader(500)
_, _ = w.Write([]byte("kaboom"))
}))
defer srv.Close()
h := webhook.New(webhook.Config{}, nil)
res, err := h.Execute(context.Background(), &harness.ExecRequest{
RunID: "r", AgentName: "a", Agent: newAgent(srv.URL, ""),
})
if err != nil {
t.Fatalf("Execute err = %v (want graceful mapping)", err)
}
if res.ExitCode == 0 {
t.Errorf("ExitCode = 0, want non-zero")
}
if !strings.Contains(res.Logs, "kaboom") {
t.Errorf("logs missing body: %q", res.Logs)
}
}
func TestWebhook_Execute_NoConfig(t *testing.T) {
h := webhook.New(webhook.Config{}, nil)
_, err := h.Execute(context.Background(), &harness.ExecRequest{
RunID: "r", AgentName: "a", Agent: &agents.Agent{Name: "a"},
})
if !errors.Is(err, webhook.ErrNoWebhookConfig) {
t.Fatalf("err = %v, want ErrNoWebhookConfig", err)
}
}
func TestWebhook_Execute_InvalidConfigJSON(t *testing.T) {
h := webhook.New(webhook.Config{}, nil)
a := &agents.Agent{Name: "a", HarnessConfigJSON: "not-json"}
_, err := h.Execute(context.Background(), &harness.ExecRequest{
RunID: "r", AgentName: "a", Agent: a,
})
if err == nil {
t.Fatal("expected parse error")
}
}
func TestWebhook_Execute_Timeout(t *testing.T) {
var attempts int32
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
atomic.AddInt32(&attempts, 1)
time.Sleep(500 * time.Millisecond)
w.WriteHeader(200)
}))
defer srv.Close()
h := webhook.New(webhook.Config{}, nil)
req := &harness.ExecRequest{
RunID: "r",
AgentName: "a",
Agent: newAgent(srv.URL, ""),
Budget: harness.Budget{MaxWallClock: 30 * time.Millisecond},
}
start := time.Now()
_, err := h.Execute(context.Background(), req)
elapsed := time.Since(start)
if err == nil {
t.Fatal("expected timeout error")
}
if elapsed > time.Second {
t.Errorf("timeout not enforced; elapsed=%s", elapsed)
}
if atomic.LoadInt32(&attempts) != 1 {
t.Errorf("attempts = %d, want 1 (no retries in harness)", attempts)
}
}
func TestWebhook_Execute_NilAgent(t *testing.T) {
h := webhook.New(webhook.Config{}, nil)
_, err := h.Execute(context.Background(), &harness.ExecRequest{RunID: "r"})
if err == nil {
t.Fatal("expected error for nil agent")
}
}
func TestWebhook_Execute_PerAgentTimeoutOverridesDefault(t *testing.T) {
var seenDeadline time.Time
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if d, ok := r.Context().Deadline(); ok {
seenDeadline = d
}
w.WriteHeader(200)
_, _ = w.Write([]byte(`{}`))
}))
defer srv.Close()
cfg, _ := json.Marshal(map[string]any{"url": srv.URL, "timeout_seconds": 2})
a := &agents.Agent{Name: "a", HarnessConfigJSON: string(cfg)}
h := webhook.New(webhook.Config{DefaultTimeout: 30 * time.Second}, nil)
_, err := h.Execute(context.Background(), &harness.ExecRequest{
RunID: "r", AgentName: "a", Agent: a,
})
if err != nil {
t.Fatalf("Execute err = %v", err)
}
if seenDeadline.IsZero() {
t.Skip("server did not capture deadline")
}
// Approximate: per-agent timeout (2s) should override the 30s default.
// We expect the deadline to be well under 10 seconds from now.
if time.Until(seenDeadline) > 10*time.Second {
t.Errorf("deadline too far out: %s", time.Until(seenDeadline))
}
}
+179
View File
@@ -0,0 +1,179 @@
// Package observability initialises OpenTelemetry tracing for SynapBus.
// It is opt-in via SYNAPBUS_OTEL_ENABLED=1 and degrades to no-op
// behaviour when disabled, so every caller can unconditionally use
// the otel.Tracer API.
package observability
import (
"context"
"fmt"
"log/slog"
"strings"
"time"
"go.opentelemetry.io/otel"
"go.opentelemetry.io/otel/exporters/otlp/otlptrace"
"go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp"
"go.opentelemetry.io/otel/propagation"
"go.opentelemetry.io/otel/sdk/resource"
sdktrace "go.opentelemetry.io/otel/sdk/trace"
semconv "go.opentelemetry.io/otel/semconv/v1.21.0"
"go.opentelemetry.io/otel/trace"
)
// TracerName is the library-level identifier used by all SynapBus
// packages when getting a tracer. Keep it stable — it becomes the
// otel.library.name attribute on every span.
const TracerName = "github.com/synapbus/synapbus"
// Config tunes the tracer provider. Zero-value config has Enabled=false
// and causes Init to return a noop shutdown func.
type Config struct {
// Enabled turns tracing on. Default: false.
Enabled bool
// Endpoint is the OTLP HTTP target (host:port). Empty means
// "otlptracehttp default" — usually localhost:4318.
Endpoint string
// Insecure switches to http:// instead of https://. Default true
// because most LAN collectors on kubic are TLS-less.
Insecure bool
// ServiceName overrides the service.name resource attribute.
// Defaults to "synapbus".
ServiceName string
// ServiceVersion adds a service.version attribute when non-empty.
ServiceVersion string
// Sampler chooses the sampling rate. Values in [0,1] are treated
// as a TraceIDRatio. 0 or negative = AlwaysSample. 1 = AlwaysSample.
Sampler float64
}
// ConfigFromEnv reads SYNAPBUS_OTEL_* variables from the environment.
func ConfigFromEnv(getenv func(string) string) Config {
if getenv == nil {
getenv = func(string) string { return "" }
}
cfg := Config{
Enabled: parseBoolEnv(getenv, "SYNAPBUS_OTEL_ENABLED", false),
Endpoint: getenv("SYNAPBUS_OTEL_ENDPOINT"),
Insecure: parseBoolEnv(getenv, "SYNAPBUS_OTEL_INSECURE", true),
ServiceName: getenv("SYNAPBUS_OTEL_SERVICE_NAME"),
ServiceVersion: getenv("SYNAPBUS_OTEL_SERVICE_VERSION"),
}
if cfg.ServiceName == "" {
cfg.ServiceName = "synapbus"
}
return cfg
}
// Init wires a TracerProvider and a W3C text-map propagator. Returns
// a shutdown func the caller must defer. When cfg.Enabled is false
// the returned shutdown func is a no-op.
func Init(ctx context.Context, cfg Config, logger *slog.Logger) (func(context.Context) error, error) {
if logger == nil {
logger = slog.Default()
}
// Always install the propagator so trace context crosses our
// HTTP boundaries even when we're not the ones exporting spans.
otel.SetTextMapPropagator(propagation.NewCompositeTextMapPropagator(
propagation.TraceContext{},
propagation.Baggage{},
))
if !cfg.Enabled {
logger.Info("otel tracing disabled (set SYNAPBUS_OTEL_ENABLED=1 to enable)")
return func(context.Context) error { return nil }, nil
}
opts := []otlptracehttp.Option{}
if cfg.Endpoint != "" {
// otlptracehttp expects host[:port], not a full URL.
clean := strings.TrimPrefix(cfg.Endpoint, "http://")
clean = strings.TrimPrefix(clean, "https://")
clean = strings.TrimSuffix(clean, "/")
opts = append(opts, otlptracehttp.WithEndpoint(clean))
}
if cfg.Insecure {
opts = append(opts, otlptracehttp.WithInsecure())
}
exporter, err := otlptrace.New(ctx, otlptracehttp.NewClient(opts...))
if err != nil {
return nil, fmt.Errorf("observability: create OTLP HTTP exporter: %w", err)
}
res, err := resource.Merge(
resource.Default(),
resource.NewWithAttributes(semconv.SchemaURL,
semconv.ServiceName(cfg.ServiceName),
semconv.ServiceVersion(cfg.ServiceVersion),
),
)
if err != nil {
return nil, fmt.Errorf("observability: build resource: %w", err)
}
tp := sdktrace.NewTracerProvider(
sdktrace.WithBatcher(exporter,
sdktrace.WithBatchTimeout(5*time.Second),
),
sdktrace.WithResource(res),
sdktrace.WithSampler(pickSampler(cfg.Sampler)),
)
otel.SetTracerProvider(tp)
logger.Info("otel tracing enabled",
"endpoint", cfg.Endpoint,
"service", cfg.ServiceName,
)
return tp.Shutdown, nil
}
// InjectTraceContext writes the W3C traceparent / tracestate headers
// from ctx into dst as env-var style keys (TRACEPARENT, TRACESTATE).
// Harnesses call this to push trace context into child processes and
// outbound HTTP requests in a single uniform shape.
func InjectTraceContext(ctx context.Context, dst map[string]string) {
if dst == nil {
return
}
carrier := propagation.MapCarrier{}
otel.GetTextMapPropagator().Inject(ctx, carrier)
for k, v := range carrier {
dst[strings.ToUpper(k)] = v
}
}
// TraceIDFromContext extracts the trace id hex from the span currently
// attached to ctx, or returns the empty string if there isn't one.
func TraceIDFromContext(ctx context.Context) string {
span := trace.SpanFromContext(ctx)
sc := span.SpanContext()
if !sc.HasTraceID() {
return ""
}
return sc.TraceID().String()
}
func parseBoolEnv(getenv func(string) string, key string, def bool) bool {
v := strings.TrimSpace(strings.ToLower(getenv(key)))
switch v {
case "":
return def
case "1", "t", "true", "yes", "on":
return true
case "0", "f", "false", "no", "off":
return false
}
return def
}
func pickSampler(rate float64) sdktrace.Sampler {
if rate <= 0 || rate >= 1 {
return sdktrace.AlwaysSample()
}
return sdktrace.ParentBased(sdktrace.TraceIDRatioBased(rate))
}
+99
View File
@@ -0,0 +1,99 @@
package observability_test
import (
"context"
"testing"
"github.com/synapbus/synapbus/internal/observability"
"go.opentelemetry.io/otel"
sdktrace "go.opentelemetry.io/otel/sdk/trace"
"go.opentelemetry.io/otel/sdk/trace/tracetest"
)
func TestConfigFromEnv_DefaultDisabled(t *testing.T) {
cfg := observability.ConfigFromEnv(func(string) string { return "" })
if cfg.Enabled {
t.Error("default Enabled should be false")
}
if cfg.ServiceName != "synapbus" {
t.Errorf("ServiceName default = %q, want synapbus", cfg.ServiceName)
}
if !cfg.Insecure {
t.Error("default Insecure should be true (LAN-friendly)")
}
}
func TestConfigFromEnv_Overrides(t *testing.T) {
env := map[string]string{
"SYNAPBUS_OTEL_ENABLED": "1",
"SYNAPBUS_OTEL_ENDPOINT": "otel.kubic.home.arpa:4318",
"SYNAPBUS_OTEL_INSECURE": "false",
"SYNAPBUS_OTEL_SERVICE_NAME": "synapbus-dev",
}
cfg := observability.ConfigFromEnv(func(k string) string { return env[k] })
if !cfg.Enabled {
t.Error("Enabled=1 not honoured")
}
if cfg.Endpoint != "otel.kubic.home.arpa:4318" {
t.Errorf("Endpoint = %q", cfg.Endpoint)
}
if cfg.Insecure {
t.Error("Insecure=false not honoured")
}
if cfg.ServiceName != "synapbus-dev" {
t.Errorf("ServiceName = %q", cfg.ServiceName)
}
}
func TestInit_DisabledIsNoopShutdown(t *testing.T) {
shutdown, err := observability.Init(context.Background(), observability.Config{Enabled: false}, nil)
if err != nil {
t.Fatalf("Init err = %v", err)
}
if shutdown == nil {
t.Fatal("shutdown is nil")
}
if err := shutdown(context.Background()); err != nil {
t.Fatalf("shutdown err = %v", err)
}
}
func TestInjectTraceContext_NoSpan(t *testing.T) {
dst := map[string]string{}
observability.InjectTraceContext(context.Background(), dst)
// No span → no traceparent written (TextMapPropagator sees empty
// span context). Assert the function didn't panic; dst may be empty
// or contain TRACEPARENT with all-zero trace id depending on setup.
}
func TestInjectTraceContext_WithSpan(t *testing.T) {
// Install a TracerProvider with an in-memory exporter so spans
// have real trace ids.
tp := sdktrace.NewTracerProvider(sdktrace.WithSyncer(tracetest.NewInMemoryExporter()))
t.Cleanup(func() { _ = tp.Shutdown(context.Background()) })
otel.SetTracerProvider(tp)
// Install the propagator (normally done by observability.Init)
_, _ = observability.Init(context.Background(), observability.Config{Enabled: false}, nil)
ctx, span := tp.Tracer("test").Start(context.Background(), "outer")
defer span.End()
dst := map[string]string{}
observability.InjectTraceContext(ctx, dst)
if _, ok := dst["TRACEPARENT"]; !ok {
t.Errorf("TRACEPARENT missing from dst: %v", dst)
}
tid := observability.TraceIDFromContext(ctx)
if tid == "" {
t.Error("TraceIDFromContext returned empty")
}
}
func TestTraceIDFromContext_NoSpan(t *testing.T) {
tid := observability.TraceIDFromContext(context.Background())
if tid != "" {
t.Errorf("TraceIDFromContext with no span = %q, want empty", tid)
}
}
+4 -1
View File
@@ -45,7 +45,10 @@ func setupTestDB(t *testing.T) *sql.DB {
k8s_image TEXT,
k8s_env_json TEXT,
k8s_resource_preset TEXT NOT NULL DEFAULT 'default',
pending_work INTEGER NOT NULL DEFAULT 0
pending_work INTEGER NOT NULL DEFAULT 0,
harness_name TEXT,
local_command TEXT,
harness_config_json TEXT
);
CREATE TABLE reactive_runs (
id INTEGER PRIMARY KEY AUTOINCREMENT,
+47
View File
@@ -0,0 +1,47 @@
-- 019: Harness-agnostic wrappers
-- Adds per-agent harness configuration and a backend-agnostic harness_runs
-- table used by all harness implementations (k8sjob, subprocess, webhook).
-- See docs/harness-otel-design.md.
-- Per-agent harness configuration ----------------------------------------
-- Explicit harness name ("k8sjob", "subprocess", "webhook", "stub").
-- When NULL, Registry.Resolve picks a backend by fallback rules.
ALTER TABLE agents ADD COLUMN harness_name TEXT;
-- JSON-encoded argv for the subprocess backend, e.g.
-- ["claude", "--print", "--max-turns", "50"]
ALTER TABLE agents ADD COLUMN local_command TEXT;
-- Per-harness opaque config blob; parsed by the backend that owns it.
ALTER TABLE agents ADD COLUMN harness_config_json TEXT;
-- Backend-agnostic harness_runs table ------------------------------------
CREATE TABLE harness_runs (
id INTEGER PRIMARY KEY AUTOINCREMENT,
run_id TEXT NOT NULL UNIQUE,
agent_name TEXT NOT NULL,
backend TEXT NOT NULL, -- 'k8sjob' | 'subprocess' | 'webhook' | 'stub'
message_id INTEGER, -- triggering message, if any
status TEXT NOT NULL, -- 'pending' | 'running' | 'success' | 'failed' | 'cancelled' | 'timeout'
exit_code INTEGER,
trace_id TEXT,
span_id TEXT,
session_id TEXT, -- backend session id for resume, if any
tokens_in INTEGER NOT NULL DEFAULT 0,
tokens_out INTEGER NOT NULL DEFAULT 0,
tokens_cached INTEGER NOT NULL DEFAULT 0,
cost_usd REAL NOT NULL DEFAULT 0,
duration_ms INTEGER,
result_json TEXT,
logs_excerpt TEXT, -- bounded; full logs live on disk
created_at DATETIME NOT NULL DEFAULT CURRENT_TIMESTAMP,
finished_at DATETIME
);
CREATE INDEX idx_harness_runs_agent ON harness_runs(agent_name, created_at DESC);
CREATE INDEX idx_harness_runs_status ON harness_runs(status, created_at DESC);
CREATE INDEX idx_harness_runs_trace ON harness_runs(trace_id);
CREATE INDEX idx_harness_runs_run_id ON harness_runs(run_id);