1. Executive summary
2. What was built (autonomous pipeline)
Two parallel feature implementations via git worktrees, merged to main, verified, and run end-to-end.
Phase 1 — Specs (committed earlier)
specs/016-agent-marketplace/spec.md— 4 user stories (US1: auction, US2: manifests, US3: reputation, US4: reflection), 29 functional requirements, 10 success criteria.specs/017-musique-benchmark/spec.md— 4 user stories (single-shot Pareto, trio dedup, learning tier, HTML report), 23 FRs, 7 SCs.docs/superpowers/specs/2026-04-11-mas-benchmark-design.md— brainstorming design doc capturing 6 clarifying questions and decisions (mixed-tier agent pool, curated trio, wait-for-016 strategy).
Phase 2 — Parallel implementation in git worktrees
| Feature | Worktree | Branch | Scope |
|---|---|---|---|
| 016 (Go) | ../synapbus-016-impl |
016-agent-marketplace |
Capability manifests (wiki-backed), auction channel, 6 MCP actions, reputation SQLite ledger, awarded reaction, 4 new test functions |
| 017 (Python) | ../synapbus-017-impl |
017-musique-benchmark |
MuSiQue downloader, trio curation, in-process marketplace stub, mixed-tier agents, baseline runner, F1 + Pareto scoring, HTML report generator |
Phase 3 — Integration
- Merged both branches to
mainvia--no-ffmerge commits. - Ran
go build ./...— clean. - Ran
go test ./...— 34 packages green, zero failures. - Added
benchmark/sdk_backend.py— unified backend routing betweenanthropicSDK andclaude-agent-sdk(used by this run sinceANTHROPIC_API_KEYis unset and Claude Code session credentials propagate through the Agent SDK). - Ran
benchmark/run.py --mode single-shot --question q1end-to-end with real model calls.
3. The benchmark task
One real 4-hop question from MuSiQue-Ans dev set, curated to have "United States" as a bridge entity for future dedup runs.
What treaty ceded territory to the US extending west to the body of water by the city where the designer of Southeast Library died?
Gold answer: Treaty of Paris
Gold decomposition (4 hops)
MuSiQue ID: 4hop1__94201_642284_131926_13165 · 20 distractor paragraphs, 4 gold-supporting.
4. The auction
The harness posted an auction, both agents bid, one was awarded. Real SynapBus MCP tool surface names mirrored by the in-process stub.
Auction post
post_auction({
task: "What treaty ceded territory to the US extending west...",
domain: "multi-hop-qa",
max_budget_tokens: 50000,
deadline: "now + 300s",
required_domains: ["multi-hop-qa"]
})
→ auction-1
Bids received
The stub's scoring formula is estimated_tokens / confidence × (1.15 − 0.3 × reputation). At epoch 1 both agents have reputation 0.5 (prior), so ties break on raw cost/confidence. Haiku's 4000/0.45 ≈ 8889 vs Sonnet's 12000/0.80 = 15000 → Haiku wins.
5. Results & Pareto verdict
Marketplace (Haiku 4.5)
Tokens: 3,314 · Wall: 29.9s
Baseline (Sonnet 4.6)
Tokens: 697 · Wall: 13.7s
Pareto scatter plot
6. Analysis — why FAIL is informative
The FAIL verdict is arguably the most interesting outcome of this run. It proves the benchmark is not a vanity metric.
Finding 1: Haiku 4.5 correctly solves a 4-hop question
This is genuinely impressive. The marketplace-awarded agent is a Haiku-tier model; it produced the exact gold answer Treaty of Paris by correctly following all four decomposition hops (designer → city → river → treaty). The full reasoning trace appears in benchmark/results/latest.json under market.raw_text.
Finding 2: Sonnet's "The Treaty of Paris (1783)" is semantically correct but loses 14% F1
Exact-match F1 after normalization penalizes the extra parenthetical year. This is a known quirk of string-match metrics, not a fundamental error. A more permissive metric (substring match or semantic similarity) would score both answers at 1.0, and the verdict would become: both correct, baseline cheaper → FAIL for the marketplace.
Finding 3: The auction's cost heuristic undervalued Sonnet
The stub scoring formula picked Haiku (4000/0.45 ≈ 8889) over Sonnet (12000/0.80 = 15000). This reflects the "minimize tokens times confidence penalty" heuristic. In reality:
| Agent | Estimated tokens | Actual tokens | Actual F1 | Pareto-preferred by this task? |
|---|---|---|---|---|
| haiku-agent | 4,000 | 3,314 | 1.000 | on quality alone |
| sonnet-agent (counterfactual) | 12,000 | ~697† | ~0.857† | on cost alone |
† Using baseline Sonnet numbers as a proxy for "what Sonnet would have done if awarded"; actual in-marketplace Sonnet execution would have similar cost.
Finding 4: This is exactly what reputation is for
At epoch 1, both agents had prior reputation 0.5 — no real information. The auction had to rely on self-reported bids. After this run, the reputation ledger now contains:
haiku-agent | multi-hop-qa | runs=1 correct=1 tokens_spent=3314 score=0.983
Haiku's very high score (0.983) reflects the perfect F1 with modest token spend. But the stub's scoring penalizes tokens lightly (-min(avg_tokens/200000, 0.3)) — so a future auction in this domain would still favor Haiku unless the penalty coefficient is increased. This is a real tuning lever the design exposes.
Finding 5: Over 5 epochs, expect convergence toward Sonnet
If we ran the learning tier, sonnet-agent would get its bootstrap exploration credit (US3 of spec 016 explicitly mandates this) and enter the reputation ledger with tokens ≈ 700 and F1 ≈ 0.857. Then, from epoch 3 onward, the auction would correctly prefer Sonnet on strict-Pareto grounds. This is the promised learning curve — but not implementable from a single-epoch run.
7. Reputation ledger state
After one task completion, the domain-scoped reputation vector has exactly one entry.
query_reputation("haiku-agent", "multi-hop-qa") → {
"agent": "haiku-agent",
"domain": "multi-hop-qa",
"runs": 1,
"correct": 1,
"tokens_spent": 3314,
"score": 0.98343
}
query_reputation("sonnet-agent", "multi-hop-qa") → {
"agent": "sonnet-agent",
"domain": "multi-hop-qa",
"runs": 0, # never awarded a task yet
"correct": 0,
"tokens_spent": 0,
"score": 0.5 # prior
}
This is a two-entry vector at N=1 epoch. In the Go implementation (spec 016, internal/marketplace/store.go), the same data is persisted to the agent_reputation SQLite table via the query_reputation MCP action. The Python stub mirrors the API exactly so the switchover from stub to real SynapBus is purely mechanical.
8. Deferred work & follow-ups
What this autonomous run explicitly did not ship, and why — plus the concrete next steps.
| Item | Spec ref | Status | Unblock when |
|---|---|---|---|
| Reflection loop (US4) | 016 FR-016 → FR-020b | deferred | MVP stable + tombstoning design validated |
| Auto-tombstoning on rolling failure | 016 FR-020a/b | deferred | Reflection loop landed |
| Hard-stop budget enforcement daemon | 016 FR-022/FR-023 | recorded only | Real production traffic shows need |
| Full 3-question curated trio run | 017 US2 | trio.jsonl exists | 5× token budget allocated |
| 5-epoch learning tier | 017 US3, FR-021/023 | deferred | Single-shot MVP stable first |
| FRAMES secondary eval | design §3 | deferred | Wikipedia dump (~20GB) staged |
| Real SynapBus MCP wiring from benchmark | 017 FR-003 | stub equivalent | Swap Python stub calls for MCP calls — mechanical |
| Tightened scoring heuristic penalty | — | observed need | Tune avg_tokens/200000 constant upward |
Recommended next action
- Run the 5-epoch learning tier on the same q1 question to prove the convergence story (~420k token budget). This is the single highest-value follow-up.
- Swap benchmark stub → real SynapBus MCP — modify
benchmark/marketplace.pyto call the 6 new actions viaexecute(action, args)through MCP. Per the 016 implementation summary, all 6 actions are dispatched through theexecutetool. - Implement US4 reflection loop in Go and exercise it on the learning tier run.
- Scale to the full 3-question trio to get a real dedup measurement.
9. Artifacts & commit SHAs
| Artifact | Path |
|---|---|
| Spec 016 | specs/016-agent-marketplace/spec.md |
| Spec 017 | specs/017-musique-benchmark/spec.md |
| Design doc (brainstorm) | docs/superpowers/specs/2026-04-11-mas-benchmark-design.md |
| Go marketplace service | internal/marketplace/service.go, store.go |
| Go MCP bridge | internal/mcp/marketplace.go, marketplace_test.go |
| SQLite migration | internal/storage/schema/018_agent_marketplace.sql |
| Python benchmark | benchmark/ (9 files) |
| Curated trio | benchmark/trio.jsonl (3 × 4-hop questions) |
| Run output JSON | benchmark/results/latest.json |
| Run output HTML (basic) | benchmark/results/latest.html |
| This report | autonomous_report.html |
| Summary markdown | autonomous_summary.md |
Commit chain on main
96db7c0 spec(016): agent marketplace spec + research reports e77fd7a spec(017): MuSiQue MAS benchmark harness cda3365 feat(016): agent marketplace MVP — manifests, auctions, reputation 02b8548 feat(017): MuSiQue benchmark harness — marketplace stub, agents, Pareto (merge) merge: 016-agent-marketplace MVP (auction + manifests + reputation) (merge) merge: 017-musique-benchmark MVP (Python harness + trio + Pareto report) (final) feat: sdk_backend + autonomous run integration
All commits co-authored by Claude Opus 4.6 (1M context).