MVP implementation of the MuSiQue multi-agent benchmark (spec 017):
- benchmark/setup.py: downloads musique_v1.0.zip from the canonical Google
Drive source (mirrors upstream download_data.sh). Idempotent.
- benchmark/curate.py: deterministic selection of 3 4-hop questions from
the dev set sharing a US pivot entity; writes benchmark/trio.jsonl.
- benchmark/marketplace.py: in-process 016-marketplace stub with
post_auction / bid / award / mark_done / query_reputation and a
domain-scoped reputation ledger. Designed for mechanical swap to real
SynapBus MCP tools.
- benchmark/agents.py: HaikuAgent + SonnetAgent, using the official
anthropic SDK (no Claude Agent SDK, no subprocesses). Models pinned
to claude-haiku-4-5-20251001 and claude-sonnet-4-6.
- benchmark/baseline.py: single Sonnet call with all 20 distractors
plus chain-of-thought.
- benchmark/score.py: SQuAD-style normalized F1 + Pareto verdict
(strictly northwest = PASS).
- benchmark/run.py: main entry. --mode single-shot, --question, --dry-run.
- benchmark/report.py: self-contained HTML with inline SVG scatter plot.
- benchmark/trio.jsonl: curated reproducible trio (all three converge on
"Treaty of Paris" US territory cession).
Verified with benchmark/run.py --dry-run end-to-end; all 8 files
py_compile clean. Real-token execution is deferred to the user's main
session.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>