Adds benchmark/sdk_backend.py that routes model calls through either
the anthropic SDK (preferred, requires ANTHROPIC_API_KEY) or the
claude-agent-sdk as a Claude Code session fallback. agents.py and
baseline.py now go through this unified backend instead of calling
anthropic directly.
Ran benchmark/run.py --mode single-shot --question q1 end-to-end
with real Claude API calls via claude-agent-sdk. Real numbers:
- Marketplace (Haiku 4.5): 3314 tokens, F1 1.000 (exact match)
- Baseline (Sonnet 4.6): 697 tokens, F1 0.857 (penalized for "1783")
- Pareto verdict: FAIL (not strictly NW; marketplace wins quality,
loses cost — informative failure per spec design).
Added autonomous_report.html (rich narrative with Pareto chart)
and autonomous_summary.md. All 34 Go packages still green.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
MVP implementation of the MuSiQue multi-agent benchmark (spec 017):
- benchmark/setup.py: downloads musique_v1.0.zip from the canonical Google
Drive source (mirrors upstream download_data.sh). Idempotent.
- benchmark/curate.py: deterministic selection of 3 4-hop questions from
the dev set sharing a US pivot entity; writes benchmark/trio.jsonl.
- benchmark/marketplace.py: in-process 016-marketplace stub with
post_auction / bid / award / mark_done / query_reputation and a
domain-scoped reputation ledger. Designed for mechanical swap to real
SynapBus MCP tools.
- benchmark/agents.py: HaikuAgent + SonnetAgent, using the official
anthropic SDK (no Claude Agent SDK, no subprocesses). Models pinned
to claude-haiku-4-5-20251001 and claude-sonnet-4-6.
- benchmark/baseline.py: single Sonnet call with all 20 distractors
plus chain-of-thought.
- benchmark/score.py: SQuAD-style normalized F1 + Pareto verdict
(strictly northwest = PASS).
- benchmark/run.py: main entry. --mode single-shot, --question, --dry-run.
- benchmark/report.py: self-contained HTML with inline SVG scatter plot.
- benchmark/trio.jsonl: curated reproducible trio (all three converge on
"Treaty of Paris" US territory cession).
Verified with benchmark/run.py --dry-run end-to-end; all 8 files
py_compile clean. Real-token execution is deferred to the user's main
session.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>