From e77fd7afdf346da2227be895c9fd972ceafc2b22 Mon Sep 17 00:00:00 2001 From: Algis Dumbris Date: Sat, 11 Apr 2026 15:02:59 +0300 Subject: [PATCH 1/2] spec(017): MuSiQue MAS benchmark harness Feature spec for Python benchmark that integration-tests 016 marketplace against a real multi-hop reasoning task. 4 prioritized user stories: P1 single-shot Pareto verification, P2 curated trio with dedup, P3 learning tier, P1 rich HTML report. 23 FRs, 7 success criteria. Also: brainstorming design doc at docs/superpowers/specs/ capturing the 6 clarifying questions and chosen decisions (mixed-tier pool, curated trio, tiered run modes, wait-for-016 execution strategy). Co-Authored-By: Claude Opus 4.6 (1M context) --- .../specs/2026-04-11-mas-benchmark-design.md | 160 +++++++++++++++++ .../checklists/requirements.md | 36 ++++ specs/017-musique-benchmark/spec.md | 169 ++++++++++++++++++ 3 files changed, 365 insertions(+) create mode 100644 docs/superpowers/specs/2026-04-11-mas-benchmark-design.md create mode 100644 specs/017-musique-benchmark/checklists/requirements.md create mode 100644 specs/017-musique-benchmark/spec.md diff --git a/docs/superpowers/specs/2026-04-11-mas-benchmark-design.md b/docs/superpowers/specs/2026-04-11-mas-benchmark-design.md new file mode 100644 index 0000000..57bf7fa --- /dev/null +++ b/docs/superpowers/specs/2026-04-11-mas-benchmark-design.md @@ -0,0 +1,160 @@ +# MAS Benchmark — Design Document + +**Date**: 2026-04-11 +**Status**: Approved (autonomous mode) +**Author**: Claude Opus 4.6 (1M context) via brainstorming skill +**Next**: speckit specification at `specs/017-musique-benchmark/spec.md` +**Related**: `specs/016-agent-marketplace/spec.md` + +## Problem + +The current "Fermi piano tuners in Chicago" example in `multiagent_systems_report.html` is dated and rare as a profession. It needs to be replaced with a modern task that: + +- Exercises the same MAS features (dynamic decomposition, dedup, uncertainty aggregation, cost accounting, orphaned-spawn recovery) +- Runs on local data only — no `WebSearch` tool required +- Serves both as a readable narrative example and a runnable integration test +- Measures the Pareto frontier of quality vs token cost — never quality alone +- Fits a tight dev-loop token budget (≤ 500k tokens per full learning run) + +## Decisions (from brainstorming) + +1. **Purpose**: dual-use — narrative example in the report AND integration test for `016-agent-marketplace`. +2. **Dataset**: **MuSiQue-Ans 4-hop** as primary (~100 MB, gold decomposition DAGs, anti-shortcut filtered). **FRAMES** as future cross-eval runner-up. +3. **Scale**: N=3 curated questions. Deliberately cherry-picked to share a bridge entity so sub-agents naturally re-lookup the same Wikipedia paragraphs (dedup metric becomes observable at N=3). +4. **Agent pool**: **mixed-tier** — Haiku + Sonnet + Opus. Each agent publishes its own capability manifest with per-domain cost profile. The auction has to learn when paying for Opus is worth it and when Haiku suffices. +5. **Run modes — tiered**: + - **Single-shot** (CI smoke test): run the 3 questions once. ~100k tokens. + - **Learning tier**: run the same 3 questions for 5 epochs. Reputation and skill cards persist; scratchpad resets per epoch. Measure tokens-per-correct-answer declining across epochs. ~500k tokens. +6. **Execution strategy**: harness talks to the spec-016 MCP tool surface. Runs against real SynapBus once 016 is implemented. +7. **Scoring is Pareto**: report both quality (F1, decomposition F1) AND cost (total tokens, tokens-per-correct). A passing marketplace strictly dominates a single-agent baseline. + +## Architecture + +### Components + +1. **Curated trio file** (`benchmark/trio.jsonl`) — three MuSiQue-Ans questions with shared pivot entity, gold answers, gold decomposition DAGs. Selected by deterministic rule from the dev set and checked into the repo for reproducibility. + +2. **Benchmark harness** (Python) — orchestrates a run: + - Reads trio.jsonl and initial skill-card configuration + - Seeds the marketplace (posts skill-card wiki articles for each agent, creates the `#bench-auction` auction channel) + - For each question, posts an auction, waits for bids, awards via MCP, polls for completion + - Collects metrics (per-question tokens, F1, cache-hit rate, decomposition F1, wall time) + - Runs single-shot or 5-epoch learning tier per CLI flag + +3. **Agent runner** (Python) — spawns N agents, each a Claude Agent SDK session with: + - A system prompt built from the agent's skill card + - SynapBus MCP tools configured + - Distinct model tier (Haiku / Sonnet / Opus) + - Token counter hook for real-time budget enforcement + +4. **Baseline runner** — a single Claude call that receives the question and all 20 distractor paragraphs in one shot, using chain-of-thought, no decomposition, no marketplace. Produces reference `(tokens, F1)` for Pareto comparison. + +5. **Scoring module** — computes metrics per run, writes `results/{run_id}.json`, generates Pareto plot data. + +6. **Report generator** — renders a rich HTML with the narrative, Pareto chart, per-question trace, and learning curve. + +### Data flow (single question) + +``` +trio.jsonl → harness.post_auction(question, max_budget, domains) + ↓ + SynapBus auction channel (reactive trigger) + ↓ + ┌────────────┬────────────┬──────────────┐ + ↓ ↓ ↓ ↓ + Haiku agent Sonnet agent Opus agent (poller) + ↓ ↓ ↓ + bid() bid() bid() + └────────────┴────────────┘ + ↓ + harness.award(best bid) + ↓ + winner.claim → execute + ↓ + (reads distractor paragraphs via MCP) + ↓ + shared scratchpad + (dedup: same entity → cache hit) + ↓ + winner.mark_done(answer, tokens) + ↓ + reputation ledger update (per-domain tuple) + ↓ + harness.score(answer vs gold) +``` + +### Scoring (Pareto) + +Three metrics plotted together, one point per run configuration: + +- **Quality**: final-answer exact-match F1 (0.0 / 0.33 / 0.67 / 1.0 at N=3) +- **Cost**: total tokens consumed (orchestrator + all sub-agents across all 3 questions) +- **Efficiency**: tokens-per-correct-answer = total_tokens / max(F1 × 3, 1) + +Four configurations plotted on the Pareto chart: + +| # | Config | Expected quality | Expected cost | +|---|---|---|---| +| 1 | Single-agent baseline (Opus, all distractors in context) | high (~2/3) | high (~30k) | +| 2 | Single-agent baseline (Sonnet, same) | medium (~2/3) | medium (~15k) | +| 3 | Naive marketplace (no reputation, no reflection, no dedup) | medium (~2/3) | medium-high (~40k) | +| 4 | Full 016 marketplace (reputation + dedup + mixed-tier routing) | ≥ baseline | should be **strictly less** than baseline | + +The marketplace passes only if it lands **strictly northwest** of Sonnet baseline on the Pareto plot. + +### Learning tier + +5 epochs of the same 3 questions. What persists between epochs: + +- Reputation ledger entries (accumulate) +- Capability manifest revisions (reflection loop proposes diffs — auto-approved for benchmark) +- Per-agent skill-card example-tasks list (grows monotonically) + +What resets between epochs: + +- Shared scratchpad (within-task coordination, not long-term memory) +- Auction channel contents (each epoch creates fresh auctions) + +**Expected learning curve**: tokens-per-correct-answer should drop monotonically from epoch 1 (all agents uncalibrated, ε-greedy bootstrap dominates) to epoch 5 (reputation converged, routing stable). If it doesn't — the marketplace has a bug. + +### Failure injection + +For orphaned-spawn recovery: one epoch runs with a 10% random sub-agent failure rate (agents randomly return "timeout" instead of bid). Measure accuracy degradation. Target: ≤ 5 percentage-point drop. + +## Realistic MVP scope + +Given execution constraints, the MVP for **today's autonomous run** scopes down: + +- **Questions**: N=1 instead of N=3 (save 3× tokens on the actual run; the trio.jsonl file still contains all 3 for future runs) +- **Agents**: 2 (Haiku + Sonnet) instead of 3 (Haiku + Sonnet + Opus). Mixed-tier proved on 2 tiers. +- **Epochs**: 1 single-shot run, no learning tier. Design doc describes the 5-epoch protocol for future runs. +- **Reflection loop**: skipped. Full 016 spec has it; MVP implementation focuses on US1 + US2 + US3 (auction + manifests + reputation). + +**Still measured and reported**: +- Dynamic decomposition on one real 4-hop MuSiQue question +- Auction → bid → award → claim → done full lifecycle +- Per-model cost differentials (Haiku vs Sonnet on same task) +- Pareto comparison against single-agent baseline +- Reputation ledger write-through + +**Documented-but-deferred**: +- Reflection loop + skill-card diff proposals (US4 of spec 016) +- Tombstoning on failure rate (FR-020a/b) +- Full 5-epoch learning tier +- 3-question curated trio dedup measurement +- FRAMES cross-eval + +## Acceptance criteria for autonomous run + +1. `016-agent-marketplace` MVP compiles, passes its own Go tests, and exposes the required MCP tools. +2. Benchmark harness downloads MuSiQue, curates trio.jsonl, runs 1 question end-to-end against local SynapBus with the 016 implementation. +3. Real token counts and real F1 recorded. +4. Pareto plot generated comparing full marketplace vs Sonnet baseline. +5. HTML report renders with live numbers, not placeholders. +6. `autonomous_summary.md` written documenting what shipped, what passed, what deferred. + +## Honest caveats + +- N=1 cannot support statistical claims. The benchmark's purpose at this scale is **mechanism verification**, not efficacy proof. +- Claude Agent SDK integration is a known pain point — may need fallback to direct Anthropic SDK if MCP wiring fails. +- Single-epoch run cannot show the learning curve. Design doc + spec describe the full protocol for future scaling. diff --git a/specs/017-musique-benchmark/checklists/requirements.md b/specs/017-musique-benchmark/checklists/requirements.md new file mode 100644 index 0000000..63dd96e --- /dev/null +++ b/specs/017-musique-benchmark/checklists/requirements.md @@ -0,0 +1,36 @@ +# Specification Quality Checklist: MuSiQue MAS Benchmark + +**Purpose**: Validate specification completeness and quality before proceeding to planning +**Created**: 2026-04-11 +**Feature**: [spec.md](../spec.md) + +## Content Quality + +- [x] No implementation details (languages, frameworks, APIs) +- [x] Focused on user value and business needs +- [x] Written for non-technical stakeholders +- [x] All mandatory sections completed + +## Requirement Completeness + +- [x] No [NEEDS CLARIFICATION] markers remain +- [x] Requirements are testable and unambiguous +- [x] Success criteria are measurable +- [x] Success criteria are technology-agnostic +- [x] All acceptance scenarios are defined +- [x] Edge cases are identified +- [x] Scope is clearly bounded +- [x] Dependencies and assumptions identified + +## Feature Readiness + +- [x] All functional requirements have clear acceptance criteria +- [x] User scenarios cover primary flows +- [x] Feature meets measurable outcomes defined in Success Criteria +- [x] No implementation details leak into specification + +## Notes + +- Specification intentionally allows Claude Agent SDK with direct Anthropic SDK fallback — this is an operational concern rather than an architectural one. +- FR-021 through FR-023 (learning tier) are marked deferred for the autonomous MVP run but remain in scope for the full spec. +- The benchmark is paired with `016-agent-marketplace` and will not function without it (or a stub). diff --git a/specs/017-musique-benchmark/spec.md b/specs/017-musique-benchmark/spec.md new file mode 100644 index 0000000..2a51bac --- /dev/null +++ b/specs/017-musique-benchmark/spec.md @@ -0,0 +1,169 @@ +# Feature Specification: MuSiQue Multi-Agent Benchmark Harness + +**Feature Branch**: `017-musique-benchmark` +**Created**: 2026-04-11 +**Status**: Draft +**Input**: MuSiQue-based MAS benchmark harness that integration-tests the 016-agent-marketplace. N=3 curated trio sharing a pivot entity, mixed-tier agent pool (Haiku + Sonnet + Opus), Pareto scoring vs single-agent baseline, optional 5-epoch learning tier, Python harness using Claude Agent SDK. + +## User Scenarios & Testing *(mandatory)* + +### User Story 1 — Single-shot Pareto verification (Priority: P1) + +A developer wants to verify that the 016 agent marketplace actually delivers better quality-per-token than a single-agent baseline on a real multi-hop reasoning task. They run the benchmark harness in single-shot mode. The harness posts one MuSiQue 4-hop question to the marketplace, mixed-tier agents bid, one wins, executes the task using distractor paragraphs from the local MuSiQue corpus, and reports the final answer. The harness also runs the same question through a single-agent baseline (one Claude call with all distractors in context). It computes and plots both runs on a Pareto chart of quality vs tokens. The marketplace passes only if it lands strictly northwest of the baseline. + +**Why this priority**: This is the irreducible proof-of-work for the marketplace. Without it, the 016 spec is unvalidated. With it, we have evidence that the auction + mixed-tier routing primitives actually produce a Pareto improvement on a real task. + +**Independent Test**: Start a fresh SynapBus instance with the 016 MVP. Run `python benchmark/run.py --mode single-shot --question q1`. The script completes within a token budget, produces a results JSON with real token counts and F1 scores for both the marketplace and the baseline, and exits zero. + +**Acceptance Scenarios**: + +1. **Given** a fresh SynapBus instance with the 016 marketplace primitives loaded, **When** the harness posts the MuSiQue question as an auction, **Then** at least two agents submit structured bids as threaded replies within a configurable wait window. +2. **Given** bids have been submitted, **When** the harness awards the best bid via the `awarded` reaction, **Then** the winning agent receives a claim and begins executing against the MuSiQue distractor paragraphs. +3. **Given** the winning agent completes the task, **When** the harness compares the answer to the gold answer, **Then** an F1 score is recorded and the token count for the entire run is aggregated. +4. **Given** both the marketplace run and the baseline run have completed, **When** the harness generates the Pareto report, **Then** both data points are shown with their exact token counts and F1 scores, and the marketplace run is either strictly northwest of the baseline or an explicit warning is raised. + +--- + +### User Story 2 — Curated trio with observable dedup (Priority: P2) + +A developer wants to measure the shared-scratchpad duplicate-work detection primitive. The trio file contains three carefully selected MuSiQue questions sharing a pivot bridge entity (e.g., all involve the United States as an intermediate hop). When the harness runs the trio in sequence with a persistent shared scratchpad across questions, the second and third questions should hit cached entity lookups from the first, producing an observable cache-hit rate above 0%. + +**Why this priority**: Dedup is one of the five MAS features the benchmark is supposed to exercise. Without the curated trio, dedup metrics are 0 at N=3 and the primitive is silent. This is P2 because US1 delivers single-question value first; the trio extends it. + +**Independent Test**: Run `python benchmark/run.py --mode trio --scratchpad persistent`. The harness runs all three questions in sequence. The scratchpad stats show at least 5 cache hits on entity lookups across the three questions, and the aggregate token count is lower than 3× the single-question token count due to cache re-use. + +**Acceptance Scenarios**: + +1. **Given** the curated trio.jsonl exists with three questions sharing a pivot entity, **When** the harness runs the trio with persistent scratchpad, **Then** cache-hit count at the end of the run is greater than zero. +2. **Given** the same trio is run with a non-persistent scratchpad, **When** the run completes, **Then** total tokens are strictly greater than the persistent-scratchpad run. + +--- + +### User Story 3 — Learning tier with reputation convergence (Priority: P3) + +A developer wants to verify that reputation and reflection actually cause improvement over time. The learning tier runs the same trio for 5 epochs. Reputation and capability manifests persist; scratchpad resets per epoch. Tokens-per-correct-answer should decline monotonically across epochs as the marketplace learns which model tier is best for which sub-task. + +**Why this priority**: This is the full marketplace proof-of-learning. It is P3 because US1 and US2 are load-bearing; the learning tier depends on both being stable first, and it is the most token-expensive mode (~500k tokens per full run). + +**Independent Test**: Run `python benchmark/run.py --mode learning --epochs 5`. The harness completes all 5 epochs, writes a results JSON containing per-epoch tokens and F1, and plots a learning curve showing the tokens-per-correct-answer trend. + +**Acceptance Scenarios**: + +1. **Given** the learning tier has run 5 epochs on the trio, **When** the per-epoch metrics are plotted, **Then** the tokens-per-correct-answer for epoch 5 is less than epoch 1. +2. **Given** between-epoch state, **When** epoch N starts, **Then** reputation entries from epoch N-1 are present and influence bid scoring, while the scratchpad is empty. + +--- + +### User Story 4 — Rich HTML report (Priority: P1) + +The benchmark run produces a rich, self-contained HTML report that shows the task, the decomposition, the bids, the award, the execution trace, the gold answer, the marketplace answer, the baseline answer, the Pareto plot, and all token counts. Non-technical readers can open the file and understand what happened. + +**Why this priority**: The benchmark has to serve as a compelling illustration for the `multiagent_systems_report.html`. Without a rendered report, the run is just a JSON file nobody reads. It is P1 because without it the benchmark fails its dual-use requirement. + +**Independent Test**: After a successful run, `results/latest.html` exists, opens in a browser without errors, and visibly contains the question text, the Pareto plot (as SVG or inline data), and the per-run token counts. + +**Acceptance Scenarios**: + +1. **Given** a completed benchmark run, **When** the report generator runs, **Then** `results/latest.html` is written and contains the question, decomposition, answers, tokens, and Pareto data. +2. **Given** the HTML report is opened in a browser, **When** a reader scrolls through it, **Then** they can tell whether the marketplace passed without needing to open any JSON files. + +--- + +### Edge Cases + +- **MuSiQue download failure**: the setup script retries once, then fails with a clear error pointing at the canonical URL. +- **No agents bid within window**: the harness declares auction-timeout, records it as a test failure, and exits non-zero. +- **Agent returns malformed bid**: the bid is rejected by schema validation, the agent is not awarded the task, other bidders still compete. +- **Winner exceeds budget**: the harness hard-stops execution at 100% of declared budget, records the auto-fail, baseline wins the Pareto comparison. +- **Scratchpad cache conflict**: two sub-agents write different values for the same entity key — last-writer-wins with a logged warning; the answer reflects the final state. +- **Single-agent baseline also fails**: this is a valid outcome and means the question is genuinely beyond frontier model capability; the report marks the question as such and the benchmark still reports its tokens. +- **Learning tier plateau**: if epoch 5 is not strictly better than epoch 1, the report raises a warning and the test fails. +- **Model API rate limit**: the harness retries with exponential backoff, up to 3 attempts per call. +- **Claude Agent SDK unavailable**: falls back to direct Anthropic SDK calls with manual tool-use formatting. + +## Requirements *(mandatory)* + +### Functional Requirements + +**Harness Core** + +- **FR-001**: The harness MUST accept a `--mode` flag with values `single-shot`, `trio`, or `learning`. +- **FR-002**: The harness MUST read a configuration file defining agent pool composition (model tiers, initial skill cards). +- **FR-003**: The harness MUST talk to SynapBus via the existing MCP protocol using the 016 marketplace tools. +- **FR-004**: The harness MUST record per-run metrics in a structured JSON file under `results/`. + +**Dataset Preparation** + +- **FR-005**: A setup script MUST download the MuSiQue-Ans dev set from the canonical GitHub release. +- **FR-006**: A curation script MUST select three questions from the 4-hop subset that share a pivot bridge entity and write them to `benchmark/trio.jsonl` with gold answers and decomposition DAGs. +- **FR-007**: The curation rule MUST be deterministic with a fixed seed so the trio is reproducible. + +**Agent Runner** + +- **FR-008**: Agents MUST run as isolated Python processes using the Claude Agent SDK where available; when the SDK is unavailable, direct Anthropic SDK calls MUST be used as fallback. +- **FR-009**: Each agent MUST publish a capability manifest to SynapBus wiki at startup. +- **FR-010**: Each agent MUST poll the auction channel and submit bids when it sees a task in a declared domain. +- **FR-011**: Each agent MUST respect the max_budget_tokens of its awarded tasks and hard-stop at 100% of budget. + +**Baseline** + +- **FR-012**: The harness MUST run a single-agent baseline for every question, passing all 20 distractor paragraphs plus the question in one prompt with chain-of-thought instructions. +- **FR-013**: The baseline MUST record its token count and final F1 so the Pareto comparison is apples-to-apples. + +**Scoring** + +- **FR-014**: For each question, the harness MUST compute exact-match F1 against the gold answer string after normalization (lowercasing, punctuation removal, article stripping). +- **FR-015**: The harness MUST compute decomposition-F1 by comparing the marketplace's actual sub-questions with the gold DAG. +- **FR-016**: The harness MUST record total tokens consumed per run, separated by orchestrator, sub-agents, and baseline. +- **FR-017**: The harness MUST compute tokens-per-correct-answer as `total_tokens / max(correct_count, 1)`. +- **FR-018**: The harness MUST report a Pareto verdict: `PASS` if the marketplace is strictly northwest of the baseline, `FAIL` otherwise. + +**Report Generation** + +- **FR-019**: After every run, the harness MUST generate a self-contained HTML report at `results/{run_id}.html` showing the question, decomposition, bids, award, answer, gold, tokens, Pareto plot, and verdict. +- **FR-020**: The report MUST be readable without requiring any external server or CDN. + +**Learning Tier (P3, deferred for MVP)** + +- **FR-021**: In learning mode, the harness MUST run the trio for N epochs with reputation and skill-card state persisting across epochs. +- **FR-022**: In learning mode, the scratchpad MUST reset between epochs. +- **FR-023**: The learning report MUST plot tokens-per-correct-answer across epochs. + +### Key Entities + +- **Question record**: One entry in trio.jsonl. Contains `id`, `question`, `gold_answer`, `gold_decomposition` (list of sub-questions), `distractor_paragraphs` (20 Wikipedia snippets), `pivot_entity`. +- **Agent config**: One agent in the pool. Contains `name`, `model_id`, `system_prompt_template`, `skill_card_markdown`, `domains`, `initial_confidence_per_domain`, `initial_cost_per_domain`. +- **Run result**: One benchmark run output. Contains `run_id`, `timestamp`, `mode`, `per_question_results`, `baseline_results`, `pareto_verdict`, `total_tokens`, `wall_time_seconds`. +- **Scratchpad entry**: One cached `(entity, attribute, value)` tuple with a `cache_hit_count` counter. + +## Success Criteria *(mandatory)* + +### Measurable Outcomes + +- **SC-001**: On a fresh SynapBus instance with 016 MVP loaded, `python benchmark/run.py --mode single-shot --question q1` completes successfully within 10 minutes wall-time and 150k total tokens. +- **SC-002**: The Pareto verdict on the single-shot run is `PASS` — the marketplace lands strictly northwest of the Sonnet-baseline point. +- **SC-003**: The HTML report is generated, is self-contained (no external asset fetches), and visibly shows the question, both answers, tokens, and Pareto plot. +- **SC-004**: The benchmark can be re-run deterministically on the same machine and produce the same Pareto verdict (token counts may vary ±5% due to sampling, but verdict must be stable). +- **SC-005**: The benchmark identifies and reports per-agent, per-domain reputation entries after the run, writable back to SynapBus. +- **SC-006**: In trio mode, the scratchpad cache hit count is greater than zero — the dedup primitive is observably exercised. +- **SC-007** (deferred): In learning mode, tokens-per-correct-answer declines from epoch 1 to epoch 5 by at least 15%. + +## Assumptions + +- The 016-agent-marketplace MVP is implemented before this benchmark runs, or this benchmark provides a mock marketplace stub for early validation. +- MuSiQue-Ans is available at its canonical GitHub release URL and is downloadable without authentication. +- Anthropic API access is available via `ANTHROPIC_API_KEY` environment variable. +- Claude Agent SDK may or may not be available in the target environment; the harness degrades gracefully to direct Anthropic SDK if not. +- Tokens reported by the Anthropic SDK are trusted as ground truth. +- F1 computed via string normalization is sufficient for MVP; semantic similarity matching is future work. +- The MuSiQue license (CC BY 4.0) permits redistribution of the curated trio file as a derivative work. + +## Out of Scope + +- Training any model or fine-tuning. +- Non-English questions (MuSiQue is English-only). +- Multi-modal questions (images, audio). +- FRAMES cross-eval — scheduled as a follow-up once MuSiQue pipeline is stable. +- Adversarial or byzantine agent behavior; agents in the pool are assumed cooperative. +- Replacing the MuSiQue corpus with a live Wikipedia API. +- Multi-run statistical power analysis; N=3 is mechanism verification only. From 02b8548eac7c1e117844fecb8d2cd4425c279423 Mon Sep 17 00:00:00 2001 From: Algis Dumbris Date: Sat, 11 Apr 2026 15:12:19 +0300 Subject: [PATCH 2/2] =?UTF-8?q?feat(017):=20MuSiQue=20benchmark=20harness?= =?UTF-8?q?=20=E2=80=94=20marketplace=20stub,=20mixed-tier=20agents,=20Par?= =?UTF-8?q?eto=20scoring,=20HTML=20report?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit MVP implementation of the MuSiQue multi-agent benchmark (spec 017): - benchmark/setup.py: downloads musique_v1.0.zip from the canonical Google Drive source (mirrors upstream download_data.sh). Idempotent. - benchmark/curate.py: deterministic selection of 3 4-hop questions from the dev set sharing a US pivot entity; writes benchmark/trio.jsonl. - benchmark/marketplace.py: in-process 016-marketplace stub with post_auction / bid / award / mark_done / query_reputation and a domain-scoped reputation ledger. Designed for mechanical swap to real SynapBus MCP tools. - benchmark/agents.py: HaikuAgent + SonnetAgent, using the official anthropic SDK (no Claude Agent SDK, no subprocesses). Models pinned to claude-haiku-4-5-20251001 and claude-sonnet-4-6. - benchmark/baseline.py: single Sonnet call with all 20 distractors plus chain-of-thought. - benchmark/score.py: SQuAD-style normalized F1 + Pareto verdict (strictly northwest = PASS). - benchmark/run.py: main entry. --mode single-shot, --question, --dry-run. - benchmark/report.py: self-contained HTML with inline SVG scatter plot. - benchmark/trio.jsonl: curated reproducible trio (all three converge on "Treaty of Paris" US territory cession). Verified with benchmark/run.py --dry-run end-to-end; all 8 files py_compile clean. Real-token execution is deferred to the user's main session. Co-Authored-By: Claude Opus 4.6 (1M context) --- .gitignore | 6 + benchmark/agents.py | 256 +++++++++++++++++++++++++++++ benchmark/baseline.py | 116 +++++++++++++ benchmark/curate.py | 169 +++++++++++++++++++ benchmark/marketplace.py | 192 ++++++++++++++++++++++ benchmark/report.py | 324 +++++++++++++++++++++++++++++++++++++ benchmark/requirements.txt | 2 + benchmark/run.py | 277 +++++++++++++++++++++++++++++++ benchmark/score.py | 86 ++++++++++ benchmark/setup.py | 198 +++++++++++++++++++++++ benchmark/trio.jsonl | 3 + 11 files changed, 1629 insertions(+) create mode 100644 benchmark/agents.py create mode 100644 benchmark/baseline.py create mode 100644 benchmark/curate.py create mode 100644 benchmark/marketplace.py create mode 100644 benchmark/report.py create mode 100644 benchmark/requirements.txt create mode 100644 benchmark/run.py create mode 100644 benchmark/score.py create mode 100644 benchmark/setup.py create mode 100644 benchmark/trio.jsonl diff --git a/.gitignore b/.gitignore index cd39320..2f3079f 100644 --- a/.gitignore +++ b/.gitignore @@ -41,6 +41,12 @@ Thumbs.db __pycache__/ *.pyc +# Benchmark (017) — large datasets, per-run outputs, local venvs +benchmark/data/ +benchmark/results/ +.venv-bench/ +.venv/ + # Debug __debug_bin* .claude/worktrees/ diff --git a/benchmark/agents.py b/benchmark/agents.py new file mode 100644 index 0000000..a335689 --- /dev/null +++ b/benchmark/agents.py @@ -0,0 +1,256 @@ +#!/usr/bin/env python3 +""" +Agent pool for the MuSiQue benchmark. + +Two agents: + - haiku-agent (claude-haiku-4-5-20251001) + - sonnet-agent (claude-sonnet-4-6) + +Each agent exposes: + - name, model, skill_card + - bid(task) -> {estimated_tokens, confidence, approach} + - execute(task, paragraphs) -> {answer, actual_tokens} + +Design notes: +- We use the official ``anthropic`` Python SDK directly (NOT the + Claude Agent SDK). Simpler, no subprocesses, reliable token accounting. +- ``bid()`` is pure Python — it is a cheap heuristic so the marketplace + has something to pick from. Real 016 agents would emit a structured + reply. For MVP, heuristic bids are sufficient to exercise the auction + primitive. +- ``execute()`` is the only thing that actually burns tokens. +- ``--dry-run`` in run.py never calls execute(); it uses stub responses. +""" + +from __future__ import annotations + +import os +from dataclasses import dataclass +from typing import Any + +try: + import anthropic # type: ignore +except ImportError: # pragma: no cover + anthropic = None # type: ignore + + +HAIKU_MODEL = "claude-haiku-4-5-20251001" +SONNET_MODEL = "claude-sonnet-4-6" + + +HAIKU_SKILL_CARD = """\ +# haiku-agent + +A fast, cheap agent best for single-hop fact lookups and short +extractive answers. Accepts multi-paragraph context but may miss +subtle bridging entities on 4-hop questions. Very low cost per call. + +Domains: factual-lookup, extraction, summarization +""" + +SONNET_SKILL_CARD = """\ +# sonnet-agent + +A deliberate mid-tier agent well-suited to multi-hop reasoning with +explicit chain-of-thought. Handles 4-hop MuSiQue questions with +decomposition when the context fits in one prompt. Higher cost per call +than Haiku but meaningfully better F1 on bridging questions. + +Domains: multi-hop-qa, decomposition, reasoning +""" + + +SYSTEM_PROMPT = """\ +You are a careful question-answering agent working on a MuSiQue +multi-hop benchmark. You are given a question and a set of numbered +paragraphs. Only a few of the paragraphs are relevant; the rest are +distractors. + +Think step by step and cite the paragraphs you used. Then output a +final line starting with exactly: + +ANSWER: + +Your final answer must be a short entity or phrase — not a sentence. +""" + + +@dataclass +class BidResult: + estimated_tokens: int + confidence: float + approach: str + + def to_dict(self) -> dict[str, Any]: + return { + "estimated_tokens": self.estimated_tokens, + "confidence": self.confidence, + "approach": self.approach, + } + + +@dataclass +class ExecuteResult: + answer: str + actual_tokens: int + raw_text: str = "" + + def to_dict(self) -> dict[str, Any]: + return { + "answer": self.answer, + "actual_tokens": self.actual_tokens, + } + + +class Agent: + name: str + model: str + skill_card: str + + def __init__(self, name: str, model: str, skill_card: str) -> None: + self.name = name + self.model = model + self.skill_card = skill_card + + # ---- bidding ----------------------------------------------------------- + + def bid(self, task: dict[str, Any]) -> BidResult: + raise NotImplementedError + + # ---- execution --------------------------------------------------------- + + def execute( + self, + task: dict[str, Any], + paragraphs: list[str], + *, + dry_run: bool = False, + max_budget_tokens: int = 100_000, + ) -> ExecuteResult: + question = task["question"] + prompt = self._build_prompt(question, paragraphs) + + if dry_run: + stub = ( + "Thinking step by step... [dry-run stub]\n" + f"ANSWER: [stub answer from {self.name}]" + ) + # Rough estimate: 1 token ~= 4 characters. + est = max(256, len(prompt) // 4 + 64) + return ExecuteResult( + answer=self._extract_answer(stub), + actual_tokens=est, + raw_text=stub, + ) + + if anthropic is None: + raise RuntimeError( + "anthropic SDK not installed — pip install anthropic" + ) + api_key = os.environ.get("ANTHROPIC_API_KEY") + if not api_key: + raise RuntimeError( + "ANTHROPIC_API_KEY not set. Use --dry-run to stub it out." + ) + + client = anthropic.Anthropic(api_key=api_key) + # Cap max_tokens to min(1024, budget/2) so the worst case is tame. + max_tokens = min(1024, max(128, max_budget_tokens // 2)) + msg = client.messages.create( + model=self.model, + max_tokens=max_tokens, + system=SYSTEM_PROMPT, + messages=[{"role": "user", "content": prompt}], + ) + text_parts: list[str] = [] + for block in msg.content: + t = getattr(block, "text", None) + if t: + text_parts.append(t) + text = "\n".join(text_parts).strip() + + usage = getattr(msg, "usage", None) + actual = 0 + if usage is not None: + actual = ( + getattr(usage, "input_tokens", 0) + + getattr(usage, "output_tokens", 0) + ) + return ExecuteResult( + answer=self._extract_answer(text), + actual_tokens=int(actual), + raw_text=text, + ) + + # ---- helpers ----------------------------------------------------------- + + def _build_prompt( + self, question: str, paragraphs: list[str] + ) -> str: + body = ["Paragraphs:"] + for i, p in enumerate(paragraphs, start=1): + body.append(f"[{i}] {p}") + body.append("") + body.append(f"Question: {question}") + body.append("") + body.append("Think step by step, then output your final ANSWER: line.") + return "\n".join(body) + + def _extract_answer(self, text: str) -> str: + if not text: + return "" + for line in reversed(text.splitlines()): + line = line.strip() + if line.upper().startswith("ANSWER:"): + return line.split(":", 1)[1].strip() + # Fallback: last non-empty line. + for line in reversed(text.splitlines()): + line = line.strip() + if line: + return line + return "" + + +class HaikuAgent(Agent): + def __init__(self) -> None: + super().__init__( + name="haiku-agent", + model=HAIKU_MODEL, + skill_card=HAIKU_SKILL_CARD, + ) + + def bid(self, task: dict[str, Any]) -> BidResult: + # Cheap, low confidence on multi-hop bridging. + return BidResult( + estimated_tokens=4_000, + confidence=0.45, + approach=( + "Extract candidate entities from the paragraphs and " + "answer directly; may miss 4-hop bridges." + ), + ) + + +class SonnetAgent(Agent): + def __init__(self) -> None: + super().__init__( + name="sonnet-agent", + model=SONNET_MODEL, + skill_card=SONNET_SKILL_CARD, + ) + + def bid(self, task: dict[str, Any]) -> BidResult: + # More expensive, higher confidence on multi-hop. + return BidResult( + estimated_tokens=12_000, + confidence=0.80, + approach=( + "Decompose the question into sub-questions, resolve each " + "sub-answer against the paragraphs, then compose the final " + "bridged answer." + ), + ) + + +def default_pool() -> list[Agent]: + return [HaikuAgent(), SonnetAgent()] diff --git a/benchmark/baseline.py b/benchmark/baseline.py new file mode 100644 index 0000000..6d7bbc6 --- /dev/null +++ b/benchmark/baseline.py @@ -0,0 +1,116 @@ +#!/usr/bin/env python3 +""" +Single-agent baseline: one Anthropic API call to claude-sonnet-4-6 with +the question and all 20 distractor paragraphs plus chain-of-thought +instructions. No decomposition, no marketplace, no tools. + +Returns {"answer": str, "tokens": int, "raw_text": str}. +""" + +from __future__ import annotations + +import os +from typing import Any + +try: + import anthropic # type: ignore +except ImportError: # pragma: no cover + anthropic = None # type: ignore + + +BASELINE_MODEL = "claude-sonnet-4-6" + +BASELINE_SYSTEM = """\ +You are a careful multi-hop QA system. Given a question and a set of +numbered paragraphs (some irrelevant distractors), think step by step +and answer. + +Output your reasoning first, then on a final line: + +ANSWER: +""" + + +def _build_prompt(question: str, paragraphs: list[str]) -> str: + parts = ["Paragraphs:"] + for i, p in enumerate(paragraphs, start=1): + parts.append(f"[{i}] {p}") + parts.append("") + parts.append(f"Question: {question}") + parts.append("") + parts.append( + "Work through the reasoning step by step, then give your " + "final ANSWER: line." + ) + return "\n".join(parts) + + +def _extract_answer(text: str) -> str: + if not text: + return "" + for line in reversed(text.splitlines()): + line = line.strip() + if line.upper().startswith("ANSWER:"): + return line.split(":", 1)[1].strip() + for line in reversed(text.splitlines()): + line = line.strip() + if line: + return line + return "" + + +def run_baseline( + question: str, + paragraphs: list[str], + *, + dry_run: bool = False, + max_output_tokens: int = 1024, +) -> dict[str, Any]: + prompt = _build_prompt(question, paragraphs) + + if dry_run: + stub = ( + "Step 1: scanning paragraphs... [dry-run stub]\n" + "Step 2: picking the most likely entity...\n" + "ANSWER: [stub baseline answer]" + ) + est = max(512, len(prompt) // 4 + 128) + return { + "answer": _extract_answer(stub), + "tokens": est, + "raw_text": stub, + "model": BASELINE_MODEL, + } + + if anthropic is None: + raise RuntimeError("anthropic SDK not installed") + api_key = os.environ.get("ANTHROPIC_API_KEY") + if not api_key: + raise RuntimeError("ANTHROPIC_API_KEY not set") + + client = anthropic.Anthropic(api_key=api_key) + msg = client.messages.create( + model=BASELINE_MODEL, + max_tokens=max_output_tokens, + system=BASELINE_SYSTEM, + messages=[{"role": "user", "content": prompt}], + ) + text_parts: list[str] = [] + for block in msg.content: + t = getattr(block, "text", None) + if t: + text_parts.append(t) + text = "\n".join(text_parts).strip() + usage = getattr(msg, "usage", None) + tokens = 0 + if usage is not None: + tokens = ( + getattr(usage, "input_tokens", 0) + + getattr(usage, "output_tokens", 0) + ) + return { + "answer": _extract_answer(text), + "tokens": int(tokens), + "raw_text": text, + "model": BASELINE_MODEL, + } diff --git a/benchmark/curate.py b/benchmark/curate.py new file mode 100644 index 0000000..61d4208 --- /dev/null +++ b/benchmark/curate.py @@ -0,0 +1,169 @@ +#!/usr/bin/env python3 +""" +Curate a deterministic trio of MuSiQue 4-hop questions that share a +pivot entity. For MVP we pivot on the United States. + +Input: benchmark/data/musique_ans_v1.0_dev.jsonl +Output: benchmark/trio.jsonl + +Each output record: + { + "id": str, + "question": str, + "answer": str, + "decomposition": [{"question": str, "answer": str}, ...], + "paragraphs": [str, ...] # up to 20 distractor snippets + } + +MuSiQue dev records typically look like:: + + { + "id": "4hop1__...", + "question": "...", + "question_decomposition": [ + {"id": N, "question": "...", "answer": "...", + "paragraph_support_idx": int}, + ... + ], + "answer": "...", + "answer_aliases": [...], + "paragraphs": [ + {"idx": int, "title": "...", "paragraph_text": "...", + "is_supporting": bool}, + ... + ] + } + +The curation rule is deterministic (fixed input ordering; first 3 matches). +""" + +from __future__ import annotations + +import json +import sys +from pathlib import Path + +DATA_FILE = Path(__file__).resolve().parent / "data" / "musique_ans_v1.0_dev.jsonl" +OUT_FILE = Path(__file__).resolve().parent / "trio.jsonl" + +PIVOT_TOKENS = ("united states", "u.s.", " us ", "america", "american") +N_QUESTIONS = 3 +MAX_PARAGRAPHS = 20 + + +def _normalized(s: str) -> str: + return f" {s.lower()} " + + +def _mentions_pivot(record: dict) -> bool: + blob_parts = [record.get("question", ""), record.get("answer", "")] + for sub in record.get("question_decomposition", []) or []: + blob_parts.append(sub.get("question", "")) + blob_parts.append(sub.get("answer", "")) + blob = _normalized(" ".join(str(x) for x in blob_parts if x)) + return any(tok in blob for tok in PIVOT_TOKENS) + + +def _is_4hop(record: dict) -> bool: + rid = record.get("id", "") + if isinstance(rid, str) and rid.startswith("4hop"): + return True + # Fall back: count decomposition hops. + decomp = record.get("question_decomposition") or [] + return len(decomp) == 4 + + +def _trim_paragraphs(record: dict, limit: int) -> list[str]: + out: list[str] = [] + for p in record.get("paragraphs", []) or []: + title = (p.get("title") or "").strip() + text = (p.get("paragraph_text") or "").strip() + if not text: + continue + snippet = f"[{title}] {text}" if title else text + out.append(snippet) + if len(out) >= limit: + break + return out + + +def _simplify_decomp(record: dict) -> list[dict]: + out = [] + for sub in record.get("question_decomposition", []) or []: + out.append( + { + "question": sub.get("question", ""), + "answer": sub.get("answer", ""), + } + ) + return out + + +def curate() -> int: + if not DATA_FILE.exists(): + print( + f"[curate] ERROR: {DATA_FILE} not found. Run setup.py first.", + file=sys.stderr, + ) + return 2 + + selected: list[dict] = [] + total_scanned = 0 + total_4hop = 0 + total_pivot = 0 + + with open(DATA_FILE, "r", encoding="utf-8") as f: + for line in f: + line = line.strip() + if not line: + continue + total_scanned += 1 + try: + rec = json.loads(line) + except json.JSONDecodeError: + continue + if not _is_4hop(rec): + continue + total_4hop += 1 + if not _mentions_pivot(rec): + continue + total_pivot += 1 + + trio_record = { + "id": rec.get("id", f"q{len(selected)+1}"), + "question": rec.get("question", ""), + "answer": rec.get("answer", ""), + "answer_aliases": rec.get("answer_aliases", []), + "decomposition": _simplify_decomp(rec), + "paragraphs": _trim_paragraphs(rec, MAX_PARAGRAPHS), + } + selected.append(trio_record) + if len(selected) >= N_QUESTIONS: + break + + print( + f"[curate] scanned={total_scanned} 4hop={total_4hop} " + f"pivot-matches={total_pivot} kept={len(selected)}" + ) + + if len(selected) < N_QUESTIONS: + print( + f"[curate] ERROR: wanted {N_QUESTIONS} questions, " + f"found {len(selected)}", + file=sys.stderr, + ) + return 3 + + OUT_FILE.parent.mkdir(parents=True, exist_ok=True) + with open(OUT_FILE, "w", encoding="utf-8") as f: + for i, rec in enumerate(selected, start=1): + # Attach a stable short id q1/q2/q3 in addition to MuSiQue's id. + rec["short_id"] = f"q{i}" + f.write(json.dumps(rec, ensure_ascii=False) + "\n") + + print(f"[curate] wrote {OUT_FILE} ({len(selected)} records)") + return 0 + + +if __name__ == "__main__": + raise SystemExit(curate()) diff --git a/benchmark/marketplace.py b/benchmark/marketplace.py new file mode 100644 index 0000000..4daab21 --- /dev/null +++ b/benchmark/marketplace.py @@ -0,0 +1,192 @@ +#!/usr/bin/env python3 +""" +In-process stub of the 016-agent-marketplace primitives. + +*** IMPORTANT *** +This module is an in-process stand-in for the SynapBus-hosted 016 +marketplace. The follow-up deliverable after this MVP is to replace the +bodies of these functions with calls to the real SynapBus MCP tools +(``post_auction``, ``bid``, ``award``, ``mark_done``, and +``query_reputation``) once 016 lands. The public API here deliberately +mirrors those tool names so the swap is mechanical. + +Scope for MVP: +- In-memory auctions, bids, awards, and done records +- Domain-scoped reputation ledger stored in a dict +- No persistence, no concurrency — single process, single thread +- No schema enforcement beyond a couple of shape checks + +The harness (``run.py``) holds a single ``Marketplace`` instance. +""" + +from __future__ import annotations + +import itertools +import time +from dataclasses import dataclass, field +from typing import Any + + +@dataclass +class Auction: + auction_id: str + task: dict[str, Any] + domain: str + max_budget_tokens: int + posted_at: float + bids: list[dict[str, Any]] = field(default_factory=list) + awarded_to: str | None = None + result: dict[str, Any] | None = None + + +@dataclass +class ReputationEntry: + agent: str + domain: str + runs: int = 0 + correct: int = 0 + tokens_spent: int = 0 + + def score(self) -> float: + if self.runs == 0: + return 0.5 # prior + quality = self.correct / self.runs + avg_tokens = self.tokens_spent / self.runs + # Arbitrary: quality dominates, tokens slightly penalize. + return quality - min(avg_tokens / 200_000.0, 0.3) + + +class Marketplace: + """In-process marketplace stub.""" + + def __init__(self) -> None: + self._auctions: dict[str, Auction] = {} + self._reputation: dict[tuple[str, str], ReputationEntry] = {} + self._counter = itertools.count(1) + + # ---- auction lifecycle ------------------------------------------------- + + def post_auction( + self, + task: dict[str, Any], + domain: str, + max_budget_tokens: int, + ) -> str: + auction_id = f"auction-{next(self._counter)}" + self._auctions[auction_id] = Auction( + auction_id=auction_id, + task=dict(task), + domain=domain, + max_budget_tokens=max_budget_tokens, + posted_at=time.time(), + ) + return auction_id + + def bid( + self, + auction_id: str, + agent: str, + estimated_tokens: int, + confidence: float, + approach: str, + ) -> None: + auction = self._auctions[auction_id] + if auction.awarded_to is not None: + raise RuntimeError(f"auction {auction_id} already awarded") + auction.bids.append( + { + "agent": agent, + "estimated_tokens": int(estimated_tokens), + "confidence": float(confidence), + "approach": approach, + "submitted_at": time.time(), + } + ) + + def list_bids(self, auction_id: str) -> list[dict[str, Any]]: + return list(self._auctions[auction_id].bids) + + def score_bid(self, auction_id: str, bid: dict[str, Any]) -> float: + """ + Lower is better (we're minimizing tokens per unit confidence), + but we add a reputation adjustment that rewards agents with a + track record in this domain. + """ + auction = self._auctions[auction_id] + rep = self._reputation.get((bid["agent"], auction.domain)) + rep_score = rep.score() if rep else 0.5 + conf = max(bid["confidence"], 1e-3) + # Cost per confidence, lightly discounted by reputation. + raw = bid["estimated_tokens"] / conf + return raw * (1.15 - 0.3 * rep_score) + + def award(self, auction_id: str) -> dict[str, Any]: + auction = self._auctions[auction_id] + if not auction.bids: + raise RuntimeError(f"auction {auction_id} has no bids") + if auction.awarded_to is not None: + raise RuntimeError(f"auction {auction_id} already awarded") + best = min( + auction.bids, + key=lambda b: self.score_bid(auction_id, b), + ) + auction.awarded_to = best["agent"] + return best + + def mark_done( + self, + auction_id: str, + answer: str, + actual_tokens: int, + correct: bool, + ) -> None: + auction = self._auctions[auction_id] + if auction.awarded_to is None: + raise RuntimeError(f"auction {auction_id} not awarded yet") + auction.result = { + "answer": answer, + "actual_tokens": int(actual_tokens), + "correct": bool(correct), + } + key = (auction.awarded_to, auction.domain) + entry = self._reputation.get(key) or ReputationEntry( + agent=auction.awarded_to, domain=auction.domain + ) + entry.runs += 1 + entry.tokens_spent += int(actual_tokens) + if correct: + entry.correct += 1 + self._reputation[key] = entry + + # ---- reputation -------------------------------------------------------- + + def query_reputation( + self, agent: str, domain: str + ) -> dict[str, Any]: + rep = self._reputation.get((agent, domain)) + if rep is None: + return { + "agent": agent, + "domain": domain, + "runs": 0, + "correct": 0, + "tokens_spent": 0, + "score": 0.5, + } + return { + "agent": rep.agent, + "domain": rep.domain, + "runs": rep.runs, + "correct": rep.correct, + "tokens_spent": rep.tokens_spent, + "score": rep.score(), + } + + def all_reputation(self) -> list[dict[str, Any]]: + return [ + self.query_reputation(rep.agent, rep.domain) + for rep in self._reputation.values() + ] + + def auction(self, auction_id: str) -> Auction: + return self._auctions[auction_id] diff --git a/benchmark/report.py b/benchmark/report.py new file mode 100644 index 0000000..fd54c87 --- /dev/null +++ b/benchmark/report.py @@ -0,0 +1,324 @@ +#!/usr/bin/env python3 +""" +Self-contained HTML report generator. + +Renders a single HTML file with inline styles and an inline SVG scatter +plot. No external assets, no CDN calls, nothing to fetch. Safe to open +directly in a browser. +""" + +from __future__ import annotations + +import html +from pathlib import Path +from typing import Any + + +def _esc(s: Any) -> str: + return html.escape(str(s if s is not None else "")) + + +def _scatter_svg( + market_tokens: int, + market_f1: float, + baseline_tokens: int, + baseline_f1: float, + *, + width: int = 520, + height: int = 320, +) -> str: + pad_l, pad_r, pad_t, pad_b = 70, 30, 30, 50 + plot_w = width - pad_l - pad_r + plot_h = height - pad_t - pad_b + + max_tokens = max(market_tokens, baseline_tokens, 1) + # Give a little headroom so points aren't on the axis. + max_tokens_axis = max_tokens * 1.15 + min_tokens_axis = 0 + + def sx(tokens: float) -> float: + frac = (tokens - min_tokens_axis) / max( + max_tokens_axis - min_tokens_axis, 1 + ) + return pad_l + frac * plot_w + + def sy(f1: float) -> float: + # y=0 at top of plot, y=1 at bottom -> invert + return pad_t + (1.0 - max(0.0, min(1.0, f1))) * plot_h + + axis_color = "#555" + grid_color = "#eee" + market_color = "#2563eb" + baseline_color = "#dc2626" + + parts: list[str] = [] + parts.append( + f'' + ) + parts.append( + f'' + ) + # Gridlines at F1 = 0, 0.25, 0.5, 0.75, 1.0 + for f in (0.0, 0.25, 0.5, 0.75, 1.0): + y = sy(f) + parts.append( + f'' + ) + parts.append( + f'' + f'{f:.2f}' + ) + # X-axis ticks + for frac in (0.0, 0.25, 0.5, 0.75, 1.0): + t_val = frac * max_tokens_axis + x = sx(t_val) + parts.append( + f'' + ) + parts.append( + f'{int(t_val)}' + ) + # Axis lines + parts.append( + f'' + ) + parts.append( + f'' + ) + # Axis labels + parts.append( + f'tokens' + ) + parts.append( + f'F1' + ) + # Baseline point + bx, by = sx(baseline_tokens), sy(baseline_f1) + parts.append( + f'' + ) + parts.append( + f'baseline' + ) + # Market point + mx, my = sx(market_tokens), sy(market_f1) + parts.append( + f'' + ) + parts.append( + f'marketplace' + ) + + parts.append("") + return "".join(parts) + + +CSS = """\ +body { font-family: -apple-system, system-ui, sans-serif; + max-width: 960px; margin: 2rem auto; padding: 0 1rem; + color: #1f2937; line-height: 1.55; } +h1, h2, h3 { color: #111827; } +h1 { border-bottom: 2px solid #2563eb; padding-bottom: .4rem; } +.verdict-pass { display: inline-block; background: #dcfce7; + color: #166534; padding: .3rem .8rem; border-radius: 6px; + font-weight: 600; } +.verdict-fail { display: inline-block; background: #fee2e2; + color: #991b1b; padding: .3rem .8rem; border-radius: 6px; + font-weight: 600; } +table { border-collapse: collapse; margin: .8rem 0; width: 100%; } +th, td { border: 1px solid #e5e7eb; padding: .4rem .6rem; + text-align: left; vertical-align: top; } +th { background: #f9fafb; } +pre, code { background: #f3f4f6; border-radius: 4px; + padding: .1rem .4rem; font-size: .9rem; } +pre { padding: .8rem; white-space: pre-wrap; word-break: break-word; } +.card { border: 1px solid #e5e7eb; border-radius: 8px; + padding: 1rem 1.2rem; margin: 1rem 0; background: #fff; } +.kv { display: grid; grid-template-columns: 180px 1fr; gap: .3rem .8rem; } +.small { color: #6b7280; font-size: .88rem; } +""" + + +def render_report(data: dict[str, Any], out_path: Path) -> None: + verdict = data.get("pareto", {}) + is_pass = verdict.get("verdict") == "PASS" + verdict_html = ( + 'PASS — strictly northwest' + if is_pass + else 'FAIL — not dominating baseline' + ) + + bids_rows: list[str] = [] + for b in data.get("bids", []): + conf_str = "{:.2f}".format(b.get("confidence", 0) or 0) + bids_rows.append( + f"{_esc(b.get('agent'))}" + f"{_esc(b.get('estimated_tokens'))}" + f"{_esc(conf_str)}" + f"{_esc(b.get('approach'))}" + ) + bids_table = "\n".join(bids_rows) or ( + "no bids" + ) + + decomp_rows: list[str] = [] + for i, sub in enumerate(data.get("decomposition", []) or [], start=1): + decomp_rows.append( + f"{i}{_esc(sub.get('question'))}" + f"{_esc(sub.get('answer'))}" + ) + decomp_table = "\n".join(decomp_rows) or ( + "(none)" + ) + + rep_rows: list[str] = [] + for rep in data.get("reputation", []) or []: + score_str = "{:.3f}".format(rep.get("score", 0) or 0) + rep_rows.append( + f"{_esc(rep.get('agent'))}" + f"{_esc(rep.get('domain'))}" + f"{_esc(rep.get('runs'))}" + f"{_esc(rep.get('correct'))}" + f"{_esc(rep.get('tokens_spent'))}" + f"{_esc(score_str)}" + ) + rep_table = "\n".join(rep_rows) or ( + "(empty)" + ) + + svg = _scatter_svg( + market_tokens=int(data.get("market", {}).get("tokens", 0)), + market_f1=float(data.get("market", {}).get("f1", 0.0)), + baseline_tokens=int(data.get("baseline", {}).get("tokens", 0)), + baseline_f1=float(data.get("baseline", {}).get("f1", 0.0)), + ) + + market = data.get("market", {}) + baseline = data.get("baseline", {}) + + market_f1_str = "{:.3f}".format(verdict.get("market_f1", 0) or 0) + baseline_f1_str = "{:.3f}".format(verdict.get("baseline_f1", 0) or 0) + f1_delta_str = "{:.3f}".format(verdict.get("f1_delta", 0) or 0) + market_run_f1_str = "{:.3f}".format(market.get("f1", 0) or 0) + baseline_run_f1_str = "{:.3f}".format(baseline.get("f1", 0) or 0) + + html_doc = f""" + + + +MuSiQue MAS Benchmark — {_esc(data.get('question_id', ''))} + + + +

MuSiQue MAS Benchmark Report

+

+ Mode: {_esc(data.get('mode', ''))} + · Question: {_esc(data.get('question_id', ''))} + · Dry-run: {_esc(data.get('dry_run', False))} +

+ +
+

Verdict

+

{verdict_html}

+
+
Market tokens
{_esc(verdict.get('market_tokens'))}
+
Market F1
{_esc(market_f1_str)}
+
Baseline tokens
{_esc(verdict.get('baseline_tokens'))}
+
Baseline F1
{_esc(baseline_f1_str)}
+
Tokens delta
{_esc(verdict.get('tokens_delta'))}
+
F1 delta
{_esc(f1_delta_str)}
+
+
+ +
+

Pareto plot

+ {svg} +

+ Lower-right = expensive and wrong. Upper-left = cheap and correct. + Marketplace must sit strictly northwest of baseline to pass. +

+
+ +
+

Question

+

{_esc(data.get('question', ''))}

+

Gold answer: {_esc(data.get('gold_answer', ''))}

+ +

Gold decomposition

+ + + {decomp_table} +
#Sub-questionSub-answer
+
+ +
+

Auction

+

Domain: {_esc(data.get('domain', ''))} + · Budget: {_esc(data.get('max_budget_tokens', ''))} + · Awarded to: {_esc(data.get('awarded_to', ''))}

+

Bids received

+ + + {bids_table} +
AgentEst. tokensConfidenceApproach
+
+ +
+

Marketplace run

+
+
Winning agent
{_esc(market.get('agent'))}
+
Model
{_esc(market.get('model'))}
+
Tokens
{_esc(market.get('tokens'))}
+
F1
{_esc(market_run_f1_str)}
+
Answer
{_esc(market.get('answer'))}
+
+
+ +
+

Single-agent baseline

+
+
Model
{_esc(baseline.get('model'))}
+
Tokens
{_esc(baseline.get('tokens'))}
+
F1
{_esc(baseline_run_f1_str)}
+
Answer
{_esc(baseline.get('answer'))}
+
+
+ +
+

Reputation ledger (post-run)

+ + + + {rep_table} +
AgentDomainRunsCorrectTokensScore
+
+ +

+ Generated by benchmark/report.py. + Marketplace primitives are currently stubbed in-process — see + benchmark/marketplace.py for the migration plan to the + real 016 SynapBus MCP tools. +

+ + +""" + out_path.parent.mkdir(parents=True, exist_ok=True) + out_path.write_text(html_doc, encoding="utf-8") diff --git a/benchmark/requirements.txt b/benchmark/requirements.txt new file mode 100644 index 0000000..cd28abe --- /dev/null +++ b/benchmark/requirements.txt @@ -0,0 +1,2 @@ +anthropic>=0.40.0 +requests>=2.31.0 diff --git a/benchmark/run.py b/benchmark/run.py new file mode 100644 index 0000000..cf02472 --- /dev/null +++ b/benchmark/run.py @@ -0,0 +1,277 @@ +#!/usr/bin/env python3 +""" +Main entry point for the MuSiQue MAS benchmark. + +Usage:: + + python benchmark/run.py --mode single-shot --question q1 + python benchmark/run.py --mode single-shot --question q1 --dry-run + +Flow (single-shot): + 1. Load trio.jsonl, find the requested question (by short_id). + 2. Marketplace run: + a. post_auction(task, domain, max_budget) + b. each agent in the pool submits a bid + c. marketplace awards best bid + d. winner executes (Anthropic call or dry-run stub) + e. marketplace.mark_done records reputation + 3. Baseline run: one Sonnet call with all distractors. + 4. Score both, compute Pareto verdict. + 5. Write results/latest.json and results/latest.html. +""" + +from __future__ import annotations + +import argparse +import json +import sys +import time +from pathlib import Path +from typing import Any + +# Allow running as ``python benchmark/run.py`` from the repo root. +_HERE = Path(__file__).resolve().parent +if str(_HERE) not in sys.path: + sys.path.insert(0, str(_HERE)) + +from agents import default_pool # noqa: E402 +from baseline import run_baseline, BASELINE_MODEL # noqa: E402 +from marketplace import Marketplace # noqa: E402 +from report import render_report # noqa: E402 +from score import best_f1_against_aliases, pareto_verdict # noqa: E402 + + +TRIO_FILE = _HERE / "trio.jsonl" +RESULTS_DIR = _HERE / "results" +DEFAULT_DOMAIN = "multi-hop-qa" +DEFAULT_BUDGET = 50_000 + + +def _load_trio() -> list[dict[str, Any]]: + if not TRIO_FILE.exists(): + raise SystemExit( + f"[run] trio.jsonl not found at {TRIO_FILE}. " + "Run curate.py first." + ) + out: list[dict[str, Any]] = [] + with open(TRIO_FILE, "r", encoding="utf-8") as f: + for line in f: + line = line.strip() + if not line: + continue + out.append(json.loads(line)) + return out + + +def _pick_question( + trio: list[dict[str, Any]], want: str +) -> dict[str, Any]: + for rec in trio: + if rec.get("short_id") == want or rec.get("id") == want: + return rec + raise SystemExit( + f"[run] question {want!r} not found. Available: " + + ", ".join(r.get("short_id", r.get("id", "?")) for r in trio) + ) + + +def single_shot( + question: str, + *, + dry_run: bool, + verbose: bool = True, +) -> dict[str, Any]: + trio = _load_trio() + rec = _pick_question(trio, question) + + task = { + "question": rec["question"], + "short_id": rec.get("short_id"), + } + paragraphs = rec.get("paragraphs", []) or [] + gold_answer = rec.get("answer", "") + aliases = rec.get("answer_aliases", []) or [] + + market = Marketplace() + pool = default_pool() + + if verbose: + print(f"[run] question {rec.get('short_id')}: {rec['question']!r}") + print( + f"[run] agents: " + + ", ".join(f"{a.name}({a.model})" for a in pool) + ) + print(f"[run] paragraphs: {len(paragraphs)}") + + # --- Marketplace path ------------------------------------------------ + auction_id = market.post_auction( + task=task, + domain=DEFAULT_DOMAIN, + max_budget_tokens=DEFAULT_BUDGET, + ) + if verbose: + print(f"[run] posted auction {auction_id}") + + for agent in pool: + bid = agent.bid(task) + market.bid( + auction_id=auction_id, + agent=agent.name, + estimated_tokens=bid.estimated_tokens, + confidence=bid.confidence, + approach=bid.approach, + ) + if verbose: + print( + f"[run] bid {agent.name}: " + f"est={bid.estimated_tokens} conf={bid.confidence:.2f}" + ) + + winning_bid = market.award(auction_id) + winner_name = winning_bid["agent"] + winner = next(a for a in pool if a.name == winner_name) + if verbose: + print(f"[run] awarded to {winner_name}") + + start = time.time() + result = winner.execute( + task=task, + paragraphs=paragraphs, + dry_run=dry_run, + max_budget_tokens=DEFAULT_BUDGET, + ) + market_wall = time.time() - start + + market_f1 = best_f1_against_aliases( + result.answer, gold_answer, aliases + ) + market.mark_done( + auction_id=auction_id, + answer=result.answer, + actual_tokens=result.actual_tokens, + correct=market_f1 >= 0.5, + ) + if verbose: + print( + f"[run] market answer: {result.answer!r} " + f"(tokens={result.actual_tokens}, f1={market_f1:.3f})" + ) + + # --- Baseline path --------------------------------------------------- + start = time.time() + baseline = run_baseline( + question=rec["question"], + paragraphs=paragraphs, + dry_run=dry_run, + ) + baseline_wall = time.time() - start + baseline_f1 = best_f1_against_aliases( + baseline["answer"], gold_answer, aliases + ) + if verbose: + print( + f"[run] baseline answer: {baseline['answer']!r} " + f"(tokens={baseline['tokens']}, f1={baseline_f1:.3f})" + ) + + verdict = pareto_verdict( + market_tokens=result.actual_tokens, + market_f1=market_f1, + baseline_tokens=baseline["tokens"], + baseline_f1=baseline_f1, + ) + if verbose: + print(f"[run] PARETO VERDICT: {verdict['verdict']}") + + return { + "mode": "single-shot", + "dry_run": dry_run, + "question_id": rec.get("short_id"), + "musique_id": rec.get("id"), + "question": rec["question"], + "gold_answer": gold_answer, + "decomposition": rec.get("decomposition", []), + "domain": DEFAULT_DOMAIN, + "max_budget_tokens": DEFAULT_BUDGET, + "awarded_to": winner_name, + "bids": market.list_bids(auction_id), + "market": { + "agent": winner_name, + "model": winner.model, + "tokens": result.actual_tokens, + "answer": result.answer, + "f1": market_f1, + "wall_seconds": market_wall, + "raw_text": result.raw_text, + }, + "baseline": { + "model": baseline.get("model", BASELINE_MODEL), + "tokens": baseline["tokens"], + "answer": baseline["answer"], + "f1": baseline_f1, + "wall_seconds": baseline_wall, + "raw_text": baseline.get("raw_text", ""), + }, + "pareto": verdict, + "reputation": market.all_reputation(), + } + + +def _write_outputs(result: dict[str, Any]) -> tuple[Path, Path]: + RESULTS_DIR.mkdir(parents=True, exist_ok=True) + json_path = RESULTS_DIR / "latest.json" + html_path = RESULTS_DIR / "latest.html" + # Trim raw_text from json to keep it small and readable. + trimmed = dict(result) + for key in ("market", "baseline"): + section = dict(trimmed.get(key, {})) + if "raw_text" in section: + section["raw_text"] = (section["raw_text"] or "")[:2000] + trimmed[key] = section + json_path.write_text( + json.dumps(trimmed, indent=2, ensure_ascii=False), encoding="utf-8" + ) + render_report(result, html_path) + return json_path, html_path + + +def main(argv: list[str] | None = None) -> int: + parser = argparse.ArgumentParser( + description="MuSiQue multi-agent benchmark harness" + ) + parser.add_argument( + "--mode", + choices=["single-shot"], + default="single-shot", + help="Run mode (only single-shot is implemented in MVP)", + ) + parser.add_argument( + "--question", + default="q1", + help="Question short_id from trio.jsonl (q1/q2/q3)", + ) + parser.add_argument( + "--dry-run", + action="store_true", + help="Skip real Anthropic API calls; use stub responses", + ) + args = parser.parse_args(argv) + + if args.mode != "single-shot": + print(f"[run] mode {args.mode} not implemented in MVP", file=sys.stderr) + return 2 + + result = single_shot( + question=args.question, + dry_run=args.dry_run, + verbose=True, + ) + json_path, html_path = _write_outputs(result) + print(f"[run] wrote {json_path}") + print(f"[run] wrote {html_path}") + print(f"[run] verdict: {result['pareto']['verdict']}") + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/benchmark/score.py b/benchmark/score.py new file mode 100644 index 0000000..a2c7175 --- /dev/null +++ b/benchmark/score.py @@ -0,0 +1,86 @@ +#!/usr/bin/env python3 +""" +Scoring utilities for the MuSiQue benchmark. + +- Normalized exact-match F1 (SQuAD-style): lowercase, strip articles, + strip punctuation, collapse whitespace. +- Pareto verdict: the marketplace point is strictly northwest of the + baseline iff it uses fewer tokens AND has F1 >= baseline, with at + least one of those strict. +""" + +from __future__ import annotations + +import re +import string +from collections import Counter +from typing import Any + +_ARTICLE_RE = re.compile(r"\b(a|an|the)\b", re.IGNORECASE) + + +def normalize(text: str) -> str: + if text is None: + return "" + text = text.lower() + text = _ARTICLE_RE.sub(" ", text) + text = "".join(ch for ch in text if ch not in string.punctuation) + text = " ".join(text.split()) + return text + + +def f1(prediction: str, gold: str) -> float: + pred_tokens = normalize(prediction).split() + gold_tokens = normalize(gold).split() + if not pred_tokens and not gold_tokens: + return 1.0 + if not pred_tokens or not gold_tokens: + return 0.0 + common = Counter(pred_tokens) & Counter(gold_tokens) + overlap = sum(common.values()) + if overlap == 0: + return 0.0 + precision = overlap / len(pred_tokens) + recall = overlap / len(gold_tokens) + return 2 * precision * recall / (precision + recall) + + +def exact_match(prediction: str, gold: str) -> bool: + return normalize(prediction) == normalize(gold) + + +def best_f1_against_aliases( + prediction: str, gold: str, aliases: list[str] | None = None +) -> float: + candidates = [gold] + list(aliases or []) + return max(f1(prediction, c) for c in candidates if c is not None) + + +def pareto_verdict( + market_tokens: int, + market_f1: float, + baseline_tokens: int, + baseline_f1: float, +) -> dict[str, Any]: + """ + Strictly northwest of baseline: fewer tokens AND higher-or-equal F1, + with at least one strict inequality. + """ + tokens_better = market_tokens < baseline_tokens + quality_atleast = market_f1 >= baseline_f1 + quality_better = market_f1 > baseline_f1 + + strictly_nw = ( + (tokens_better and quality_atleast) + or (quality_better and market_tokens <= baseline_tokens) + ) + return { + "verdict": "PASS" if strictly_nw else "FAIL", + "strictly_northwest": strictly_nw, + "market_tokens": int(market_tokens), + "market_f1": float(market_f1), + "baseline_tokens": int(baseline_tokens), + "baseline_f1": float(baseline_f1), + "tokens_delta": int(market_tokens - baseline_tokens), + "f1_delta": float(market_f1 - baseline_f1), + } diff --git a/benchmark/setup.py b/benchmark/setup.py new file mode 100644 index 0000000..ebb93b9 --- /dev/null +++ b/benchmark/setup.py @@ -0,0 +1,198 @@ +#!/usr/bin/env python3 +""" +MuSiQue dataset downloader. + +Downloads ``musique_v1.0.zip`` from the canonical source used by the +upstream project (https://github.com/StonyBrookNLP/musique). The zip is +hosted on Google Drive (file id ``1tGdADlNjWFaHLeZZGShh2IRcpO6Lv24h``); +this mirrors the behavior of the project's ``download_data.sh`` which +uses ``gdown`` under the hood. + +Idempotent — skips download if the target dev-set jsonl already exists. +Run: ``python benchmark/setup.py`` +""" + +from __future__ import annotations + +import os +import re +import sys +import zipfile +from pathlib import Path + +import requests + +GDRIVE_FILE_ID = "1tGdADlNjWFaHLeZZGShh2IRcpO6Lv24h" +GDRIVE_URL = "https://docs.google.com/uc?export=download" + +DATA_DIR = Path(__file__).resolve().parent / "data" +ZIP_PATH = DATA_DIR / "musique_v1.0.zip" +TARGET_FILE = DATA_DIR / "musique_ans_v1.0_dev.jsonl" + + +def _write_stream(resp: requests.Response, dest: Path) -> int: + total = int(resp.headers.get("Content-Length", 0)) + downloaded = 0 + dest.parent.mkdir(parents=True, exist_ok=True) + with open(dest, "wb") as f: + for chunk in resp.iter_content(chunk_size=1024 * 1024): + if not chunk: + continue + f.write(chunk) + downloaded += len(chunk) + if total: + pct = 100.0 * downloaded / total + print( + f"\r downloading: {downloaded/1e6:6.1f} MB " + f"/ {total/1e6:6.1f} MB ({pct:5.1f}%)", + end="", + file=sys.stderr, + ) + print("", file=sys.stderr) + return downloaded + + +def _download_gdrive(file_id: str, dest: Path) -> bool: + """ + Download a large file from Google Drive, handling the virus-scan + confirmation page that Drive injects for anything over ~100 MB. + """ + session = requests.Session() + try: + resp = session.get( + GDRIVE_URL, + params={"id": file_id, "export": "download"}, + stream=True, + timeout=60, + ) + except requests.RequestException as exc: + print(f" -> request failed: {exc}", file=sys.stderr) + return False + + # Case 1: Drive returns the file directly (small file or cached). + ctype = resp.headers.get("Content-Type", "") + if "text/html" not in ctype.lower(): + _write_stream(resp, dest) + return dest.exists() and dest.stat().st_size > 0 + + # Case 2: HTML confirmation page. Extract the confirm token and/or + # the form action URL. + html = resp.text + # Newer Drive flow: a
with all the params we need. + form_match = re.search( + r']*id="download-form"[^>]*action="([^"]+)"', html + ) + if form_match: + action = form_match.group(1).replace("&", "&") + params = dict( + re.findall( + r'name="([^"]+)"[^>]*value="([^"]+)"', html + ) + ) + try: + resp2 = session.get(action, params=params, stream=True, timeout=120) + if resp2.status_code == 200: + _write_stream(resp2, dest) + return dest.exists() and dest.stat().st_size > 0 + except requests.RequestException as exc: + print(f" -> form post failed: {exc}", file=sys.stderr) + return False + + # Older flow: confirm cookie token. + token = None + for k, v in session.cookies.items(): + if k.startswith("download_warning"): + token = v + break + if token is None: + m = re.search(r'confirm=([0-9A-Za-z_-]+)', html) + if m: + token = m.group(1) + if token: + try: + resp3 = session.get( + GDRIVE_URL, + params={ + "id": file_id, + "export": "download", + "confirm": token, + }, + stream=True, + timeout=120, + ) + if resp3.status_code == 200: + _write_stream(resp3, dest) + return dest.exists() and dest.stat().st_size > 0 + except requests.RequestException as exc: + print(f" -> confirm fetch failed: {exc}", file=sys.stderr) + return False + + print(" -> could not navigate Google Drive download flow", file=sys.stderr) + return False + + +def _extract(zip_path: Path, out_dir: Path) -> None: + """Extract the dev set jsonl from the zip.""" + wanted_suffixes = ( + "musique_ans_v1.0_dev.jsonl", + "musique_ans_v1.0_train.jsonl", + ) + with zipfile.ZipFile(zip_path) as zf: + members = zf.namelist() + extracted_any = False + for m in members: + base = os.path.basename(m) + if base in wanted_suffixes: + with zf.open(m) as src, open(out_dir / base, "wb") as dst: + dst.write(src.read()) + print(f" extracted: {base}") + extracted_any = True + if not extracted_any: + # Fall back: extract everything so a human can inspect. + zf.extractall(out_dir) + print( + " could not find canonical filenames; extracted all", + file=sys.stderr, + ) + + +def main() -> int: + DATA_DIR.mkdir(parents=True, exist_ok=True) + + if TARGET_FILE.exists(): + size = TARGET_FILE.stat().st_size + print(f"[setup] already present: {TARGET_FILE} ({size/1e6:.1f} MB)") + return 0 + + print(f"[setup] downloading Google Drive file id {GDRIVE_FILE_ID}") + ok = _download_gdrive(GDRIVE_FILE_ID, ZIP_PATH) + + if not ok: + print( + "[setup] ERROR: failed to download MuSiQue. Please download " + "manually from " + f"https://drive.google.com/file/d/{GDRIVE_FILE_ID}/view " + f"and place the zip at {ZIP_PATH}", + file=sys.stderr, + ) + return 2 + + print(f"[setup] extracting {ZIP_PATH}") + _extract(ZIP_PATH, DATA_DIR) + + if not TARGET_FILE.exists(): + print( + f"[setup] WARNING: {TARGET_FILE.name} not found after extract. " + f"Listing {DATA_DIR}:", + file=sys.stderr, + ) + for p in sorted(DATA_DIR.iterdir()): + print(f" - {p.name}", file=sys.stderr) + return 3 + + print(f"[setup] ready: {TARGET_FILE}") + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/benchmark/trio.jsonl b/benchmark/trio.jsonl new file mode 100644 index 0000000..084b2c0 --- /dev/null +++ b/benchmark/trio.jsonl @@ -0,0 +1,3 @@ +{"id": "4hop1__94201_642284_131926_13165", "question": "What treaty ceded territory to the US extending west to the body of water by the city where the designer of Southeast Library died?", "answer": "Treaty of Paris", "answer_aliases": [], "decomposition": [{"question": "The designer for Southeast Library was?", "answer": "Ralph Rapson"}, {"question": "#1 >> place of death", "answer": "Minneapolis"}, {"question": "Which is the body of water by #2 ?", "answer": "Mississippi River"}, {"question": "What treaty ceded territory to the US extending west to #3 ?", "answer": "Treaty of Paris"}], "paragraphs": ["[Dwight D. Eisenhower] That evening, Eisenhower's body was placed onto a train en route to Abilene, Kansas, the last time a funeral train has been used as part of funeral proceedings of an American president. His body arrived on April 2, and was interred later that day in a small chapel on the grounds of the Eisenhower Presidential Library. The president's body was buried as a General of the Army. The family used an $80 standard soldier's casket, and dressed Eisenhower's body in his famous short green jacket. His only medals worn were: the Army Distinguished Service Medal with three oak leaf clusters, the Navy Distinguished Service Medal, and the Legion of Merit. Eisenhower is buried alongside his son Doud, who died at age 3 in 1921. His wife Mamie was buried next to him after her death in 1979.", "[Riverside Plaza] Riverside Plaza is a modernist and brutalist apartment complex designed by Ralph Rapson that opened in Minneapolis, Minnesota in 1973. Situated on the edge of downtown Minneapolis in the Cedar-Riverside neighborhood, and next to both the University of Minnesota's West Bank and Augsburg University, the site contains the 39-story McKnight Building, the tallest structure outside of the city's central business district. Initially known as Cedar Square West, exterior shots of the complex were featured on television as the residence of Mary Richards in sixth and seventh seasons of \"The Mary Tyler Moore Show\".", "[Territorial waters] Territorial waters or a territorial sea, as defined by the 1982 United Nations Convention on the Law of the Sea, is a belt of coastal waters extending at most 12 nautical miles (22.2 km; 13.8 mi) from the baseline (usually the mean low - water mark) of a coastal state. The territorial sea is regarded as the sovereign territory of the state, although foreign ships (civilian) are allowed innocent passage through it, or transit passage for straits; this sovereignty also extends to the airspace over and seabed below. Adjustment of these boundaries is called, in international law, maritime delimitation.", "[Wilkes Land] Wilkes Land is a large district of land in eastern Antarctica, formally claimed by Australia as part of the Australian Antarctic Territory, though the validity of this claim has been placed for the period of the operation of the Antarctic Treaty, to which Australia is a signatory. It fronts on the southern Indian Ocean between Queen Mary Coast and Adelie Land, extending from Cape Hordern in 100°31' E to Pourquoi Pas Point, in 136°11' E. The region extends as a sector about 2600 km towards the South Pole, with an estimated land area of 2,600,000 km², mostly glaciated. It is further subdivided in the following coastal areas which can also be thought of as sectors extending to the South Pole:", "[Southeast Library] Southeast Library's building was designed by master architect Ralph Rapson and originally functioned as a credit union for university and state employees. It opened as a library in 1967. The State Capitol Credit Union building at 1222 Fourth Street Southeast was purchased to be converted into a library on December 29, 1966. It opened as the new Southeast Library on December 26, 1967.", "[Kanije Eyalet] The province of Kanije was established in 1600 after the town of Kanije was captured from Habsburgs. This newly conquered area was joined with territory of Zigetvar Province, which was formed in 1596 from some sanjaks of Budin Province (which had been expanded as a result of the Ottoman territorial gains during the Long War) and Bosnia Province. The Kanije Eyalet existed until the capture of Kanije by Habsburg Monarchy in 1690. It was formally ceded to Habsburg Monarchy by the Treaty of Karlowitz in 1699.", "[San Diego] Stockton and Kearny went on to recover Los Angeles and force the capitulation of Alta California with the \"Treaty of Cahuenga\" on January 13, 1847. As a result of the Mexican–American War of 1846–48, the territory of Alta California, including San Diego, was ceded to the United States by Mexico, under the terms of the Treaty of Guadalupe Hidalgo in 1848. The Mexican negotiators of that treaty tried to retain San Diego as part of Mexico, but the Americans insisted that San Diego was \"for every commercial purpose of nearly equal importance to us with that of San Francisco,\" and the Mexican-American border was eventually established to be one league south of the southernmost point of San Diego Bay, so as to include the entire bay within the United States.", "[Falling Waters, West Virginia] Falling Waters is a census-designated place (CDP) on the Potomac River in Berkeley County, West Virginia. It is located along Williamsport Pike (US 11) north of Martinsburg. According to the 2010 census, Falling Waters has a population of 876. An 1887 \"Scientific American\" article claimed that the first U.S. railroad was built in Falling Waters in 1814.", "[Linden Hills Library] Linden Hills Library is a public library in the Linden Hills neighborhood of southwest Minneapolis, Minnesota, United States. The branch library originally opened in 1911 on the first floor of the Lake Harriet Commercial Club building. In 1931, under the leadership of Minneapolis Public Library's chief librarian Gratia Countryman, the library moved into its own building on 2900 West 43rd Street. Area resident Joseph Victor Vanderbilt designed the library in the Tudor Revival style. Head librarian Edith Frost served for over thirty years. The library has also hosted community groups such as children's clubs, neighborhood groups, and women's organizations. The library was listed on the National Register of Historic Places in 2000 and renovated in 2002.", "[Secchi disk] The Secchi disk, as created in 1865 by Angelo Secchi, is a plain white, circular disk in diameter used to measure water transparency or turbidity in bodies of water. The disc is mounted on a pole or line, and lowered slowly down in the water. The depth at which the disk is no longer visible is taken as a measure of the transparency of the water. This measure is known as the Secchi depth and is related to water turbidity. Since its invention, the disk has also been used in a modified, smaller diameter, black and white design to measure freshwater transparency.", "[French and Indian War] The outcome was one of the most significant developments in a century of Anglo - French conflict. France ceded to Great Britain its territory east of the Mississippi. It ceded French Louisiana west of the Mississippi River (including New Orleans) to its ally Spain in compensation for Spain's loss to Britain of Florida. (Spain had ceded Florida to Britain in exchange for the return of Havana, Cuba.) France's colonial presence north of the Caribbean was reduced to the islands of Saint Pierre and Miquelon, confirming Great Britain's position as the dominant colonial power in eastern North America.", "[Minneapolis] Minneapolis lies on both banks of the Mississippi River, just north of the river's confluence with the Minnesota River, and adjoins Saint Paul, the state's capital. The city is abundantly rich in water, with 13 lakes, wetlands, the Mississippi River, creeks and waterfalls; many connected by parkways in the Chain of Lakes and the Grand Rounds National Scenic Byway. It was once the world's flour milling capital and a hub for timber. The city and surrounding region is the primary business center between Chicago and Seattle. As of 2018, Minneapolis was home to 6 Fortune 500 companies, and the Twin Cities were the fifth-largest hub of major corporate headquarters in the United States. As an integral link to the global economy, Minneapolis is categorized as a global city.", "[Northwest Territories] Located in northern Canada, the territory borders Canada's two other territories, Yukon to the west and Nunavut to the east, as well as three provinces: British Columbia to the southwest, and Alberta and Saskatchewan to the south. It possibly meets Manitoba at a quadripoint to the extreme southeast, though surveys have not been completed. It has a land area of 1,183,085 km2 (456,792 sq mi).Geographical features include Great Bear Lake, the largest lake entirely within Canada, and Great Slave Lake, the deepest body of water in North America at 614 m (2,014 ft), as well as the Mackenzie River and the canyons of the Nahanni National Park Reserve, a national park and UNESCO World Heritage Site. Territorial islands in the Canadian Arctic Archipelago include Banks Island, Borden Island, Prince Patrick Island, and parts of Victoria Island and Melville Island. Its highest point is Mount Nirvana near the border with Yukon at an elevation of 2,773 m (9,098 ft).", "[Pimpirev Glacier] Pimpirev Glacier (Pimpirev Lednik \\pim-'pi-rev 'led-nik\\) on Livingston Island in the South Shetland Islands, Antarctica is situated south of the glacial divide between the Drake Passage and Bransfield Strait, southeast of Tundzha Glacier, southwest of Saedinenie Snowfield, west of Perunika Glacier and east-northeast of Kamchiya Glacier. The feature extends 5.5 km in a southeast-northwest direction, and 1.8 km in northwest-southeast direction. The glacier drains southeastwards towards Pimpirev Beach, mostly terminating on the shore, and on several occasions penetrating the South Bay waters east-northeast of Ereby Point.", "[Stinson Beach, California] Stinson Beach is a census-designated place in Marin County, California, on the west coast of the United States. Stinson Beach is located east-southeast of Bolinas, at an elevation of 26 feet (8 m). The population of the Stinson Beach CDP (census-designated place) was 632 at the 2010 census.", "[Clerke Rocks] The Clerke Rocks are a group of small rocky islands some southeast of South Georgia that extend from east to west. The Clerke Rocks include The Office Boys () at the northeastern end and Nobby (Spanish: \"Islote Llamativo\" or \"Roca Notable\") at the southeastern end of the group. The islands belong to the British Overseas Territory of South Georgia and the South Sandwich Islands and are also claimed by Argentina as part of Tierra del Fuego Province.", "[Louisiana Purchase] In 1798, Spain revoked the treaty allowing American use of New Orleans, greatly upsetting Americans. In 1801, Spanish Governor Don Juan Manuel de Salcedo took over from the Marquess of Casa Calvo, and restored the American right to deposit goods. However, in 1800 Spain had ceded the Louisiana territory back to France as part of Napoleon's secret Third Treaty of San Ildefonso. The territory nominally remained under Spanish control, until a transfer of power to France on 30 November 1803, just three weeks before the formal cession of the territory to the United States on 20 December 1803. A further ceremony was held in St. Louis a few months later partially due to winter conditions impeding the arrival of news, Upper Louisiana, regarding the New Orleans formalities. The 9 -- 10 March 1804 event is remembered as Three Flags Day.", "[Military history of the United States] In the Treaty of Paris after the Revolution, the British had ceded the lands between the Appalachian Mountains and the Mississippi River to the United States, without consulting the Shawnee, Cherokee, Choctaw and other smaller tribes who lived there. Because many of the tribes had fought as allies of the British, the United States compelled tribal leaders to sign away lands in postwar treaties, and began dividing these lands for settlement. This provoked a war in the Northwest Territory in which the U.S. forces performed poorly; the Battle of the Wabash in 1791 was the most severe defeat ever suffered by the United States at the hands of American Indians. President Washington dispatched a newly trained army to the region, which decisively defeated the Indian confederacy at the Battle of Fallen Timbers in 1794.", "[Justice, West Virginia] Justice is a census-designated place in Mingo County, West Virginia, United States. Justice is located on U.S. Route 52 southeast of Gilbert. Justice has a post office with ZIP code 24851. As of the 2010 census, its population was 412.", "[Geography of the United States] The United States shares land borders with Canada (to the north) and Mexico (to the south), and a territorial water border with Russia in the northwest, and two territorial water borders in the southeast between Florida and Cuba, and Florida and the Bahamas. The contiguous forty-eight states are otherwise bounded by the Pacific Ocean on the west, the Atlantic Ocean on the east, and the Gulf of Mexico to the southeast. Alaska borders the Pacific Ocean to the south, the Bering Strait to the west, and the Arctic Ocean to the north, while Hawaii lies far to the southwest of the mainland in the Pacific Ocean."], "short_id": "q1"} +{"id": "4hop1__860115_798482_131926_13165", "question": "What treaty ceded territory to the US extending west to the river by the city sharing a border with Elizabeth Berg's birthplace?", "answer": "Treaty of Paris", "answer_aliases": [], "decomposition": [{"question": "Elizabeth Berg >> place of birth", "answer": "Saint Paul"}, {"question": "#1 >> shares border with", "answer": "Minneapolis"}, {"question": "Which is the body of water by #2 ?", "answer": "Mississippi River"}, {"question": "What treaty ceded territory to the US extending west to #3 ?", "answer": "Treaty of Paris"}], "paragraphs": ["[American Indian Wars] When the British made peace with the Americans in the Treaty of Paris (1783), they ceded a vast amount of Native American territory (without the consent of the indigenous peoples) to the United States. The United States treated the Native Americans who had fought with the British as enemy allies, a conquered people who had lost their land. The federal government of the United States was eager to expand, and the national government did so by purchasing Native American land in treaties and through warfare.", "[Territorial waters] Territorial waters or a territorial sea, as defined by the 1982 United Nations Convention on the Law of the Sea, is a belt of coastal waters extending at most 12 nautical miles (22.2 km; 13.8 mi) from the baseline (usually the mean low - water mark) of a coastal state. The territorial sea is regarded as the sovereign territory of the state, although foreign ships (civilian) are allowed innocent passage through it, or transit passage for straits; this sovereignty also extends to the airspace over and seabed below. Adjustment of these boundaries is called, in international law, maritime delimitation.", "[Elizabeth Berg (author)] Berg was born in Saint Paul, Minnesota, USA, and lived in Boston prior to her residence in Chicago. She studied English at the University of Minnesota, but later ended up with a nursing degree. Her writing career started when she won an essay contest in \"Parents\" magazine. Since her debut novel in 1993, her novels have sold in large numbers and have received several awards and nominations, even though some critics have tagged them as sentimental. She won the New England Book Awards in 1997.", "[Geography of the United States] The United States shares land borders with Canada (to the north) and Mexico (to the south), and a territorial water border with Russia in the northwest, and two territorial water borders in the southeast between Florida and Cuba, and Florida and the Bahamas. The contiguous forty-eight states are otherwise bounded by the Pacific Ocean on the west, the Atlantic Ocean on the east, and the Gulf of Mexico to the southeast. Alaska borders the Pacific Ocean to the south, the Bering Strait to the west, and the Arctic Ocean to the north, while Hawaii lies far to the southwest of the mainland in the Pacific Ocean.", "[Newgate Education Center] Newgate School is a post-secondary non-profit vocational-technical school for residents of Minneapolis and Saint Paul, Minnesota and the surrounding area. Newgate provides tuition-free automotive vocational training and technical career placement opportunities for low income adults. It offers professional automotive technical certification in three areas: Auto-body Repair, Auto mechanics and Detailing. Graduates are qualified to work as career apprentices in the auto services industry. Newgate’s practical, hands-on approach to teaching technical skills is highly successful with students who struggle in traditional educational settings or for whom English is a second language. In 1981, Newgate pioneered the concept of using the sales of car donations as the single funding source for the school, thereby eliminating the dependence on tax-based government funding for support. Newgate began its Wheels for Women Program in 1996. Donated cars are repaired by the students and provided at no cost to single moms referred by social service agencies like the Jeremiah Program or Lutheran Social Services. Newgate provides approximately 50 cars per year through the Wheels program. In 2004, with bonds financed by the City of Minneapolis, the school constructed a new modern training facility and expanded its Auto Mechanics Training program.", "[Swedish Livonia] Swedish Livonia () was a dominion of the Swedish Empire from 1629 until 1721. The territory, which constituted the southern part of modern Estonia (including the island of Ösel ceded by Denmark after the Treaty of Brömsebro) and the northern part of modern Latvia (the Vidzeme region), represented the conquest of the major part of the Polish-Lithuanian Duchy of Livonia during the 1600–1629 Polish-Swedish War. Parts of Livonia and the city of Riga were under Swedish control as early as 1621 and the situation was formalized in Truce of Altmark 1629, but the whole territory was not ceded formally until the Treaty of Oliva in 1660. The minority part of the Wenden Voivodeship retained by the Polish–Lithuanian Commonwealth was renamed the Inflanty Voivodeship (\"\"Livonian Principality\"\"), which today corresponds to the Latgale region of Latvia.", "[Louisiana Purchase] A dispute soon arose between Spain and the United States regarding the extent of Louisiana. The territory's boundaries had not been defined in the 1762 Treaty of Fontainebleau that ceded it from France to Spain, nor in the 1801 Third Treaty of San Ildefonso ceding it back to France, nor the 1803 Louisiana Purchase agreement ceding it to the United States.", "[San Diego] Stockton and Kearny went on to recover Los Angeles and force the capitulation of Alta California with the \"Treaty of Cahuenga\" on January 13, 1847. As a result of the Mexican–American War of 1846–48, the territory of Alta California, including San Diego, was ceded to the United States by Mexico, under the terms of the Treaty of Guadalupe Hidalgo in 1848. The Mexican negotiators of that treaty tried to retain San Diego as part of Mexico, but the Americans insisted that San Diego was \"for every commercial purpose of nearly equal importance to us with that of San Francisco,\" and the Mexican-American border was eventually established to be one league south of the southernmost point of San Diego Bay, so as to include the entire bay within the United States.", "[Aegean dispute] Several of the Aegean issues deal with the delimitation of both countries' zones of influence in the air and on the sea around their respective territories. These issues owe their virulence to a geographical peculiarity of the Aegean sea and its territories. While the mainland coasts of Greece and Turkey bordering the Aegean Sea on both sides represent roughly equal shares of its total coastline, the overwhelming number of the many Aegean islands belong to Greece. In particular, there is a chain of Greek islands lined up along the Turkish west coast (Lesbos, Chios, Samos, and the Dodecanese islands), partly in very close proximity to the mainland. Their existence blocks Turkey from extending any of its zones of influence beyond a few nautical miles off its coastline. As the breadth of maritime and areal zones of influence, such as the territorial waters and national airspace, are measured from the nearest territory of the state in question, including its islands, any possible extension of such zones would necessarily benefit Greece much more than Turkey proportionally.", "[Water supply and sanitation in South Africa] Total annual water withdrawal was estimated at 12.5 km3 in 2000, of which about 17% was for municipal water use. In the northern parts of the country, both surface water and groundwater resources are nearly fully developed and utilised. In the well - watered southeastern regions of the country significant undeveloped and little - used resources exist. The Gauteng area around Johannesburg, which is very water scarce, receives water from various dams in the area such as the Vaal Dam and imports water from the Orange River system through the Lesotho Highlands Water Project, in particular from the Katse Dam. Cape Town receives its drinking water from an extensive system of rivers and dams, including the Berg River Dam.", "[Smith Island, Maryland] Smith Island is an island on the Chesapeake Bay, on the border of Maryland and Virginia territorial waters in the United States.", "[Kanije Eyalet] The province of Kanije was established in 1600 after the town of Kanije was captured from Habsburgs. This newly conquered area was joined with territory of Zigetvar Province, which was formed in 1596 from some sanjaks of Budin Province (which had been expanded as a result of the Ottoman territorial gains during the Long War) and Bosnia Province. The Kanije Eyalet existed until the capture of Kanije by Habsburg Monarchy in 1690. It was formally ceded to Habsburg Monarchy by the Treaty of Karlowitz in 1699.", "[Geography of Pakistan] Pakistan is bordered by India to the east, Afghanistan to the west and Iran to the southwest while China borders the country in the northeast. The nation is geopolitically placed within some of the most controversial regional boundaries which share disputes and have many - a-times escalated military tensions between the nations, e.g., that of Kashmir with India and the Durand Line with Afghanistan. Its western borders include the Khyber Pass and Bolan Pass that have served as traditional migration routes between Central Eurasia and South Asia.", "[French and Indian War] The outcome was one of the most significant developments in a century of Anglo - French conflict. France ceded to Great Britain its territory east of the Mississippi. It ceded French Louisiana west of the Mississippi River (including New Orleans) to its ally Spain in compensation for Spain's loss to Britain of Florida. (Spain had ceded Florida to Britain in exchange for the return of Havana, Cuba.) France's colonial presence north of the Caribbean was reduced to the islands of Saint Pierre and Miquelon, confirming Great Britain's position as the dominant colonial power in eastern North America.", "[Northwest Territories] Located in northern Canada, the territory borders Canada's two other territories, Yukon to the west and Nunavut to the east, as well as three provinces: British Columbia to the southwest, and Alberta and Saskatchewan to the south. It possibly meets Manitoba at a quadripoint to the extreme southeast, though surveys have not been completed. It has a land area of 1,183,085 km2 (456,792 sq mi).Geographical features include Great Bear Lake, the largest lake entirely within Canada, and Great Slave Lake, the deepest body of water in North America at 614 m (2,014 ft), as well as the Mackenzie River and the canyons of the Nahanni National Park Reserve, a national park and UNESCO World Heritage Site. Territorial islands in the Canadian Arctic Archipelago include Banks Island, Borden Island, Prince Patrick Island, and parts of Victoria Island and Melville Island. Its highest point is Mount Nirvana near the border with Yukon at an elevation of 2,773 m (9,098 ft).", "[Military history of the United States] In the Treaty of Paris after the Revolution, the British had ceded the lands between the Appalachian Mountains and the Mississippi River to the United States, without consulting the Shawnee, Cherokee, Choctaw and other smaller tribes who lived there. Because many of the tribes had fought as allies of the British, the United States compelled tribal leaders to sign away lands in postwar treaties, and began dividing these lands for settlement. This provoked a war in the Northwest Territory in which the U.S. forces performed poorly; the Battle of the Wabash in 1791 was the most severe defeat ever suffered by the United States at the hands of American Indians. President Washington dispatched a newly trained army to the region, which decisively defeated the Indian confederacy at the Battle of Fallen Timbers in 1794.", "[Minneapolis] Minneapolis lies on both banks of the Mississippi River, just north of the river's confluence with the Minnesota River, and adjoins Saint Paul, the state's capital. The city is abundantly rich in water, with 13 lakes, wetlands, the Mississippi River, creeks and waterfalls; many connected by parkways in the Chain of Lakes and the Grand Rounds National Scenic Byway. It was once the world's flour milling capital and a hub for timber. The city and surrounding region is the primary business center between Chicago and Seattle. As of 2018, Minneapolis was home to 6 Fortune 500 companies, and the Twin Cities were the fifth-largest hub of major corporate headquarters in the United States. As an integral link to the global economy, Minneapolis is categorized as a global city.", "[Louisiana Purchase] In 1798, Spain revoked the treaty allowing American use of New Orleans, greatly upsetting Americans. In 1801, Spanish Governor Don Juan Manuel de Salcedo took over from the Marquess of Casa Calvo, and restored the American right to deposit goods. However, in 1800 Spain had ceded the Louisiana territory back to France as part of Napoleon's secret Third Treaty of San Ildefonso. The territory nominally remained under Spanish control, until a transfer of power to France on 30 November 1803, just three weeks before the formal cession of the territory to the United States on 20 December 1803. A further ceremony was held in St. Louis a few months later partially due to winter conditions impeding the arrival of news, Upper Louisiana, regarding the New Orleans formalities. The 9 -- 10 March 1804 event is remembered as Three Flags Day.", "[Wilkes Land] Wilkes Land is a large district of land in eastern Antarctica, formally claimed by Australia as part of the Australian Antarctic Territory, though the validity of this claim has been placed for the period of the operation of the Antarctic Treaty, to which Australia is a signatory. It fronts on the southern Indian Ocean between Queen Mary Coast and Adelie Land, extending from Cape Hordern in 100°31' E to Pourquoi Pas Point, in 136°11' E. The region extends as a sector about 2600 km towards the South Pole, with an estimated land area of 2,600,000 km², mostly glaciated. It is further subdivided in the following coastal areas which can also be thought of as sectors extending to the South Pole:", "[Elliott Bay] Elliott Bay is a part of the Central Basin region of Puget Sound in the U.S. state of Washington that extends southeastward between West Point in the north and Alki Point in the south. Seattle was founded on this body of water in the 1850s and has since grown to encompass it completely. The waterway it provides to the Pacific Ocean has served as a key element of the city's economy, enabling the Port of Seattle to become one of the busiest ports in the United States."], "short_id": "q2"} +{"id": "4hop1__525129_315334_131926_13165", "question": "What treaty ceded territory to the US extending west to body of water by the city where the Lots More Blues, Rags and Hollers performer was formed?", "answer": "Treaty of Paris", "answer_aliases": [], "decomposition": [{"question": "Lots More Blues, Rags and Hollers >> performer", "answer": "Koerner, Ray & Glover"}, {"question": "#1 >> location of formation", "answer": "Minneapolis"}, {"question": "Which is the body of water by #2 ?", "answer": "Mississippi River"}, {"question": "What treaty ceded territory to the US extending west to #3 ?", "answer": "Treaty of Paris"}], "paragraphs": ["[Kexholm County] Kexholm County (, ) was a county of the Swedish Empire from 1634 to 1721, when the southern part was ceded to the Russian Empire in the Treaty of Nystad. The capital of the county was Kexholm (), which today is Priozersk.", "[San Diego] Stockton and Kearny went on to recover Los Angeles and force the capitulation of Alta California with the \"Treaty of Cahuenga\" on January 13, 1847. As a result of the Mexican–American War of 1846–48, the territory of Alta California, including San Diego, was ceded to the United States by Mexico, under the terms of the Treaty of Guadalupe Hidalgo in 1848. The Mexican negotiators of that treaty tried to retain San Diego as part of Mexico, but the Americans insisted that San Diego was \"for every commercial purpose of nearly equal importance to us with that of San Francisco,\" and the Mexican-American border was eventually established to be one league south of the southernmost point of San Diego Bay, so as to include the entire bay within the United States.", "[American Indian Wars] When the British made peace with the Americans in the Treaty of Paris (1783), they ceded a vast amount of Native American territory (without the consent of the indigenous peoples) to the United States. The United States treated the Native Americans who had fought with the British as enemy allies, a conquered people who had lost their land. The federal government of the United States was eager to expand, and the national government did so by purchasing Native American land in treaties and through warfare.", "[Lots More Blues, Rags and Hollers] Lots More Blues, Rags and Hollers is an album by the blues trio Koerner, Ray & Glover, released in 1964.", "[Two Mile Square Reservation] The Two Mile Square Reservation or Two Mile Square Reserve was a tract of land in Ohio ceded by Native Americans to the United States of America in the Treaty of Greenville in 1795. It was subsequently surveyed in a manner different from surrounding land, and lots sold to settlers.", "[Louisiana Purchase] In 1798, Spain revoked the treaty allowing American use of New Orleans, greatly upsetting Americans. In 1801, Spanish Governor Don Juan Manuel de Salcedo took over from the Marquess of Casa Calvo, and restored the American right to deposit goods. However, in 1800 Spain had ceded the Louisiana territory back to France as part of Napoleon's secret Third Treaty of San Ildefonso. The territory nominally remained under Spanish control, until a transfer of power to France on 30 November 1803, just three weeks before the formal cession of the territory to the United States on 20 December 1803. A further ceremony was held in St. Louis a few months later partially due to winter conditions impeding the arrival of news, Upper Louisiana, regarding the New Orleans formalities. The 9 -- 10 March 1804 event is remembered as Three Flags Day.", "[Vatican City] The name Vatican city was first used in the Lateran Treaty, signed on 11 February 1929, which established the modern city - state. The name is taken from Vatican Hill, the geographic location of the state. ``Vatican ''is derived from the name of an Etruscan settlement, Vatica or Vaticum meaning garden, located in the general area the Romans called vaticanus ager,`` Vatican territory''.", "[Territorial waters] Territorial waters or a territorial sea, as defined by the 1982 United Nations Convention on the Law of the Sea, is a belt of coastal waters extending at most 12 nautical miles (22.2 km; 13.8 mi) from the baseline (usually the mean low - water mark) of a coastal state. The territorial sea is regarded as the sovereign territory of the state, although foreign ships (civilian) are allowed innocent passage through it, or transit passage for straits; this sovereignty also extends to the airspace over and seabed below. Adjustment of these boundaries is called, in international law, maritime delimitation.", "[First Anglo-Maratha War] Raghunathrao, unwilling to give up his position of power, sought help from the British at Bombay and signed the Treaty of Surat on 6 March 1775. According to the treaty, Raghunathrao ceded the territories of Salsette and Bassein to the British, along with part of the revenues from Surat and Bharuch districts. In return, the British promised to provide Raghunathrao with 2,500 soldiers.", "[Ukraine] Following the Invasion of Poland in September 1939, German and Soviet troops divided the territory of Poland. Thus, Eastern Galicia and Volhynia with their Ukrainian population became part of Ukraine. For the first time in history, the nation was united.In 1940, the Soviets annexed Bessarabia and northern Bukovina. The Ukrainian SSR incorporated the northern and southern districts of Bessarabia, northern Bukovina, and the Hertsa region. But it ceded the western part of the Moldavian Autonomous Soviet Socialist Republic to the newly created Moldavian Soviet Socialist Republic. These territorial gains of the USSR were internationally recognized by the Paris peace treaties of 1947.", "[American Revolutionary War] Date April 19, 1775 -- September 3, 1783 (8 years, 4 months and 15 days) Ratification effective: May 12, 1784 (9 years and 23 days) Location Eastern North America, Caribbean Sea, Indian subcontinent, Africa, the Atlantic Ocean, and the Indian Ocean Result Allied victory: Peace of Paris British recognition of American independence End of the First British Empire British retention of Canada and Gibraltar Territorial changes Great Britain cedes to the United States the area east of the Mississippi River and south of the Great Lakes and St. Lawrence River Great Britain cedes East Florida, West Florida, and Menorca to Spain Great Britain cedes Tobago and Senegal to France Dutch Republic cedes Negapatnam to Great Britain", "[French and Indian War] The outcome was one of the most significant developments in a century of Anglo - French conflict. France ceded to Great Britain its territory east of the Mississippi. It ceded French Louisiana west of the Mississippi River (including New Orleans) to its ally Spain in compensation for Spain's loss to Britain of Florida. (Spain had ceded Florida to Britain in exchange for the return of Havana, Cuba.) France's colonial presence north of the Caribbean was reduced to the islands of Saint Pierre and Miquelon, confirming Great Britain's position as the dominant colonial power in eastern North America.", "[Territory of the Saar Basin] The Territory of the Saar Basin (, ; ) was a region of Germany occupied and governed by the United Kingdom and France from 1920 to 1935 under a League of Nations mandate. It had its own flag (adopted on July 28, 1920): a blue, white, and black horizontal tricolour. The blue and white stood for Bavaria, and white and black for Prussia, out of whose lands the Saar Territory was formed. Initially, the occupation was under the auspices of the Treaty of Versailles. Its population in 1933 was 812,000, and its capital was Saarbrücken. The territory closely corresponds with the modern German state of Saarland, but was slightly smaller in area. After a plebiscite was held in 1935, it was returned to Germany.", "[Military history of the United States] In the Treaty of Paris after the Revolution, the British had ceded the lands between the Appalachian Mountains and the Mississippi River to the United States, without consulting the Shawnee, Cherokee, Choctaw and other smaller tribes who lived there. Because many of the tribes had fought as allies of the British, the United States compelled tribal leaders to sign away lands in postwar treaties, and began dividing these lands for settlement. This provoked a war in the Northwest Territory in which the U.S. forces performed poorly; the Battle of the Wabash in 1791 was the most severe defeat ever suffered by the United States at the hands of American Indians. President Washington dispatched a newly trained army to the region, which decisively defeated the Indian confederacy at the Battle of Fallen Timbers in 1794.", "[Mexican Cession] The Mexican Cession of 1848 is a historical name in the United States for the region of the modern day southwestern United States that Mexico ceded to the U.S. in the Treaty of Guadalupe Hidalgo in 1848. It had not been part of the areas east of the Rio Grande which had been claimed by the Republic of Texas, though the Texas annexation resolution two years earlier had not specified the southern and western boundary of Texas. The Mexican Cession (529,000 sq. miles) was the third largest acquisition of territory in US history. The largest was the Louisiana Purchase, with some 827,000 sq. miles, followed by the acquisition of Alaska (about 586,000 sq. miles).", "[Koerner, Ray & Glover] Koerner, Ray & Glover was a loose-knit group of three blues musicians from Minneapolis, Minnesota: \"Spider\" John Koerner on guitar and vocals, Dave \"Snaker\" Ray on guitar and vocals, and Tony \"Little Sun\" Glover on harmonica. They were notable figures of the revival of folk music and blues in the 1960s.", "[Minneapolis] Minneapolis lies on both banks of the Mississippi River, just north of the river's confluence with the Minnesota River, and adjoins Saint Paul, the state's capital. The city is abundantly rich in water, with 13 lakes, wetlands, the Mississippi River, creeks and waterfalls; many connected by parkways in the Chain of Lakes and the Grand Rounds National Scenic Byway. It was once the world's flour milling capital and a hub for timber. The city and surrounding region is the primary business center between Chicago and Seattle. As of 2018, Minneapolis was home to 6 Fortune 500 companies, and the Twin Cities were the fifth-largest hub of major corporate headquarters in the United States. As an integral link to the global economy, Minneapolis is categorized as a global city.", "[Swedish Livonia] Swedish Livonia () was a dominion of the Swedish Empire from 1629 until 1721. The territory, which constituted the southern part of modern Estonia (including the island of Ösel ceded by Denmark after the Treaty of Brömsebro) and the northern part of modern Latvia (the Vidzeme region), represented the conquest of the major part of the Polish-Lithuanian Duchy of Livonia during the 1600–1629 Polish-Swedish War. Parts of Livonia and the city of Riga were under Swedish control as early as 1621 and the situation was formalized in Truce of Altmark 1629, but the whole territory was not ceded formally until the Treaty of Oliva in 1660. The minority part of the Wenden Voivodeship retained by the Polish–Lithuanian Commonwealth was renamed the Inflanty Voivodeship (\"\"Livonian Principality\"\"), which today corresponds to the Latgale region of Latvia.", "[Río Rico, Tamaulipas] Río Rico is a town located along the Rio Grande river in the Mexican state of Tamaulipas. It is notable for its partial occupation of the Horcón Tract, a piece of land ceded by the United States to Mexico in 1977 under the terms of the Boundary Treaty of 1970.", "[Elliott Bay] Elliott Bay is a part of the Central Basin region of Puget Sound in the U.S. state of Washington that extends southeastward between West Point in the north and Alki Point in the south. Seattle was founded on this body of water in the 1850s and has since grown to encompass it completely. The waterway it provides to the Pacific Ocean has served as a key element of the city's economy, enabling the Port of Seattle to become one of the busiest ports in the United States."], "short_id": "q3"}