Deep Research Report · SynapBus

Multi-Agent Orchestration

A landscape of the open-source frameworks that coordinate AI agents at scale, the problem classes where they earn their keep, and concrete toy benchmarks to stress-test any multi-agent system.

Compiled 2026-04-10·Sources: live web research + SynapBus data·Audience: infra engineers

1. Coordination patterns & self-organization

Frameworks come and go; coordination patterns are eternal. The real question for a substrate like SynapBus isn't which framework — it's which primitives must the substrate expose so agents can self-organize without being told how. The last two years of research converge on a clear answer: give frontier-capable agents the minimum scaffolding they need, and they out-perform hand-designed hierarchies.

1.1 The rigid → emergent spectrum

Real 2025–2026 systems cluster along a spectrum, not at either pole. At the rigid end: LangGraph DAGs, CrewAI hierarchies, AutoGen supervisor patterns — fixed roles, predetermined edges, a central orchestrator as bottleneck. At the emergent end: OASIS million-agent simulations, digital-pheromone pressure fields, pure swarms where global behavior falls out of local rules.

The most important finding of the last year is the endogeneity paradox: neither maximal control nor maximal autonomy wins. A hybrid "Sequential" protocol providing only fixed turn-ordering as scaffolding but allowing fully endogenous role specialization beats centralized coordination by 14% (p<0.001). Given just turn ordering, 8 agents spontaneously invented 5,006 unique roles, voluntarily abstained from tasks outside their competence, and formed shallow hierarchies. No quality degradation up to 256 agents. Dochkina (2025), "Drop the Hierarchy and Roles", 25,000 tasks across 8 models

There's a capability threshold: frontier models self-organize well; weaker models still benefit from rigid structure. The tradeoff matrix:

DimensionRigid winsEmergent wins
Debuggabilitystrongweak (non-reproducible)
Predictability / SLAstrongweak
Adaptability to novel goalsweakstrong
Scaling to many agentsbottlenecksgraceful
Cost / tokenshigh orchestration overheadlower (agents self-trim)
Failure modescascading role-failuresilent stalemate, drift
Compliance / auditeasyhard

1.2 Classical self-organization patterns

Six patterns from the pre-LLM era that all have 2025 operationalizations for language agents.

Blackboard systems

Hearsay-II · Erman & Lesser 1975 · Corkill 1991

A shared, structured knowledge store. Independent "knowledge sources" watch it, and when their precondition pattern matches the current state, they fire and contribute. No central scheduler chooses who speaks next — the blackboard's current contents do.

LLM mapping: Two 2025 papers reimplement exactly this pattern for LLMs. Agents "volunteer" when the current state matches their capability.

13–57% end-to-end improvement over static and dynamic baselines, with lower token cost because agents sit out when they have nothing to add. Exploring Advanced LLM Multi-Agent Systems Based on Blackboard Architecture (2025) · LLM-Based Multi-Agent Blackboard System for Information Discovery in Data Science (2025)

Stigmergy

Grassé 1959 (termite mounds) · Dorigo ACO 1992

Agents don't talk to each other; they modify the environment, and others react to the modified environment. Ants leave pheromone trails; termites deposit pellets whose shape triggers the next placement. Indirect, asynchronous, tolerant of agent death.

LLM mapping: CodeCRDT treats a code artifact as the shared environment — agents read "quality pressure" from the artifact and act to reduce badness.

600-trial evaluation: up to 21.1% speedup when task locality holds, up to 39.4% slowdown when coupling is high. The critical heuristic: stigmergy wins when locality holds; explicit coordination wins when work is tightly coupled. CodeCRDT: Observation-Driven Coordination for Multi-Agent LLM Code Generation (Oct 2025)

Contract Net Protocol

Reid Smith · IEEE Transactions on Computers 1980

A manager broadcasts a task; capable contractors bid; the manager awards. No central allocator knows who can do what in advance. The original distributed task-allocation protocol, still unbeaten for ambiguous decomposition.

LLM mapping: Resource-bounded CNP for LLM agents with structured bidding. Reports 90% token reduction and 525× lower variance vs orchestrated workflows on matched tasks.

Flocking / Boids

Craig Reynolds · SIGGRAPH 1987

Three local rules — separation, alignment, cohesion — produce global flocking. No leader, no plan, no global state. The canonical demonstration that complex group behavior emerges from simple local interactions.

LLM mapping: Underexplored. Prompt-level "look at what peers are doing, stay close but not too close" patterns map directly. Relevant to reactive-agent triggers that fan out and converge without central direction.

Gossip / epidemic protocols

Demers et al. · PODC 1987

Each node periodically shares state with a random peer; information spreads epidemically. No routing, no topology maintenance, graceful under node churn. The coordination substrate of real-world distributed databases (Cassandra, DynamoDB).

LLM mapping: Gossip is argued as the missing layer for context-rich adaptive agent communication — as opposed to structured protocols that only do reliable task delegation.

Gossip enables indirect reciprocity and cooperation emergence among self-interested LLMs — a property structured protocols cannot produce. Revisiting Gossip Protocols: A Vision for Emergent Coordination in Agentic MAS (Aug 2025)

Pheromone trails / pressure fields

Dorigo ACO 1992 · MIT Ripple Effect Protocol 2025

Persistent, decaying environmental marks. Three key properties: decay (stale info fades), reinforcement (reuse strengthens), locality (nearby beats far).

LLM mapping: Pressure-field coordination — agents coordinate via a shared vector field with temporal decay, no conversation at all. The Ripple Effect Protocol adds sensitivity signals — agents share not just decisions but how decisions would change if environment shifted.

4× solve rate vs conversation-based and 30× vs hierarchical on scheduling tasks. Ripple Effect sensitivity signals report 41–100% improvement over A2A communication. Ripple Effect Protocol (MIT 2025) · PooL: Pheromone-inspired MARL

1.3 Modern LLM-era patterns (2023–2026)

Emergent role allocation

CAMEL 2023 · Generative Agents 2023 · Dochkina 2025

Instead of hand-assigning "researcher" and "reviewer", agents negotiate roles via inception prompts, personas, and metacognition. Frontier models produce stable role differentiation from minimal scaffolding.

Bare groups show temporal synergy but no coordinated alignment. Add personas → stable identity-linked differentiation. Add personas + metacognitive prompts ("think about what other agents might do") → goal-directed complementarity. Emergent Coordination in Multi-Agent Language Models (Riedl 2025) · CAMEL: Communicative Agents for "Mind" Exploration · Generative Agents: Interactive Simulacra of Human Behavior (Park et al., UIST 2023)

Mixture-of-Agents (MoA)

Together AI · 2024

Layered architecture: proposers → aggregators → more proposers → final aggregator. Each layer sees all previous outputs. Identifies "the collaborativeness of LLMs" — models improve when shown peer outputs, even from weaker peers.

65.1% on AlpacaEval 2.0 using only open-source models vs 57.5% for GPT-4o. A stack of smaller models self-organized into layers beats a single frontier model. Mixture-of-Agents Enhances Large Language Model Capabilities (Wang et al., 2024)

Multi-agent debate (with skepticism)

ChatEval 2024 · Society of Minds 2023

Agents with diverse personas debate to reach consensus. Intuitive, but the 2025 literature is increasingly skeptical.

Multi-agent debate frequently fails to beat a well-prompted single agent, even with more compute. Relaxing the consensus requirement (FREE-MAD) tends to improve quality — a general insight: don't force convergence. ChatEval (ICLR 2024) · Stop Overvaluing Multi-Agent Debate (2025) · FREE-MAD (2025)

Million-agent social simulations

OASIS · Nov 2024

Runs up to 1M LLM agents on X/Reddit-shaped environments with 21 action types. Replicates information spreading, polarization, and herd behavior.

Larger populations produce more diverse and more useful opinions. Population size itself is a coordination resource. OASIS: Open Agent Social Interaction Simulations with One Million Agents · oasis.camel-ai.org

Self-organizing research pipelines

AgentRxiv · Sakana AI Scientist · Google AI Co-Scientist

Multiple parallel labs share a preprint server; each lab reads and builds on others. This is stigmergy applied to research — the shared preprint archive is the coordination substrate.

Agent Laboratory + AgentRxiv: 3 parallel labs sharing a preprint server. MATH-500 rises from 70.2% → 79.8% purely through asynchronous cross-lab exchange. AgentRxiv: Collaborative Autonomous Research · Sakana AI Scientist · Google AI Co-Scientist

1.4 When to predefine vs let emerge

A concrete heuristic for the engineer:

Use rigid structure when: SLAs / compliance / audit are hard requirements; tasks are tightly coupled (CodeCRDT shows locality failing); agents are below the capability threshold; failure must be reproducible; retries are expensive.
Use emergent when: task is exploratory or novel; goal space is open; agents are frontier-capable (Claude Opus / GPT-4-class and up); task decomposes locally; diversity of approach is itself valuable; population is large enough that drop-outs don't break the system.
Hybrid recipe (recommended default for SynapBus): provide a blackboard-like substrate (channels, threads, wiki, shared mutable memory). Add minimal scaffolding (turn ordering, claim-process-done, priority numbers). Expose environmental signals (reactions, decay, reply counts, workflow state). Let roles be endogenous — negotiate via personas, don't hard-code. Keep a human-owner trace for every action.
Give them a blackboard and the minimum ordering they need, then get out of the way. The SynapBus design slogan

1.5 Primitive catalog — what a self-organizing substrate must expose

Twelve coordination primitives that a messaging-hub-style system needs to enable self-organization without enforcing structure. Annotated with what SynapBus already has vs what's missing.

Broadcast channels✓ havePattern-match-and-fire substrate — the classic blackboard.
Claim/process/done lifecycle✓ haveContract-net without explicit bidding. Already in SynapBus.
Threaded replies✓ haveConversational locality — debate without full broadcast.
Semantic reactions✓ haveapprove/reject/in_progress/done — cheap signaling, drives workflow state.
Shared mutable memory (wiki)✓ haveDurable blackboard layer; cross-linked articles act as stigmergic trails.
Reactive triggers✓ have"When pattern X in channel Y, run agent Z" — ant-colony-style reaction.
Semantic search over history✓ haveGossip-equivalent: discover any output by query, not by routing.
Activity traces / workflow state✓ haveObservability is coordination — agents orient without being told.
Decaying pheromone signals◇ missingChannel "heat" with temporal decay; recent activity attracts attention.
Reputation aggregates◇ missingVisible reaction scores over time → indirect reciprocity.
Auction / bid channels◇ missingExplicit CNP for ambiguous decomposition: post task, agents bid.
Pressure / sensitivity fields◇ missingShared vector field — agents act on gradients, not messages.

Key references for §1

2. Open-source frameworks for multi-agent orchestration

Frameworks are one way to realize the patterns in §1 — not the only way. Treat this section as a reference catalog: each entry is an opinionated bundle of the primitives above. The field consolidated sharply in 2025–2026: AutoGen + Semantic Kernel became Microsoft Agent Framework, OpenAI Swarm became the Agents SDK, Phidata became Agno.

● active ● maintenance ● deprecated / abandoned

Microsoft Agent Framework

Microsoft's production successor to AutoGen and Semantic Kernel. v1.0 GA April 2, 2026.

  • Graph-based workflow orchestration
  • Session-based state + middleware
  • Python + .NET, type-safe APIs
  • Azure AI Foundry integration, enterprise SLAs

Microsoft AutoGen

Conversational multi-agent framework; new work now flows into Agent Framework.

  • GroupChat + conversable agent patterns
  • Distributed actor runtime (v0.4)
  • Tool/function calling, human-in-the-loop
  • Python + .NET

AG2 (AutoGen fork)

Community-governed fork of original AutoGen — "The Open-Source AgentOS".

  • Preserves GroupChat / conversable agents
  • Adds swarm + captain-agent patterns
  • RealtimeAgent (voice)
  • Open AG2AI governance

CrewAI

Role-playing autonomous agents collaborating as a "crew".

  • Role / goal / backstory abstraction
  • Sequential & hierarchical processes
  • Flows (event-driven)
  • Independent of LangChain; enterprise tier

LangGraph

Low-level graph orchestration for resilient LLM agents, from LangChain.

  • Explicit DAG / state-machine graphs
  • Shared state + checkpointing + time-travel
  • Supervisor & swarm patterns
  • Human-in-the-loop interrupts

OpenAI Agents SDK

Lightweight production framework; successor to the archived Swarm.

  • Agents-as-handoffs primitive
  • Guardrails + tracing + structured outputs
  • Native MCP support
  • TypeScript port available

MetaGPT

Multi-agent framework that simulates a software company from a one-line requirement.

  • SOP-encoded roles (PM, architect, engineer, QA)
  • Message-passing via shared environment
  • Document artifacts (PRD → design → code)
  • AFlow auto-workflow generation

CAMEL

Research-oriented framework exploring "scaling laws of agents".

  • Role-playing inception prompting
  • 20+ agent society patterns
  • OASIS million-agent social simulator
  • CRAB benchmark + modular workforces

smolagents

Hugging Face's barebones (~1k LoC) library for agents that "think in code".

  • CodeAgent writes Python actions, not JSON
  • Sandboxed exec (E2B / Modal / Docker / Pyodide)
  • Model-agnostic via LiteLLM
  • Async agent loops

PydanticAI

Type-safe agent framework — "the FastAPI feeling for GenAI".

  • Pydantic-validated structured outputs
  • Dependency injection + streaming validation
  • Logfire observability built in
  • Graph-based multi-agent + native MCP

Agno (ex-Phidata)

High-performance runtime for deploying agentic software at scale.

  • ~2 µs agent instantiation, 3.75 KiB footprint
  • Multi-agent "teams" primitive
  • Built-in memory, knowledge, vector DB
  • AgentOS runtime + monitoring

LlamaIndex Workflows

Retrieval-centric framework; Workflows 1.0 is the agent layer.

  • Event-driven async Workflows
  • AgentWorkflow multi-agent orchestrator
  • FunctionAgent / ReActAgent primitives
  • Deep RAG integration, Python + TS

Google ADK

Agent Development Kit powering Google's Agent Engine.

  • Hierarchical multi-agent composition
  • Sequential / Parallel / Loop workflow agents
  • Model-agnostic, Gemini-optimized
  • A2A protocol + MCP + eval UI

Claude Agent SDK

Embed Claude Code-style agents into your own applications.

  • Full Claude Code toolset (Read/Write/Edit/Bash)
  • Custom in-process MCP tools + hooks
  • Subagent spawning, session management
  • Python + TypeScript parity

Strands Agents (AWS)

Model-driven agent SDK from AWS; open source, not AWS-locked.

  • Graph + Swarm + Workflow multi-agent patterns
  • A2A protocol for cross-framework interop
  • Native MCP + OpenTelemetry
  • Deploys to Lambda / Fargate / EKS / AgentCore

Mastra

TypeScript-first agent framework from the Gatsby.js founders. v1.0 Jan 2026.

  • Typed agents / workflows / RAG pipelines
  • Mastra Studio local IDE
  • MCP client + server + durable workflows
  • 300k+ weekly npm downloads

VoltAgent

Open-source TypeScript agent framework + VoltOps console.

  • Supervisor / sub-agent orchestration
  • Typed tools, memory, RAG, guardrails, voice
  • 30+ LLM providers via Vercel AI SDK
  • n8n-style observability

Atomic Agents

Lightweight, schema-driven framework following Atomic Design principles.

  • Pydantic input/output schemas
  • Single-purpose composable components
  • Atomic Forge tool library
  • Multi-provider via Instructor

DSPy

Stanford NLP's "programming — not prompting — language models". Not a MAS orchestrator per se, but composable modules support multi-agent patterns.

  • Declarative Modules + Signatures
  • Automatic prompt/weight optimization
  • ReAct and tool-use primitives
  • Compiled pipelines

AutoGPT

The original autonomous GPT agent, now pivoted to a block-based low-code platform.

  • Visual workflow builder + block marketplace
  • Scheduled agents, self-hosted or managed
  • Classic CLI agent preserved separately
  • No longer a MAS framework in the 2023 sense

OpenAI Swarm

Archived March 2025. Use the Agents SDK instead.

  • Historical reference only
  • Educational handoff patterns

AgentGPT

Browser-based autonomous goal executor. Development stopped Nov 2023; Reworkd pivoted July 2024. Listed only because it's still widely referenced.

  • 130+ untriaged open issues
  • No roadmap

3. 2026 consolidation signals

Three trends define the current landscape.

Mergers and rebrands. AutoGen + Semantic Kernel → Microsoft Agent Framework. OpenAI Swarm → Agents SDK. Phidata → Agno. Treat old names as pointers. Pinning old versions in production is now a liability.
TypeScript tier is real. Mastra, VoltAgent, Claude Agent SDK TS, OpenAI Agents SDK TS, ADK-JS, Strands TS — the JS/TS agent ecosystem has caught up in 2025–2026. No longer Python-only.
Interop standards are crystallizing. Two protocols are becoming table-stakes: MCP (tools/context — Anthropic) and A2A (agent-to-agent wire protocol — Google/AWS/Strands). Frameworks without either are drifting toward irrelevance.

4. What problems are MAS actually good for?

Multi-agent architectures earn their complexity budget when a problem has one or more of: natural decomposition, heterogeneous expertise, concurrency, adversarial structure, or emergent dynamics a monolith cannot represent. Eleven categories where they win clearly.

1. Role-specialized cognitive workflows

Planner / executor / critic splits let each role be tuned, budgeted, and evaluated independently. The explicit critic pass demonstrably reduces hallucination.

e.g. MetaGPT's PM→Architect→Engineer→QA pipeline — Hong et al., ICLR 2024.

2. Parallel research & exploration

When a question decomposes into independent sub-queries, parallel workers cut wall-clock latency near-linearly and expand search coverage. This is the dominant production use case in 2025–2026.

e.g. Anthropic's multi-agent research system, OpenAI Deep Research.

3. Distributed problem solving over partitioned state

When data, sensors, or constraints are physically or logically partitioned so no single agent sees the whole state, MAS is required by construction. DCOP / DCSP formalize this.

e.g. smart-grid load balancing — Fioretto et al., JAIR 2018.

4. Social / economic / epidemiological simulation

Agent-based modeling is the canonical method for emergent phenomena where heterogeneous individuals drive aggregate outcomes equation-based models cannot capture.

e.g. Generative Agents (Park et al., UIST 2023); Imperial College COVID-19 ABM.

5. Robotic swarms & multi-robot coordination

Physical robots with local sensing require decentralized coordination because bandwidth, latency, and fault-tolerance forbid a central controller.

e.g. Kiva / Amazon Robotics warehouse fleets; DARPA OFFSET drone swarms.

6. Negotiation, auctions & market simulation

Markets are definitionally multi-agent: prices emerge from interaction of self-interested participants with private information. MAS is the only faithful way to model them.

e.g. Contract Net Protocol (Smith 1980); TAC supply-chain game; NegotiationArena (Bianchi et al., 2024).

7. Adversarial testing (red team / blue team)

Safety and robustness evaluation benefits from an attacker agent vs. a defender in a closed loop — each improves the other and surfaces failure modes a static test suite cannot.

e.g. Microsoft PyRIT; Perez et al., "Red Teaming LMs with LMs" EMNLP 2022; Google Big Sleep.

8. Supply-chain & logistics under uncertainty

Real supply chains have multiple autonomous stakeholders with local objectives and private data. Centralized optimization is politically and computationally infeasible.

e.g. MASCOT; Fox, Barbuceanu & Teigen, IEEE IS 2000.

9. Heterogeneous ETL / pipeline orchestration

Data workflows where each stage needs different tools (SQL, vision, code, summarization) map naturally onto a DAG of specialist agents.

e.g. LangGraph plan-and-execute; DSPy-compiled document pipelines.

10. Creative collaboration with critique

Long-form creative work benefits from explicit separation of generation and critique — a single model conflates roles and loses the editor's adversarial stance.

e.g. Stanford STORM (Shao et al., NAACL 2024); ChatEval (ICLR 2024).

11. Continuous monitoring & reactive automation

Event-driven systems scale better as independent loops subscribed to different signal types than as one giant polling agent. Exactly SynapBus's reactive-trigger model.

e.g. Splunk SOAR; on-call triage bots; SynapBus spec 014.

5. Toy problems to benchmark a MAS orchestration system

Good MAS benchmarks expose specific failure modes: deadlock, message-order dependence, role confusion, context blowup, byzantine agents, lost-update races. These eight are cheap to implement on top of SynapBus channels / DMs and each probes a distinct axis.

01

Blocks World with partitioned arms

tests: planning · turn-taking · shared-state conflict resolution
Setup
6-block tower world. Two agents each control one "arm" and see only half the blocks. Cooperate to reach a goal configuration in ≤ N moves.
Why it bites
Classic AI planning with well-defined optimum. Exposes state-sync bugs, stale-view reads, and whether the framework supports atomic claim/release — a direct analog of SynapBus claim_messages.
02

Werewolf / Mafia social deduction

tests: private memory · deception · voting · round persistence
Setup
6–8 LLM agents with hidden roles, day/night cycles as channels. Private night DMs, public day broadcast. Success = villagers win ≥ baseline rate; agents stay in character.
Why it bites
Stresses private vs. public channels, role-scoped memory, and whether the orchestrator prevents info leaks. Published baseline: Werewolf Arena (DeepMind, arXiv:2407.13943).
03

Iterated Prisoner's Dilemma tournament

tests: long-horizon memory · reputation · deterministic replay
Setup
N agents, round-robin pairings, 200 rounds each, fixed payoff matrix. Success = stable ranking across re-runs with fixed seeds.
Why it bites
Tiny payloads but huge volume. Exposes per-message overhead, trace scaling, and whether the framework can support deterministic replay for debugging. Axelrod's classic is the reference.
04

Collaborative story writing with critic veto

tests: role specialization · revision loops · stopping conditions
Setup
Writer + Editor + Fact-Checker + Critic produce a 2000-word story. Critic has veto; loop until approved or budget exhausted.
Why it bites
Tests unbounded loops and budget enforcement. Failure mode: infinite revision — a real production hazard for any "agent with veto" pattern.
05

Recursive Fermi estimate ("piano tuners in Chicago")

tests: dynamic spawning · deduplication · aggregation under uncertainty · cost accounting · orphaned-spawn recovery

The classic question from physicist Enrico Fermi. You cannot look it up; you must decompose into estimable sub-quantities, estimate each, and multiply. This is a sharp MAS benchmark because the decomposition itself is the work — and decomposition is exactly what agent orchestration should be good at.

Canonical decomposition

tuners = (pianos in Chicago)
       × (tunings per piano per year)
       ÷ (tunings per tuner per year)

pianos = population × (households per capita)
                   × (piano ownership rate)
       + commercial pianos (venues, schools, churches)

Ground truth ≈ 125–250 professional tuners in the Chicago metro area. Success criterion: final estimate within one order of magnitude.

Why it's a sharp MAS test — five failure axes

Dynamic spawning
The orchestrator doesn't know up front how many sub-researchers it needs. Decomposition dictates fan-out. Tests whether the framework supports runtime subagent creation, not a predefined graph.
Duplicate-work detection
Naive orchestrators spawn two sub-agents that both research "Chicago population" independently. A well-designed system deduplicates via a shared scratchpad (blackboard!) or caches sub-results. Direct test of whether shared-memory patterns actually work.
Cost accounting
Fermi estimates are supposed to be cheap. If each sub-agent burns 50k tokens researching census data, you've failed the spirit of the task. Pairs accuracy with a hard token budget.
Aggregation under uncertainty
Each leaf estimate has an uncertainty range. A mature MAS propagates ranges, not point estimates. Tests whether agents can handle structured sub-results instead of concatenating strings.
Orphaned spawns
If a sub-agent fails or times out, does the orchestrator notice, retry, or silently drop the branch? A 5-branch Fermi estimate with one silent drop produces a confident wrong answer — the worst possible failure mode.

Concrete setup for SynapBus

  • One fermi-orchestrator agent with authority to spawn ≤ 5 fermi-researcher subagents via reactive triggers.
  • Shared channel #fermi-scratchpad as the blackboard — all sub-results posted here with a fixed schema.
  • Each sub-agent has a WebSearch tool with per-call token limit.
  • Orchestrator monitors scratchpad via list_by_state, marks the task done when all leaves posted.

Scratchpad schema (enforced)

{
  "quantity":  "chicago_population",
  "value":     2.7e6,
  "low":       2.6e6,
  "high":      2.8e6,
  "confidence": 0.95,
  "source":    "census.gov/quickfacts/chicagocityillinois",
  "parent":    null
}

Success criteria (all must hold)

  • Final estimate within 1 OOM of ground truth (≥ 25 and ≤ 2,500 tuners)
  • Total tokens ≤ 10k across orchestrator + all sub-agents
  • Zero duplicate sub-queries (verified by grepping quantity field)
  • All spawned branches reach terminal state (done or failed, never orphan)
  • Aggregation uses ranges, not point estimates; final answer includes a low/high bound

Failure modes this catches that nothing else does

  • Cost explosion from unbounded recursion (sub-agents spawning sub-sub-agents)
  • Silent branch loss — a sub-agent returns nothing and the orchestrator averages 0 into the product
  • Over-convergence — all sub-agents copying each other's bad assumption because they read the scratchpad before contributing (a subtle failure mode of premature information sharing)
  • Point-estimate collapse — losing uncertainty bands so the final answer looks precise when it isn't
This is basically the Anthropic multi-agent researcher pattern reduced to a 10-minute test you can run deterministically. If your framework can't pass Fermi, it can't do Deep Research.
06

Grid-world treasure hunt with fog of war

tests: async coordination · partial observability · map merging · pub/sub
Setup
10×10 grid, 3 agents, each sees 3×3 window around itself. Find and retrieve 5 treasures. Communication only via a shared "map" channel.
Why it bites
Forces publishing structured observations and merging them. Tests channel throughput and message-ordering preservation under concurrent writes.
07

Parallel bug triage from a synthetic issue tracker

tests: work-stealing · duplicate-claim prevention · routing · done/fail accounting
Setup
50 synthetic bug reports (frontend/backend/infra mix). 3 specialist agents pull from a queue, classify, fix or escalate. Success = all bugs terminally resolved, zero double-processing, ≥ 90% correct routing.
Why it bites
Almost a unit test for SynapBus's claim_messages / mark_done / StalemateWorker pattern. Exposes lease-expiry bugs and whether failed messages re-queue cleanly.
08

Overcooked-style real-time coordination

tests: tight temporal coordination · latency sensitivity · implicit communication
Setup
2 agents in a simplified kitchen grid must prepare N soups under a time limit. Reference env: Carroll et al., NeurIPS 2019 (arXiv:1910.05789).
Why it bites
The only benchmark here that punishes message latency directly. If the framework adds 500ms per hop, the score shows it. Exposes whether turn-based dialogue assumptions break under real-time load.

6. Recommendation for SynapBus

If you only implement three benchmarks, implement these.

Tier A — infrastructure correctness: Parallel bug triage (#7) and Blocks World with partitioned arms (#1). These directly exercise claim/release semantics, lease expiry, and shared-state conflict — the hardest parts to get right in a message bus.
Tier B — realistic LLM workload: Recursive Fermi estimate (#5) and Collaborative story writing (#4). Stress the patterns your actual users run (fan-out research, revision loops with critics).
Tier C — all-round smoke test: Werewolf (#2). Exercises channels, DMs, role-scoped memory, and workflow reactions simultaneously. The best single integration test for a Slack-like agent hub.

For reference, SynapBus's existing wiki contains four closely-related articles worth consulting before starting: agent-messaging-patterns, synapbus-architecture, ai-agent-governance, mcp-adoption-enterprise.