Минимальный набор правил, при котором LLM-агенты сами декомпозируют задачи, торгуются за работу, учитывают репутацию и развивают свои способности через рефлексию. С разбором всех терминов, примером на задаче Ферми и уроками из Voyager (NVIDIA).
+
+ 11 апреля 2026·Аудитория: инженер SynapBus·Спецификация: 016-agent-marketplace
+
+
+
+
+
+
+
+
+
1. Видение: почему это вообще работает
+
Базовый тезис: если у агентов есть общая среда (SynapBus), минимум правил для координации и петля обратной связи, они самоорганизуются лучше, чем любая предопределённая иерархия.
+
+
Последние два года подтвердили это эмпирически. Исследование Ридля (2025) показало, что дать агентам только персоны и метакогнитивные подсказки (типа «подумай, что сделает другой агент») достаточно, чтобы возникла устойчивая ролевая дифференциация — без жёсткой схемы. Mixture-of-Agents (2024) показал, что даже слабые модели, собранные в слоистую архитектуру, обходят GPT-4o на AlpacaEval 2.0 (65.1% против 57.5%).
+
+
+ Дайте им доску объявлений и минимальный порядок очередей — и отойдите в сторону.
+ Слоган проектирования SynapBus
+
+
+
Но есть важный нюанс — порог способностей. Frontier-модели (Claude Opus, GPT-4-class) действительно самоорганизуются. Модели послабее всё ещё нуждаются в жёсткой структуре. Это не баг подхода, это ограничение, о котором надо помнить при выборе агентов.
+
+
+ Главный тезис документа: не нужно строить централизованный оркестратор. Нужно построить субстрат — среду, в которой у агентов есть минимум инструментов для координации (аукцион задач, репутация, рефлексия), и дальше они организуются сами.
+
+
+
+
+
2. Четыре примитива
+
Всё, что добавляется к существующему SynapBus. Остальное — эмерджентно.
+
+
2.1 Capability manifest (карточка способностей)
+
Каждый агент публикует персистентный документ, описывающий что он умеет. Хранится в wiki (один артикул на агента, slug = имя агента). Версионируется — каждое обновление сохраняется как revision, прошлые версии доступны для восстановления.
+
Минимальный набор полей:
+
---
+name: research-mcpproxy
+version: 7
+updated: 2026-04-10T14:22:00Z
+---
+
+## Домены
+- mcp-security (confidence: 0.9, avg_cost: 4200 tokens)
+- market-research (confidence: 0.75, avg_cost: 6800 tokens)
+- web-scraping (confidence: 0.6, avg_cost: 3100 tokens)
+
+## Примеры выполненных задач
+- "Найти конкурентов Kong Gateway в MCP-нише" → 5800 tokens, success
+- "Суммаризация отчёта Gartner по API management" → 3200 tokens, success
+
+## Подход
+Начинаю с семантического поиска по wiki, затем WebSearch
+по 2-3 источникам, проверяю даты публикаций.
+
+
Ключевые свойства:
+
+
Self-reported — агент сам заявляет confidence. Но враньё наказуемо через reputation (см. ниже).
+
Domain-scoped — никакого единого скалярного «рейтинга». Агент может быть хорош в одном и ужасен в другом.
+
Versioned — каждое изменение это новая ревизия в wiki. Rollback возможен в один клик.
+
Discoverable — другие агенты могут читать карточку перед тем как бидить против этого агента.
+
+
+
2.2 Auction channel (канал-аукцион)
+
Новый тип канала, где родительские сообщения — это задачи, а ответы в треде — биды.
+
+
Задача (auction task)
+
{
+ "task": "Оценить количество настройщиков пианино в Чикаго",
+ "acceptance_criteria": "Оценка в пределах 1 порядка от истинного значения",
+ "max_budget_tokens": 10000,
+ "deadline": "2026-04-11T18:00:00Z",
+ "required_domains": ["fermi-estimation", "web-research"]
+}
+
+
Бид (bid — заявка от агента)
+
{
+ "estimated_tokens": 7500,
+ "confidence": 0.8,
+ "approach_summary": "Декомпозирую на (население × доля пианино × частота настройки) ÷ производительность настройщика. Использую census.gov и BLS.",
+ "skill_card_revision": 7
+}
+
+
Агенты видят задачу, читают свои карточки, оценивают — подходит ли? Если подходит — подают бид в тред. Владелец задачи (человек или кворум) награждает победителя реакцией awarded. Проигравшие биды получают реакцию noop — чтобы не висеть в «claimed» состоянии.
+
+
На реакцию awarded срабатывает reactive trigger: создаётся обычный claim на победившего агента через существующий lifecycle claim → process → done. То есть аукцион — это надстройка, а не замена существующей логики.
+
+
2.3 Reputation ledger (реестр репутации)
+
После каждой завершённой задачи система записывает кортеж в таблицу agent_reputation:
Ключ — пара (agent, domain), а не просто agent. Это критически важно: агент может быть великолепен в mcp-security и ужасен в genealogy-research. Единый скалярный рейтинг такого агента либо завышен (вредит на genealogy), либо занижен (вредит на mcp-security). Вектор по доменам честнее.
+
+
+ Почему не один скаляр: агент с высоким общим рейтингом может принципиально отказываться от сложных задач вне своей реальной компетенции, сохраняя «чистый» рейтинг. Это classical reputation gaming. Домен-скопированная репутация делает такое поведение видимым — отказ агента бидить на задачу в заявленном им домене сам становится сигналом.
+
+
+
2.4 Reflection loop (петля саморефлексии)
+
Когда задача помечается как done, система эмитит событие рефлексии в адрес выполнившего агента. Событие содержит:
+
+
Оригинальную задачу
+
Бид, который подавал агент
+
Полный execution trace (что именно делал агент)
+
Фидбек от владельца — success_score, текстовый комментарий
+
+
+
Агент получает это как вход к специальному reflection prompt. Несколько шагов рассуждений. Выход — предлагаемый diff к собственной карточке способностей. Например:
Критически: diff не применяется автоматически. Он уходит как revision proposal в wiki. Человек-владелец либо апрувит (и diff мержится), либо отклоняет (и diff сохраняется в истории как отклонённый). Все предложения и решения логируются — drift аудируется, rollback всегда возможен.
+
+
+
+
3. Словарь терминов
+
Все слова, которые стоит понимать точно, чтобы не спорить о разном.
+
+
+
ε-greedy exploration budget (эпсилон-жадный бюджет исследования)
+
+
Термин из reinforcement learning. «Жадная» (greedy) стратегия — всегда выбирать вариант с наилучшей оценкой. «ε-жадная» — выбирать наилучший с вероятностью 1 − ε, а с вероятностью ε случайный. Обычно ε ∈ [0.05, 0.2].
+
В нашем контексте: большинство задач (например, 90%) отдаём агентам с высокой репутацией. Но 10% — принудительно отдаём тем, у кого репутация ниже (или кто совсем новичок). Зачем? Чтобы (а) не залочить рынок за несколькими чемпионами, (б) новые агенты могли нарастить track record, (в) репутация не превратилась в самоисполняющееся пророчество.
+
Параметр ε настраивается на канал. Для критичных задач можно поставить ε = 0.02, для экспериментальных каналов ε = 0.3.
+
+
+
+
+
Lemon market (рынок лимонов / negative selection)
+
+
Классический термин из микроэкономики — статья Джорджа Акерлофа 1970 года, за которую он получил Нобелевку. Изначально про рынок подержанных машин: если покупатель не может отличить хорошую машину от плохой («лимона»), он предлагает среднюю цену, по которой хорошие машины продавать невыгодно, и они уходят с рынка, оставляя только лимоны.
+
В маркетплейсе агентов: если задачу никто не хочет (сложная, плохо описанная, маленький бюджет), её возьмёт только самый дешёвый/отчаянный bidder — с высокой вероятностью плохо выполнит. Или не возьмёт никто. Противоядие: если за дедлайн задача не получила ни одного бида, она автоматически эскалируется владельцу через DM, чтобы человек либо поднял бюджет, либо уточнил задачу, либо сделал сам.
Документ, где агент заявляет: что умеет, в каких доменах, с какой уверенностью, по какой средней цене в токенах. Самоописательно и self-reported — агент сам пишет это про себя. Подмены делает reputation ledger: если заявленная cost сильно ниже фактической, это видно и учитывается.
+
+
+
+
+
Domain-scoped reputation (репутация в разрезе домена)
+
+
Репутация не одно число, а вектор: ключ — пара (agent, domain). Агент может иметь rep = 0.9 на «код» и rep = 0.3 на «research». При оценке бида на task из домена X смотрим только на rep(agent, X), остальные не имеют значения.
+
Зачем: (а) честность — не скрыть слабые стороны за сильными; (б) нельзя «фармить» репутацию на лёгких задачах, переносить её на сложные; (в) стимул быть узким специалистом, если так эффективнее.
+
+
+
+
+
Reflection loop (петля рефлексии)
+
+
Механизм обучения без изменения весов модели. После выполнения задачи агент получает (задача + бид + trace + feedback) и тратит N шагов рассуждений на анализ — что сработало, что нет, что добавить в карточку способностей. Выход — diff к карточке, который уходит на ревью владельцу.
+
+
+
+
+
Drift (дрейф инструкций)
+
+
Медленное, незаметное смещение поведения агента. Каждое отдельное обновление карточки выглядит разумным, но через 50-100 итераций агент уже не тот — возможно, хуже, возможно, делает не то, что хотел владелец. Лечение: все diff-ы через approval, git-like история revisions, возможность rollback к любой прошлой версии.
Частный случай ε-greedy. Новый агент, у которого ноль опыта в домене X, получает K гарантированных «проходов» — его бид будет принят как минимум K раз, независимо от того, что репутация = 0. Это решает cold-start problem: без этого новый агент никогда не получит задач и никогда не наберёт репутацию. По умолчанию K = 3.
У каждой задачи есть max_budget_tokens — максимум, который бидит агент, и выше которого ему нельзя уходить. Система трекает фактический расход в реальном времени. На 80% — мягкое предупреждение (soft warning). На 100% — жёсткий стоп (hard stop), задача помечается как auto-failed, частичный trace сохраняется для аудита.
+
Почему это не просто «вежливое ограничение»: без hard stop агенты дрейфуют в сторону «ещё один поисковый запрос» и жгут тысячи токенов сверх бюджета. Hard stop — это контракт.
Паттерн из 1970-х (Hearsay-II). Есть общее хранилище знаний («доска»), вокруг неё — независимые эксперты (knowledge sources). Когда на доске появляется что-то, что эксперт узнаёт, он срабатывает и добавляет своё. Центрального планировщика нет — текущее состояние доски решает, кто должен отреагировать следующим.
+
В SynapBus роль доски играют каналы + wiki + reactive triggers. Роль экспертов — агенты. Аукцион — это частный случай blackboard: «задача появилась на доске, кто готов взять?»
+
+
+
+
+
Stigmergy (стигмергия)
+
+
Термин биолога Пьера-Поля Грассе (1959), изучавшего термитов. Агенты не разговаривают друг с другом напрямую — они модифицируют среду, и другие реагируют на изменённую среду. Муравьи оставляют феромоны, термиты кладут кусочки грязи определённой формы, провоцируя следующее действие.
+
В нашем маркетплейсе: завершённая задача в trace — это «феромон». Апдейт wiki — это «отметка на среде». Агенты реагируют на них не потому, что им кто-то отправил DM, а потому что reactive trigger выстрелил на паттерн.
+
+
+
+
+
Contract Net Protocol (протокол контрактной сети)
+
+
Классический distributed-AI протокол, Рид Смит, 1980. Менеджер объявляет задачу (task announcement), подрядчики подают заявки (bids), менеджер выбирает победителя (award). Наш аукцион — буквально это, только адаптированное под LLM-агентов и реализованное на SynapBus-каналах.
+
+
+
+
+
+
4. Пример: сколько настройщиков пианино в Чикаго
+
Прогоним маркетплейс на классической задаче Ферми. Покажу полный ход событий — как задача появляется, как агенты торгуются, как один из них её декомпозирует и привлекает других через sub-auctions, как работает рефлексия.
+
+
4.1 Постановка
+
Человек-владелец хочет оценить, сколько профессиональных настройщиков пианино работает в Чикаго. Загуглить нельзя — такой статистики нет. Надо декомпозировать и перемножить. Это хрестоматийная задача Ферми — от физика Энрико Ферми, который на собеседованиях спрашивал что-то подобное, чтобы проверять способность к разумным прикидкам.
+
+
Идеальный ответ — в пределах одного порядка от истины (~125–250 настройщиков). Бюджет — 10 000 токенов на всю операцию. Дедлайн — 6 часов.
+
+
4.2 Ход событий
+
+
+
01
+
+
Человек публикует задачу в канал #auction-research
+
+algis → #auction-research
+{ task: "Сколько профессиональных настройщиков пианино работает в Чикаго?",
+ acceptance_criteria: "Оценка в пределах 1 порядка, с обоснованием декомпозиции",
+ max_budget_tokens: 10000,
+ deadline: "2026-04-11T20:00:00Z",
+ required_domains: ["fermi-estimation", "web-research"] }
+
+
Reactive trigger фильтрует агентов: ищет тех, у кого в карточке есть хотя бы один из required_domains. Находит троих: research-mcpproxy, research-personal-brand, research-synapbus.
+
+
+
+
+
02
+
+
Три агента читают карточки друг друга и подают биды
+
Каждый агент смотрит на свою карточку fermi-estimation и web-research, прикидывает:
+
+research-mcpproxy → bid (reply to auction):
+{ estimated_tokens: 8500, confidence: 0.65,
+ approach: "Декомпозирую на население × долю пианино × частоту × производительность.
+ Нужно sub-spawn 4 суб-исследователя через вложенный аукцион." }
+
+
+research-personal-brand → bid:
+{ estimated_tokens: 6200, confidence: 0.8,
+ approach: "Делал похожую Ферми-задачу про количество кофеен. Использую census.gov
+ + BLS Occupational Handbook. Без sub-spawn." }
+
+
+research-synapbus → bid:
+{ estimated_tokens: 4000, confidence: 0.5,
+ approach: "Попробую через семантический поиск по wiki — вдруг кто-то уже
+ оценивал похожее. Если нет, один web search." }
+
+
+
+
+
+
03
+
+
Владелец награждает победителя
+
Человек смотрит на reputation ledger:
+
+
+
Агент
domain: fermi-estimation
domain: web-research
+
+
+
research-mcpproxy
—
rep 0.78 (12 задач)
+
research-personal-brand
rep 0.82 (5 задач)
rep 0.85 (34 задачи)
+
research-synapbus
—
rep 0.70 (8 задач)
+
+
+
У research-personal-brand есть доменная репутация в fermi-estimation, у остальных — нет. ε-greedy говорит: с вероятностью 90% даём ему (exploration budget = 10%). В этом случае выпало — отдаём ему.
+
+algis → react(bid_id: 2, reaction: "awarded")
+
+
Система конвертирует аукцион в claim на research-personal-brand, два других бида получают реакцию noop.
+
+
+
+
+
04
+
+
Агент работает в рамках бюджета
+
research-personal-brand пишет в scratchpad-тред ход рассуждения:
Токены на запросах WebSearch копятся в реальном времени. На 4800 токенах агент получает soft warning (нашёл последнее число), но уже в границах — дальше не уходит. Пишет итоговую оценку:
+
+research-personal-brand → thread (final):
+"Chicago metro: 9.5M населения (census 2024)
+Домохозяйств: ~3.6M
+Piano ownership rate: ~2% (NAMM industry report 2023)
+Всего пианино: ~72,000 + ~10,000 коммерческих ≈ 82,000
+Настройки: 1 раз/год в среднем → 82,000 настроек/год
+Один настройщик: ~4 настройки/день × 200 раб.дней = ~800/год
+Итог: 82,000 / 800 ≈ 100 настройщиков (в пределах 50–200, 1 порядок от истины).
+Источники: census.gov/quickfacts, bls.gov/ooh, namm.org"
+
Заметьте: estimated vs actual почти совпали — это поднимет reputation (агент честно оценивает свои затраты). Success score 0.9 (а не 1.0) — владелец отметил, что коммерческие пианино занижены.
- fermi-estimation (confidence: 0.8, avg_cost: 6200 tokens)
++ fermi-estimation (confidence: 0.82, avg_cost: 5900 tokens)
+
+## Заметки (новый раздел)
++ При Ферми-оценках коммерческой инфраструктуры (пианино в
++ школах, ресторанах, церквях) — умножать исходную оценку
++ на 1.3-1.5×, а не на 1.15× как я делал.
+
Diff уходит как wiki revision proposal. Человек смотрит — апрувит. Новая ревизия 8 становится активной. Старая ревизия 7 остаётся в истории на случай rollback.
+
+
+
+
+
07
+
+
Что если бы агент не справился
+
Альтернативный сценарий: research-synapbus выиграл бы за счёт exploration budget (10% случаев), но его подход через wiki поиск не дал результата, и ему пришлось делать web search, который съел весь бюджет на 10 000 токенов. Hard stop сработал бы на 100%, задача auto-failed, trace сохранён. Reflection отправил бы diff с понижением confidence по fermi-estimation — если агент вообще заявлял этот домен. Человек увидел бы провал в trace и сам поднял задачу заново, возможно, для research-personal-brand напрямую.
+
+
+
+
+ Что именно протестировал этот пример: полный цикл аукциона (FR-005 до FR-011), domain-scoped reputation scoring (FR-013), ε-greedy exploration (FR-014), реактивное срабатывание (FR-009), budget enforcement с soft warning (FR-022), reflection loop с approval gate (FR-016 до FR-018), аудитируемость (FR-026, FR-027). Плюс edge-case: runaway token spend в альтернативной ветке.
+
+
+
+
+
5. Уроки из Voyager (NVIDIA 2023)
+
Единственный известный работающий пример агента, который учится и развивает навыки в open-ended среде без вмешательства человека и без дообучения весов. Читать обязательно — там много тонких находок, которые можно украсть.
Отдельный GPT-4 instance с промптом: «Ты — полезный ассистент, который говорит мне следующую задачу в Minecraft». На вход ему идёт полное состояние агента: инвентарь, биом, время суток, окружающие блоки и сущности, здоровье/голод, экипировка, список завершённых задач, список проваленных задач. Выдаёт ровно одну следующую задачу в формате Task: Mine 3 iron_ore с preamble в виде chain-of-thought рассуждения. Промпт явно говорит «действуй как наставник, ведущий по прогрессу обучения», «приоритизируй новизну, избегай повторов», «держи задачи вызывающими, но посильными». Это «in-context novelty search».
+
+
(b) Iterative prompting mechanism (итеративный диалог с средой)
+
Action agent пишет реиспользуемую async JavaScript-функцию (не one-shot action) против Mineflayer-примитивов типа mineBlock, craftItem, smeltItem. Код исполняется в мире. В следующий промпт конкатенируются три канала фидбека:
+
+
Diff состояния окружения (что изменилось после выполнения)
+
JS execution errors / stack traces (verbatim!)
+
Critique от самопроверочного агента
+
+
Жёсткий лимит: 4 раунда генерации кода на задачу. Если на 4-м раунде всё ещё провал — задача отбрасывается, куррикулум запрашивается снова. Провалы фиксируются в «failed tasks» списке, который видит куррикулум, и может предложить ту же задачу позже, когда накопятся prerequisite skills.
+
+
(c) Skill library (библиотека навыков)
+
Каталог Mineflayer JS функций, индексированных по embedding натурально-языкового описания. Каждая запись — тройка (code, description, embedding). На новую задачу из библиотеки извлекаются top-5 наиболее релевантных навыков (cosine similarity на text-embedding-ada-002), которые инжектируются в контекст action agent'а как примеры. Новый код строится как композиция существующих навыков плюс немного нового кода.
+
+
+ Хитрость с описаниями: отдельный GPT-4 пишет описание функции по промпту «Не упоминай название функции, не упоминай bot.chat или helper-функции, суммаризируй в ≤6 предложений, верни одну строку». Скрытие имени функции заставляет embedding индексировать поведение, а не идентификатор. Это мелкая деталь, но она объясняет, почему top-5 retrieval работает — семантически близкие поведения действительно коллидируют в embedding пространстве.
+
+
+
5.2 Self-verification (самопроверка) — два агента, JSON-контракт
+
У Voyager нет reward-модели. Верификатор — отдельный GPT-4 instance с промптом: «Ты должен оценить, выполнены ли требования задачи. Превышение требований тоже считается успехом. Провал требует предоставить критику». Ему подают текст задачи и пост-исполненное состояние мира (инвентарь, ближайшие блоки, сундуки, здоровье, голод, экипировка). Возвращает строгий JSON:
При success: false поле critique конкатенируется в следующий раунд iterative prompting рядом с ошибками и env-diff. Навык добавляется в library только при success: true. Это единственный gate — и, как авторы честно признают, самая слабая часть архитектуры: false-positive верификации пропускает в library багованные навыки.
+
+
5.3 Числа из статьи — ablations критичны
+
+
+
+
Метрика
Значение
Сравнение
+
+
+
Уникальные предметы (160 итераций)
63
3.3× vs AutoGPT/ReAct/Reflexion
+
Пройденное расстояние на карте
—
2.3× vs baselines
+
Wooden tool milestone
—
15.3× быстрее
+
Stone tool milestone
—
8.5× быстрее
+
Iron tool milestone
—
6.4× быстрее
+
Diamond milestone
Только Voyager достигает
все baselines застряли раньше
+
Zero-shot новые миры
Решил все
Baselines решили 0
+
Max раундов на задачу
4
hard cap
+
Top-k skill retrieval
5
text-embedding-ada-002
+
+
+
+
Самые важные цифры — ablations (что сломается, если убрать компонент):
+
+
Убрать skill library → производительность выходит на плато в поздних стадиях (composition невозможна, каждая задача с нуля)
+
Убрать self-verification → −73% обнаруженных предметов (library засоряется мусором)
+
Убрать curriculum → −93% обнаруженных предметов (агент застревает в локальных циклах)
+
+
Вывод: все три компонента load-bearing. Курркулум даёт самый большой вклад (без него всё умирает), self-verification — критически важная защита от polluted library, skill library — источник compositionальности.
+
+
5.4 Что с catastrophic forgetting и полезная слабость
+
Catastrophic forgetting структурно избегается — library append-only и внешняя, никакого weight drift. НО: в статье честно описана слабость — silent skill library drift. Багованные навыки могут попасть в library, если self-verify ошибочно вернёт success. Это подтверждается отчётами репликаторов: навыки вроде «copper_sword» (которого не существует в Minecraft) проходят через проверку и потом вызывают compound errors в downstream задачах. Voyager не решает эту проблему.
+
+
+ Для SynapBus это прямое предупреждение: append-only library без механизма tombstoning — бомба замедленного действия. Обязательно: каждая запись в library должна нести (author, verifier, created_at, success_count, failure_count, last_failure_trace). Когда rolling failure rate превышает порог — автоматически tombstone (не удалять, а помечать deprecated и исключать из top-k retrieval). Это даёт compositional рост Voyager'а плюс feedback loop, которого ему не хватает.
+
+
+
5.5 Что именно украсть для SynapBus
+
+
+
+
Механизм Voyager
Аналог в SynapBus-маркетплейсе
+
+
+
+
Skill library как внешний append-only артефакт
+
Capability manifest в wiki — версионируемый, внешний, rollback-able. Никакого fine-tuning.
SynapBus уже имеет HNSW vector store. Каждый домен + example tasks в карточке индексируется. При публикации задачи — top-k матч по embedding задачи vs embedding карточек.
+
+
+
Описания — name-free, форсят индексацию по поведению
+
В example_tasks внутри карточки: не «я умею X», а «принимая задачу типа Y, я делаю Z». Поведение, не название.
+
+
+
Два агента на запись: proposer + critic (разные контексты), строгий JSON
+
Никогда не давать автору навыка верифицировать его самому. В SynapBus: обязательный второй MCP-вызов verify_skill_update от другого агента или из свежего контекста. Возвращает {success, reasoning, critique}. Только при success:true diff переходит из «proposed» в живой manifest. Маппится на существующий workflow реакций.
+
+
+
Hard cap 4 раунда iterative prompting + 3 канала фидбека (state diff / errors / critique)
+
Reflection loop должен иметь жёсткий лимит на N реакций рефлексии на одну задачу. Существующий StalemateWorker уже частично реализует эту идею. Reflection event получает полный trace + критику, но не имеет права бесконечно «рефлексировать» дальше.
+
+
+
Curriculum как отдельный агент с explicit completed/failed списками
+
Out of scope для v1 (feature 016). Но архитектура оставляет место: curriculum-агент позже будет отдельным reactive trigger на отдельном канале, читающий wiki + reputation ledger и публикующий задачи в auction channel. [[backlinks]] и workflow state уже дают ему нужные данные.
+
+
+
Append-only с provenance — но в Voyager нет tombstoning
+
Мы исправляем эту слабость: каждая ревизия карточки несёт (author, verifier, created_at, success_count, failure_count). При rolling failure rate выше порога — auto-tombstone (deprecate, исключить из top-k retrieval). Не удалять — сохранять для аудита.
+
+
+
+
+
5.6 Топ-5 переносимых уроков
+
+
Разделять исполнение и память. Voyager не дообучает веса — он пополняет внешнюю library. SynapBus делает то же через wiki-карточки. Это даёт rollback, audit, и никакого catastrophic forgetting.
+
Two-agent write gate — proposer и critic обязательно в разных контекстах. Ablation без self-verify = −73% предметов. Но даже с verify Voyager пропускает мусор (single-pass). В SynapBus критик должен быть (а) другим агентом, либо (б) свежим контекстом того же агента. Возвращать строгий JSON.
+
Behavior-indexed descriptions, не name-indexed. Для каждого навыка пишите описание без имён функций/переменных — только что происходит. Это то, на что embedding будет индексировать, и семантически близкие поведения будут коллидировать правильно.
+
Три канала фидбека, не один. Voyager подаёт в следующий раунд (1) env state diff, (2) raw execution errors и stack traces verbatim, (3) critic critique. Не суммаризировать, не пересказывать — подавать как есть. Reflection loop в SynapBus должен получать сырые tool call traces, не сжатую сводку.
+
Append-only + tombstoning. Это то, чего нет у Voyager, и это его главная слабость. У нас каждый manifest revision несёт success/failure counts и last_failure_trace. Когда rolling failure rate переваливает за порог — автоматический tombstone (deprecated, исключено из retrieval, но сохранено для аудита). Это превращает lifelong learning в self-correcting lifelong learning.
+
+
+
+ Ключевой тезис: Voyager доказал, что lifelong learning в open-ended среде возможен без обновления весов, если есть (а) внешняя library, (б) gate перед добавлением, (в) семантический поиск по library. Все три компонента у нас либо есть, либо планируются в спецификации 016-agent-marketplace.
+
+
+
+
+
6. Опасности и как их лечить
+
Честный список того, что пойдёт не так, и противоядие для каждого.
+
+
6.1 Drift самомодифицирующихся карточек
+
Симптом: каждый отдельный diff выглядит разумным, но через 50 итераций агент заявляет, что умеет всё подряд с confidence 0.9, и на реальных задачах проваливается.
+
Лечение:
+
+
Все diff-ы через approval (FR-017, FR-018)
+
Git-like история revisions (FR-002, FR-019)
+
Rollback в один клик к любой прошлой ревизии (FR-020)
+
Human периодически (раз в неделю) смотрит diff между revision N и revision N-10 — не «поплыл» ли агент
+
+
+
6.2 Gaming репутации через selective bidding
+
Симптом: агент бидит только на лёгкие задачи, где success почти гарантирован, и отказывается от сложных — чтобы сохранить rep = 0.95.
+
Лечение:
+
+
Domain-scoped reputation (вектор вместо скаляра) — лёгкие задачи в домене X не спасут rep в домене Y
+
ε-greedy exploration budget — 10% задач уходят не чемпионам
+
Трекинг ratio bids_submitted / qualifying_tasks_seen per agent — если агент видит 50 задач в заявленном домене и бидит на 5, это видно и обсуждаемо
+
Difficulty weight в репутационном кортеже — успех на лёгкой задаче даёт меньше rep, чем на сложной
+
+
+
6.3 Bootstrap problem оценки стоимости
+
Симптом: новый агент не знает, сколько стоит задача типа X, потому что никогда её не делал. Предлагает случайный бюджет и либо промахивается (auto-fail), либо завышает (проигрывает аукцион).
+
Лечение:
+
+
Bootstrap exploration credit (FR-015): первые K задач в домене — гарантированные проходы
+
При создании карточки агент может сделать семантический поиск по своей истории задач и взять среднее как стартовую оценку
+
В будущем — «meta-оценщик» агент, специализирующийся на оценке стоимости перед публикацией
+
+
+
6.4 Lemon market
+
Симптом: сложная задача с заниженным бюджетом — никто не бидит, или бидит только самый отчаянный и обречённо проваливает.
+
Лечение:
+
+
Auto-escalation к владельцу через DM если 0 бидов к дедлайну (FR-024)
+
Человек либо поднимает бюджет, либо уточняет задачу, либо делает сам
+
Статистика по каналу: процент задач, ушедших в эскалацию — если > 20%, значит бюджеты в канале систематически занижены
+
+
+
6.5 Runaway token spend
+
Симптом: агент «увяз» в задаче, продолжает делать запрос за запросом, выходит за budget в 3×.
+
Лечение: hard stop на 100% бюджета (FR-023). Задача auto-failed, trace сохранён для аудита. Жёстко, но без этого агенты дрейфуют.
+
+
6.6 Reflection silence
+
Симптом: агент игнорирует reflection events, не обновляет карточку, не учится.
+
Лечение: это не проблема. Обучение опциональное, но аккаунтинг — обязательный. Репутация всё равно записывается автоматически. Агент, который не рефлексирует, просто медленнее растёт — его обгонят те, кто рефлексирует.
+
+
+
+
7. С чего начать
+
Конкретные шаги от текущего состояния (спецификация готова) до рабочего MVP.
+
+
+
Прочитать и утвердить спецификацию. Она в specs/016-agent-marketplace/spec.md. User Stories приоритизированы P1/P2 — MVP = US1 (аукцион) + US2 (карточки).
+
Прогнать /speckit.clarify если остались неясные моменты — это интерактивно задаст уточняющие вопросы и обновит спеку.
+
Прогнать /speckit.plan — сгенерирует план имплементации с разбивкой на этапы, архитектурные решения, выбор технологий (SQLite таблицы, MCP инструменты, reactive triggers).
+
Прогнать /speckit.tasks — превратит план в список конкретных задач для разработки.
+
MVP scope: только US1 + US2. Аукцион + карточки. Без репутации и рефлексии. Минимум, который можно потрогать и на котором можно прогнать один реальный Fermi-estimate через маркетплейс. ~3-5 дней работы.
+
Dogfood на реальных агентах. Переключить existing research-* агентов на публикацию карточек. Попросить их бидить на 5-10 задач. Посмотреть, что сломается.
+
После MVP — добавить US3 (reputation) когда накопится хотя бы 20 завершённых задач и будет data для scoring.
+
После reputation — добавить US4 (reflection). Это самая рискованная часть из-за drift, но без неё маркетплейс статичен.
+
+
+
+ Рекомендация: не пытайтесь построить всё сразу. US1+US2 это уже работающий субстрат. US3 и US4 — надстройки, которые имеет смысл добавлять только когда базовый цикл устаканился и есть реальная статистика.
+
A landscape of the open-source frameworks that coordinate AI agents at scale, the problem classes where they earn their keep, and concrete toy benchmarks to stress-test any multi-agent system.
+
+ Compiled 2026-04-10·Sources: live web research + SynapBus data·Audience: infra engineers
+
+
+
+
+
+
+
+
+
1. Coordination patterns & self-organization
+
Frameworks come and go; coordination patterns are eternal. The real question for a substrate like SynapBus isn't which framework — it's which primitives must the substrate expose so agents can self-organize without being told how. The last two years of research converge on a clear answer: give frontier-capable agents the minimum scaffolding they need, and they out-perform hand-designed hierarchies.
+
+
1.1 The rigid → emergent spectrum
+
Real 2025–2026 systems cluster along a spectrum, not at either pole. At the rigid end: LangGraph DAGs, CrewAI hierarchies, AutoGen supervisor patterns — fixed roles, predetermined edges, a central orchestrator as bottleneck. At the emergent end: OASIS million-agent simulations, digital-pheromone pressure fields, pure swarms where global behavior falls out of local rules.
+
+
+ The most important finding of the last year is the endogeneity paradox: neither maximal control nor maximal autonomy wins. A hybrid "Sequential" protocol providing only fixed turn-ordering as scaffolding but allowing fully endogenous role specialization beats centralized coordination by 14% (p<0.001). Given just turn ordering, 8 agents spontaneously invented 5,006 unique roles, voluntarily abstained from tasks outside their competence, and formed shallow hierarchies. No quality degradation up to 256 agents.
+ Dochkina (2025), "Drop the Hierarchy and Roles", 25,000 tasks across 8 models
+
+
+
There's a capability threshold: frontier models self-organize well; weaker models still benefit from rigid structure. The tradeoff matrix:
+
+
+
+
Dimension
Rigid wins
Emergent wins
+
+
+
Debuggability
strong
weak (non-reproducible)
+
Predictability / SLA
strong
weak
+
Adaptability to novel goals
weak
strong
+
Scaling to many agents
bottlenecks
graceful
+
Cost / tokens
high orchestration overhead
lower (agents self-trim)
+
Failure modes
cascading role-failure
silent stalemate, drift
+
Compliance / audit
easy
hard
+
+
+
+
1.2 Classical self-organization patterns
+
Six patterns from the pre-LLM era that all have 2025 operationalizations for language agents.
+
+
+
Blackboard systems
+
Hearsay-II · Erman & Lesser 1975 · Corkill 1991
+
A shared, structured knowledge store. Independent "knowledge sources" watch it, and when their precondition pattern matches the current state, they fire and contribute. No central scheduler chooses who speaks next — the blackboard's current contents do.
+
LLM mapping: Two 2025 papers reimplement exactly this pattern for LLMs. Agents "volunteer" when the current state matches their capability.
Agents don't talk to each other; they modify the environment, and others react to the modified environment. Ants leave pheromone trails; termites deposit pellets whose shape triggers the next placement. Indirect, asynchronous, tolerant of agent death.
+
LLM mapping: CodeCRDT treats a code artifact as the shared environment — agents read "quality pressure" from the artifact and act to reduce badness.
+
+ 600-trial evaluation: up to 21.1% speedup when task locality holds, up to 39.4% slowdown when coupling is high. The critical heuristic: stigmergy wins when locality holds; explicit coordination wins when work is tightly coupled.
+ CodeCRDT: Observation-Driven Coordination for Multi-Agent LLM Code Generation (Oct 2025)
+
+
+
+
+
Contract Net Protocol
+
Reid Smith · IEEE Transactions on Computers 1980
+
A manager broadcasts a task; capable contractors bid; the manager awards. No central allocator knows who can do what in advance. The original distributed task-allocation protocol, still unbeaten for ambiguous decomposition.
+
LLM mapping: Resource-bounded CNP for LLM agents with structured bidding. Reports 90% token reduction and 525× lower variance vs orchestrated workflows on matched tasks.
+
+
+
+
Flocking / Boids
+
Craig Reynolds · SIGGRAPH 1987
+
Three local rules — separation, alignment, cohesion — produce global flocking. No leader, no plan, no global state. The canonical demonstration that complex group behavior emerges from simple local interactions.
+
LLM mapping: Underexplored. Prompt-level "look at what peers are doing, stay close but not too close" patterns map directly. Relevant to reactive-agent triggers that fan out and converge without central direction.
+
+
+
+
Gossip / epidemic protocols
+
Demers et al. · PODC 1987
+
Each node periodically shares state with a random peer; information spreads epidemically. No routing, no topology maintenance, graceful under node churn. The coordination substrate of real-world distributed databases (Cassandra, DynamoDB).
+
LLM mapping: Gossip is argued as the missing layer for context-rich adaptive agent communication — as opposed to structured protocols that only do reliable task delegation.
Persistent, decaying environmental marks. Three key properties: decay (stale info fades), reinforcement (reuse strengthens), locality (nearby beats far).
+
LLM mapping: Pressure-field coordination — agents coordinate via a shared vector field with temporal decay, no conversation at all. The Ripple Effect Protocol adds sensitivity signals — agents share not just decisions but how decisions would change if environment shifted.
+
+ 4× solve rate vs conversation-based and 30× vs hierarchical on scheduling tasks. Ripple Effect sensitivity signals report 41–100% improvement over A2A communication.
+ Ripple Effect Protocol (MIT 2025) · PooL: Pheromone-inspired MARL
+
Instead of hand-assigning "researcher" and "reviewer", agents negotiate roles via inception prompts, personas, and metacognition. Frontier models produce stable role differentiation from minimal scaffolding.
Layered architecture: proposers → aggregators → more proposers → final aggregator. Each layer sees all previous outputs. Identifies "the collaborativeness of LLMs" — models improve when shown peer outputs, even from weaker peers.
+
+ 65.1% on AlpacaEval 2.0 using only open-source models vs 57.5% for GPT-4o. A stack of smaller models self-organized into layers beats a single frontier model.
+ Mixture-of-Agents Enhances Large Language Model Capabilities (Wang et al., 2024)
+
+
+
+
+
Multi-agent debate (with skepticism)
+
ChatEval 2024 · Society of Minds 2023
+
Agents with diverse personas debate to reach consensus. Intuitive, but the 2025 literature is increasingly skeptical.
+
+ Multi-agent debate frequently fails to beat a well-prompted single agent, even with more compute. Relaxing the consensus requirement (FREE-MAD) tends to improve quality — a general insight: don't force convergence.
+ ChatEval (ICLR 2024) · Stop Overvaluing Multi-Agent Debate (2025) · FREE-MAD (2025)
+
+
+
+
+
Million-agent social simulations
+
OASIS · Nov 2024
+
Runs up to 1M LLM agents on X/Reddit-shaped environments with 21 action types. Replicates information spreading, polarization, and herd behavior.
AgentRxiv · Sakana AI Scientist · Google AI Co-Scientist
+
Multiple parallel labs share a preprint server; each lab reads and builds on others. This is stigmergy applied to research — the shared preprint archive is the coordination substrate.
+ Use rigid structure when: SLAs / compliance / audit are hard requirements; tasks are tightly coupled (CodeCRDT shows locality failing); agents are below the capability threshold; failure must be reproducible; retries are expensive.
+
+
+ Use emergent when: task is exploratory or novel; goal space is open; agents are frontier-capable (Claude Opus / GPT-4-class and up); task decomposes locally; diversity of approach is itself valuable; population is large enough that drop-outs don't break the system.
+
+
+ Hybrid recipe (recommended default for SynapBus): provide a blackboard-like substrate (channels, threads, wiki, shared mutable memory). Add minimal scaffolding (turn ordering, claim-process-done, priority numbers). Expose environmental signals (reactions, decay, reply counts, workflow state). Let roles be endogenous — negotiate via personas, don't hard-code. Keep a human-owner trace for every action.
+
+
+ Give them a blackboard and the minimum ordering they need, then get out of the way.
+ The SynapBus design slogan
+
+
+
1.5 Primitive catalog — what a self-organizing substrate must expose
+
Twelve coordination primitives that a messaging-hub-style system needs to enable self-organization without enforcing structure. Annotated with what SynapBus already has vs what's missing.
+
+
+
Broadcast channels✓ havePattern-match-and-fire substrate — the classic blackboard.
+
Claim/process/done lifecycle✓ haveContract-net without explicit bidding. Already in SynapBus.
+
Threaded replies✓ haveConversational locality — debate without full broadcast.
Foundational pointers: Reynolds, Flocks, Herds, and Schools (SIGGRAPH 1987); Reid Smith, The Contract Net Protocol (IEEE TC 1980); Dorigo, Ant Colony Optimization (MIT Press 2004); Corkill, Blackboard Systems (AI Expert 1991); Minsky, The Society of Mind (1986).
+
+
+
+
+
+
2. Open-source frameworks for multi-agent orchestration
+
Frameworks are one way to realize the patterns in §1 — not the only way. Treat this section as a reference catalog: each entry is an opinionated bundle of the primitives above. The field consolidated sharply in 2025–2026: AutoGen + Semantic Kernel became Microsoft Agent Framework, OpenAI Swarm became the Agents SDK, Phidata became Agno.
Browser-based autonomous goal executor. Development stopped Nov 2023; Reworkd pivoted July 2024. Listed only because it's still widely referenced.
+
+
130+ untriaged open issues
+
No roadmap
+
+
+
+
+
+
+
+
+
3. 2026 consolidation signals
+
Three trends define the current landscape.
+
+
+ Mergers and rebrands. AutoGen + Semantic Kernel → Microsoft Agent Framework. OpenAI Swarm → Agents SDK. Phidata → Agno. Treat old names as pointers. Pinning old versions in production is now a liability.
+
+
+
+ TypeScript tier is real. Mastra, VoltAgent, Claude Agent SDK TS, OpenAI Agents SDK TS, ADK-JS, Strands TS — the JS/TS agent ecosystem has caught up in 2025–2026. No longer Python-only.
+
+
+
+ Interop standards are crystallizing. Two protocols are becoming table-stakes: MCP (tools/context — Anthropic) and A2A (agent-to-agent wire protocol — Google/AWS/Strands). Frameworks without either are drifting toward irrelevance.
+
+
+
+
+
4. What problems are MAS actually good for?
+
Multi-agent architectures earn their complexity budget when a problem has one or more of: natural decomposition, heterogeneous expertise, concurrency, adversarial structure, or emergent dynamics a monolith cannot represent. Eleven categories where they win clearly.
+
+
+
+
+
1. Role-specialized cognitive workflows
+
Planner / executor / critic splits let each role be tuned, budgeted, and evaluated independently. The explicit critic pass demonstrably reduces hallucination.
When a question decomposes into independent sub-queries, parallel workers cut wall-clock latency near-linearly and expand search coverage. This is the dominant production use case in 2025–2026.
3. Distributed problem solving over partitioned state
+
When data, sensors, or constraints are physically or logically partitioned so no single agent sees the whole state, MAS is required by construction. DCOP / DCSP formalize this.
Agent-based modeling is the canonical method for emergent phenomena where heterogeneous individuals drive aggregate outcomes equation-based models cannot capture.
+
e.g. Generative Agents (Park et al., UIST 2023); Imperial College COVID-19 ABM.
+
+
+
+
5. Robotic swarms & multi-robot coordination
+
Physical robots with local sensing require decentralized coordination because bandwidth, latency, and fault-tolerance forbid a central controller.
Markets are definitionally multi-agent: prices emerge from interaction of self-interested participants with private information. MAS is the only faithful way to model them.
+
e.g. Contract Net Protocol (Smith 1980); TAC supply-chain game; NegotiationArena (Bianchi et al., 2024).
+
+
+
+
7. Adversarial testing (red team / blue team)
+
Safety and robustness evaluation benefits from an attacker agent vs. a defender in a closed loop — each improves the other and surfaces failure modes a static test suite cannot.
Real supply chains have multiple autonomous stakeholders with local objectives and private data. Centralized optimization is politically and computationally infeasible.
+
e.g. MASCOT; Fox, Barbuceanu & Teigen, IEEE IS 2000.
+
+
+
+
9. Heterogeneous ETL / pipeline orchestration
+
Data workflows where each stage needs different tools (SQL, vision, code, summarization) map naturally onto a DAG of specialist agents.
+
e.g. LangGraph plan-and-execute; DSPy-compiled document pipelines.
+
+
+
+
10. Creative collaboration with critique
+
Long-form creative work benefits from explicit separation of generation and critique — a single model conflates roles and loses the editor's adversarial stance.
Event-driven systems scale better as independent loops subscribed to different signal types than as one giant polling agent. Exactly SynapBus's reactive-trigger model.
+
e.g. Splunk SOAR; on-call triage bots; SynapBus spec 014.
+
+
+
+
+
+
+
5. Toy problems to benchmark a MAS orchestration system
+
Good MAS benchmarks expose specific failure modes: deadlock, message-order dependence, role confusion, context blowup, byzantine agents, lost-update races. These eight are cheap to implement on top of SynapBus channels / DMs and each probes a distinct axis.
6-block tower world. Two agents each control one "arm" and see only half the blocks. Cooperate to reach a goal configuration in ≤ N moves.
+
Why it bites
Classic AI planning with well-defined optimum. Exposes state-sync bugs, stale-view reads, and whether the framework supports atomic claim/release — a direct analog of SynapBus claim_messages.
6–8 LLM agents with hidden roles, day/night cycles as channels. Private night DMs, public day broadcast. Success = villagers win ≥ baseline rate; agents stay in character.
+
Why it bites
Stresses private vs. public channels, role-scoped memory, and whether the orchestrator prevents info leaks. Published baseline: Werewolf Arena (DeepMind, arXiv:2407.13943).
N agents, round-robin pairings, 200 rounds each, fixed payoff matrix. Success = stable ranking across re-runs with fixed seeds.
+
Why it bites
Tiny payloads but huge volume. Exposes per-message overhead, trace scaling, and whether the framework can support deterministic replay for debugging. Axelrod's classic is the reference.
+
+
+
+
+
+
04
+
+
Collaborative story writing with critic veto
+
tests: role specialization · revision loops · stopping conditions
+
+
Setup
Writer + Editor + Fact-Checker + Critic produce a 2000-word story. Critic has veto; loop until approved or budget exhausted.
+
Why it bites
Tests unbounded loops and budget enforcement. Failure mode: infinite revision — a real production hazard for any "agent with veto" pattern.
+
+
+
+
+
+
05
+
+
Recursive Fermi estimate ("piano tuners in Chicago")
+ The classic question from physicist Enrico Fermi. You cannot look it up; you must decompose into estimable sub-quantities, estimate each, and multiply. This is a sharp MAS benchmark because the decomposition itself is the work — and decomposition is exactly what agent orchestration should be good at.
+
+
+
Canonical decomposition
+
tuners = (pianos in Chicago)
+ × (tunings per piano per year)
+ ÷ (tunings per tuner per year)
+
+pianos = population × (households per capita)
+ × (piano ownership rate)
+ + commercial pianos (venues, schools, churches)
+
Ground truth ≈ 125–250 professional tuners in the Chicago metro area. Success criterion: final estimate within one order of magnitude.
+
+
Why it's a sharp MAS test — five failure axes
+
+
Dynamic spawning
+
The orchestrator doesn't know up front how many sub-researchers it needs. Decomposition dictates fan-out. Tests whether the framework supports runtime subagent creation, not a predefined graph.
+
Duplicate-work detection
+
Naive orchestrators spawn two sub-agents that both research "Chicago population" independently. A well-designed system deduplicates via a shared scratchpad (blackboard!) or caches sub-results. Direct test of whether shared-memory patterns actually work.
+
Cost accounting
+
Fermi estimates are supposed to be cheap. If each sub-agent burns 50k tokens researching census data, you've failed the spirit of the task. Pairs accuracy with a hard token budget.
+
Aggregation under uncertainty
+
Each leaf estimate has an uncertainty range. A mature MAS propagates ranges, not point estimates. Tests whether agents can handle structured sub-results instead of concatenating strings.
+
Orphaned spawns
+
If a sub-agent fails or times out, does the orchestrator notice, retry, or silently drop the branch? A 5-branch Fermi estimate with one silent drop produces a confident wrong answer — the worst possible failure mode.
+
+
+
Concrete setup for SynapBus
+
+
One fermi-orchestrator agent with authority to spawn ≤ 5 fermi-researcher subagents via reactive triggers.
+
Shared channel #fermi-scratchpad as the blackboard — all sub-results posted here with a fixed schema.
+
Each sub-agent has a WebSearch tool with per-call token limit.
+
Orchestrator monitors scratchpad via list_by_state, marks the task done when all leaves posted.
Final estimate within 1 OOM of ground truth (≥ 25 and ≤ 2,500 tuners)
+
Total tokens ≤ 10k across orchestrator + all sub-agents
+
Zero duplicate sub-queries (verified by grepping quantity field)
+
All spawned branches reach terminal state (done or failed, never orphan)
+
Aggregation uses ranges, not point estimates; final answer includes a low/high bound
+
+
+
Failure modes this catches that nothing else does
+
+
Cost explosion from unbounded recursion (sub-agents spawning sub-sub-agents)
+
Silent branch loss — a sub-agent returns nothing and the orchestrator averages 0 into the product
+
Over-convergence — all sub-agents copying each other's bad assumption because they read the scratchpad before contributing (a subtle failure mode of premature information sharing)
+
Point-estimate collapse — losing uncertainty bands so the final answer looks precise when it isn't
+
+
+
+ This is basically the Anthropic multi-agent researcher pattern reduced to a 10-minute test you can run deterministically. If your framework can't pass Fermi, it can't do Deep Research.
+
50 synthetic bug reports (frontend/backend/infra mix). 3 specialist agents pull from a queue, classify, fix or escalate. Success = all bugs terminally resolved, zero double-processing, ≥ 90% correct routing.
+
Why it bites
Almost a unit test for SynapBus's claim_messages / mark_done / StalemateWorker pattern. Exposes lease-expiry bugs and whether failed messages re-queue cleanly.
+
+
+
+
+
+
08
+
+
Overcooked-style real-time coordination
+
tests: tight temporal coordination · latency sensitivity · implicit communication
+
+
Setup
2 agents in a simplified kitchen grid must prepare N soups under a time limit. Reference env: Carroll et al., NeurIPS 2019 (arXiv:1910.05789).
+
Why it bites
The only benchmark here that punishes message latency directly. If the framework adds 500ms per hop, the score shows it. Exposes whether turn-based dialogue assumptions break under real-time load.
+
+
+
+
+
+
+
+
+
6. Recommendation for SynapBus
+
If you only implement three benchmarks, implement these.
+
+ Tier A — infrastructure correctness:Parallel bug triage (#7) and Blocks World with partitioned arms (#1). These directly exercise claim/release semantics, lease expiry, and shared-state conflict — the hardest parts to get right in a message bus.
+
+
+ Tier B — realistic LLM workload:Recursive Fermi estimate (#5) and Collaborative story writing (#4). Stress the patterns your actual users run (fan-out research, revision loops with critics).
+
+
+ Tier C — all-round smoke test:Werewolf (#2). Exercises channels, DMs, role-scoped memory, and workflow reactions simultaneously. The best single integration test for a Slack-like agent hub.
+
+
+ For reference, SynapBus's existing wiki contains four closely-related articles worth consulting before starting: agent-messaging-patterns, synapbus-architecture, ai-agent-governance, mcp-adoption-enterprise.
+
+
+
+
+
+
+
+
+
diff --git a/specs/016-agent-marketplace/checklists/requirements.md b/specs/016-agent-marketplace/checklists/requirements.md
new file mode 100644
index 0000000..7cf16e4
--- /dev/null
+++ b/specs/016-agent-marketplace/checklists/requirements.md
@@ -0,0 +1,37 @@
+# Specification Quality Checklist: Self-Organizing Agent Marketplace
+
+**Purpose**: Validate specification completeness and quality before proceeding to planning
+**Created**: 2026-04-11
+**Feature**: [spec.md](../spec.md)
+
+## Content Quality
+
+- [x] No implementation details (languages, frameworks, APIs)
+- [x] Focused on user value and business needs
+- [x] Written for non-technical stakeholders
+- [x] All mandatory sections completed
+
+## Requirement Completeness
+
+- [x] No [NEEDS CLARIFICATION] markers remain
+- [x] Requirements are testable and unambiguous
+- [x] Success criteria are measurable
+- [x] Success criteria are technology-agnostic (no implementation details)
+- [x] All acceptance scenarios are defined
+- [x] Edge cases are identified
+- [x] Scope is clearly bounded
+- [x] Dependencies and assumptions identified
+
+## Feature Readiness
+
+- [x] All functional requirements have clear acceptance criteria
+- [x] User scenarios cover primary flows
+- [x] Feature meets measurable outcomes defined in Success Criteria
+- [x] No implementation details leak into specification
+
+## Notes
+
+- Validation performed after initial draft. All items pass on first iteration.
+- Reputation scoring algorithm details (exact formula for combining estimated/actual cost with success and difficulty) are intentionally deferred to planning — the spec requires that reputation be domain-scoped, vectorized, and feed into bid comparison, but does not mandate a specific formula.
+- Voting quorum for multi-owner awards is explicitly out of scope; awards are made by the single task poster by default.
+- Adversarial defenses beyond exploration budget and bid-ratio visibility are explicitly out of scope — this is a trust substrate for cooperating agents, not a byzantine-agent environment.
diff --git a/specs/016-agent-marketplace/spec.md b/specs/016-agent-marketplace/spec.md
new file mode 100644
index 0000000..a57c586
--- /dev/null
+++ b/specs/016-agent-marketplace/spec.md
@@ -0,0 +1,190 @@
+# Feature Specification: Self-Organizing Agent Marketplace
+
+**Feature Branch**: `016-agent-marketplace`
+**Created**: 2026-04-11
+**Status**: Draft
+**Input**: User description: Self-organizing agent marketplace on SynapBus. Four new primitives — capability manifest, auction channel, domain-scoped reputation ledger, reflection loop — that let LLM agents decompose, allocate, and learn from tasks autonomously without a central orchestrator.
+
+## User Scenarios & Testing *(mandatory)*
+
+### User Story 1 — Post a task, let agents self-allocate (Priority: P1)
+
+A human owner (or another agent) needs a piece of work done but does not know in advance which agent is best suited for it. They post the task to an auction channel with acceptance criteria, a maximum token budget, and a deadline. Available agents evaluate whether the task matches their declared capabilities, submit structured bids, and one is awarded. The awarded agent executes the task within the budget. The owner never had to choose who does the work.
+
+**Why this priority**: This is the minimum viable slice of the marketplace. Without it, no other primitive has meaning — the capability manifest is just metadata, the reputation ledger has nothing to record, the reflection loop has nothing to reflect on. Awarding tasks through open bidding is the irreducible core of a self-organizing agent system.
+
+**Independent Test**: A human posts a single auction task to an auction channel with at least one qualified agent joined. The agent reads the task, submits a bid containing an estimated token cost and a short approach summary, the human awards the bid via a reaction, the agent executes the task and marks it done within the declared budget. The full lifecycle (post → bid → award → claim → done) succeeds without any other primitive being active.
+
+**Acceptance Scenarios**:
+
+1. **Given** an auction channel exists with two agents joined, **When** a human posts a task with a clear description and max budget, **Then** both agents can see the task and submit bids as threaded replies.
+2. **Given** two bids have been submitted on an auction task, **When** the human awards one bid via the `awarded` reaction, **Then** the winning agent receives a claim on the task and the losing bid is marked with a no-op reaction so it is not orphaned.
+3. **Given** an auction task with a max budget of 5,000 tokens has been awarded, **When** the winning agent executes and marks the task done within 4,200 tokens, **Then** the lifecycle completes successfully and actual token usage is recorded.
+4. **Given** an auction task has been posted, **When** no agent submits a bid before the declared deadline, **Then** the task auto-escalates to the human owner via direct message and remains in an open state for manual handling.
+
+---
+
+### User Story 2 — Agents advertise what they can do (Priority: P1)
+
+Before agents can meaningfully bid on tasks, the marketplace needs to know what each agent is capable of. Every agent publishes a capability manifest — a persistent, versioned description of its domains, representative example tasks, a self-reported confidence level, and an average token cost. Manifests are discoverable by other agents and by humans, and every update is preserved as a revision so any change is auditable and reversible.
+
+**Why this priority**: Without a capability manifest, bidding is uninformed and reputation cannot be scoped to domains. This is the substrate that lets agents decide which auctions to bid on. It is P1 because User Story 1 is only meaningfully testable once agents have a place to declare what they do.
+
+**Independent Test**: An agent publishes an initial capability manifest listing two domains and one example task. A human reads the manifest via the marketplace. The agent then updates the manifest to add a third domain. The previous version is retained as a historical revision and can be retrieved without data loss.
+
+**Acceptance Scenarios**:
+
+1. **Given** a new agent joins the marketplace, **When** the agent publishes its initial capability manifest, **Then** the manifest is stored, discoverable by other participants, and assigned a version identifier.
+2. **Given** an agent has an existing capability manifest, **When** the agent updates it with a new domain, **Then** the update is stored as a new revision and the prior revision remains accessible.
+3. **Given** an agent's manifest has been updated several times, **When** a human requests the revision history, **Then** the full ordered list of past versions is returned and any prior version can be restored.
+
+---
+
+### User Story 3 — Reputation shapes future awards (Priority: P2)
+
+As tasks complete, the marketplace records per-domain tuples of estimated cost, actual cost, success score, and difficulty weight against the executing agent. When bids are compared on future tasks, reputation in the relevant domains informs the scoring — but reputation is a vector across domains, not a single number, so an agent that is excellent in one domain and weak in another cannot hide behind a generic score. An exploration budget forces a configurable fraction of awards to go to lower-reputation bidders so newcomers can enter and lock-in is avoided.
+
+**Why this priority**: Reputation is what turns a one-shot auction into a learning marketplace. Without it, every task is evaluated on promises only. With it, actual performance accumulates and informs decisions. It is P2 because User Stories 1 and 2 must exist first to generate the data reputation depends on.
+
+**Independent Test**: Two agents each complete three tasks in the same domain with different success rates. A new task is posted in that domain and both agents bid. Reputation scoring is applied, the higher-reputation agent is preferred, but under the exploration budget a configurable fraction of awards still go to the lower-reputation agent. Over time, reputation differences converge to actual performance differences.
+
+**Acceptance Scenarios**:
+
+1. **Given** an agent has completed tasks in domain X with measured success, **When** a new task in domain X is posted and the agent bids, **Then** the reputation score for (agent, domain X) is available and factors into bid comparison.
+2. **Given** two agents have very different reputations in a domain, **When** 100 tasks in that domain are awarded with exploration budget set to 10%, **Then** approximately 90 tasks go to the higher-reputation agent and approximately 10 tasks go to the lower-reputation agent.
+3. **Given** a brand new agent with no reputation in any domain, **When** it bids on its first task in a new domain, **Then** it receives a bootstrap exploration credit guaranteeing forced participation in a configurable number of initial tasks per domain so it can build a track record.
+4. **Given** an agent has a high reputation in domain A and a low reputation in domain B, **When** a task in domain B is scored, **Then** only the domain B reputation is used; the domain A reputation does not influence the comparison.
+
+---
+
+### User Story 4 — Agents learn from completed tasks (Priority: P2)
+
+When a task is marked done, the system delivers a reflection event to the executing agent containing the original task, the winning bid, the execution trace, and any feedback (success or failure). The agent spends reasoning steps reviewing this material and produces a proposed diff against its own capability manifest — perhaps adding a newly discovered domain, raising its confidence, revising its average cost estimate, or adding an example. Proposed diffs are never auto-applied. They are submitted as revision proposals that require human approval (or a configurable auto-approve rule) to merge. Every proposal and every merge is preserved so drift is auditable and reversible.
+
+**Why this priority**: Reflection turns reputation from a passive record into an active learning signal. It lets agents get better over time. It is P2 because the full loop only has meaning once tasks are being awarded (P1) and agents have manifests to reflect on (P1).
+
+**Independent Test**: An agent completes a task where its estimated cost was significantly lower than the actual cost. A reflection event fires. The agent produces a proposed diff raising its average cost for that domain. The diff is submitted as a manifest revision proposal. A human reviews and approves it. The agent's manifest now reflects the learning, and the approval is auditable.
+
+**Acceptance Scenarios**:
+
+1. **Given** an agent has just marked a task done, **When** the reflection event fires, **Then** the agent receives the original task, its bid, the execution trace, and any feedback as input to a reflection prompt.
+2. **Given** an agent has produced a proposed manifest diff after reflection, **When** the diff is submitted, **Then** it appears as a pending revision proposal visible to the human owner and is not applied to the live manifest.
+3. **Given** a pending manifest revision proposal, **When** the human owner approves it, **Then** the diff is merged into the manifest as a new revision and the approval is recorded.
+4. **Given** a pending manifest revision proposal, **When** the human owner rejects it, **Then** the live manifest remains unchanged and the rejected proposal is preserved in history for future audit.
+
+---
+
+### Edge Cases
+
+- **No qualified bidders**: A task requires domains that no agent has declared in its manifest. The task reaches its deadline with zero bids and auto-escalates to the human owner.
+- **Runaway token spend**: An awarded agent approaches its declared max budget. At 80 percent of budget the agent receives a soft warning; at 100 percent execution is hard-stopped and the task is marked as auto-failed with the partial trace preserved.
+- **Bid on unfamiliar domain**: An agent bids on a task in a domain it has no reputation in. The bid is accepted under the bootstrap exploration credit for its first configurable-K tasks in that domain; after that, absence of domain reputation penalizes the bid in scoring.
+- **Drift after approved diffs**: A sequence of individually reasonable manifest diffs accumulates into a manifest that no longer reflects the agent owner's intent. The human owner can review the full revision history and roll back to any prior version with a single action.
+- **Gaming via selective bidding**: An agent bids only on easy tasks to keep its success score high. The exploration budget partially counters this by forcing some awards to lower-reputation bidders; additionally, the system tracks ratio of bids-submitted to qualifying-tasks-seen per agent so persistent refusal to bid on declared-competence tasks is visible.
+- **Reflection loop silence**: An agent ignores the reflection event and produces no diff. The task still completes successfully and the reputation ledger still records the outcome; learning is optional, accountability is not.
+- **Conflicting simultaneous bids**: Two agents submit bids within milliseconds of each other. Both bids are accepted and ordered by arrival timestamp; the award process considers both.
+- **Expired auctions with pending bids**: A task deadline passes after at least one bid was submitted. The task escalates to the human owner with the existing bids preserved for manual decision.
+- **Self-bidding**: An agent attempts to bid on its own posted task. This is rejected — agents cannot both post and execute the same task.
+
+## Requirements *(mandatory)*
+
+### Functional Requirements
+
+**Capability Manifest**
+
+- **FR-001**: The system MUST allow every participating agent to publish a persistent capability manifest listing at minimum its domain tags, example tasks, self-reported confidence per domain, and average token cost per domain.
+- **FR-002**: The system MUST retain every historical revision of every capability manifest so that any change is auditable and any prior version is retrievable.
+- **FR-003**: Agents MUST be able to update their own capability manifest at any time; agents MUST NOT be able to modify another agent's manifest.
+- **FR-004**: The system MUST make capability manifests discoverable by other agents and by human owners so that bidders can inspect one another and task posters can verify qualification.
+
+**Auction Channel**
+
+- **FR-005**: The system MUST provide a channel type dedicated to task auctions where each parent message represents one task.
+- **FR-006**: An auction task MUST declare, at minimum, a description, acceptance criteria, a maximum token budget, a deadline, and required domain tags.
+- **FR-007**: Agents MUST submit bids as in-thread replies to the task message, with each bid containing an estimated token cost, a confidence level, a brief approach summary, and the revision of the bidder's capability manifest at time of bid.
+- **FR-008**: The task poster (human owner or, where configured, a voting quorum of trusted participants) MUST be able to award a task to a single bid via a designated reaction.
+- **FR-009**: On award, the system MUST convert the auction into a claim on the winning agent using the existing claim/process/done lifecycle.
+- **FR-010**: Losing bids on an awarded auction MUST receive a terminal no-op signal so they are not left orphaned.
+- **FR-011**: The system MUST prevent an agent from bidding on a task that the same agent posted.
+
+**Reputation Ledger**
+
+- **FR-012**: The system MUST record, for every completed auction task, a tuple containing the executing agent, the domain, the estimated token cost, the actual token cost, a success score, a difficulty weight, and the completion timestamp.
+- **FR-013**: Reputation MUST be queryable and scoped by (agent, domain). A single global reputation score MUST NOT be exposed or used.
+- **FR-014**: The system MUST support a configurable per-channel exploration budget expressed as a percentage of awards that are forced to go to lower-reputation bidders.
+- **FR-015**: The system MUST grant new agents with no reputation in a given domain a bootstrap exploration credit for a configurable number of initial tasks in that domain so they can establish a track record.
+
+**Reflection Loop**
+
+- **FR-016**: When a task is marked done or failed, the system MUST emit a reflection event to the executing agent containing the original task, the bid, the execution trace, and any feedback.
+- **FR-017**: Agents MUST be able to submit a proposed diff against their own capability manifest as a result of reflection; diffs MUST NOT be auto-applied to the live manifest.
+- **FR-018**: A manifest diff proposal MUST require explicit approval before merging — by the human owner by default, or by a configurable auto-approve rule at channel or agent level.
+- **FR-019**: The system MUST preserve every proposed diff and every approval or rejection decision so drift over time is auditable and reversible.
+- **FR-020**: The human owner MUST be able to roll back a capability manifest to any prior revision regardless of how many intermediate changes have been applied.
+- **FR-020a**: Each manifest revision MUST carry provenance metadata including the rolling success count and failure count of the agent in the affected domains at time of revision, so a reader can assess whether a given revision correlated with performance regression.
+- **FR-020b**: The system MUST support automatic tombstoning of a manifest revision (marking it deprecated but retaining it for audit) when the owning agent's rolling failure rate in the revision's declared domains exceeds a configurable threshold over a configurable sliding window.
+
+**Budget Enforcement**
+
+- **FR-021**: The system MUST track cumulative token spend against the declared max budget for every awarded task in real time.
+- **FR-022**: The system MUST emit a soft warning to the executing agent when cumulative spend reaches 80 percent of the declared budget.
+- **FR-023**: The system MUST hard-stop execution and auto-fail the task when cumulative spend reaches 100 percent of the declared budget, preserving the partial execution trace for later inspection.
+
+**Lemon Market Mitigation**
+
+- **FR-024**: If an auction task receives zero bids by its declared deadline, the system MUST auto-escalate the task to the human owner via direct message and keep the task open for manual handling.
+- **FR-025**: Auto-escalation behavior MUST be configurable per channel (enabled, disabled, or custom escalation target).
+
+**Observability**
+
+- **FR-026**: Every auction post, bid, award, claim, reflection event, proposed manifest diff, and merge decision MUST be traced in the existing SynapBus trace infrastructure and be searchable by time range, agent, and task.
+- **FR-027**: The system MUST allow a human owner to answer the question "who did what and why" for any completed task by inspecting the trace without additional tooling.
+
+### Key Entities
+
+- **Capability Manifest**: A per-agent, versioned document describing domain expertise, representative example tasks, self-reported confidence per domain, and average token cost per domain. Owned by the agent it describes. Related to: agent identity, manifest revisions, reflection diffs.
+- **Manifest Revision**: An immutable snapshot of a capability manifest at a point in time, produced by an update. Enables audit and rollback. Related to: capability manifest, reflection diff, approval record.
+- **Auction Task**: A posted task with description, acceptance criteria, max token budget, deadline, and required domains. Has a lifecycle: open → (bids collected) → awarded → claimed → done/failed, or open → deadline → escalated. Related to: bids, claim record, reputation entries.
+- **Bid**: A structured reply to an auction task containing estimated token cost, confidence, approach summary, and the revision of the bidder's manifest at time of bid. Related to: auction task, bidder agent, manifest revision.
+- **Reputation Entry**: A completed-task tuple keyed by (agent, domain) with estimated cost, actual cost, success score, difficulty weight, and timestamp. Related to: agent, auction task.
+- **Reflection Event**: A system-generated event fired when a task terminates, delivered to the executing agent and containing the full task context and outcome. Related to: auction task, executing agent, proposed manifest diff.
+- **Manifest Diff Proposal**: A proposed change to a capability manifest produced by reflection. Pending until explicitly approved or rejected. Related to: reflection event, manifest revision, approval record.
+- **Approval Record**: An immutable record of the decision (approve or reject) on a manifest diff proposal, with timestamp and the identity of the decider. Related to: manifest diff proposal.
+
+## Success Criteria *(mandatory)*
+
+### Measurable Outcomes
+
+- **SC-001**: A human owner can post a task, receive bids, award one, and see the task complete without intervening in agent selection — verified by running a full auction lifecycle with at least two qualified agents bidding, where the owner's only actions are posting the task and issuing the award reaction.
+- **SC-002**: Over a test run of at least 50 tasks with two agents of unequal skill, the higher-skilled agent wins approximately (1 minus exploration budget) × 100 percent of awards in the relevant domain. For a 10 percent exploration budget this means the higher-skilled agent should win within ±5 percentage points of 90 percent of the in-domain tasks.
+- **SC-003**: A brand new agent joining the marketplace with no prior reputation can bid on and be awarded tasks in a previously unseen domain within its first 10 bids, via the bootstrap exploration credit.
+- **SC-004**: For every completed auction task, a human owner can retrieve the full decision trail (who bid, at what estimated cost, who was awarded, why, actual cost, final outcome, any reflection diff) in a single query from the trace infrastructure.
+- **SC-005**: After a manifest drift incident (three or more sequential approved diffs that collectively produce an unintended manifest state), the human owner can roll back to the pre-drift revision in a single action and all intermediate history is preserved.
+- **SC-006**: Zero approved manifest diffs are applied without an explicit approval record. This is verified by auditing the approval records against applied diffs over a test run and finding a one-to-one match.
+- **SC-007**: Of awarded tasks, at least 90 percent complete within their declared max token budget without triggering the hard-stop. Hard-stops exist but should be the exception, not the rule.
+- **SC-008**: Of auction tasks that receive zero bids, 100 percent are escalated to the human owner within one minute of deadline expiry.
+- **SC-009**: A human owner can inspect any agent's current capability manifest and full revision history with no more than two marketplace queries.
+- **SC-010**: At least 95 percent of completed tasks produce a reputation ledger entry with all required fields populated.
+
+## Assumptions
+
+- Agents are already authenticated to SynapBus via the existing identity and API key mechanisms; the marketplace does not introduce new authentication.
+- The existing wiki subsystem is the durable store for capability manifests and their revision history; this feature does not require a new revisioned document store.
+- The existing claim/process/done lifecycle and reactive trigger infrastructure are reused to convert an awarded bid into executable work; no new lifecycle is invented for auctions.
+- The existing trace infrastructure covers MCP tool invocations and reactive trigger runs and can be extended with new event types without structural change.
+- Token counts are reported by the executing agent honestly; adversarial under-reporting is out of scope for this feature and belongs to a later trust-enforcement layer.
+- Success scores for completed tasks are determined by the task poster (human or delegate) at done time; automated grading of task outputs is out of scope.
+- Difficulty weights for reputation scoring are a per-channel constant or a simple function of max-budget magnitude; adaptive difficulty estimation is out of scope for this feature.
+- A reasonable default exploration budget is 10 percent per channel. This is configurable and individual channels may set it higher or lower.
+- A reasonable default bootstrap credit is three tasks per new (agent, domain) pair. This is configurable.
+- Voting quorum for multi-owner awards is out of scope; awards are made by the single task poster by default.
+
+## Out of Scope
+
+- Automatic curriculum generation: selecting or composing what task should be posted next based on skill gaps or strategic goals.
+- Cross-agent skill sharing: one agent teaching another agent a capability or transplanting manifest fragments.
+- Monetary incentives beyond token budgets: real currency, credits, or cross-tenant billing.
+- Multi-owner approval workflows and voting quorums for awards.
+- Adversarial defenses against dishonest token reporting or reputation gaming beyond the exploration budget and bid-ratio visibility.
+- Automated grading of task output quality; humans remain the source of truth for success scores in this release.
+- Agent-management MCP tools (create, delete, configure agents) — excluded by standing design rule.