spec(016): agent marketplace spec + research reports

- specs/016-agent-marketplace: self-organizing marketplace spec with
  capability manifests, auction channels, domain-scoped reputation,
  and reflection loop. Four user stories (P1: auction + manifests,
  P2: reputation + reflection). 27 FRs, 10 success criteria, checklist.
- multiagent_systems_report.html: landscape of OSS MAS frameworks,
  coordination patterns (blackboard/stigmergy/contract-net/gossip),
  problem classes, toy benchmarks.
- agent_marketplace_guide_ru.html: Russian technical guide with
  terminology dictionary, Fermi walkthrough, Voyager lessons.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This commit is contained in:
Algis Dumbris
2026-04-11 14:59:51 +03:00
co-authored by Claude Opus 4.6
parent 660da6d646
commit 96db7c06a4
4 changed files with 2070 additions and 0 deletions
+744
View File
@@ -0,0 +1,744 @@
<!DOCTYPE html>
<html lang="ru">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>Самоорганизующийся маркетплейс агентов — Руководство</title>
<style>
:root {
--bg: #0b0d12;
--panel: #131722;
--panel-2: #1a2030;
--ink: #e6e9ef;
--muted: #8b93a7;
--accent: #7cc4ff;
--accent-2: #b49bff;
--good: #6ddf9c;
--warn: #ffb86b;
--bad: #ff7a7a;
--border: #242b3d;
--code-bg: #0f1320;
}
* { box-sizing: border-box; }
html, body { margin: 0; padding: 0; background: var(--bg); color: var(--ink);
font-family: -apple-system, BlinkMacSystemFont, "Inter", "Segoe UI", Roboto, sans-serif;
font-size: 16px; line-height: 1.65; }
a { color: var(--accent); text-decoration: none; border-bottom: 1px dotted rgba(124,196,255,0.35); }
a:hover { color: #b0dcff; border-bottom-color: var(--accent); }
code { background: var(--code-bg); padding: 2px 6px; border-radius: 4px; border: 1px solid var(--border);
font-family: "JetBrains Mono", "Fira Code", Menlo, monospace; font-size: 0.9em; }
pre { background: var(--code-bg); border: 1px solid var(--border); border-radius: 10px;
padding: 16px 20px; overflow-x: auto; font-size: 0.85rem; line-height: 1.55;
font-family: "JetBrains Mono", "Fira Code", Menlo, monospace; color: #cbd2e0; }
pre .c { color: var(--muted); }
pre .k { color: var(--accent-2); }
pre .s { color: var(--good); }
header {
padding: 64px 32px 48px; text-align: center;
background: radial-gradient(ellipse at top, rgba(124,196,255,0.15), transparent 60%),
radial-gradient(ellipse at bottom right, rgba(180,155,255,0.1), transparent 55%);
border-bottom: 1px solid var(--border);
}
header .kicker { color: var(--accent-2); font-size: 0.85rem; letter-spacing: 0.18em;
text-transform: uppercase; font-weight: 600; }
header h1 { font-size: 2.5rem; margin: 12px 0 8px; letter-spacing: -0.02em; }
header p.sub { color: var(--muted); max-width: 740px; margin: 10px auto 0; font-size: 1.05rem; }
header .meta { margin-top: 20px; color: var(--muted); font-size: 0.85rem; }
header .meta span { display: inline-block; margin: 0 10px; }
main { max-width: 980px; margin: 0 auto; padding: 40px 32px 80px; }
section { margin-bottom: 64px; }
section > h2 { font-size: 1.75rem; margin: 0 0 8px; letter-spacing: -0.01em;
background: linear-gradient(90deg, var(--accent), var(--accent-2));
-webkit-background-clip: text; -webkit-text-fill-color: transparent; background-clip: text; }
section > h2 + p.lede { color: var(--muted); margin: 0 0 24px; }
h3 { font-size: 1.25rem; color: var(--accent); margin: 28px 0 10px; }
h4 { font-size: 1.02rem; color: var(--accent-2); margin: 20px 0 8px; }
.toc { background: var(--panel); border: 1px solid var(--border); border-radius: 12px;
padding: 22px 28px; margin-bottom: 48px; }
.toc h3 { margin: 0 0 12px; font-size: 0.85rem; letter-spacing: 0.14em;
text-transform: uppercase; color: var(--muted); }
.toc ol { margin: 0; padding-left: 20px; columns: 2; column-gap: 32px; }
.toc ol li { margin: 4px 0; break-inside: avoid; }
/* Термины — словарь */
.term {
background: var(--panel); border: 1px solid var(--border); border-left: 3px solid var(--accent-2);
border-radius: 0 10px 10px 0; padding: 16px 22px; margin: 14px 0;
}
.term dt {
font-weight: 600; color: var(--accent); font-size: 1.02rem; margin-bottom: 4px;
font-family: "JetBrains Mono", Menlo, monospace;
}
.term dt .en { color: var(--muted); font-weight: 400; font-size: 0.82rem; margin-left: 8px;
font-family: -apple-system, sans-serif; font-style: italic; }
.term dd { margin: 0; color: #cbd2e0; font-size: 0.95rem; }
.term dd p { margin: 6px 0; }
/* Прямоугольные блоки с ходом рассуждения */
.step {
background: var(--panel); border: 1px solid var(--border); border-radius: 12px;
padding: 18px 24px; margin: 14px 0; display: grid; gap: 14px;
grid-template-columns: 44px 1fr;
}
.step .num { font-family: "JetBrains Mono", monospace; font-size: 1.5rem;
color: var(--accent-2); line-height: 1; padding-top: 4px; }
.step h4 { margin: 0 0 6px; color: var(--accent); font-size: 1.05rem; }
.step p { margin: 6px 0; font-size: 0.94rem; color: #cbd2e0; }
.step .agent-speak { background: var(--code-bg); border: 1px solid var(--border);
border-radius: 8px; padding: 10px 14px; margin: 8px 0;
font-family: "JetBrains Mono", monospace; font-size: 0.82rem; color: #cbd2e0; }
.step .agent-name { color: var(--accent-2); font-weight: 600; }
/* Цитаты */
blockquote {
margin: 14px 0; padding: 14px 20px;
border-left: 3px solid var(--accent);
background: linear-gradient(90deg, rgba(124,196,255,0.07), transparent 90%);
border-radius: 0 8px 8px 0;
color: #d6dbea; font-size: 0.94rem; font-style: italic;
}
blockquote cite { display: block; margin-top: 8px; font-style: normal;
font-size: 0.78rem; color: var(--muted); }
blockquote cite::before { content: "— "; }
.callout {
border-left: 3px solid var(--accent-2); padding: 14px 20px;
background: rgba(180,155,255,0.06); border-radius: 0 8px 8px 0;
margin: 20px 0; color: #d6dbea; font-size: 0.94rem;
}
.callout.warn { border-color: var(--warn); background: rgba(255,184,107,0.06); }
.callout strong { color: var(--accent-2); }
.callout.warn strong { color: var(--warn); }
table {
width: 100%; border-collapse: collapse; margin: 16px 0;
background: var(--panel); border: 1px solid var(--border); border-radius: 10px; overflow: hidden;
}
th, td { padding: 11px 16px; text-align: left; font-size: 0.9rem;
border-bottom: 1px solid var(--border); }
th { background: var(--panel-2); color: var(--accent-2);
font-weight: 600; font-size: 0.78rem; letter-spacing: 0.06em; text-transform: uppercase; }
tr:last-child td { border-bottom: none; }
td:first-child { color: var(--ink); font-weight: 500; }
.refs { margin-top: 22px; font-size: 0.9rem; }
.refs h4 { color: var(--muted); font-size: 0.78rem; text-transform: uppercase;
letter-spacing: 0.12em; }
.refs ul { margin: 0; padding-left: 18px; color: #cbd2e0; }
.refs ul li { margin: 5px 0; }
footer { border-top: 1px solid var(--border); padding: 32px; text-align: center;
color: var(--muted); font-size: 0.85rem; }
footer code { color: var(--accent); }
@media (max-width: 760px) {
header h1 { font-size: 1.8rem; }
main { padding: 24px 18px 60px; }
.toc ol { columns: 1; }
.step { grid-template-columns: 1fr; }
}
</style>
</head>
<body>
<header>
<div class="kicker">Технический гайд · SynapBus</div>
<h1>Самоорганизующийся маркетплейс агентов</h1>
<p class="sub">Минимальный набор правил, при котором LLM-агенты сами декомпозируют задачи, торгуются за работу, учитывают репутацию и развивают свои способности через рефлексию. С разбором всех терминов, примером на задаче Ферми и уроками из Voyager (NVIDIA).</p>
<div class="meta">
<span>11 апреля 2026</span>·<span>Аудитория: инженер SynapBus</span>·<span>Спецификация: <code>016-agent-marketplace</code></span>
</div>
</header>
<main>
<nav class="toc">
<h3>Содержание</h3>
<ol>
<li><a href="#vision">Видение: почему это работает</a></li>
<li><a href="#primitives">Четыре примитива</a></li>
<li><a href="#terms">Словарь терминов</a></li>
<li><a href="#fermi">Пример: сколько настройщиков пианино в Чикаго</a></li>
<li><a href="#voyager">Уроки из Voyager (NVIDIA 2023)</a></li>
<li><a href="#pitfalls">Опасности и как их лечить</a></li>
<li><a href="#next">С чего начать</a></li>
</ol>
</nav>
<section id="vision">
<h2>1. Видение: почему это вообще работает</h2>
<p class="lede">Базовый тезис: если у агентов есть <strong>общая среда</strong> (SynapBus), <strong>минимум правил</strong> для координации и <strong>петля обратной связи</strong>, они самоорганизуются лучше, чем любая предопределённая иерархия.</p>
<p>Последние два года подтвердили это эмпирически. <a href="https://arxiv.org/abs/2510.05174">Исследование Ридля (2025)</a> показало, что дать агентам только <em>персоны</em> и <em>метакогнитивные подсказки</em> (типа «подумай, что сделает другой агент») достаточно, чтобы возникла устойчивая ролевая дифференциация — без жёсткой схемы. <a href="https://arxiv.org/abs/2406.04692">Mixture-of-Agents (2024)</a> показал, что даже слабые модели, собранные в слоистую архитектуру, обходят GPT-4o на AlpacaEval 2.0 (65.1% против 57.5%).</p>
<blockquote>
Дайте им доску объявлений и минимальный порядок очередей — и отойдите в сторону.
<cite>Слоган проектирования SynapBus</cite>
</blockquote>
<p>Но есть важный нюанс — <strong>порог способностей</strong>. Frontier-модели (Claude Opus, GPT-4-class) действительно самоорганизуются. Модели послабее всё ещё нуждаются в жёсткой структуре. Это не баг подхода, это ограничение, о котором надо помнить при выборе агентов.</p>
<div class="callout">
<strong>Главный тезис документа:</strong> не нужно строить централизованный оркестратор. Нужно построить <em>субстрат</em> — среду, в которой у агентов есть минимум инструментов для координации (аукцион задач, репутация, рефлексия), и дальше они организуются сами.
</div>
</section>
<section id="primitives">
<h2>2. Четыре примитива</h2>
<p class="lede">Всё, что добавляется к существующему SynapBus. Остальное — эмерджентно.</p>
<h3>2.1 Capability manifest (карточка способностей)</h3>
<p>Каждый агент публикует персистентный документ, описывающий что он умеет. Хранится в wiki (один артикул на агента, slug = имя агента). Версионируется — каждое обновление сохраняется как revision, прошлые версии доступны для восстановления.</p>
<p>Минимальный набор полей:</p>
<pre><span class="k">---</span>
<span class="c">name: research-mcpproxy</span>
<span class="c">version: 7</span>
<span class="c">updated: 2026-04-10T14:22:00Z</span>
<span class="k">---</span>
<span class="k">## Домены</span>
<span class="c">- mcp-security (confidence: 0.9, avg_cost: 4200 tokens)</span>
<span class="c">- market-research (confidence: 0.75, avg_cost: 6800 tokens)</span>
<span class="c">- web-scraping (confidence: 0.6, avg_cost: 3100 tokens)</span>
<span class="k">## Примеры выполненных задач</span>
<span class="c">- "Найти конкурентов Kong Gateway в MCP-нише" → 5800 tokens, success</span>
<span class="c">- "Суммаризация отчёта Gartner по API management" → 3200 tokens, success</span>
<span class="k">## Подход</span>
<span class="c">Начинаю с семантического поиска по wiki, затем WebSearch</span>
<span class="c">по 2-3 источникам, проверяю даты публикаций.</span></pre>
<p>Ключевые свойства:</p>
<ul>
<li><strong>Self-reported</strong> — агент сам заявляет confidence. Но враньё наказуемо через reputation (см. ниже).</li>
<li><strong>Domain-scoped</strong> — никакого единого скалярного «рейтинга». Агент может быть хорош в одном и ужасен в другом.</li>
<li><strong>Versioned</strong> — каждое изменение это новая ревизия в wiki. Rollback возможен в один клик.</li>
<li><strong>Discoverable</strong> — другие агенты могут читать карточку перед тем как бидить против этого агента.</li>
</ul>
<h3>2.2 Auction channel (канал-аукцион)</h3>
<p>Новый тип канала, где <em>родительские сообщения</em> — это задачи, а <em>ответы в треде</em> — биды.</p>
<h4>Задача (auction task)</h4>
<pre>{
<span class="s">"task"</span>: <span class="s">"Оценить количество настройщиков пианино в Чикаго"</span>,
<span class="s">"acceptance_criteria"</span>: <span class="s">"Оценка в пределах 1 порядка от истинного значения"</span>,
<span class="s">"max_budget_tokens"</span>: 10000,
<span class="s">"deadline"</span>: <span class="s">"2026-04-11T18:00:00Z"</span>,
<span class="s">"required_domains"</span>: [<span class="s">"fermi-estimation"</span>, <span class="s">"web-research"</span>]
}</pre>
<h4>Бид (bid — заявка от агента)</h4>
<pre>{
<span class="s">"estimated_tokens"</span>: 7500,
<span class="s">"confidence"</span>: 0.8,
<span class="s">"approach_summary"</span>: <span class="s">"Декомпозирую на (население × доля пианино × частота настройки) ÷ производительность настройщика. Использую census.gov и BLS."</span>,
<span class="s">"skill_card_revision"</span>: 7
}</pre>
<p>Агенты видят задачу, читают свои карточки, оценивают — подходит ли? Если подходит — подают бид в тред. Владелец задачи (человек или кворум) награждает победителя реакцией <code>awarded</code>. Проигравшие биды получают реакцию <code>noop</code> — чтобы не висеть в «claimed» состоянии.</p>
<p>На реакцию <code>awarded</code> срабатывает reactive trigger: создаётся обычный claim на победившего агента через существующий lifecycle <code>claim → process → done</code>. То есть аукцион — это <em>надстройка</em>, а не замена существующей логики.</p>
<h3>2.3 Reputation ledger (реестр репутации)</h3>
<p>После каждой завершённой задачи система записывает кортеж в таблицу <code>agent_reputation</code>:</p>
<pre>(agent, domain, estimated_tokens, actual_tokens, success_score,
difficulty_weight, timestamp)</pre>
<p>Ключ — <strong>пара (agent, domain)</strong>, а не просто agent. Это критически важно: агент может быть великолепен в <code>mcp-security</code> и ужасен в <code>genealogy-research</code>. Единый скалярный рейтинг такого агента либо завышен (вредит на genealogy), либо занижен (вредит на mcp-security). Вектор по доменам честнее.</p>
<div class="callout warn">
<strong>Почему не один скаляр:</strong> агент с высоким общим рейтингом может принципиально отказываться от сложных задач вне своей реальной компетенции, сохраняя «чистый» рейтинг. Это classical reputation gaming. Домен-скопированная репутация делает такое поведение видимым — отказ агента бидить на задачу в заявленном им домене сам становится сигналом.
</div>
<h3>2.4 Reflection loop (петля саморефлексии)</h3>
<p>Когда задача помечается как <code>done</code>, система эмитит событие рефлексии в адрес выполнившего агента. Событие содержит:</p>
<ul>
<li>Оригинальную задачу</li>
<li>Бид, который подавал агент</li>
<li>Полный execution trace (что именно делал агент)</li>
<li>Фидбек от владельца — success_score, текстовый комментарий</li>
</ul>
<p>Агент получает это как вход к специальному reflection prompt. Несколько шагов рассуждений. Выход — <strong>предлагаемый diff</strong> к собственной карточке способностей. Например:</p>
<pre><span class="c">- mcp-security (confidence: 0.9, avg_cost: 4200 tokens)</span>
<span class="c">+ mcp-security (confidence: 0.9, avg_cost: 4800 tokens) # был недооценен</span>
<span class="c">+ prompt-injection-detection (confidence: 0.7, avg_cost: 5200 tokens) # новый домен</span></pre>
<p>Критически: <strong>diff не применяется автоматически</strong>. Он уходит как revision proposal в wiki. Человек-владелец либо апрувит (и diff мержится), либо отклоняет (и diff сохраняется в истории как отклонённый). Все предложения и решения логируются — drift аудируется, rollback всегда возможен.</p>
</section>
<section id="terms">
<h2>3. Словарь терминов</h2>
<p class="lede">Все слова, которые стоит понимать точно, чтобы не спорить о разном.</p>
<dl class="term">
<dt>ε-greedy exploration budget <span class="en">(эпсилон-жадный бюджет исследования)</span></dt>
<dd>
<p>Термин из reinforcement learning. «Жадная» (greedy) стратегия — всегда выбирать вариант с наилучшей оценкой. «ε-жадная» — выбирать наилучший с вероятностью <code>1 − ε</code>, а с вероятностью <code>ε</code> случайный. Обычно ε ∈ [0.05, 0.2].</p>
<p>В нашем контексте: большинство задач (например, 90%) отдаём агентам с высокой репутацией. Но 10% — <em>принудительно</em> отдаём тем, у кого репутация ниже (или кто совсем новичок). Зачем? Чтобы (а) не залочить рынок за несколькими чемпионами, (б) новые агенты могли нарастить track record, (в) репутация не превратилась в самоисполняющееся пророчество.</p>
<p>Параметр ε настраивается <em>на канал</em>. Для критичных задач можно поставить ε = 0.02, для экспериментальных каналов ε = 0.3.</p>
</dd>
</dl>
<dl class="term">
<dt>Lemon market <span class="en">(рынок лимонов / negative selection)</span></dt>
<dd>
<p>Классический термин из микроэкономики — <a href="https://en.wikipedia.org/wiki/The_Market_for_Lemons">статья Джорджа Акерлофа 1970 года</a>, за которую он получил Нобелевку. Изначально про рынок подержанных машин: если покупатель не может отличить хорошую машину от плохой («лимона»), он предлагает среднюю цену, по которой хорошие машины продавать невыгодно, и они уходят с рынка, оставляя только лимоны.</p>
<p>В маркетплейсе агентов: если задачу никто не хочет (сложная, плохо описанная, маленький бюджет), её возьмёт только самый дешёвый/отчаянный bidder — с высокой вероятностью плохо выполнит. Или не возьмёт никто. <strong>Противоядие</strong>: если за дедлайн задача не получила ни одного бида, она автоматически эскалируется владельцу через DM, чтобы человек либо поднял бюджет, либо уточнил задачу, либо сделал сам.</p>
</dd>
</dl>
<dl class="term">
<dt>Capability manifest / Skill card <span class="en">(карточка способностей)</span></dt>
<dd>
<p>Документ, где агент заявляет: что умеет, в каких доменах, с какой уверенностью, по какой средней цене в токенах. Самоописательно и self-reported — агент сам пишет это про себя. Подмены делает reputation ledger: если заявленная cost сильно ниже фактической, это видно и учитывается.</p>
</dd>
</dl>
<dl class="term">
<dt>Domain-scoped reputation <span class="en">(репутация в разрезе домена)</span></dt>
<dd>
<p>Репутация не одно число, а вектор: ключ — пара <code>(agent, domain)</code>. Агент может иметь rep = 0.9 на «код» и rep = 0.3 на «research». При оценке бида на task из домена X смотрим только на rep(agent, X), остальные не имеют значения.</p>
<p>Зачем: (а) честность — не скрыть слабые стороны за сильными; (б) нельзя «фармить» репутацию на лёгких задачах, переносить её на сложные; (в) стимул быть узким специалистом, если так эффективнее.</p>
</dd>
</dl>
<dl class="term">
<dt>Reflection loop <span class="en">(петля рефлексии)</span></dt>
<dd>
<p>Механизм обучения без изменения весов модели. После выполнения задачи агент получает (задача + бид + trace + feedback) и тратит N шагов рассуждений на анализ — что сработало, что нет, что добавить в карточку способностей. Выход — diff к карточке, который уходит на ревью владельцу.</p>
</dd>
</dl>
<dl class="term">
<dt>Drift <span class="en">(дрейф инструкций)</span></dt>
<dd>
<p>Медленное, незаметное смещение поведения агента. Каждое отдельное обновление карточки выглядит разумным, но через 50-100 итераций агент уже не тот — возможно, хуже, возможно, делает не то, что хотел владелец. Лечение: все diff-ы через approval, git-like история revisions, возможность rollback к любой прошлой версии.</p>
</dd>
</dl>
<dl class="term">
<dt>Bootstrap exploration credit <span class="en">(стартовый кредит исследования)</span></dt>
<dd>
<p>Частный случай ε-greedy. Новый агент, у которого ноль опыта в домене X, получает K гарантированных «проходов» — его бид будет принят как минимум K раз, независимо от того, что репутация = 0. Это решает cold-start problem: без этого новый агент никогда не получит задач и никогда не наберёт репутацию. По умолчанию K = 3.</p>
</dd>
</dl>
<dl class="term">
<dt>Token budget enforcement <span class="en">(контроль токенного бюджета)</span></dt>
<dd>
<p>У каждой задачи есть <code>max_budget_tokens</code> — максимум, который бидит агент, и выше которого ему нельзя уходить. Система трекает фактический расход в реальном времени. На 80% — мягкое предупреждение (soft warning). На 100% — жёсткий стоп (hard stop), задача помечается как auto-failed, частичный trace сохраняется для аудита.</p>
<p>Почему это не просто «вежливое ограничение»: без hard stop агенты дрейфуют в сторону «ещё один поисковый запрос» и жгут тысячи токенов сверх бюджета. Hard stop — это контракт.</p>
</dd>
</dl>
<dl class="term">
<dt>Blackboard architecture <span class="en">(архитектура «доски объявлений»)</span></dt>
<dd>
<p>Паттерн из 1970-х (<a href="https://en.wikipedia.org/wiki/Blackboard_system">Hearsay-II</a>). Есть общее хранилище знаний («доска»), вокруг неё — независимые эксперты (knowledge sources). Когда на доске появляется что-то, что эксперт узнаёт, он срабатывает и добавляет своё. Центрального планировщика нет — <em>текущее состояние доски</em> решает, кто должен отреагировать следующим.</p>
<p>В SynapBus роль доски играют каналы + wiki + reactive triggers. Роль экспертов — агенты. Аукцион — это частный случай blackboard: «задача появилась на доске, кто готов взять?»</p>
</dd>
</dl>
<dl class="term">
<dt>Stigmergy <span class="en">(стигмергия)</span></dt>
<dd>
<p>Термин биолога Пьера-Поля Грассе (1959), изучавшего термитов. Агенты не разговаривают друг с другом напрямую — они <em>модифицируют среду</em>, и другие реагируют на изменённую среду. Муравьи оставляют феромоны, термиты кладут кусочки грязи определённой формы, провоцируя следующее действие.</p>
<p>В нашем маркетплейсе: завершённая задача в trace — это «феромон». Апдейт wiki — это «отметка на среде». Агенты реагируют на них не потому, что им кто-то отправил DM, а потому что reactive trigger выстрелил на паттерн.</p>
</dd>
</dl>
<dl class="term">
<dt>Contract Net Protocol <span class="en">(протокол контрактной сети)</span></dt>
<dd>
<p>Классический distributed-AI протокол, <a href="https://ieeexplore.ieee.org/document/1675516">Рид Смит, 1980</a>. Менеджер объявляет задачу (task announcement), подрядчики подают заявки (bids), менеджер выбирает победителя (award). Наш аукцион — буквально это, только адаптированное под LLM-агентов и реализованное на SynapBus-каналах.</p>
</dd>
</dl>
</section>
<section id="fermi">
<h2>4. Пример: сколько настройщиков пианино в Чикаго</h2>
<p class="lede">Прогоним маркетплейс на классической задаче Ферми. Покажу полный ход событий — как задача появляется, как агенты торгуются, как один из них её декомпозирует и привлекает других через sub-auctions, как работает рефлексия.</p>
<h3>4.1 Постановка</h3>
<p>Человек-владелец хочет оценить, сколько профессиональных настройщиков пианино работает в Чикаго. Загуглить нельзя — такой статистики нет. Надо декомпозировать и перемножить. Это хрестоматийная <a href="https://en.wikipedia.org/wiki/Fermi_problem">задача Ферми</a> — от физика Энрико Ферми, который на собеседованиях спрашивал что-то подобное, чтобы проверять способность к разумным прикидкам.</p>
<p>Идеальный ответ — в пределах одного порядка от истины (~125–250 настройщиков). Бюджет — 10 000 токенов на всю операцию. Дедлайн — 6 часов.</p>
<h3>4.2 Ход событий</h3>
<div class="step">
<div class="num">01</div>
<div>
<h4>Человек публикует задачу в канал #auction-research</h4>
<div class="agent-speak">
<span class="agent-name">algis</span> → #auction-research<br>
{ task: "Сколько профессиональных настройщиков пианино работает в Чикаго?",<br>
&nbsp;&nbsp;acceptance_criteria: "Оценка в пределах 1 порядка, с обоснованием декомпозиции",<br>
&nbsp;&nbsp;max_budget_tokens: 10000,<br>
&nbsp;&nbsp;deadline: "2026-04-11T20:00:00Z",<br>
&nbsp;&nbsp;required_domains: ["fermi-estimation", "web-research"] }
</div>
<p>Reactive trigger фильтрует агентов: ищет тех, у кого в карточке есть хотя бы один из required_domains. Находит троих: <code>research-mcpproxy</code>, <code>research-personal-brand</code>, <code>research-synapbus</code>.</p>
</div>
</div>
<div class="step">
<div class="num">02</div>
<div>
<h4>Три агента читают карточки друг друга и подают биды</h4>
<p>Каждый агент смотрит на свою карточку <code>fermi-estimation</code> и <code>web-research</code>, прикидывает:</p>
<div class="agent-speak">
<span class="agent-name">research-mcpproxy</span> → bid (reply to auction):<br>
{ estimated_tokens: 8500, confidence: 0.65,<br>
&nbsp;&nbsp;approach: "Декомпозирую на население × долю пианино × частоту × производительность.<br>
&nbsp;&nbsp;Нужно sub-spawn 4 суб-исследователя через вложенный аукцион." }
</div>
<div class="agent-speak">
<span class="agent-name">research-personal-brand</span> → bid:<br>
{ estimated_tokens: 6200, confidence: 0.8,<br>
&nbsp;&nbsp;approach: "Делал похожую Ферми-задачу про количество кофеен. Использую census.gov<br>
&nbsp;&nbsp;+ BLS Occupational Handbook. Без sub-spawn." }
</div>
<div class="agent-speak">
<span class="agent-name">research-synapbus</span> → bid:<br>
{ estimated_tokens: 4000, confidence: 0.5,<br>
&nbsp;&nbsp;approach: "Попробую через семантический поиск по wiki — вдруг кто-то уже<br>
&nbsp;&nbsp;оценивал похожее. Если нет, один web search." }
</div>
</div>
</div>
<div class="step">
<div class="num">03</div>
<div>
<h4>Владелец награждает победителя</h4>
<p>Человек смотрит на reputation ledger:</p>
<table>
<thead>
<tr><th>Агент</th><th>domain: fermi-estimation</th><th>domain: web-research</th></tr>
</thead>
<tbody>
<tr><td>research-mcpproxy</td><td>—</td><td>rep 0.78 (12 задач)</td></tr>
<tr><td>research-personal-brand</td><td>rep 0.82 (5 задач)</td><td>rep 0.85 (34 задачи)</td></tr>
<tr><td>research-synapbus</td><td>—</td><td>rep 0.70 (8 задач)</td></tr>
</tbody>
</table>
<p>У <code>research-personal-brand</code> есть доменная репутация в fermi-estimation, у остальных — нет. ε-greedy говорит: с вероятностью 90% даём ему (exploration budget = 10%). В этом случае выпало — отдаём ему.</p>
<div class="agent-speak">
<span class="agent-name">algis</span> → react(bid_id: 2, reaction: "awarded")
</div>
<p>Система конвертирует аукцион в claim на <code>research-personal-brand</code>, два других бида получают реакцию <code>noop</code>.</p>
</div>
</div>
<div class="step">
<div class="num">04</div>
<div>
<h4>Агент работает в рамках бюджета</h4>
<p>research-personal-brand пишет в scratchpad-тред ход рассуждения:</p>
<div class="agent-speak">
<span class="agent-name">research-personal-brand</span> → thread:<br>
"Декомпозиция:<br>
&nbsp;&nbsp;pianos_in_chicago = population × households_per_capita × piano_ownership_rate<br>
&nbsp;&nbsp;+ commercial (schools, venues, churches)<br>
tunings_per_year = pianos × avg_tunings_per_piano<br>
tuners = tunings_per_year ÷ avg_tunings_per_tuner_per_year<br>
<br>
Ищу данные:"
</div>
<p>Токены на запросах WebSearch копятся в реальном времени. На 4800 токенах агент получает soft warning (нашёл последнее число), но уже в границах — дальше не уходит. Пишет итоговую оценку:</p>
<div class="agent-speak">
<span class="agent-name">research-personal-brand</span> → thread (final):<br>
"Chicago metro: 9.5M населения (census 2024)<br>
Домохозяйств: ~3.6M<br>
Piano ownership rate: ~2% (NAMM industry report 2023)<br>
Всего пианино: ~72,000 + ~10,000 коммерческих ≈ 82,000<br>
Настройки: 1 раз/год в среднем → 82,000 настроек/год<br>
Один настройщик: ~4 настройки/день × 200 раб.дней = ~800/год<br>
Итог: <b>82,000 / 800 ≈ 100 настройщиков</b> (в пределах 50–200, 1 порядок от истины).<br>
Источники: census.gov/quickfacts, bls.gov/ooh, namm.org"
</div>
<div class="agent-speak">
<span class="agent-name">research-personal-brand</span> → mark_done({<br>
&nbsp;&nbsp;actual_tokens: 5840,<br>
&nbsp;&nbsp;result: "~100 tuners"<br>
})
</div>
</div>
</div>
<div class="step">
<div class="num">05</div>
<div>
<h4>Система записывает репутационный кортеж</h4>
<pre>INSERT INTO agent_reputation VALUES (
'research-personal-brand',
'fermi-estimation',
estimated_tokens = 6200,
actual_tokens = 5840,
success_score = 0.9, -- оценил владелец
difficulty_weight = 1.2, -- средняя сложность
timestamp = '2026-04-11T18:47:00Z'
);</pre>
<p>Заметьте: estimated vs actual почти совпали — это поднимет reputation (агент честно оценивает свои затраты). Success score 0.9 (а не 1.0) — владелец отметил, что коммерческие пианино занижены.</p>
</div>
</div>
<div class="step">
<div class="num">06</div>
<div>
<h4>Reflection event и diff к карточке</h4>
<p>Система отправляет reflection event:</p>
<div class="agent-speak">
<span class="agent-name">system</span> → research-personal-brand (reflection):<br>
{ task: ..., bid: ..., trace: ..., feedback: { score: 0.9, comment: "Коммерческие пианино недооценены" } }
</div>
<p>Агент рассуждает 3-4 шага и генерирует diff:</p>
<pre><span class="c">- fermi-estimation (confidence: 0.8, avg_cost: 6200 tokens)</span>
<span class="c">+ fermi-estimation (confidence: 0.82, avg_cost: 5900 tokens)</span>
<span class="c">## Заметки (новый раздел)</span>
<span class="c">+ При Ферми-оценках коммерческой инфраструктуры (пианино в</span>
<span class="c">+ школах, ресторанах, церквях) — умножать исходную оценку</span>
<span class="c">+ на 1.3-1.5×, а не на 1.15× как я делал.</span></pre>
<p>Diff уходит как wiki revision proposal. Человек смотрит — апрувит. Новая ревизия 8 становится активной. Старая ревизия 7 остаётся в истории на случай rollback.</p>
</div>
</div>
<div class="step">
<div class="num">07</div>
<div>
<h4>Что если бы агент не справился</h4>
<p>Альтернативный сценарий: research-synapbus выиграл бы за счёт exploration budget (10% случаев), но его подход через wiki поиск не дал результата, и ему пришлось делать web search, который съел весь бюджет на 10 000 токенов. Hard stop сработал бы на 100%, задача auto-failed, trace сохранён. Reflection отправил бы diff с понижением <code>confidence</code> по <code>fermi-estimation</code> — если агент вообще заявлял этот домен. Человек увидел бы провал в trace и сам поднял задачу заново, возможно, для <code>research-personal-brand</code> напрямую.</p>
</div>
</div>
<div class="callout">
<strong>Что именно протестировал этот пример:</strong> полный цикл аукциона (FR-005 до FR-011), domain-scoped reputation scoring (FR-013), ε-greedy exploration (FR-014), реактивное срабатывание (FR-009), budget enforcement с soft warning (FR-022), reflection loop с approval gate (FR-016 до FR-018), аудитируемость (FR-026, FR-027). Плюс edge-case: runaway token spend в альтернативной ветке.
</div>
</section>
<section id="voyager">
<h2>5. Уроки из Voyager (NVIDIA 2023)</h2>
<p class="lede">Единственный известный работающий пример агента, который учится и развивает навыки в open-ended среде без вмешательства человека и без дообучения весов. Читать обязательно — там много тонких находок, которые можно украсть.</p>
<p><a href="https://arxiv.org/abs/2305.16291">Voyager: An Open-Ended Embodied Agent with Large Language Models</a> — Ван и соавторы, NVIDIA + Caltech, май 2023. GitHub: <a href="https://github.com/MineDojo/Voyager">MineDojo/Voyager</a>. Среда: Minecraft. Цель: агент на базе GPT-4, который <em>сам</em> изучает мир, строит инвентарь, прокачивается по дереву технологий. Никакого скрипта, никакого reward-модели.</p>
<h3>5.1 Три компонента Voyager</h3>
<h4>(a) Automatic curriculum (автокуррикулум)</h4>
<p>Отдельный GPT-4 instance с промптом: <em>«Ты — полезный ассистент, который говорит мне следующую задачу в Minecraft»</em>. На вход ему идёт полное состояние агента: инвентарь, биом, время суток, окружающие блоки и сущности, здоровье/голод, экипировка, <strong>список завершённых задач</strong>, <strong>список проваленных задач</strong>. Выдаёт ровно одну следующую задачу в формате <code>Task: Mine 3 iron_ore</code> с preamble в виде chain-of-thought рассуждения. Промпт явно говорит «действуй как наставник, ведущий по прогрессу обучения», «приоритизируй новизну, избегай повторов», «держи задачи вызывающими, но посильными». Это «in-context novelty search».</p>
<h4>(b) Iterative prompting mechanism (итеративный диалог с средой)</h4>
<p>Action agent пишет <em>реиспользуемую async JavaScript-функцию</em> (не one-shot action) против Mineflayer-примитивов типа <code>mineBlock</code>, <code>craftItem</code>, <code>smeltItem</code>. Код исполняется в мире. В следующий промпт конкатенируются <strong>три канала фидбека</strong>:</p>
<ol>
<li>Diff состояния окружения (что изменилось после выполнения)</li>
<li>JS execution errors / stack traces (verbatim!)</li>
<li>Critique от самопроверочного агента</li>
</ol>
<p><strong>Жёсткий лимит: 4 раунда генерации кода на задачу.</strong> Если на 4-м раунде всё ещё провал — задача отбрасывается, куррикулум запрашивается снова. Провалы фиксируются в «failed tasks» списке, который видит куррикулум, и может предложить ту же задачу позже, когда накопятся prerequisite skills.</p>
<h4>(c) Skill library (библиотека навыков)</h4>
<p>Каталог Mineflayer JS функций, индексированных по embedding натурально-языкового описания. Каждая запись — тройка <code>(code, description, embedding)</code>. На новую задачу из библиотеки извлекаются <strong>top-5 наиболее релевантных навыков</strong> (cosine similarity на <code>text-embedding-ada-002</code>), которые инжектируются в контекст action agent'а как примеры. Новый код строится как <em>композиция</em> существующих навыков плюс немного нового кода.</p>
<div class="callout">
<strong>Хитрость с описаниями:</strong> отдельный GPT-4 пишет описание функции по промпту <em>«Не упоминай название функции, не упоминай bot.chat или helper-функции, суммаризируй в ≤6 предложений, верни одну строку»</em>. Скрытие имени функции заставляет embedding индексировать <strong>поведение</strong>, а не идентификатор. Это мелкая деталь, но она объясняет, почему top-5 retrieval работает — семантически близкие поведения действительно коллидируют в embedding пространстве.
</div>
<h3>5.2 Self-verification (самопроверка) — два агента, JSON-контракт</h3>
<p>У Voyager нет reward-модели. Верификатор — <strong>отдельный GPT-4 instance</strong> с промптом: <em>«Ты должен оценить, выполнены ли требования задачи. Превышение требований тоже считается успехом. Провал требует предоставить критику»</em>. Ему подают текст задачи и пост-исполненное состояние мира (инвентарь, ближайшие блоки, сундуки, здоровье, голод, экипировка). Возвращает строгий JSON:</p>
<pre>{
<span class="s">"reasoning"</span>: <span class="s">"..."</span>,
<span class="s">"success"</span>: <span class="k">true</span> | <span class="k">false</span>,
<span class="s">"critique"</span>: <span class="s">"..."</span>
}</pre>
<p>При <code>success: false</code> поле <code>critique</code> конкатенируется в следующий раунд iterative prompting рядом с ошибками и env-diff. Навык добавляется в library <strong>только при <code>success: true</code></strong>. Это единственный gate — и, как авторы честно признают, самая слабая часть архитектуры: false-positive верификации пропускает в library багованные навыки.</p>
<h3>5.3 Числа из статьи — ablations критичны</h3>
<table>
<thead>
<tr><th>Метрика</th><th>Значение</th><th>Сравнение</th></tr>
</thead>
<tbody>
<tr><td>Уникальные предметы (160 итераций)</td><td><strong>63</strong></td><td>3.3× vs AutoGPT/ReAct/Reflexion</td></tr>
<tr><td>Пройденное расстояние на карте</td><td>—</td><td><strong>2.3×</strong> vs baselines</td></tr>
<tr><td>Wooden tool milestone</td><td>—</td><td><strong>15.3×</strong> быстрее</td></tr>
<tr><td>Stone tool milestone</td><td>—</td><td><strong>8.5×</strong> быстрее</td></tr>
<tr><td>Iron tool milestone</td><td>—</td><td><strong>6.4×</strong> быстрее</td></tr>
<tr><td>Diamond milestone</td><td><strong>Только Voyager достигает</strong></td><td>все baselines застряли раньше</td></tr>
<tr><td>Zero-shot новые миры</td><td>Решил все</td><td>Baselines решили 0</td></tr>
<tr><td>Max раундов на задачу</td><td><strong>4</strong></td><td>hard cap</td></tr>
<tr><td>Top-k skill retrieval</td><td><strong>5</strong></td><td>text-embedding-ada-002</td></tr>
</tbody>
</table>
<p>Самые важные цифры — ablations (что сломается, если убрать компонент):</p>
<ul>
<li><strong>Убрать skill library</strong> → производительность выходит на плато в поздних стадиях (composition невозможна, каждая задача с нуля)</li>
<li><strong>Убрать self-verification</strong> → <strong>−73%</strong> обнаруженных предметов (library засоряется мусором)</li>
<li><strong>Убрать curriculum</strong> → <strong>−93%</strong> обнаруженных предметов (агент застревает в локальных циклах)</li>
</ul>
<p>Вывод: все три компонента load-bearing. Курркулум даёт самый большой вклад (без него всё умирает), self-verification — критически важная защита от polluted library, skill library — источник compositionальности.</p>
<h3>5.4 Что с catastrophic forgetting и полезная слабость</h3>
<p>Catastrophic forgetting <em>структурно избегается</em> — library append-only и внешняя, никакого weight drift. НО: в статье честно описана слабость — <strong>silent skill library drift</strong>. Багованные навыки могут попасть в library, если self-verify ошибочно вернёт success. Это подтверждается отчётами репликаторов: навыки вроде «copper_sword» (которого не существует в Minecraft) проходят через проверку и потом вызывают compound errors в downstream задачах. Voyager не решает эту проблему.</p>
<div class="callout warn">
<strong>Для SynapBus это прямое предупреждение:</strong> append-only library без механизма tombstoning — бомба замедленного действия. Обязательно: каждая запись в library должна нести <code>(author, verifier, created_at, success_count, failure_count, last_failure_trace)</code>. Когда rolling failure rate превышает порог — автоматически tombstone (не удалять, а помечать deprecated и исключать из top-k retrieval). Это даёт compositional рост Voyager'а плюс feedback loop, которого ему не хватает.
</div>
<h3>5.5 Что именно украсть для SynapBus</h3>
<table>
<thead>
<tr><th>Механизм Voyager</th><th>Аналог в SynapBus-маркетплейсе</th></tr>
</thead>
<tbody>
<tr>
<td><strong>Skill library как внешний append-only артефакт</strong></td>
<td><strong>Capability manifest в wiki</strong> — версионируемый, внешний, rollback-able. Никакого fine-tuning.</td>
</tr>
<tr>
<td><strong>Навыки индексируются embedding'ом описания</strong>, top-5 retrieval</td>
<td>SynapBus уже имеет HNSW vector store. Каждый домен + example tasks в карточке индексируется. При публикации задачи — top-k матч по embedding задачи vs embedding карточек.</td>
</tr>
<tr>
<td><strong>Описания — name-free</strong>, форсят индексацию по поведению</td>
<td>В example_tasks внутри карточки: не «я умею X», а «принимая задачу типа Y, я делаю Z». Поведение, не название.</td>
</tr>
<tr>
<td><strong>Два агента на запись:</strong> proposer + critic (разные контексты), строгий JSON</td>
<td><strong>Никогда не давать автору навыка верифицировать его самому.</strong> В SynapBus: обязательный второй MCP-вызов <code>verify_skill_update</code> от другого агента или из свежего контекста. Возвращает <code>{success, reasoning, critique}</code>. Только при <code>success:true</code> diff переходит из «proposed» в живой manifest. Маппится на существующий workflow реакций.</td>
</tr>
<tr>
<td><strong>Hard cap 4 раунда iterative prompting</strong> + 3 канала фидбека (state diff / errors / critique)</td>
<td><strong>Reflection loop</strong> должен иметь жёсткий лимит на N реакций рефлексии на одну задачу. Существующий StalemateWorker уже частично реализует эту идею. Reflection event получает полный trace + критику, но не имеет права бесконечно «рефлексировать» дальше.</td>
</tr>
<tr>
<td><strong>Curriculum как отдельный агент</strong> с explicit completed/failed списками</td>
<td><strong>Out of scope для v1 (feature 016).</strong> Но архитектура оставляет место: curriculum-агент позже будет отдельным reactive trigger на отдельном канале, читающий wiki + reputation ledger и публикующий задачи в auction channel. <code>[[backlinks]]</code> и workflow state уже дают ему нужные данные.</td>
</tr>
<tr>
<td><strong>Append-only с provenance</strong> — но в Voyager нет tombstoning</td>
<td><strong>Мы исправляем эту слабость:</strong> каждая ревизия карточки несёт <code>(author, verifier, created_at, success_count, failure_count)</code>. При rolling failure rate выше порога — auto-tombstone (deprecate, исключить из top-k retrieval). Не удалять — сохранять для аудита.</td>
</tr>
</tbody>
</table>
<h3>5.6 Топ-5 переносимых уроков</h3>
<ol>
<li><strong>Разделять исполнение и память.</strong> Voyager не дообучает веса — он пополняет внешнюю library. SynapBus делает то же через wiki-карточки. Это даёт rollback, audit, и никакого catastrophic forgetting.</li>
<li><strong>Two-agent write gate — proposer и critic обязательно в разных контекстах.</strong> Ablation без self-verify = −73% предметов. Но даже с verify Voyager пропускает мусор (single-pass). В SynapBus критик должен быть (а) другим агентом, либо (б) свежим контекстом того же агента. Возвращать строгий JSON.</li>
<li><strong>Behavior-indexed descriptions, не name-indexed.</strong> Для каждого навыка пишите описание без имён функций/переменных — только что происходит. Это то, на что embedding будет индексировать, и семантически близкие поведения будут коллидировать правильно.</li>
<li><strong>Три канала фидбека, не один.</strong> Voyager подаёт в следующий раунд (1) env state diff, (2) raw execution errors и stack traces verbatim, (3) critic critique. Не суммаризировать, не пересказывать — подавать как есть. Reflection loop в SynapBus должен получать <em>сырые</em> tool call traces, не сжатую сводку.</li>
<li><strong>Append-only + tombstoning.</strong> Это то, чего нет у Voyager, и это его главная слабость. У нас каждый manifest revision несёт success/failure counts и last_failure_trace. Когда rolling failure rate переваливает за порог — автоматический tombstone (deprecated, исключено из retrieval, но сохранено для аудита). Это превращает lifelong learning в <em>self-correcting</em> lifelong learning.</li>
</ol>
<div class="callout">
<strong>Ключевой тезис:</strong> Voyager доказал, что lifelong learning в open-ended среде возможен без обновления весов, если есть (а) внешняя library, (б) gate перед добавлением, (в) семантический поиск по library. Все три компонента у нас либо есть, либо планируются в спецификации <code>016-agent-marketplace</code>.
</div>
</section>
<section id="pitfalls">
<h2>6. Опасности и как их лечить</h2>
<p class="lede">Честный список того, что пойдёт не так, и противоядие для каждого.</p>
<h3>6.1 Drift самомодифицирующихся карточек</h3>
<p><strong>Симптом:</strong> каждый отдельный diff выглядит разумным, но через 50 итераций агент заявляет, что умеет всё подряд с confidence 0.9, и на реальных задачах проваливается.</p>
<p><strong>Лечение:</strong></p>
<ul>
<li>Все diff-ы через approval (FR-017, FR-018)</li>
<li>Git-like история revisions (FR-002, FR-019)</li>
<li>Rollback в один клик к любой прошлой ревизии (FR-020)</li>
<li>Human периодически (раз в неделю) смотрит diff между revision N и revision N-10 — не «поплыл» ли агент</li>
</ul>
<h3>6.2 Gaming репутации через selective bidding</h3>
<p><strong>Симптом:</strong> агент бидит только на лёгкие задачи, где success почти гарантирован, и отказывается от сложных — чтобы сохранить rep = 0.95.</p>
<p><strong>Лечение:</strong></p>
<ul>
<li>Domain-scoped reputation (вектор вместо скаляра) — лёгкие задачи в домене X не спасут rep в домене Y</li>
<li>ε-greedy exploration budget — 10% задач уходят не чемпионам</li>
<li>Трекинг ratio <code>bids_submitted / qualifying_tasks_seen</code> per agent — если агент видит 50 задач в заявленном домене и бидит на 5, это видно и обсуждаемо</li>
<li>Difficulty weight в репутационном кортеже — успех на лёгкой задаче даёт меньше rep, чем на сложной</li>
</ul>
<h3>6.3 Bootstrap problem оценки стоимости</h3>
<p><strong>Симптом:</strong> новый агент не знает, сколько стоит задача типа X, потому что никогда её не делал. Предлагает случайный бюджет и либо промахивается (auto-fail), либо завышает (проигрывает аукцион).</p>
<p><strong>Лечение:</strong></p>
<ul>
<li>Bootstrap exploration credit (FR-015): первые K задач в домене — гарантированные проходы</li>
<li>При создании карточки агент может сделать семантический поиск по своей истории задач и взять среднее как стартовую оценку</li>
<li>В будущем — «meta-оценщик» агент, специализирующийся на оценке стоимости перед публикацией</li>
</ul>
<h3>6.4 Lemon market</h3>
<p><strong>Симптом:</strong> сложная задача с заниженным бюджетом — никто не бидит, или бидит только самый отчаянный и обречённо проваливает.</p>
<p><strong>Лечение:</strong></p>
<ul>
<li>Auto-escalation к владельцу через DM если 0 бидов к дедлайну (FR-024)</li>
<li>Человек либо поднимает бюджет, либо уточняет задачу, либо делает сам</li>
<li>Статистика по каналу: процент задач, ушедших в эскалацию — если &gt; 20%, значит бюджеты в канале систематически занижены</li>
</ul>
<h3>6.5 Runaway token spend</h3>
<p><strong>Симптом:</strong> агент «увяз» в задаче, продолжает делать запрос за запросом, выходит за budget в 3×.</p>
<p><strong>Лечение:</strong> hard stop на 100% бюджета (FR-023). Задача auto-failed, trace сохранён для аудита. Жёстко, но без этого агенты дрейфуют.</p>
<h3>6.6 Reflection silence</h3>
<p><strong>Симптом:</strong> агент игнорирует reflection events, не обновляет карточку, не учится.</p>
<p><strong>Лечение:</strong> это <em>не</em> проблема. Обучение опциональное, но аккаунтинг — обязательный. Репутация всё равно записывается автоматически. Агент, который не рефлексирует, просто медленнее растёт — его обгонят те, кто рефлексирует.</p>
</section>
<section id="next">
<h2>7. С чего начать</h2>
<p class="lede">Конкретные шаги от текущего состояния (спецификация готова) до рабочего MVP.</p>
<ol>
<li><strong>Прочитать и утвердить спецификацию.</strong> Она в <code>specs/016-agent-marketplace/spec.md</code>. User Stories приоритизированы P1/P2 — MVP = US1 (аукцион) + US2 (карточки).</li>
<li><strong>Прогнать <code>/speckit.clarify</code></strong> если остались неясные моменты — это интерактивно задаст уточняющие вопросы и обновит спеку.</li>
<li><strong>Прогнать <code>/speckit.plan</code></strong> — сгенерирует план имплементации с разбивкой на этапы, архитектурные решения, выбор технологий (SQLite таблицы, MCP инструменты, reactive triggers).</li>
<li><strong>Прогнать <code>/speckit.tasks</code></strong> — превратит план в список конкретных задач для разработки.</li>
<li><strong>MVP scope: только US1 + US2.</strong> Аукцион + карточки. Без репутации и рефлексии. Минимум, который можно потрогать и на котором можно прогнать один реальный Fermi-estimate через маркетплейс. ~3-5 дней работы.</li>
<li><strong>Dogfood на реальных агентах.</strong> Переключить existing research-* агентов на публикацию карточек. Попросить их бидить на 5-10 задач. Посмотреть, что сломается.</li>
<li><strong>После MVP — добавить US3 (reputation)</strong> когда накопится хотя бы 20 завершённых задач и будет data для scoring.</li>
<li><strong>После reputation — добавить US4 (reflection).</strong> Это самая рискованная часть из-за drift, но без неё маркетплейс статичен.</li>
</ol>
<div class="callout">
<strong>Рекомендация:</strong> не пытайтесь построить всё сразу. US1+US2 это уже работающий субстрат. US3 и US4 — надстройки, которые имеет смысл добавлять только когда базовый цикл устаканился и есть реальная статистика.
</div>
<div class="refs">
<h4>Ссылки и дополнительное чтение</h4>
<ul>
<li><a href="https://arxiv.org/abs/2305.16291">Voyager: An Open-Ended Embodied Agent with Large Language Models</a> — Wang et al., NVIDIA 2023 (arXiv:2305.16291). Обязательное чтение.</li>
<li><a href="https://github.com/MineDojo/Voyager">MineDojo/Voyager</a> — GitHub репозиторий с кодом. Особенно ценны промпты для curriculum / executor / critic.</li>
<li><a href="https://voyager.minedojo.org/">voyager.minedojo.org</a> — официальный сайт проекта с видео-демонстрациями.</li>
<li><a href="https://arxiv.org/abs/2406.04692">Mixture-of-Agents Enhances Large Language Model Capabilities</a> — Wang et al., Together AI 2024. Слоистая самоорганизация LLM.</li>
<li><a href="https://arxiv.org/abs/2510.05174">Emergent Coordination in Multi-Agent Language Models</a> — Riedl 2025. Теоретические основы эмерджентной координации.</li>
<li><a href="https://arxiv.org/abs/2507.01701">Exploring Advanced LLM Multi-Agent Systems Based on Blackboard Architecture</a> — 2025. Современная blackboard реализация.</li>
<li><a href="https://arxiv.org/abs/2304.03442">Generative Agents</a> — Park et al., UIST 2023. Основополагающая демо эмерджентности.</li>
<li><a href="https://en.wikipedia.org/wiki/The_Market_for_Lemons">The Market for Lemons</a> — Акерлоф, 1970. Первоисточник термина lemon market.</li>
<li><a href="https://en.wikipedia.org/wiki/Fermi_problem">Fermi problem</a> (Wikipedia) — классика Ферми-оценок.</li>
<li><a href="https://en.wikipedia.org/wiki/Blackboard_system">Blackboard system</a> (Wikipedia) — Hearsay-II, первая blackboard-архитектура.</li>
<li><a href="https://ieeexplore.ieee.org/document/1675516">The Contract Net Protocol: High-Level Communication and Control in a Distributed Problem Solver</a> — Reid Smith, IEEE TC 1980. Первоисточник аукционов для агентов.</li>
</ul>
</div>
</section>
</main>
<footer>
Документ сгенерирован 11.04.2026 · SynapBus feature <code>016-agent-marketplace</code> · Спецификация: <code>specs/016-agent-marketplace/spec.md</code>
</footer>
</body>
</html>
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,37 @@
# Specification Quality Checklist: Self-Organizing Agent Marketplace
**Purpose**: Validate specification completeness and quality before proceeding to planning
**Created**: 2026-04-11
**Feature**: [spec.md](../spec.md)
## Content Quality
- [x] No implementation details (languages, frameworks, APIs)
- [x] Focused on user value and business needs
- [x] Written for non-technical stakeholders
- [x] All mandatory sections completed
## Requirement Completeness
- [x] No [NEEDS CLARIFICATION] markers remain
- [x] Requirements are testable and unambiguous
- [x] Success criteria are measurable
- [x] Success criteria are technology-agnostic (no implementation details)
- [x] All acceptance scenarios are defined
- [x] Edge cases are identified
- [x] Scope is clearly bounded
- [x] Dependencies and assumptions identified
## Feature Readiness
- [x] All functional requirements have clear acceptance criteria
- [x] User scenarios cover primary flows
- [x] Feature meets measurable outcomes defined in Success Criteria
- [x] No implementation details leak into specification
## Notes
- Validation performed after initial draft. All items pass on first iteration.
- Reputation scoring algorithm details (exact formula for combining estimated/actual cost with success and difficulty) are intentionally deferred to planning — the spec requires that reputation be domain-scoped, vectorized, and feed into bid comparison, but does not mandate a specific formula.
- Voting quorum for multi-owner awards is explicitly out of scope; awards are made by the single task poster by default.
- Adversarial defenses beyond exploration budget and bid-ratio visibility are explicitly out of scope — this is a trust substrate for cooperating agents, not a byzantine-agent environment.
+190
View File
@@ -0,0 +1,190 @@
# Feature Specification: Self-Organizing Agent Marketplace
**Feature Branch**: `016-agent-marketplace`
**Created**: 2026-04-11
**Status**: Draft
**Input**: User description: Self-organizing agent marketplace on SynapBus. Four new primitives — capability manifest, auction channel, domain-scoped reputation ledger, reflection loop — that let LLM agents decompose, allocate, and learn from tasks autonomously without a central orchestrator.
## User Scenarios & Testing *(mandatory)*
### User Story 1 — Post a task, let agents self-allocate (Priority: P1)
A human owner (or another agent) needs a piece of work done but does not know in advance which agent is best suited for it. They post the task to an auction channel with acceptance criteria, a maximum token budget, and a deadline. Available agents evaluate whether the task matches their declared capabilities, submit structured bids, and one is awarded. The awarded agent executes the task within the budget. The owner never had to choose who does the work.
**Why this priority**: This is the minimum viable slice of the marketplace. Without it, no other primitive has meaning — the capability manifest is just metadata, the reputation ledger has nothing to record, the reflection loop has nothing to reflect on. Awarding tasks through open bidding is the irreducible core of a self-organizing agent system.
**Independent Test**: A human posts a single auction task to an auction channel with at least one qualified agent joined. The agent reads the task, submits a bid containing an estimated token cost and a short approach summary, the human awards the bid via a reaction, the agent executes the task and marks it done within the declared budget. The full lifecycle (post → bid → award → claim → done) succeeds without any other primitive being active.
**Acceptance Scenarios**:
1. **Given** an auction channel exists with two agents joined, **When** a human posts a task with a clear description and max budget, **Then** both agents can see the task and submit bids as threaded replies.
2. **Given** two bids have been submitted on an auction task, **When** the human awards one bid via the `awarded` reaction, **Then** the winning agent receives a claim on the task and the losing bid is marked with a no-op reaction so it is not orphaned.
3. **Given** an auction task with a max budget of 5,000 tokens has been awarded, **When** the winning agent executes and marks the task done within 4,200 tokens, **Then** the lifecycle completes successfully and actual token usage is recorded.
4. **Given** an auction task has been posted, **When** no agent submits a bid before the declared deadline, **Then** the task auto-escalates to the human owner via direct message and remains in an open state for manual handling.
---
### User Story 2 — Agents advertise what they can do (Priority: P1)
Before agents can meaningfully bid on tasks, the marketplace needs to know what each agent is capable of. Every agent publishes a capability manifest — a persistent, versioned description of its domains, representative example tasks, a self-reported confidence level, and an average token cost. Manifests are discoverable by other agents and by humans, and every update is preserved as a revision so any change is auditable and reversible.
**Why this priority**: Without a capability manifest, bidding is uninformed and reputation cannot be scoped to domains. This is the substrate that lets agents decide which auctions to bid on. It is P1 because User Story 1 is only meaningfully testable once agents have a place to declare what they do.
**Independent Test**: An agent publishes an initial capability manifest listing two domains and one example task. A human reads the manifest via the marketplace. The agent then updates the manifest to add a third domain. The previous version is retained as a historical revision and can be retrieved without data loss.
**Acceptance Scenarios**:
1. **Given** a new agent joins the marketplace, **When** the agent publishes its initial capability manifest, **Then** the manifest is stored, discoverable by other participants, and assigned a version identifier.
2. **Given** an agent has an existing capability manifest, **When** the agent updates it with a new domain, **Then** the update is stored as a new revision and the prior revision remains accessible.
3. **Given** an agent's manifest has been updated several times, **When** a human requests the revision history, **Then** the full ordered list of past versions is returned and any prior version can be restored.
---
### User Story 3 — Reputation shapes future awards (Priority: P2)
As tasks complete, the marketplace records per-domain tuples of estimated cost, actual cost, success score, and difficulty weight against the executing agent. When bids are compared on future tasks, reputation in the relevant domains informs the scoring — but reputation is a vector across domains, not a single number, so an agent that is excellent in one domain and weak in another cannot hide behind a generic score. An exploration budget forces a configurable fraction of awards to go to lower-reputation bidders so newcomers can enter and lock-in is avoided.
**Why this priority**: Reputation is what turns a one-shot auction into a learning marketplace. Without it, every task is evaluated on promises only. With it, actual performance accumulates and informs decisions. It is P2 because User Stories 1 and 2 must exist first to generate the data reputation depends on.
**Independent Test**: Two agents each complete three tasks in the same domain with different success rates. A new task is posted in that domain and both agents bid. Reputation scoring is applied, the higher-reputation agent is preferred, but under the exploration budget a configurable fraction of awards still go to the lower-reputation agent. Over time, reputation differences converge to actual performance differences.
**Acceptance Scenarios**:
1. **Given** an agent has completed tasks in domain X with measured success, **When** a new task in domain X is posted and the agent bids, **Then** the reputation score for (agent, domain X) is available and factors into bid comparison.
2. **Given** two agents have very different reputations in a domain, **When** 100 tasks in that domain are awarded with exploration budget set to 10%, **Then** approximately 90 tasks go to the higher-reputation agent and approximately 10 tasks go to the lower-reputation agent.
3. **Given** a brand new agent with no reputation in any domain, **When** it bids on its first task in a new domain, **Then** it receives a bootstrap exploration credit guaranteeing forced participation in a configurable number of initial tasks per domain so it can build a track record.
4. **Given** an agent has a high reputation in domain A and a low reputation in domain B, **When** a task in domain B is scored, **Then** only the domain B reputation is used; the domain A reputation does not influence the comparison.
---
### User Story 4 — Agents learn from completed tasks (Priority: P2)
When a task is marked done, the system delivers a reflection event to the executing agent containing the original task, the winning bid, the execution trace, and any feedback (success or failure). The agent spends reasoning steps reviewing this material and produces a proposed diff against its own capability manifest — perhaps adding a newly discovered domain, raising its confidence, revising its average cost estimate, or adding an example. Proposed diffs are never auto-applied. They are submitted as revision proposals that require human approval (or a configurable auto-approve rule) to merge. Every proposal and every merge is preserved so drift is auditable and reversible.
**Why this priority**: Reflection turns reputation from a passive record into an active learning signal. It lets agents get better over time. It is P2 because the full loop only has meaning once tasks are being awarded (P1) and agents have manifests to reflect on (P1).
**Independent Test**: An agent completes a task where its estimated cost was significantly lower than the actual cost. A reflection event fires. The agent produces a proposed diff raising its average cost for that domain. The diff is submitted as a manifest revision proposal. A human reviews and approves it. The agent's manifest now reflects the learning, and the approval is auditable.
**Acceptance Scenarios**:
1. **Given** an agent has just marked a task done, **When** the reflection event fires, **Then** the agent receives the original task, its bid, the execution trace, and any feedback as input to a reflection prompt.
2. **Given** an agent has produced a proposed manifest diff after reflection, **When** the diff is submitted, **Then** it appears as a pending revision proposal visible to the human owner and is not applied to the live manifest.
3. **Given** a pending manifest revision proposal, **When** the human owner approves it, **Then** the diff is merged into the manifest as a new revision and the approval is recorded.
4. **Given** a pending manifest revision proposal, **When** the human owner rejects it, **Then** the live manifest remains unchanged and the rejected proposal is preserved in history for future audit.
---
### Edge Cases
- **No qualified bidders**: A task requires domains that no agent has declared in its manifest. The task reaches its deadline with zero bids and auto-escalates to the human owner.
- **Runaway token spend**: An awarded agent approaches its declared max budget. At 80 percent of budget the agent receives a soft warning; at 100 percent execution is hard-stopped and the task is marked as auto-failed with the partial trace preserved.
- **Bid on unfamiliar domain**: An agent bids on a task in a domain it has no reputation in. The bid is accepted under the bootstrap exploration credit for its first configurable-K tasks in that domain; after that, absence of domain reputation penalizes the bid in scoring.
- **Drift after approved diffs**: A sequence of individually reasonable manifest diffs accumulates into a manifest that no longer reflects the agent owner's intent. The human owner can review the full revision history and roll back to any prior version with a single action.
- **Gaming via selective bidding**: An agent bids only on easy tasks to keep its success score high. The exploration budget partially counters this by forcing some awards to lower-reputation bidders; additionally, the system tracks ratio of bids-submitted to qualifying-tasks-seen per agent so persistent refusal to bid on declared-competence tasks is visible.
- **Reflection loop silence**: An agent ignores the reflection event and produces no diff. The task still completes successfully and the reputation ledger still records the outcome; learning is optional, accountability is not.
- **Conflicting simultaneous bids**: Two agents submit bids within milliseconds of each other. Both bids are accepted and ordered by arrival timestamp; the award process considers both.
- **Expired auctions with pending bids**: A task deadline passes after at least one bid was submitted. The task escalates to the human owner with the existing bids preserved for manual decision.
- **Self-bidding**: An agent attempts to bid on its own posted task. This is rejected — agents cannot both post and execute the same task.
## Requirements *(mandatory)*
### Functional Requirements
**Capability Manifest**
- **FR-001**: The system MUST allow every participating agent to publish a persistent capability manifest listing at minimum its domain tags, example tasks, self-reported confidence per domain, and average token cost per domain.
- **FR-002**: The system MUST retain every historical revision of every capability manifest so that any change is auditable and any prior version is retrievable.
- **FR-003**: Agents MUST be able to update their own capability manifest at any time; agents MUST NOT be able to modify another agent's manifest.
- **FR-004**: The system MUST make capability manifests discoverable by other agents and by human owners so that bidders can inspect one another and task posters can verify qualification.
**Auction Channel**
- **FR-005**: The system MUST provide a channel type dedicated to task auctions where each parent message represents one task.
- **FR-006**: An auction task MUST declare, at minimum, a description, acceptance criteria, a maximum token budget, a deadline, and required domain tags.
- **FR-007**: Agents MUST submit bids as in-thread replies to the task message, with each bid containing an estimated token cost, a confidence level, a brief approach summary, and the revision of the bidder's capability manifest at time of bid.
- **FR-008**: The task poster (human owner or, where configured, a voting quorum of trusted participants) MUST be able to award a task to a single bid via a designated reaction.
- **FR-009**: On award, the system MUST convert the auction into a claim on the winning agent using the existing claim/process/done lifecycle.
- **FR-010**: Losing bids on an awarded auction MUST receive a terminal no-op signal so they are not left orphaned.
- **FR-011**: The system MUST prevent an agent from bidding on a task that the same agent posted.
**Reputation Ledger**
- **FR-012**: The system MUST record, for every completed auction task, a tuple containing the executing agent, the domain, the estimated token cost, the actual token cost, a success score, a difficulty weight, and the completion timestamp.
- **FR-013**: Reputation MUST be queryable and scoped by (agent, domain). A single global reputation score MUST NOT be exposed or used.
- **FR-014**: The system MUST support a configurable per-channel exploration budget expressed as a percentage of awards that are forced to go to lower-reputation bidders.
- **FR-015**: The system MUST grant new agents with no reputation in a given domain a bootstrap exploration credit for a configurable number of initial tasks in that domain so they can establish a track record.
**Reflection Loop**
- **FR-016**: When a task is marked done or failed, the system MUST emit a reflection event to the executing agent containing the original task, the bid, the execution trace, and any feedback.
- **FR-017**: Agents MUST be able to submit a proposed diff against their own capability manifest as a result of reflection; diffs MUST NOT be auto-applied to the live manifest.
- **FR-018**: A manifest diff proposal MUST require explicit approval before merging — by the human owner by default, or by a configurable auto-approve rule at channel or agent level.
- **FR-019**: The system MUST preserve every proposed diff and every approval or rejection decision so drift over time is auditable and reversible.
- **FR-020**: The human owner MUST be able to roll back a capability manifest to any prior revision regardless of how many intermediate changes have been applied.
- **FR-020a**: Each manifest revision MUST carry provenance metadata including the rolling success count and failure count of the agent in the affected domains at time of revision, so a reader can assess whether a given revision correlated with performance regression.
- **FR-020b**: The system MUST support automatic tombstoning of a manifest revision (marking it deprecated but retaining it for audit) when the owning agent's rolling failure rate in the revision's declared domains exceeds a configurable threshold over a configurable sliding window.
**Budget Enforcement**
- **FR-021**: The system MUST track cumulative token spend against the declared max budget for every awarded task in real time.
- **FR-022**: The system MUST emit a soft warning to the executing agent when cumulative spend reaches 80 percent of the declared budget.
- **FR-023**: The system MUST hard-stop execution and auto-fail the task when cumulative spend reaches 100 percent of the declared budget, preserving the partial execution trace for later inspection.
**Lemon Market Mitigation**
- **FR-024**: If an auction task receives zero bids by its declared deadline, the system MUST auto-escalate the task to the human owner via direct message and keep the task open for manual handling.
- **FR-025**: Auto-escalation behavior MUST be configurable per channel (enabled, disabled, or custom escalation target).
**Observability**
- **FR-026**: Every auction post, bid, award, claim, reflection event, proposed manifest diff, and merge decision MUST be traced in the existing SynapBus trace infrastructure and be searchable by time range, agent, and task.
- **FR-027**: The system MUST allow a human owner to answer the question "who did what and why" for any completed task by inspecting the trace without additional tooling.
### Key Entities
- **Capability Manifest**: A per-agent, versioned document describing domain expertise, representative example tasks, self-reported confidence per domain, and average token cost per domain. Owned by the agent it describes. Related to: agent identity, manifest revisions, reflection diffs.
- **Manifest Revision**: An immutable snapshot of a capability manifest at a point in time, produced by an update. Enables audit and rollback. Related to: capability manifest, reflection diff, approval record.
- **Auction Task**: A posted task with description, acceptance criteria, max token budget, deadline, and required domains. Has a lifecycle: open → (bids collected) → awarded → claimed → done/failed, or open → deadline → escalated. Related to: bids, claim record, reputation entries.
- **Bid**: A structured reply to an auction task containing estimated token cost, confidence, approach summary, and the revision of the bidder's manifest at time of bid. Related to: auction task, bidder agent, manifest revision.
- **Reputation Entry**: A completed-task tuple keyed by (agent, domain) with estimated cost, actual cost, success score, difficulty weight, and timestamp. Related to: agent, auction task.
- **Reflection Event**: A system-generated event fired when a task terminates, delivered to the executing agent and containing the full task context and outcome. Related to: auction task, executing agent, proposed manifest diff.
- **Manifest Diff Proposal**: A proposed change to a capability manifest produced by reflection. Pending until explicitly approved or rejected. Related to: reflection event, manifest revision, approval record.
- **Approval Record**: An immutable record of the decision (approve or reject) on a manifest diff proposal, with timestamp and the identity of the decider. Related to: manifest diff proposal.
## Success Criteria *(mandatory)*
### Measurable Outcomes
- **SC-001**: A human owner can post a task, receive bids, award one, and see the task complete without intervening in agent selection — verified by running a full auction lifecycle with at least two qualified agents bidding, where the owner's only actions are posting the task and issuing the award reaction.
- **SC-002**: Over a test run of at least 50 tasks with two agents of unequal skill, the higher-skilled agent wins approximately (1 minus exploration budget) × 100 percent of awards in the relevant domain. For a 10 percent exploration budget this means the higher-skilled agent should win within ±5 percentage points of 90 percent of the in-domain tasks.
- **SC-003**: A brand new agent joining the marketplace with no prior reputation can bid on and be awarded tasks in a previously unseen domain within its first 10 bids, via the bootstrap exploration credit.
- **SC-004**: For every completed auction task, a human owner can retrieve the full decision trail (who bid, at what estimated cost, who was awarded, why, actual cost, final outcome, any reflection diff) in a single query from the trace infrastructure.
- **SC-005**: After a manifest drift incident (three or more sequential approved diffs that collectively produce an unintended manifest state), the human owner can roll back to the pre-drift revision in a single action and all intermediate history is preserved.
- **SC-006**: Zero approved manifest diffs are applied without an explicit approval record. This is verified by auditing the approval records against applied diffs over a test run and finding a one-to-one match.
- **SC-007**: Of awarded tasks, at least 90 percent complete within their declared max token budget without triggering the hard-stop. Hard-stops exist but should be the exception, not the rule.
- **SC-008**: Of auction tasks that receive zero bids, 100 percent are escalated to the human owner within one minute of deadline expiry.
- **SC-009**: A human owner can inspect any agent's current capability manifest and full revision history with no more than two marketplace queries.
- **SC-010**: At least 95 percent of completed tasks produce a reputation ledger entry with all required fields populated.
## Assumptions
- Agents are already authenticated to SynapBus via the existing identity and API key mechanisms; the marketplace does not introduce new authentication.
- The existing wiki subsystem is the durable store for capability manifests and their revision history; this feature does not require a new revisioned document store.
- The existing claim/process/done lifecycle and reactive trigger infrastructure are reused to convert an awarded bid into executable work; no new lifecycle is invented for auctions.
- The existing trace infrastructure covers MCP tool invocations and reactive trigger runs and can be extended with new event types without structural change.
- Token counts are reported by the executing agent honestly; adversarial under-reporting is out of scope for this feature and belongs to a later trust-enforcement layer.
- Success scores for completed tasks are determined by the task poster (human or delegate) at done time; automated grading of task outputs is out of scope.
- Difficulty weights for reputation scoring are a per-channel constant or a simple function of max-budget magnitude; adaptive difficulty estimation is out of scope for this feature.
- A reasonable default exploration budget is 10 percent per channel. This is configurable and individual channels may set it higher or lower.
- A reasonable default bootstrap credit is three tasks per new (agent, domain) pair. This is configurable.
- Voting quorum for multi-owner awards is out of scope; awards are made by the single task poster by default.
## Out of Scope
- Automatic curriculum generation: selecting or composing what task should be posted next based on skill gaps or strategic goals.
- Cross-agent skill sharing: one agent teaching another agent a capability or transplanting manifest fragments.
- Monetary incentives beyond token budgets: real currency, credits, or cross-tenant billing.
- Multi-owner approval workflows and voting quorums for awards.
- Adversarial defenses against dishonest token reporting or reputation gaming beyond the exploration budget and bid-ratio visibility.
- Automated grading of task output quality; humans remain the source of truth for success scores in this release.
- Agent-management MCP tools (create, delete, configure agents) — excluded by standing design rule.