Algis Dumbris and Claude Opus 4.6
ff5d0c49f4
feat(018): dynamic agent spawning — primitives + doc-gardener demo
...
Ships the MVP slice of spec 018 (dynamic agent spawning):
- 5 new SQLite migrations (021-025): goals + goal_tasks + agent_proposals
+ reputation_evidence + secrets + harness_runs.task_id. The legacy
`tasks` table (channel auctions) and `agent_trust` table (reactions
workflow) are left untouched — the new schema coexists.
- 4 new internal packages, fully tested:
- internal/goals: Goal struct + store + service, slug collision dedup,
backing-channel auto-create via ChannelCreator adapter
- internal/goaltasks: goal_tasks table with denormalized 16 KB
ancestry snapshots, single-statement optimistic-lock atomic claim,
recursive-CTE cost rollup, state machine, per-billing-code rollup
- internal/secrets: NaCl-secretbox encrypted blobs, user/agent/task
scope precedence, sanitized env injection, master-key bootstrap
- internal/trust additions: ConfigHash (deterministic SHA-256 of
model + prompt + tools + skills + mcp + subagents, sorted),
DelegationCap (tier + tool-scope + budget + depth enforcement),
append-only Ledger with exponential time-decay rolling score and
70%-of-parent child seeding. Existing trust package unchanged.
- Critical invariants under test:
- 50-goroutine concurrent claim race → exactly one winner per round
- ConfigHash stable under shuffled array inputs, sensitive to
capability changes
- DelegationCap full tier × tool-scope matrix
- Ledger time-decay + parent seed at 70 % ± 1 %
- Secret name sanitization, scope precedence, plaintext never
returned via MCP-equivalent paths
- internal/agents/types.go extended with dynamic-spawning columns
(config_hash, parent_agent_id, spawn_depth, system_prompt,
autonomy_tier, tool_scope_json, quarantined_at). Existing tests
still pass.
- cmd/docgardener: self-contained demo binary driving the end-to-end
flow. `docgardener run` creates a goal, builds a task tree with
denormalized ancestry, spawns 3 specialists (each going through
real delegation-cap validation and config-hash computation and
70 %-of-parent reputation seeding), claims tasks atomically, runs
them through the state machine, records reputation evidence.
`docgardener report` queries all of that back out and renders a
rich dark-mode HTML report (header, spend metrics, task tree,
spawned-agent cards with reputation bars, cost breakdown, artifacts,
timeline).
- examples/doc-gardener: start.sh / run_task.sh / report.sh / stop.sh
mirroring the cold-topic-explainer pattern. Launches an isolated
synapbus instance on port 18089, drives the demo, renders
report.html, cleans up. Full README documenting what's real vs
deferred, plus examples/README.md listing both examples.
- specs/018: tasks.md updated with MVP completion status; legacy tasks
naming collision noted.
Deferred (marked explicitly in example README):
- Real LLM-driven coordinator (needs MCP tool wiring + prompt
iteration)
- Real subprocess runs (needs reactor integration with task_id on
ExecRequest)
- Full MCP tool surface (contracts are written at
specs/018-dynamic-agent-spawning/contracts/mcp-tools.md)
- Svelte /goals UI (REST endpoints remain a follow-up)
- Full budget race + quarantine auto-trigger wiring
- Full resource-request → secrets fulfill reaction-workflow path
Cross-compiles clean for linux/amd64 and darwin/arm64 with no CGO
(SC-010). All new package tests pass (SC-004, SC-005, SC-007).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-14 15:29:21 +03:00
Algis Dumbris and Claude Opus 4.6
8df22457ab
feat: trust scores, claim semantics, state-change webhooks (011-trust-claims-triggers)
...
Trust scores: per (agent, action_type) pair, stored in agent_trust
table. Auto-adjusts when human reacts to AI agent messages (approve
+0.05, reject -0.1). Scores clamped [0.0, 1.0]. MCP get_trust action
+ REST API /api/trust/{agent}. Web UI shows trust progress bars on
agent detail pages.
Claim semantics: only one in_progress reaction per message enforced.
First agent to claim wins, duplicates rejected with clear error.
State-change webhooks: StateChangeNotifier interface fires
workflow.state_changed events through existing webhook infrastructure
when reactions change a message's derived workflow state.
Channel thresholds: publish_threshold and approve_threshold fields
on channels for configuring autonomy gates.
Migration 014_trust_claims.sql. 17 new test cases across trust
model + store.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-03-18 21:41:54 +02:00