feat: initial SynapBus project scaffolding
Bootstrap the SynapBus project — a local-first, MCP-native agent-to-agent messaging service written in Go. Includes: - Project constitution (10 architectural principles) - 10 feature specs with implementation tasks: 001 Core Messaging, 002 Agent Registry, 003 Human Auth (OAuth 2.1), 004 Channels, 005 Web UI, 006 MCP Server, 007 Trace Logging, 008 Semantic Search, 009 Attachments, 010 Swarm Patterns - Go project scaffold (cobra CLI, SQLite schema, Makefile) - Speckit templates and commands Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,184 @@
|
||||
---
|
||||
description: Perform a non-destructive cross-artifact consistency and quality analysis across spec.md, plan.md, and tasks.md after task generation.
|
||||
---
|
||||
|
||||
## User Input
|
||||
|
||||
```text
|
||||
$ARGUMENTS
|
||||
```
|
||||
|
||||
You **MUST** consider the user input before proceeding (if not empty).
|
||||
|
||||
## Goal
|
||||
|
||||
Identify inconsistencies, duplications, ambiguities, and underspecified items across the three core artifacts (`spec.md`, `plan.md`, `tasks.md`) before implementation. This command MUST run only after `/speckit.tasks` has successfully produced a complete `tasks.md`.
|
||||
|
||||
## Operating Constraints
|
||||
|
||||
**STRICTLY READ-ONLY**: Do **not** modify any files. Output a structured analysis report. Offer an optional remediation plan (user must explicitly approve before any follow-up editing commands would be invoked manually).
|
||||
|
||||
**Constitution Authority**: The project constitution (`.specify/memory/constitution.md`) is **non-negotiable** within this analysis scope. Constitution conflicts are automatically CRITICAL and require adjustment of the spec, plan, or tasks—not dilution, reinterpretation, or silent ignoring of the principle. If a principle itself needs to change, that must occur in a separate, explicit constitution update outside `/speckit.analyze`.
|
||||
|
||||
## Execution Steps
|
||||
|
||||
### 1. Initialize Analysis Context
|
||||
|
||||
Run `.specify/scripts/bash/check-prerequisites.sh --json --require-tasks --include-tasks` once from repo root and parse JSON for FEATURE_DIR and AVAILABLE_DOCS. Derive absolute paths:
|
||||
|
||||
- SPEC = FEATURE_DIR/spec.md
|
||||
- PLAN = FEATURE_DIR/plan.md
|
||||
- TASKS = FEATURE_DIR/tasks.md
|
||||
|
||||
Abort with an error message if any required file is missing (instruct the user to run missing prerequisite command).
|
||||
For single quotes in args like "I'm Groot", use escape syntax: e.g 'I'\''m Groot' (or double-quote if possible: "I'm Groot").
|
||||
|
||||
### 2. Load Artifacts (Progressive Disclosure)
|
||||
|
||||
Load only the minimal necessary context from each artifact:
|
||||
|
||||
**From spec.md:**
|
||||
|
||||
- Overview/Context
|
||||
- Functional Requirements
|
||||
- Non-Functional Requirements
|
||||
- User Stories
|
||||
- Edge Cases (if present)
|
||||
|
||||
**From plan.md:**
|
||||
|
||||
- Architecture/stack choices
|
||||
- Data Model references
|
||||
- Phases
|
||||
- Technical constraints
|
||||
|
||||
**From tasks.md:**
|
||||
|
||||
- Task IDs
|
||||
- Descriptions
|
||||
- Phase grouping
|
||||
- Parallel markers [P]
|
||||
- Referenced file paths
|
||||
|
||||
**From constitution:**
|
||||
|
||||
- Load `.specify/memory/constitution.md` for principle validation
|
||||
|
||||
### 3. Build Semantic Models
|
||||
|
||||
Create internal representations (do not include raw artifacts in output):
|
||||
|
||||
- **Requirements inventory**: Each functional + non-functional requirement with a stable key (derive slug based on imperative phrase; e.g., "User can upload file" → `user-can-upload-file`)
|
||||
- **User story/action inventory**: Discrete user actions with acceptance criteria
|
||||
- **Task coverage mapping**: Map each task to one or more requirements or stories (inference by keyword / explicit reference patterns like IDs or key phrases)
|
||||
- **Constitution rule set**: Extract principle names and MUST/SHOULD normative statements
|
||||
|
||||
### 4. Detection Passes (Token-Efficient Analysis)
|
||||
|
||||
Focus on high-signal findings. Limit to 50 findings total; aggregate remainder in overflow summary.
|
||||
|
||||
#### A. Duplication Detection
|
||||
|
||||
- Identify near-duplicate requirements
|
||||
- Mark lower-quality phrasing for consolidation
|
||||
|
||||
#### B. Ambiguity Detection
|
||||
|
||||
- Flag vague adjectives (fast, scalable, secure, intuitive, robust) lacking measurable criteria
|
||||
- Flag unresolved placeholders (TODO, TKTK, ???, `<placeholder>`, etc.)
|
||||
|
||||
#### C. Underspecification
|
||||
|
||||
- Requirements with verbs but missing object or measurable outcome
|
||||
- User stories missing acceptance criteria alignment
|
||||
- Tasks referencing files or components not defined in spec/plan
|
||||
|
||||
#### D. Constitution Alignment
|
||||
|
||||
- Any requirement or plan element conflicting with a MUST principle
|
||||
- Missing mandated sections or quality gates from constitution
|
||||
|
||||
#### E. Coverage Gaps
|
||||
|
||||
- Requirements with zero associated tasks
|
||||
- Tasks with no mapped requirement/story
|
||||
- Non-functional requirements not reflected in tasks (e.g., performance, security)
|
||||
|
||||
#### F. Inconsistency
|
||||
|
||||
- Terminology drift (same concept named differently across files)
|
||||
- Data entities referenced in plan but absent in spec (or vice versa)
|
||||
- Task ordering contradictions (e.g., integration tasks before foundational setup tasks without dependency note)
|
||||
- Conflicting requirements (e.g., one requires Next.js while other specifies Vue)
|
||||
|
||||
### 5. Severity Assignment
|
||||
|
||||
Use this heuristic to prioritize findings:
|
||||
|
||||
- **CRITICAL**: Violates constitution MUST, missing core spec artifact, or requirement with zero coverage that blocks baseline functionality
|
||||
- **HIGH**: Duplicate or conflicting requirement, ambiguous security/performance attribute, untestable acceptance criterion
|
||||
- **MEDIUM**: Terminology drift, missing non-functional task coverage, underspecified edge case
|
||||
- **LOW**: Style/wording improvements, minor redundancy not affecting execution order
|
||||
|
||||
### 6. Produce Compact Analysis Report
|
||||
|
||||
Output a Markdown report (no file writes) with the following structure:
|
||||
|
||||
## Specification Analysis Report
|
||||
|
||||
| ID | Category | Severity | Location(s) | Summary | Recommendation |
|
||||
|----|----------|----------|-------------|---------|----------------|
|
||||
| A1 | Duplication | HIGH | spec.md:L120-134 | Two similar requirements ... | Merge phrasing; keep clearer version |
|
||||
|
||||
(Add one row per finding; generate stable IDs prefixed by category initial.)
|
||||
|
||||
**Coverage Summary Table:**
|
||||
|
||||
| Requirement Key | Has Task? | Task IDs | Notes |
|
||||
|-----------------|-----------|----------|-------|
|
||||
|
||||
**Constitution Alignment Issues:** (if any)
|
||||
|
||||
**Unmapped Tasks:** (if any)
|
||||
|
||||
**Metrics:**
|
||||
|
||||
- Total Requirements
|
||||
- Total Tasks
|
||||
- Coverage % (requirements with >=1 task)
|
||||
- Ambiguity Count
|
||||
- Duplication Count
|
||||
- Critical Issues Count
|
||||
|
||||
### 7. Provide Next Actions
|
||||
|
||||
At end of report, output a concise Next Actions block:
|
||||
|
||||
- If CRITICAL issues exist: Recommend resolving before `/speckit.implement`
|
||||
- If only LOW/MEDIUM: User may proceed, but provide improvement suggestions
|
||||
- Provide explicit command suggestions: e.g., "Run /speckit.specify with refinement", "Run /speckit.plan to adjust architecture", "Manually edit tasks.md to add coverage for 'performance-metrics'"
|
||||
|
||||
### 8. Offer Remediation
|
||||
|
||||
Ask the user: "Would you like me to suggest concrete remediation edits for the top N issues?" (Do NOT apply them automatically.)
|
||||
|
||||
## Operating Principles
|
||||
|
||||
### Context Efficiency
|
||||
|
||||
- **Minimal high-signal tokens**: Focus on actionable findings, not exhaustive documentation
|
||||
- **Progressive disclosure**: Load artifacts incrementally; don't dump all content into analysis
|
||||
- **Token-efficient output**: Limit findings table to 50 rows; summarize overflow
|
||||
- **Deterministic results**: Rerunning without changes should produce consistent IDs and counts
|
||||
|
||||
### Analysis Guidelines
|
||||
|
||||
- **NEVER modify files** (this is read-only analysis)
|
||||
- **NEVER hallucinate missing sections** (if absent, report them accurately)
|
||||
- **Prioritize constitution violations** (these are always CRITICAL)
|
||||
- **Use examples over exhaustive rules** (cite specific instances, not generic patterns)
|
||||
- **Report zero issues gracefully** (emit success report with coverage statistics)
|
||||
|
||||
## Context
|
||||
|
||||
$ARGUMENTS
|
||||
@@ -0,0 +1,295 @@
|
||||
---
|
||||
description: Generate a custom checklist for the current feature based on user requirements.
|
||||
---
|
||||
|
||||
## Checklist Purpose: "Unit Tests for English"
|
||||
|
||||
**CRITICAL CONCEPT**: Checklists are **UNIT TESTS FOR REQUIREMENTS WRITING** - they validate the quality, clarity, and completeness of requirements in a given domain.
|
||||
|
||||
**NOT for verification/testing**:
|
||||
|
||||
- ❌ NOT "Verify the button clicks correctly"
|
||||
- ❌ NOT "Test error handling works"
|
||||
- ❌ NOT "Confirm the API returns 200"
|
||||
- ❌ NOT checking if code/implementation matches the spec
|
||||
|
||||
**FOR requirements quality validation**:
|
||||
|
||||
- ✅ "Are visual hierarchy requirements defined for all card types?" (completeness)
|
||||
- ✅ "Is 'prominent display' quantified with specific sizing/positioning?" (clarity)
|
||||
- ✅ "Are hover state requirements consistent across all interactive elements?" (consistency)
|
||||
- ✅ "Are accessibility requirements defined for keyboard navigation?" (coverage)
|
||||
- ✅ "Does the spec define what happens when logo image fails to load?" (edge cases)
|
||||
|
||||
**Metaphor**: If your spec is code written in English, the checklist is its unit test suite. You're testing whether the requirements are well-written, complete, unambiguous, and ready for implementation - NOT whether the implementation works.
|
||||
|
||||
## User Input
|
||||
|
||||
```text
|
||||
$ARGUMENTS
|
||||
```
|
||||
|
||||
You **MUST** consider the user input before proceeding (if not empty).
|
||||
|
||||
## Execution Steps
|
||||
|
||||
1. **Setup**: Run `.specify/scripts/bash/check-prerequisites.sh --json` from repo root and parse JSON for FEATURE_DIR and AVAILABLE_DOCS list.
|
||||
- All file paths must be absolute.
|
||||
- For single quotes in args like "I'm Groot", use escape syntax: e.g 'I'\''m Groot' (or double-quote if possible: "I'm Groot").
|
||||
|
||||
2. **Clarify intent (dynamic)**: Derive up to THREE initial contextual clarifying questions (no pre-baked catalog). They MUST:
|
||||
- Be generated from the user's phrasing + extracted signals from spec/plan/tasks
|
||||
- Only ask about information that materially changes checklist content
|
||||
- Be skipped individually if already unambiguous in `$ARGUMENTS`
|
||||
- Prefer precision over breadth
|
||||
|
||||
Generation algorithm:
|
||||
1. Extract signals: feature domain keywords (e.g., auth, latency, UX, API), risk indicators ("critical", "must", "compliance"), stakeholder hints ("QA", "review", "security team"), and explicit deliverables ("a11y", "rollback", "contracts").
|
||||
2. Cluster signals into candidate focus areas (max 4) ranked by relevance.
|
||||
3. Identify probable audience & timing (author, reviewer, QA, release) if not explicit.
|
||||
4. Detect missing dimensions: scope breadth, depth/rigor, risk emphasis, exclusion boundaries, measurable acceptance criteria.
|
||||
5. Formulate questions chosen from these archetypes:
|
||||
- Scope refinement (e.g., "Should this include integration touchpoints with X and Y or stay limited to local module correctness?")
|
||||
- Risk prioritization (e.g., "Which of these potential risk areas should receive mandatory gating checks?")
|
||||
- Depth calibration (e.g., "Is this a lightweight pre-commit sanity list or a formal release gate?")
|
||||
- Audience framing (e.g., "Will this be used by the author only or peers during PR review?")
|
||||
- Boundary exclusion (e.g., "Should we explicitly exclude performance tuning items this round?")
|
||||
- Scenario class gap (e.g., "No recovery flows detected—are rollback / partial failure paths in scope?")
|
||||
|
||||
Question formatting rules:
|
||||
- If presenting options, generate a compact table with columns: Option | Candidate | Why It Matters
|
||||
- Limit to A–E options maximum; omit table if a free-form answer is clearer
|
||||
- Never ask the user to restate what they already said
|
||||
- Avoid speculative categories (no hallucination). If uncertain, ask explicitly: "Confirm whether X belongs in scope."
|
||||
|
||||
Defaults when interaction impossible:
|
||||
- Depth: Standard
|
||||
- Audience: Reviewer (PR) if code-related; Author otherwise
|
||||
- Focus: Top 2 relevance clusters
|
||||
|
||||
Output the questions (label Q1/Q2/Q3). After answers: if ≥2 scenario classes (Alternate / Exception / Recovery / Non-Functional domain) remain unclear, you MAY ask up to TWO more targeted follow‑ups (Q4/Q5) with a one-line justification each (e.g., "Unresolved recovery path risk"). Do not exceed five total questions. Skip escalation if user explicitly declines more.
|
||||
|
||||
3. **Understand user request**: Combine `$ARGUMENTS` + clarifying answers:
|
||||
- Derive checklist theme (e.g., security, review, deploy, ux)
|
||||
- Consolidate explicit must-have items mentioned by user
|
||||
- Map focus selections to category scaffolding
|
||||
- Infer any missing context from spec/plan/tasks (do NOT hallucinate)
|
||||
|
||||
4. **Load feature context**: Read from FEATURE_DIR:
|
||||
- spec.md: Feature requirements and scope
|
||||
- plan.md (if exists): Technical details, dependencies
|
||||
- tasks.md (if exists): Implementation tasks
|
||||
|
||||
**Context Loading Strategy**:
|
||||
- Load only necessary portions relevant to active focus areas (avoid full-file dumping)
|
||||
- Prefer summarizing long sections into concise scenario/requirement bullets
|
||||
- Use progressive disclosure: add follow-on retrieval only if gaps detected
|
||||
- If source docs are large, generate interim summary items instead of embedding raw text
|
||||
|
||||
5. **Generate checklist** - Create "Unit Tests for Requirements":
|
||||
- Create `FEATURE_DIR/checklists/` directory if it doesn't exist
|
||||
- Generate unique checklist filename:
|
||||
- Use short, descriptive name based on domain (e.g., `ux.md`, `api.md`, `security.md`)
|
||||
- Format: `[domain].md`
|
||||
- File handling behavior:
|
||||
- If file does NOT exist: Create new file and number items starting from CHK001
|
||||
- If file exists: Append new items to existing file, continuing from the last CHK ID (e.g., if last item is CHK015, start new items at CHK016)
|
||||
- Never delete or replace existing checklist content - always preserve and append
|
||||
|
||||
**CORE PRINCIPLE - Test the Requirements, Not the Implementation**:
|
||||
Every checklist item MUST evaluate the REQUIREMENTS THEMSELVES for:
|
||||
- **Completeness**: Are all necessary requirements present?
|
||||
- **Clarity**: Are requirements unambiguous and specific?
|
||||
- **Consistency**: Do requirements align with each other?
|
||||
- **Measurability**: Can requirements be objectively verified?
|
||||
- **Coverage**: Are all scenarios/edge cases addressed?
|
||||
|
||||
**Category Structure** - Group items by requirement quality dimensions:
|
||||
- **Requirement Completeness** (Are all necessary requirements documented?)
|
||||
- **Requirement Clarity** (Are requirements specific and unambiguous?)
|
||||
- **Requirement Consistency** (Do requirements align without conflicts?)
|
||||
- **Acceptance Criteria Quality** (Are success criteria measurable?)
|
||||
- **Scenario Coverage** (Are all flows/cases addressed?)
|
||||
- **Edge Case Coverage** (Are boundary conditions defined?)
|
||||
- **Non-Functional Requirements** (Performance, Security, Accessibility, etc. - are they specified?)
|
||||
- **Dependencies & Assumptions** (Are they documented and validated?)
|
||||
- **Ambiguities & Conflicts** (What needs clarification?)
|
||||
|
||||
**HOW TO WRITE CHECKLIST ITEMS - "Unit Tests for English"**:
|
||||
|
||||
❌ **WRONG** (Testing implementation):
|
||||
- "Verify landing page displays 3 episode cards"
|
||||
- "Test hover states work on desktop"
|
||||
- "Confirm logo click navigates home"
|
||||
|
||||
✅ **CORRECT** (Testing requirements quality):
|
||||
- "Are the exact number and layout of featured episodes specified?" [Completeness]
|
||||
- "Is 'prominent display' quantified with specific sizing/positioning?" [Clarity]
|
||||
- "Are hover state requirements consistent across all interactive elements?" [Consistency]
|
||||
- "Are keyboard navigation requirements defined for all interactive UI?" [Coverage]
|
||||
- "Is the fallback behavior specified when logo image fails to load?" [Edge Cases]
|
||||
- "Are loading states defined for asynchronous episode data?" [Completeness]
|
||||
- "Does the spec define visual hierarchy for competing UI elements?" [Clarity]
|
||||
|
||||
**ITEM STRUCTURE**:
|
||||
Each item should follow this pattern:
|
||||
- Question format asking about requirement quality
|
||||
- Focus on what's WRITTEN (or not written) in the spec/plan
|
||||
- Include quality dimension in brackets [Completeness/Clarity/Consistency/etc.]
|
||||
- Reference spec section `[Spec §X.Y]` when checking existing requirements
|
||||
- Use `[Gap]` marker when checking for missing requirements
|
||||
|
||||
**EXAMPLES BY QUALITY DIMENSION**:
|
||||
|
||||
Completeness:
|
||||
- "Are error handling requirements defined for all API failure modes? [Gap]"
|
||||
- "Are accessibility requirements specified for all interactive elements? [Completeness]"
|
||||
- "Are mobile breakpoint requirements defined for responsive layouts? [Gap]"
|
||||
|
||||
Clarity:
|
||||
- "Is 'fast loading' quantified with specific timing thresholds? [Clarity, Spec §NFR-2]"
|
||||
- "Are 'related episodes' selection criteria explicitly defined? [Clarity, Spec §FR-5]"
|
||||
- "Is 'prominent' defined with measurable visual properties? [Ambiguity, Spec §FR-4]"
|
||||
|
||||
Consistency:
|
||||
- "Do navigation requirements align across all pages? [Consistency, Spec §FR-10]"
|
||||
- "Are card component requirements consistent between landing and detail pages? [Consistency]"
|
||||
|
||||
Coverage:
|
||||
- "Are requirements defined for zero-state scenarios (no episodes)? [Coverage, Edge Case]"
|
||||
- "Are concurrent user interaction scenarios addressed? [Coverage, Gap]"
|
||||
- "Are requirements specified for partial data loading failures? [Coverage, Exception Flow]"
|
||||
|
||||
Measurability:
|
||||
- "Are visual hierarchy requirements measurable/testable? [Acceptance Criteria, Spec §FR-1]"
|
||||
- "Can 'balanced visual weight' be objectively verified? [Measurability, Spec §FR-2]"
|
||||
|
||||
**Scenario Classification & Coverage** (Requirements Quality Focus):
|
||||
- Check if requirements exist for: Primary, Alternate, Exception/Error, Recovery, Non-Functional scenarios
|
||||
- For each scenario class, ask: "Are [scenario type] requirements complete, clear, and consistent?"
|
||||
- If scenario class missing: "Are [scenario type] requirements intentionally excluded or missing? [Gap]"
|
||||
- Include resilience/rollback when state mutation occurs: "Are rollback requirements defined for migration failures? [Gap]"
|
||||
|
||||
**Traceability Requirements**:
|
||||
- MINIMUM: ≥80% of items MUST include at least one traceability reference
|
||||
- Each item should reference: spec section `[Spec §X.Y]`, or use markers: `[Gap]`, `[Ambiguity]`, `[Conflict]`, `[Assumption]`
|
||||
- If no ID system exists: "Is a requirement & acceptance criteria ID scheme established? [Traceability]"
|
||||
|
||||
**Surface & Resolve Issues** (Requirements Quality Problems):
|
||||
Ask questions about the requirements themselves:
|
||||
- Ambiguities: "Is the term 'fast' quantified with specific metrics? [Ambiguity, Spec §NFR-1]"
|
||||
- Conflicts: "Do navigation requirements conflict between §FR-10 and §FR-10a? [Conflict]"
|
||||
- Assumptions: "Is the assumption of 'always available podcast API' validated? [Assumption]"
|
||||
- Dependencies: "Are external podcast API requirements documented? [Dependency, Gap]"
|
||||
- Missing definitions: "Is 'visual hierarchy' defined with measurable criteria? [Gap]"
|
||||
|
||||
**Content Consolidation**:
|
||||
- Soft cap: If raw candidate items > 40, prioritize by risk/impact
|
||||
- Merge near-duplicates checking the same requirement aspect
|
||||
- If >5 low-impact edge cases, create one item: "Are edge cases X, Y, Z addressed in requirements? [Coverage]"
|
||||
|
||||
**🚫 ABSOLUTELY PROHIBITED** - These make it an implementation test, not a requirements test:
|
||||
- ❌ Any item starting with "Verify", "Test", "Confirm", "Check" + implementation behavior
|
||||
- ❌ References to code execution, user actions, system behavior
|
||||
- ❌ "Displays correctly", "works properly", "functions as expected"
|
||||
- ❌ "Click", "navigate", "render", "load", "execute"
|
||||
- ❌ Test cases, test plans, QA procedures
|
||||
- ❌ Implementation details (frameworks, APIs, algorithms)
|
||||
|
||||
**✅ REQUIRED PATTERNS** - These test requirements quality:
|
||||
- ✅ "Are [requirement type] defined/specified/documented for [scenario]?"
|
||||
- ✅ "Is [vague term] quantified/clarified with specific criteria?"
|
||||
- ✅ "Are requirements consistent between [section A] and [section B]?"
|
||||
- ✅ "Can [requirement] be objectively measured/verified?"
|
||||
- ✅ "Are [edge cases/scenarios] addressed in requirements?"
|
||||
- ✅ "Does the spec define [missing aspect]?"
|
||||
|
||||
6. **Structure Reference**: Generate the checklist following the canonical template in `.specify/templates/checklist-template.md` for title, meta section, category headings, and ID formatting. If template is unavailable, use: H1 title, purpose/created meta lines, `##` category sections containing `- [ ] CHK### <requirement item>` lines with globally incrementing IDs starting at CHK001.
|
||||
|
||||
7. **Report**: Output full path to checklist file, item count, and summarize whether the run created a new file or appended to an existing one. Summarize:
|
||||
- Focus areas selected
|
||||
- Depth level
|
||||
- Actor/timing
|
||||
- Any explicit user-specified must-have items incorporated
|
||||
|
||||
**Important**: Each `/speckit.checklist` command invocation uses a short, descriptive checklist filename and either creates a new file or appends to an existing one. This allows:
|
||||
|
||||
- Multiple checklists of different types (e.g., `ux.md`, `test.md`, `security.md`)
|
||||
- Simple, memorable filenames that indicate checklist purpose
|
||||
- Easy identification and navigation in the `checklists/` folder
|
||||
|
||||
To avoid clutter, use descriptive types and clean up obsolete checklists when done.
|
||||
|
||||
## Example Checklist Types & Sample Items
|
||||
|
||||
**UX Requirements Quality:** `ux.md`
|
||||
|
||||
Sample items (testing the requirements, NOT the implementation):
|
||||
|
||||
- "Are visual hierarchy requirements defined with measurable criteria? [Clarity, Spec §FR-1]"
|
||||
- "Is the number and positioning of UI elements explicitly specified? [Completeness, Spec §FR-1]"
|
||||
- "Are interaction state requirements (hover, focus, active) consistently defined? [Consistency]"
|
||||
- "Are accessibility requirements specified for all interactive elements? [Coverage, Gap]"
|
||||
- "Is fallback behavior defined when images fail to load? [Edge Case, Gap]"
|
||||
- "Can 'prominent display' be objectively measured? [Measurability, Spec §FR-4]"
|
||||
|
||||
**API Requirements Quality:** `api.md`
|
||||
|
||||
Sample items:
|
||||
|
||||
- "Are error response formats specified for all failure scenarios? [Completeness]"
|
||||
- "Are rate limiting requirements quantified with specific thresholds? [Clarity]"
|
||||
- "Are authentication requirements consistent across all endpoints? [Consistency]"
|
||||
- "Are retry/timeout requirements defined for external dependencies? [Coverage, Gap]"
|
||||
- "Is versioning strategy documented in requirements? [Gap]"
|
||||
|
||||
**Performance Requirements Quality:** `performance.md`
|
||||
|
||||
Sample items:
|
||||
|
||||
- "Are performance requirements quantified with specific metrics? [Clarity]"
|
||||
- "Are performance targets defined for all critical user journeys? [Coverage]"
|
||||
- "Are performance requirements under different load conditions specified? [Completeness]"
|
||||
- "Can performance requirements be objectively measured? [Measurability]"
|
||||
- "Are degradation requirements defined for high-load scenarios? [Edge Case, Gap]"
|
||||
|
||||
**Security Requirements Quality:** `security.md`
|
||||
|
||||
Sample items:
|
||||
|
||||
- "Are authentication requirements specified for all protected resources? [Coverage]"
|
||||
- "Are data protection requirements defined for sensitive information? [Completeness]"
|
||||
- "Is the threat model documented and requirements aligned to it? [Traceability]"
|
||||
- "Are security requirements consistent with compliance obligations? [Consistency]"
|
||||
- "Are security failure/breach response requirements defined? [Gap, Exception Flow]"
|
||||
|
||||
## Anti-Examples: What NOT To Do
|
||||
|
||||
**❌ WRONG - These test implementation, not requirements:**
|
||||
|
||||
```markdown
|
||||
- [ ] CHK001 - Verify landing page displays 3 episode cards [Spec §FR-001]
|
||||
- [ ] CHK002 - Test hover states work correctly on desktop [Spec §FR-003]
|
||||
- [ ] CHK003 - Confirm logo click navigates to home page [Spec §FR-010]
|
||||
- [ ] CHK004 - Check that related episodes section shows 3-5 items [Spec §FR-005]
|
||||
```
|
||||
|
||||
**✅ CORRECT - These test requirements quality:**
|
||||
|
||||
```markdown
|
||||
- [ ] CHK001 - Are the number and layout of featured episodes explicitly specified? [Completeness, Spec §FR-001]
|
||||
- [ ] CHK002 - Are hover state requirements consistently defined for all interactive elements? [Consistency, Spec §FR-003]
|
||||
- [ ] CHK003 - Are navigation requirements clear for all clickable brand elements? [Clarity, Spec §FR-010]
|
||||
- [ ] CHK004 - Is the selection criteria for related episodes documented? [Gap, Spec §FR-005]
|
||||
- [ ] CHK005 - Are loading state requirements defined for asynchronous episode data? [Gap]
|
||||
- [ ] CHK006 - Can "visual hierarchy" requirements be objectively measured? [Measurability, Spec §FR-001]
|
||||
```
|
||||
|
||||
**Key Differences:**
|
||||
|
||||
- Wrong: Tests if the system works correctly
|
||||
- Correct: Tests if the requirements are written correctly
|
||||
- Wrong: Verification of behavior
|
||||
- Correct: Validation of requirement quality
|
||||
- Wrong: "Does it do X?"
|
||||
- Correct: "Is X clearly specified?"
|
||||
@@ -0,0 +1,181 @@
|
||||
---
|
||||
description: Identify underspecified areas in the current feature spec by asking up to 5 highly targeted clarification questions and encoding answers back into the spec.
|
||||
handoffs:
|
||||
- label: Build Technical Plan
|
||||
agent: speckit.plan
|
||||
prompt: Create a plan for the spec. I am building with...
|
||||
---
|
||||
|
||||
## User Input
|
||||
|
||||
```text
|
||||
$ARGUMENTS
|
||||
```
|
||||
|
||||
You **MUST** consider the user input before proceeding (if not empty).
|
||||
|
||||
## Outline
|
||||
|
||||
Goal: Detect and reduce ambiguity or missing decision points in the active feature specification and record the clarifications directly in the spec file.
|
||||
|
||||
Note: This clarification workflow is expected to run (and be completed) BEFORE invoking `/speckit.plan`. If the user explicitly states they are skipping clarification (e.g., exploratory spike), you may proceed, but must warn that downstream rework risk increases.
|
||||
|
||||
Execution steps:
|
||||
|
||||
1. Run `.specify/scripts/bash/check-prerequisites.sh --json --paths-only` from repo root **once** (combined `--json --paths-only` mode / `-Json -PathsOnly`). Parse minimal JSON payload fields:
|
||||
- `FEATURE_DIR`
|
||||
- `FEATURE_SPEC`
|
||||
- (Optionally capture `IMPL_PLAN`, `TASKS` for future chained flows.)
|
||||
- If JSON parsing fails, abort and instruct user to re-run `/speckit.specify` or verify feature branch environment.
|
||||
- For single quotes in args like "I'm Groot", use escape syntax: e.g 'I'\''m Groot' (or double-quote if possible: "I'm Groot").
|
||||
|
||||
2. Load the current spec file. Perform a structured ambiguity & coverage scan using this taxonomy. For each category, mark status: Clear / Partial / Missing. Produce an internal coverage map used for prioritization (do not output raw map unless no questions will be asked).
|
||||
|
||||
Functional Scope & Behavior:
|
||||
- Core user goals & success criteria
|
||||
- Explicit out-of-scope declarations
|
||||
- User roles / personas differentiation
|
||||
|
||||
Domain & Data Model:
|
||||
- Entities, attributes, relationships
|
||||
- Identity & uniqueness rules
|
||||
- Lifecycle/state transitions
|
||||
- Data volume / scale assumptions
|
||||
|
||||
Interaction & UX Flow:
|
||||
- Critical user journeys / sequences
|
||||
- Error/empty/loading states
|
||||
- Accessibility or localization notes
|
||||
|
||||
Non-Functional Quality Attributes:
|
||||
- Performance (latency, throughput targets)
|
||||
- Scalability (horizontal/vertical, limits)
|
||||
- Reliability & availability (uptime, recovery expectations)
|
||||
- Observability (logging, metrics, tracing signals)
|
||||
- Security & privacy (authN/Z, data protection, threat assumptions)
|
||||
- Compliance / regulatory constraints (if any)
|
||||
|
||||
Integration & External Dependencies:
|
||||
- External services/APIs and failure modes
|
||||
- Data import/export formats
|
||||
- Protocol/versioning assumptions
|
||||
|
||||
Edge Cases & Failure Handling:
|
||||
- Negative scenarios
|
||||
- Rate limiting / throttling
|
||||
- Conflict resolution (e.g., concurrent edits)
|
||||
|
||||
Constraints & Tradeoffs:
|
||||
- Technical constraints (language, storage, hosting)
|
||||
- Explicit tradeoffs or rejected alternatives
|
||||
|
||||
Terminology & Consistency:
|
||||
- Canonical glossary terms
|
||||
- Avoided synonyms / deprecated terms
|
||||
|
||||
Completion Signals:
|
||||
- Acceptance criteria testability
|
||||
- Measurable Definition of Done style indicators
|
||||
|
||||
Misc / Placeholders:
|
||||
- TODO markers / unresolved decisions
|
||||
- Ambiguous adjectives ("robust", "intuitive") lacking quantification
|
||||
|
||||
For each category with Partial or Missing status, add a candidate question opportunity unless:
|
||||
- Clarification would not materially change implementation or validation strategy
|
||||
- Information is better deferred to planning phase (note internally)
|
||||
|
||||
3. Generate (internally) a prioritized queue of candidate clarification questions (maximum 5). Do NOT output them all at once. Apply these constraints:
|
||||
- Maximum of 5 total questions across the whole session.
|
||||
- Each question must be answerable with EITHER:
|
||||
- A short multiple‑choice selection (2–5 distinct, mutually exclusive options), OR
|
||||
- A one-word / short‑phrase answer (explicitly constrain: "Answer in <=5 words").
|
||||
- Only include questions whose answers materially impact architecture, data modeling, task decomposition, test design, UX behavior, operational readiness, or compliance validation.
|
||||
- Ensure category coverage balance: attempt to cover the highest impact unresolved categories first; avoid asking two low-impact questions when a single high-impact area (e.g., security posture) is unresolved.
|
||||
- Exclude questions already answered, trivial stylistic preferences, or plan-level execution details (unless blocking correctness).
|
||||
- Favor clarifications that reduce downstream rework risk or prevent misaligned acceptance tests.
|
||||
- If more than 5 categories remain unresolved, select the top 5 by (Impact * Uncertainty) heuristic.
|
||||
|
||||
4. Sequential questioning loop (interactive):
|
||||
- Present EXACTLY ONE question at a time.
|
||||
- For multiple‑choice questions:
|
||||
- **Analyze all options** and determine the **most suitable option** based on:
|
||||
- Best practices for the project type
|
||||
- Common patterns in similar implementations
|
||||
- Risk reduction (security, performance, maintainability)
|
||||
- Alignment with any explicit project goals or constraints visible in the spec
|
||||
- Present your **recommended option prominently** at the top with clear reasoning (1-2 sentences explaining why this is the best choice).
|
||||
- Format as: `**Recommended:** Option [X] - <reasoning>`
|
||||
- Then render all options as a Markdown table:
|
||||
|
||||
| Option | Description |
|
||||
|--------|-------------|
|
||||
| A | <Option A description> |
|
||||
| B | <Option B description> |
|
||||
| C | <Option C description> (add D/E as needed up to 5) |
|
||||
| Short | Provide a different short answer (<=5 words) (Include only if free-form alternative is appropriate) |
|
||||
|
||||
- After the table, add: `You can reply with the option letter (e.g., "A"), accept the recommendation by saying "yes" or "recommended", or provide your own short answer.`
|
||||
- For short‑answer style (no meaningful discrete options):
|
||||
- Provide your **suggested answer** based on best practices and context.
|
||||
- Format as: `**Suggested:** <your proposed answer> - <brief reasoning>`
|
||||
- Then output: `Format: Short answer (<=5 words). You can accept the suggestion by saying "yes" or "suggested", or provide your own answer.`
|
||||
- After the user answers:
|
||||
- If the user replies with "yes", "recommended", or "suggested", use your previously stated recommendation/suggestion as the answer.
|
||||
- Otherwise, validate the answer maps to one option or fits the <=5 word constraint.
|
||||
- If ambiguous, ask for a quick disambiguation (count still belongs to same question; do not advance).
|
||||
- Once satisfactory, record it in working memory (do not yet write to disk) and move to the next queued question.
|
||||
- Stop asking further questions when:
|
||||
- All critical ambiguities resolved early (remaining queued items become unnecessary), OR
|
||||
- User signals completion ("done", "good", "no more"), OR
|
||||
- You reach 5 asked questions.
|
||||
- Never reveal future queued questions in advance.
|
||||
- If no valid questions exist at start, immediately report no critical ambiguities.
|
||||
|
||||
5. Integration after EACH accepted answer (incremental update approach):
|
||||
- Maintain in-memory representation of the spec (loaded once at start) plus the raw file contents.
|
||||
- For the first integrated answer in this session:
|
||||
- Ensure a `## Clarifications` section exists (create it just after the highest-level contextual/overview section per the spec template if missing).
|
||||
- Under it, create (if not present) a `### Session YYYY-MM-DD` subheading for today.
|
||||
- Append a bullet line immediately after acceptance: `- Q: <question> → A: <final answer>`.
|
||||
- Then immediately apply the clarification to the most appropriate section(s):
|
||||
- Functional ambiguity → Update or add a bullet in Functional Requirements.
|
||||
- User interaction / actor distinction → Update User Stories or Actors subsection (if present) with clarified role, constraint, or scenario.
|
||||
- Data shape / entities → Update Data Model (add fields, types, relationships) preserving ordering; note added constraints succinctly.
|
||||
- Non-functional constraint → Add/modify measurable criteria in Non-Functional / Quality Attributes section (convert vague adjective to metric or explicit target).
|
||||
- Edge case / negative flow → Add a new bullet under Edge Cases / Error Handling (or create such subsection if template provides placeholder for it).
|
||||
- Terminology conflict → Normalize term across spec; retain original only if necessary by adding `(formerly referred to as "X")` once.
|
||||
- If the clarification invalidates an earlier ambiguous statement, replace that statement instead of duplicating; leave no obsolete contradictory text.
|
||||
- Save the spec file AFTER each integration to minimize risk of context loss (atomic overwrite).
|
||||
- Preserve formatting: do not reorder unrelated sections; keep heading hierarchy intact.
|
||||
- Keep each inserted clarification minimal and testable (avoid narrative drift).
|
||||
|
||||
6. Validation (performed after EACH write plus final pass):
|
||||
- Clarifications session contains exactly one bullet per accepted answer (no duplicates).
|
||||
- Total asked (accepted) questions ≤ 5.
|
||||
- Updated sections contain no lingering vague placeholders the new answer was meant to resolve.
|
||||
- No contradictory earlier statement remains (scan for now-invalid alternative choices removed).
|
||||
- Markdown structure valid; only allowed new headings: `## Clarifications`, `### Session YYYY-MM-DD`.
|
||||
- Terminology consistency: same canonical term used across all updated sections.
|
||||
|
||||
7. Write the updated spec back to `FEATURE_SPEC`.
|
||||
|
||||
8. Report completion (after questioning loop ends or early termination):
|
||||
- Number of questions asked & answered.
|
||||
- Path to updated spec.
|
||||
- Sections touched (list names).
|
||||
- Coverage summary table listing each taxonomy category with Status: Resolved (was Partial/Missing and addressed), Deferred (exceeds question quota or better suited for planning), Clear (already sufficient), Outstanding (still Partial/Missing but low impact).
|
||||
- If any Outstanding or Deferred remain, recommend whether to proceed to `/speckit.plan` or run `/speckit.clarify` again later post-plan.
|
||||
- Suggested next command.
|
||||
|
||||
Behavior rules:
|
||||
|
||||
- If no meaningful ambiguities found (or all potential questions would be low-impact), respond: "No critical ambiguities detected worth formal clarification." and suggest proceeding.
|
||||
- If spec file missing, instruct user to run `/speckit.specify` first (do not create a new spec here).
|
||||
- Never exceed 5 total asked questions (clarification retries for a single question do not count as new questions).
|
||||
- Avoid speculative tech stack questions unless the absence blocks functional clarity.
|
||||
- Respect user early termination signals ("stop", "done", "proceed").
|
||||
- If no questions asked due to full coverage, output a compact coverage summary (all categories Clear) then suggest advancing.
|
||||
- If quota reached with unresolved high-impact categories remaining, explicitly flag them under Deferred with rationale.
|
||||
|
||||
Context for prioritization: $ARGUMENTS
|
||||
@@ -0,0 +1,84 @@
|
||||
---
|
||||
description: Create or update the project constitution from interactive or provided principle inputs, ensuring all dependent templates stay in sync.
|
||||
handoffs:
|
||||
- label: Build Specification
|
||||
agent: speckit.specify
|
||||
prompt: Implement the feature specification based on the updated constitution. I want to build...
|
||||
---
|
||||
|
||||
## User Input
|
||||
|
||||
```text
|
||||
$ARGUMENTS
|
||||
```
|
||||
|
||||
You **MUST** consider the user input before proceeding (if not empty).
|
||||
|
||||
## Outline
|
||||
|
||||
You are updating the project constitution at `.specify/memory/constitution.md`. This file is a TEMPLATE containing placeholder tokens in square brackets (e.g. `[PROJECT_NAME]`, `[PRINCIPLE_1_NAME]`). Your job is to (a) collect/derive concrete values, (b) fill the template precisely, and (c) propagate any amendments across dependent artifacts.
|
||||
|
||||
**Note**: If `.specify/memory/constitution.md` does not exist yet, it should have been initialized from `.specify/templates/constitution-template.md` during project setup. If it's missing, copy the template first.
|
||||
|
||||
Follow this execution flow:
|
||||
|
||||
1. Load the existing constitution at `.specify/memory/constitution.md`.
|
||||
- Identify every placeholder token of the form `[ALL_CAPS_IDENTIFIER]`.
|
||||
**IMPORTANT**: The user might require less or more principles than the ones used in the template. If a number is specified, respect that - follow the general template. You will update the doc accordingly.
|
||||
|
||||
2. Collect/derive values for placeholders:
|
||||
- If user input (conversation) supplies a value, use it.
|
||||
- Otherwise infer from existing repo context (README, docs, prior constitution versions if embedded).
|
||||
- For governance dates: `RATIFICATION_DATE` is the original adoption date (if unknown ask or mark TODO), `LAST_AMENDED_DATE` is today if changes are made, otherwise keep previous.
|
||||
- `CONSTITUTION_VERSION` must increment according to semantic versioning rules:
|
||||
- MAJOR: Backward incompatible governance/principle removals or redefinitions.
|
||||
- MINOR: New principle/section added or materially expanded guidance.
|
||||
- PATCH: Clarifications, wording, typo fixes, non-semantic refinements.
|
||||
- If version bump type ambiguous, propose reasoning before finalizing.
|
||||
|
||||
3. Draft the updated constitution content:
|
||||
- Replace every placeholder with concrete text (no bracketed tokens left except intentionally retained template slots that the project has chosen not to define yet—explicitly justify any left).
|
||||
- Preserve heading hierarchy and comments can be removed once replaced unless they still add clarifying guidance.
|
||||
- Ensure each Principle section: succinct name line, paragraph (or bullet list) capturing non‑negotiable rules, explicit rationale if not obvious.
|
||||
- Ensure Governance section lists amendment procedure, versioning policy, and compliance review expectations.
|
||||
|
||||
4. Consistency propagation checklist (convert prior checklist into active validations):
|
||||
- Read `.specify/templates/plan-template.md` and ensure any "Constitution Check" or rules align with updated principles.
|
||||
- Read `.specify/templates/spec-template.md` for scope/requirements alignment—update if constitution adds/removes mandatory sections or constraints.
|
||||
- Read `.specify/templates/tasks-template.md` and ensure task categorization reflects new or removed principle-driven task types (e.g., observability, versioning, testing discipline).
|
||||
- Read each command file in `.specify/templates/commands/*.md` (including this one) to verify no outdated references (agent-specific names like CLAUDE only) remain when generic guidance is required.
|
||||
- Read any runtime guidance docs (e.g., `README.md`, `docs/quickstart.md`, or agent-specific guidance files if present). Update references to principles changed.
|
||||
|
||||
5. Produce a Sync Impact Report (prepend as an HTML comment at top of the constitution file after update):
|
||||
- Version change: old → new
|
||||
- List of modified principles (old title → new title if renamed)
|
||||
- Added sections
|
||||
- Removed sections
|
||||
- Templates requiring updates (✅ updated / ⚠ pending) with file paths
|
||||
- Follow-up TODOs if any placeholders intentionally deferred.
|
||||
|
||||
6. Validation before final output:
|
||||
- No remaining unexplained bracket tokens.
|
||||
- Version line matches report.
|
||||
- Dates ISO format YYYY-MM-DD.
|
||||
- Principles are declarative, testable, and free of vague language ("should" → replace with MUST/SHOULD rationale where appropriate).
|
||||
|
||||
7. Write the completed constitution back to `.specify/memory/constitution.md` (overwrite).
|
||||
|
||||
8. Output a final summary to the user with:
|
||||
- New version and bump rationale.
|
||||
- Any files flagged for manual follow-up.
|
||||
- Suggested commit message (e.g., `docs: amend constitution to vX.Y.Z (principle additions + governance update)`).
|
||||
|
||||
Formatting & Style Requirements:
|
||||
|
||||
- Use Markdown headings exactly as in the template (do not demote/promote levels).
|
||||
- Wrap long rationale lines to keep readability (<100 chars ideally) but do not hard enforce with awkward breaks.
|
||||
- Keep a single blank line between sections.
|
||||
- Avoid trailing whitespace.
|
||||
|
||||
If the user supplies partial updates (e.g., only one principle revision), still perform validation and version decision steps.
|
||||
|
||||
If critical info missing (e.g., ratification date truly unknown), insert `TODO(<FIELD_NAME>): explanation` and include in the Sync Impact Report under deferred items.
|
||||
|
||||
Do not create a new template; always operate on the existing `.specify/memory/constitution.md` file.
|
||||
@@ -0,0 +1,198 @@
|
||||
---
|
||||
description: Execute the implementation plan by processing and executing all tasks defined in tasks.md
|
||||
---
|
||||
|
||||
## User Input
|
||||
|
||||
```text
|
||||
$ARGUMENTS
|
||||
```
|
||||
|
||||
You **MUST** consider the user input before proceeding (if not empty).
|
||||
|
||||
## Pre-Execution Checks
|
||||
|
||||
**Check for extension hooks (before implementation)**:
|
||||
- Check if `.specify/extensions.yml` exists in the project root.
|
||||
- If it exists, read it and look for entries under the `hooks.before_implement` key
|
||||
- If the YAML cannot be parsed or is invalid, skip hook checking silently and continue normally
|
||||
- Filter to only hooks where `enabled: true`
|
||||
- For each remaining hook, do **not** attempt to interpret or evaluate hook `condition` expressions:
|
||||
- If the hook has no `condition` field, or it is null/empty, treat the hook as executable
|
||||
- If the hook defines a non-empty `condition`, skip the hook and leave condition evaluation to the HookExecutor implementation
|
||||
- For each executable hook, output the following based on its `optional` flag:
|
||||
- **Optional hook** (`optional: true`):
|
||||
```
|
||||
## Extension Hooks
|
||||
|
||||
**Optional Pre-Hook**: {extension}
|
||||
Command: `/{command}`
|
||||
Description: {description}
|
||||
|
||||
Prompt: {prompt}
|
||||
To execute: `/{command}`
|
||||
```
|
||||
- **Mandatory hook** (`optional: false`):
|
||||
```
|
||||
## Extension Hooks
|
||||
|
||||
**Automatic Pre-Hook**: {extension}
|
||||
Executing: `/{command}`
|
||||
EXECUTE_COMMAND: {command}
|
||||
|
||||
Wait for the result of the hook command before proceeding to the Outline.
|
||||
```
|
||||
- If no hooks are registered or `.specify/extensions.yml` does not exist, skip silently
|
||||
|
||||
## Outline
|
||||
|
||||
1. Run `.specify/scripts/bash/check-prerequisites.sh --json --require-tasks --include-tasks` from repo root and parse FEATURE_DIR and AVAILABLE_DOCS list. All paths must be absolute. For single quotes in args like "I'm Groot", use escape syntax: e.g 'I'\''m Groot' (or double-quote if possible: "I'm Groot").
|
||||
|
||||
2. **Check checklists status** (if FEATURE_DIR/checklists/ exists):
|
||||
- Scan all checklist files in the checklists/ directory
|
||||
- For each checklist, count:
|
||||
- Total items: All lines matching `- [ ]` or `- [X]` or `- [x]`
|
||||
- Completed items: Lines matching `- [X]` or `- [x]`
|
||||
- Incomplete items: Lines matching `- [ ]`
|
||||
- Create a status table:
|
||||
|
||||
```text
|
||||
| Checklist | Total | Completed | Incomplete | Status |
|
||||
|-----------|-------|-----------|------------|--------|
|
||||
| ux.md | 12 | 12 | 0 | ✓ PASS |
|
||||
| test.md | 8 | 5 | 3 | ✗ FAIL |
|
||||
| security.md | 6 | 6 | 0 | ✓ PASS |
|
||||
```
|
||||
|
||||
- Calculate overall status:
|
||||
- **PASS**: All checklists have 0 incomplete items
|
||||
- **FAIL**: One or more checklists have incomplete items
|
||||
|
||||
- **If any checklist is incomplete**:
|
||||
- Display the table with incomplete item counts
|
||||
- **STOP** and ask: "Some checklists are incomplete. Do you want to proceed with implementation anyway? (yes/no)"
|
||||
- Wait for user response before continuing
|
||||
- If user says "no" or "wait" or "stop", halt execution
|
||||
- If user says "yes" or "proceed" or "continue", proceed to step 3
|
||||
|
||||
- **If all checklists are complete**:
|
||||
- Display the table showing all checklists passed
|
||||
- Automatically proceed to step 3
|
||||
|
||||
3. Load and analyze the implementation context:
|
||||
- **REQUIRED**: Read tasks.md for the complete task list and execution plan
|
||||
- **REQUIRED**: Read plan.md for tech stack, architecture, and file structure
|
||||
- **IF EXISTS**: Read data-model.md for entities and relationships
|
||||
- **IF EXISTS**: Read contracts/ for API specifications and test requirements
|
||||
- **IF EXISTS**: Read research.md for technical decisions and constraints
|
||||
- **IF EXISTS**: Read quickstart.md for integration scenarios
|
||||
|
||||
4. **Project Setup Verification**:
|
||||
- **REQUIRED**: Create/verify ignore files based on actual project setup:
|
||||
|
||||
**Detection & Creation Logic**:
|
||||
- Check if the following command succeeds to determine if the repository is a git repo (create/verify .gitignore if so):
|
||||
|
||||
```sh
|
||||
git rev-parse --git-dir 2>/dev/null
|
||||
```
|
||||
|
||||
- Check if Dockerfile* exists or Docker in plan.md → create/verify .dockerignore
|
||||
- Check if .eslintrc* exists → create/verify .eslintignore
|
||||
- Check if eslint.config.* exists → ensure the config's `ignores` entries cover required patterns
|
||||
- Check if .prettierrc* exists → create/verify .prettierignore
|
||||
- Check if .npmrc or package.json exists → create/verify .npmignore (if publishing)
|
||||
- Check if terraform files (*.tf) exist → create/verify .terraformignore
|
||||
- Check if .helmignore needed (helm charts present) → create/verify .helmignore
|
||||
|
||||
**If ignore file already exists**: Verify it contains essential patterns, append missing critical patterns only
|
||||
**If ignore file missing**: Create with full pattern set for detected technology
|
||||
|
||||
**Common Patterns by Technology** (from plan.md tech stack):
|
||||
- **Node.js/JavaScript/TypeScript**: `node_modules/`, `dist/`, `build/`, `*.log`, `.env*`
|
||||
- **Python**: `__pycache__/`, `*.pyc`, `.venv/`, `venv/`, `dist/`, `*.egg-info/`
|
||||
- **Java**: `target/`, `*.class`, `*.jar`, `.gradle/`, `build/`
|
||||
- **C#/.NET**: `bin/`, `obj/`, `*.user`, `*.suo`, `packages/`
|
||||
- **Go**: `*.exe`, `*.test`, `vendor/`, `*.out`
|
||||
- **Ruby**: `.bundle/`, `log/`, `tmp/`, `*.gem`, `vendor/bundle/`
|
||||
- **PHP**: `vendor/`, `*.log`, `*.cache`, `*.env`
|
||||
- **Rust**: `target/`, `debug/`, `release/`, `*.rs.bk`, `*.rlib`, `*.prof*`, `.idea/`, `*.log`, `.env*`
|
||||
- **Kotlin**: `build/`, `out/`, `.gradle/`, `.idea/`, `*.class`, `*.jar`, `*.iml`, `*.log`, `.env*`
|
||||
- **C++**: `build/`, `bin/`, `obj/`, `out/`, `*.o`, `*.so`, `*.a`, `*.exe`, `*.dll`, `.idea/`, `*.log`, `.env*`
|
||||
- **C**: `build/`, `bin/`, `obj/`, `out/`, `*.o`, `*.a`, `*.so`, `*.exe`, `*.dll`, `autom4te.cache/`, `config.status`, `config.log`, `.idea/`, `*.log`, `.env*`
|
||||
- **Swift**: `.build/`, `DerivedData/`, `*.swiftpm/`, `Packages/`
|
||||
- **R**: `.Rproj.user/`, `.Rhistory`, `.RData`, `.Ruserdata`, `*.Rproj`, `packrat/`, `renv/`
|
||||
- **Universal**: `.DS_Store`, `Thumbs.db`, `*.tmp`, `*.swp`, `.vscode/`, `.idea/`
|
||||
|
||||
**Tool-Specific Patterns**:
|
||||
- **Docker**: `node_modules/`, `.git/`, `Dockerfile*`, `.dockerignore`, `*.log*`, `.env*`, `coverage/`
|
||||
- **ESLint**: `node_modules/`, `dist/`, `build/`, `coverage/`, `*.min.js`
|
||||
- **Prettier**: `node_modules/`, `dist/`, `build/`, `coverage/`, `package-lock.json`, `yarn.lock`, `pnpm-lock.yaml`
|
||||
- **Terraform**: `.terraform/`, `*.tfstate*`, `*.tfvars`, `.terraform.lock.hcl`
|
||||
- **Kubernetes/k8s**: `*.secret.yaml`, `secrets/`, `.kube/`, `kubeconfig*`, `*.key`, `*.crt`
|
||||
|
||||
5. Parse tasks.md structure and extract:
|
||||
- **Task phases**: Setup, Tests, Core, Integration, Polish
|
||||
- **Task dependencies**: Sequential vs parallel execution rules
|
||||
- **Task details**: ID, description, file paths, parallel markers [P]
|
||||
- **Execution flow**: Order and dependency requirements
|
||||
|
||||
6. Execute implementation following the task plan:
|
||||
- **Phase-by-phase execution**: Complete each phase before moving to the next
|
||||
- **Respect dependencies**: Run sequential tasks in order, parallel tasks [P] can run together
|
||||
- **Follow TDD approach**: Execute test tasks before their corresponding implementation tasks
|
||||
- **File-based coordination**: Tasks affecting the same files must run sequentially
|
||||
- **Validation checkpoints**: Verify each phase completion before proceeding
|
||||
|
||||
7. Implementation execution rules:
|
||||
- **Setup first**: Initialize project structure, dependencies, configuration
|
||||
- **Tests before code**: If you need to write tests for contracts, entities, and integration scenarios
|
||||
- **Core development**: Implement models, services, CLI commands, endpoints
|
||||
- **Integration work**: Database connections, middleware, logging, external services
|
||||
- **Polish and validation**: Unit tests, performance optimization, documentation
|
||||
|
||||
8. Progress tracking and error handling:
|
||||
- Report progress after each completed task
|
||||
- Halt execution if any non-parallel task fails
|
||||
- For parallel tasks [P], continue with successful tasks, report failed ones
|
||||
- Provide clear error messages with context for debugging
|
||||
- Suggest next steps if implementation cannot proceed
|
||||
- **IMPORTANT** For completed tasks, make sure to mark the task off as [X] in the tasks file.
|
||||
|
||||
9. Completion validation:
|
||||
- Verify all required tasks are completed
|
||||
- Check that implemented features match the original specification
|
||||
- Validate that tests pass and coverage meets requirements
|
||||
- Confirm the implementation follows the technical plan
|
||||
- Report final status with summary of completed work
|
||||
|
||||
Note: This command assumes a complete task breakdown exists in tasks.md. If tasks are incomplete or missing, suggest running `/speckit.tasks` first to regenerate the task list.
|
||||
|
||||
10. **Check for extension hooks**: After completion validation, check if `.specify/extensions.yml` exists in the project root.
|
||||
- If it exists, read it and look for entries under the `hooks.after_implement` key
|
||||
- If the YAML cannot be parsed or is invalid, skip hook checking silently and continue normally
|
||||
- Filter to only hooks where `enabled: true`
|
||||
- For each remaining hook, do **not** attempt to interpret or evaluate hook `condition` expressions:
|
||||
- If the hook has no `condition` field, or it is null/empty, treat the hook as executable
|
||||
- If the hook defines a non-empty `condition`, skip the hook and leave condition evaluation to the HookExecutor implementation
|
||||
- For each executable hook, output the following based on its `optional` flag:
|
||||
- **Optional hook** (`optional: true`):
|
||||
```
|
||||
## Extension Hooks
|
||||
|
||||
**Optional Hook**: {extension}
|
||||
Command: `/{command}`
|
||||
Description: {description}
|
||||
|
||||
Prompt: {prompt}
|
||||
To execute: `/{command}`
|
||||
```
|
||||
- **Mandatory hook** (`optional: false`):
|
||||
```
|
||||
## Extension Hooks
|
||||
|
||||
**Automatic Hook**: {extension}
|
||||
Executing: `/{command}`
|
||||
EXECUTE_COMMAND: {command}
|
||||
```
|
||||
- If no hooks are registered or `.specify/extensions.yml` does not exist, skip silently
|
||||
@@ -0,0 +1,90 @@
|
||||
---
|
||||
description: Execute the implementation planning workflow using the plan template to generate design artifacts.
|
||||
handoffs:
|
||||
- label: Create Tasks
|
||||
agent: speckit.tasks
|
||||
prompt: Break the plan into tasks
|
||||
send: true
|
||||
- label: Create Checklist
|
||||
agent: speckit.checklist
|
||||
prompt: Create a checklist for the following domain...
|
||||
---
|
||||
|
||||
## User Input
|
||||
|
||||
```text
|
||||
$ARGUMENTS
|
||||
```
|
||||
|
||||
You **MUST** consider the user input before proceeding (if not empty).
|
||||
|
||||
## Outline
|
||||
|
||||
1. **Setup**: Run `.specify/scripts/bash/setup-plan.sh --json` from repo root and parse JSON for FEATURE_SPEC, IMPL_PLAN, SPECS_DIR, BRANCH. For single quotes in args like "I'm Groot", use escape syntax: e.g 'I'\''m Groot' (or double-quote if possible: "I'm Groot").
|
||||
|
||||
2. **Load context**: Read FEATURE_SPEC and `.specify/memory/constitution.md`. Load IMPL_PLAN template (already copied).
|
||||
|
||||
3. **Execute plan workflow**: Follow the structure in IMPL_PLAN template to:
|
||||
- Fill Technical Context (mark unknowns as "NEEDS CLARIFICATION")
|
||||
- Fill Constitution Check section from constitution
|
||||
- Evaluate gates (ERROR if violations unjustified)
|
||||
- Phase 0: Generate research.md (resolve all NEEDS CLARIFICATION)
|
||||
- Phase 1: Generate data-model.md, contracts/, quickstart.md
|
||||
- Phase 1: Update agent context by running the agent script
|
||||
- Re-evaluate Constitution Check post-design
|
||||
|
||||
4. **Stop and report**: Command ends after Phase 2 planning. Report branch, IMPL_PLAN path, and generated artifacts.
|
||||
|
||||
## Phases
|
||||
|
||||
### Phase 0: Outline & Research
|
||||
|
||||
1. **Extract unknowns from Technical Context** above:
|
||||
- For each NEEDS CLARIFICATION → research task
|
||||
- For each dependency → best practices task
|
||||
- For each integration → patterns task
|
||||
|
||||
2. **Generate and dispatch research agents**:
|
||||
|
||||
```text
|
||||
For each unknown in Technical Context:
|
||||
Task: "Research {unknown} for {feature context}"
|
||||
For each technology choice:
|
||||
Task: "Find best practices for {tech} in {domain}"
|
||||
```
|
||||
|
||||
3. **Consolidate findings** in `research.md` using format:
|
||||
- Decision: [what was chosen]
|
||||
- Rationale: [why chosen]
|
||||
- Alternatives considered: [what else evaluated]
|
||||
|
||||
**Output**: research.md with all NEEDS CLARIFICATION resolved
|
||||
|
||||
### Phase 1: Design & Contracts
|
||||
|
||||
**Prerequisites:** `research.md` complete
|
||||
|
||||
1. **Extract entities from feature spec** → `data-model.md`:
|
||||
- Entity name, fields, relationships
|
||||
- Validation rules from requirements
|
||||
- State transitions if applicable
|
||||
|
||||
2. **Define interface contracts** (if project has external interfaces) → `/contracts/`:
|
||||
- Identify what interfaces the project exposes to users or other systems
|
||||
- Document the contract format appropriate for the project type
|
||||
- Examples: public APIs for libraries, command schemas for CLI tools, endpoints for web services, grammars for parsers, UI contracts for applications
|
||||
- Skip if project is purely internal (build scripts, one-off tools, etc.)
|
||||
|
||||
3. **Agent context update**:
|
||||
- Run `.specify/scripts/bash/update-agent-context.sh claude`
|
||||
- These scripts detect which AI agent is in use
|
||||
- Update the appropriate agent-specific context file
|
||||
- Add only new technology from current plan
|
||||
- Preserve manual additions between markers
|
||||
|
||||
**Output**: data-model.md, /contracts/*, quickstart.md, agent-specific file
|
||||
|
||||
## Key rules
|
||||
|
||||
- Use absolute paths
|
||||
- ERROR on gate failures or unresolved clarifications
|
||||
@@ -0,0 +1,239 @@
|
||||
---
|
||||
description: Create or update the feature specification from a natural language feature description.
|
||||
handoffs:
|
||||
- label: Build Technical Plan
|
||||
agent: speckit.plan
|
||||
prompt: Create a plan for the spec. I am building with...
|
||||
- label: Clarify Spec Requirements
|
||||
agent: speckit.clarify
|
||||
prompt: Clarify specification requirements
|
||||
send: true
|
||||
---
|
||||
|
||||
## User Input
|
||||
|
||||
```text
|
||||
$ARGUMENTS
|
||||
```
|
||||
|
||||
You **MUST** consider the user input before proceeding (if not empty).
|
||||
|
||||
## Outline
|
||||
|
||||
The text the user typed after `/speckit.specify` in the triggering message **is** the feature description. Assume you always have it available in this conversation even if `$ARGUMENTS` appears literally below. Do not ask the user to repeat it unless they provided an empty command.
|
||||
|
||||
Given that feature description, do this:
|
||||
|
||||
1. **Generate a concise short name** (2-4 words) for the branch:
|
||||
- Analyze the feature description and extract the most meaningful keywords
|
||||
- Create a 2-4 word short name that captures the essence of the feature
|
||||
- Use action-noun format when possible (e.g., "add-user-auth", "fix-payment-bug")
|
||||
- Preserve technical terms and acronyms (OAuth2, API, JWT, etc.)
|
||||
- Keep it concise but descriptive enough to understand the feature at a glance
|
||||
- Examples:
|
||||
- "I want to add user authentication" → "user-auth"
|
||||
- "Implement OAuth2 integration for the API" → "oauth2-api-integration"
|
||||
- "Create a dashboard for analytics" → "analytics-dashboard"
|
||||
- "Fix payment processing timeout bug" → "fix-payment-timeout"
|
||||
|
||||
2. **Create the feature branch** by running the script with `--short-name` (and `--json`), and do NOT pass `--number` (the script auto-detects the next globally available number across all branches and spec directories):
|
||||
|
||||
- Bash example: `.specify/scripts/bash/create-new-feature.sh "$ARGUMENTS" --json --short-name "user-auth" "Add user authentication"`
|
||||
- PowerShell example: `.specify/scripts/bash/create-new-feature.sh "$ARGUMENTS" -Json -ShortName "user-auth" "Add user authentication"`
|
||||
|
||||
**IMPORTANT**:
|
||||
- Do NOT pass `--number` — the script determines the correct next number automatically
|
||||
- Always include the JSON flag (`--json` for Bash, `-Json` for PowerShell) so the output can be parsed reliably
|
||||
- You must only ever run this script once per feature
|
||||
- The JSON is provided in the terminal as output - always refer to it to get the actual content you're looking for
|
||||
- The JSON output will contain BRANCH_NAME and SPEC_FILE paths
|
||||
- For single quotes in args like "I'm Groot", use escape syntax: e.g 'I'\''m Groot' (or double-quote if possible: "I'm Groot")
|
||||
|
||||
3. Load `.specify/templates/spec-template.md` to understand required sections.
|
||||
|
||||
4. Follow this execution flow:
|
||||
|
||||
1. Parse user description from Input
|
||||
If empty: ERROR "No feature description provided"
|
||||
2. Extract key concepts from description
|
||||
Identify: actors, actions, data, constraints
|
||||
3. For unclear aspects:
|
||||
- Make informed guesses based on context and industry standards
|
||||
- Only mark with [NEEDS CLARIFICATION: specific question] if:
|
||||
- The choice significantly impacts feature scope or user experience
|
||||
- Multiple reasonable interpretations exist with different implications
|
||||
- No reasonable default exists
|
||||
- **LIMIT: Maximum 3 [NEEDS CLARIFICATION] markers total**
|
||||
- Prioritize clarifications by impact: scope > security/privacy > user experience > technical details
|
||||
4. Fill User Scenarios & Testing section
|
||||
If no clear user flow: ERROR "Cannot determine user scenarios"
|
||||
5. Generate Functional Requirements
|
||||
Each requirement must be testable
|
||||
Use reasonable defaults for unspecified details (document assumptions in Assumptions section)
|
||||
6. Define Success Criteria
|
||||
Create measurable, technology-agnostic outcomes
|
||||
Include both quantitative metrics (time, performance, volume) and qualitative measures (user satisfaction, task completion)
|
||||
Each criterion must be verifiable without implementation details
|
||||
7. Identify Key Entities (if data involved)
|
||||
8. Return: SUCCESS (spec ready for planning)
|
||||
|
||||
5. Write the specification to SPEC_FILE using the template structure, replacing placeholders with concrete details derived from the feature description (arguments) while preserving section order and headings.
|
||||
|
||||
6. **Specification Quality Validation**: After writing the initial spec, validate it against quality criteria:
|
||||
|
||||
a. **Create Spec Quality Checklist**: Generate a checklist file at `FEATURE_DIR/checklists/requirements.md` using the checklist template structure with these validation items:
|
||||
|
||||
```markdown
|
||||
# Specification Quality Checklist: [FEATURE NAME]
|
||||
|
||||
**Purpose**: Validate specification completeness and quality before proceeding to planning
|
||||
**Created**: [DATE]
|
||||
**Feature**: [Link to spec.md]
|
||||
|
||||
## Content Quality
|
||||
|
||||
- [ ] No implementation details (languages, frameworks, APIs)
|
||||
- [ ] Focused on user value and business needs
|
||||
- [ ] Written for non-technical stakeholders
|
||||
- [ ] All mandatory sections completed
|
||||
|
||||
## Requirement Completeness
|
||||
|
||||
- [ ] No [NEEDS CLARIFICATION] markers remain
|
||||
- [ ] Requirements are testable and unambiguous
|
||||
- [ ] Success criteria are measurable
|
||||
- [ ] Success criteria are technology-agnostic (no implementation details)
|
||||
- [ ] All acceptance scenarios are defined
|
||||
- [ ] Edge cases are identified
|
||||
- [ ] Scope is clearly bounded
|
||||
- [ ] Dependencies and assumptions identified
|
||||
|
||||
## Feature Readiness
|
||||
|
||||
- [ ] All functional requirements have clear acceptance criteria
|
||||
- [ ] User scenarios cover primary flows
|
||||
- [ ] Feature meets measurable outcomes defined in Success Criteria
|
||||
- [ ] No implementation details leak into specification
|
||||
|
||||
## Notes
|
||||
|
||||
- Items marked incomplete require spec updates before `/speckit.clarify` or `/speckit.plan`
|
||||
```
|
||||
|
||||
b. **Run Validation Check**: Review the spec against each checklist item:
|
||||
- For each item, determine if it passes or fails
|
||||
- Document specific issues found (quote relevant spec sections)
|
||||
|
||||
c. **Handle Validation Results**:
|
||||
|
||||
- **If all items pass**: Mark checklist complete and proceed to step 6
|
||||
|
||||
- **If items fail (excluding [NEEDS CLARIFICATION])**:
|
||||
1. List the failing items and specific issues
|
||||
2. Update the spec to address each issue
|
||||
3. Re-run validation until all items pass (max 3 iterations)
|
||||
4. If still failing after 3 iterations, document remaining issues in checklist notes and warn user
|
||||
|
||||
- **If [NEEDS CLARIFICATION] markers remain**:
|
||||
1. Extract all [NEEDS CLARIFICATION: ...] markers from the spec
|
||||
2. **LIMIT CHECK**: If more than 3 markers exist, keep only the 3 most critical (by scope/security/UX impact) and make informed guesses for the rest
|
||||
3. For each clarification needed (max 3), present options to user in this format:
|
||||
|
||||
```markdown
|
||||
## Question [N]: [Topic]
|
||||
|
||||
**Context**: [Quote relevant spec section]
|
||||
|
||||
**What we need to know**: [Specific question from NEEDS CLARIFICATION marker]
|
||||
|
||||
**Suggested Answers**:
|
||||
|
||||
| Option | Answer | Implications |
|
||||
|--------|--------|--------------|
|
||||
| A | [First suggested answer] | [What this means for the feature] |
|
||||
| B | [Second suggested answer] | [What this means for the feature] |
|
||||
| C | [Third suggested answer] | [What this means for the feature] |
|
||||
| Custom | Provide your own answer | [Explain how to provide custom input] |
|
||||
|
||||
**Your choice**: _[Wait for user response]_
|
||||
```
|
||||
|
||||
4. **CRITICAL - Table Formatting**: Ensure markdown tables are properly formatted:
|
||||
- Use consistent spacing with pipes aligned
|
||||
- Each cell should have spaces around content: `| Content |` not `|Content|`
|
||||
- Header separator must have at least 3 dashes: `|--------|`
|
||||
- Test that the table renders correctly in markdown preview
|
||||
5. Number questions sequentially (Q1, Q2, Q3 - max 3 total)
|
||||
6. Present all questions together before waiting for responses
|
||||
7. Wait for user to respond with their choices for all questions (e.g., "Q1: A, Q2: Custom - [details], Q3: B")
|
||||
8. Update the spec by replacing each [NEEDS CLARIFICATION] marker with the user's selected or provided answer
|
||||
9. Re-run validation after all clarifications are resolved
|
||||
|
||||
d. **Update Checklist**: After each validation iteration, update the checklist file with current pass/fail status
|
||||
|
||||
7. Report completion with branch name, spec file path, checklist results, and readiness for the next phase (`/speckit.clarify` or `/speckit.plan`).
|
||||
|
||||
**NOTE:** The script creates and checks out the new branch and initializes the spec file before writing.
|
||||
|
||||
## General Guidelines
|
||||
|
||||
## Quick Guidelines
|
||||
|
||||
- Focus on **WHAT** users need and **WHY**.
|
||||
- Avoid HOW to implement (no tech stack, APIs, code structure).
|
||||
- Written for business stakeholders, not developers.
|
||||
- DO NOT create any checklists that are embedded in the spec. That will be a separate command.
|
||||
|
||||
### Section Requirements
|
||||
|
||||
- **Mandatory sections**: Must be completed for every feature
|
||||
- **Optional sections**: Include only when relevant to the feature
|
||||
- When a section doesn't apply, remove it entirely (don't leave as "N/A")
|
||||
|
||||
### For AI Generation
|
||||
|
||||
When creating this spec from a user prompt:
|
||||
|
||||
1. **Make informed guesses**: Use context, industry standards, and common patterns to fill gaps
|
||||
2. **Document assumptions**: Record reasonable defaults in the Assumptions section
|
||||
3. **Limit clarifications**: Maximum 3 [NEEDS CLARIFICATION] markers - use only for critical decisions that:
|
||||
- Significantly impact feature scope or user experience
|
||||
- Have multiple reasonable interpretations with different implications
|
||||
- Lack any reasonable default
|
||||
4. **Prioritize clarifications**: scope > security/privacy > user experience > technical details
|
||||
5. **Think like a tester**: Every vague requirement should fail the "testable and unambiguous" checklist item
|
||||
6. **Common areas needing clarification** (only if no reasonable default exists):
|
||||
- Feature scope and boundaries (include/exclude specific use cases)
|
||||
- User types and permissions (if multiple conflicting interpretations possible)
|
||||
- Security/compliance requirements (when legally/financially significant)
|
||||
|
||||
**Examples of reasonable defaults** (don't ask about these):
|
||||
|
||||
- Data retention: Industry-standard practices for the domain
|
||||
- Performance targets: Standard web/mobile app expectations unless specified
|
||||
- Error handling: User-friendly messages with appropriate fallbacks
|
||||
- Authentication method: Standard session-based or OAuth2 for web apps
|
||||
- Integration patterns: Use project-appropriate patterns (REST/GraphQL for web services, function calls for libraries, CLI args for tools, etc.)
|
||||
|
||||
### Success Criteria Guidelines
|
||||
|
||||
Success criteria must be:
|
||||
|
||||
1. **Measurable**: Include specific metrics (time, percentage, count, rate)
|
||||
2. **Technology-agnostic**: No mention of frameworks, languages, databases, or tools
|
||||
3. **User-focused**: Describe outcomes from user/business perspective, not system internals
|
||||
4. **Verifiable**: Can be tested/validated without knowing implementation details
|
||||
|
||||
**Good examples**:
|
||||
|
||||
- "Users can complete checkout in under 3 minutes"
|
||||
- "System supports 10,000 concurrent users"
|
||||
- "95% of searches return results in under 1 second"
|
||||
- "Task completion rate improves by 40%"
|
||||
|
||||
**Bad examples** (implementation-focused):
|
||||
|
||||
- "API response time is under 200ms" (too technical, use "Users see results instantly")
|
||||
- "Database can handle 1000 TPS" (implementation detail, use user-facing metric)
|
||||
- "React components render efficiently" (framework-specific)
|
||||
- "Redis cache hit rate above 80%" (technology-specific)
|
||||
@@ -0,0 +1,200 @@
|
||||
---
|
||||
description: Generate an actionable, dependency-ordered tasks.md for the feature based on available design artifacts.
|
||||
handoffs:
|
||||
- label: Analyze For Consistency
|
||||
agent: speckit.analyze
|
||||
prompt: Run a project analysis for consistency
|
||||
send: true
|
||||
- label: Implement Project
|
||||
agent: speckit.implement
|
||||
prompt: Start the implementation in phases
|
||||
send: true
|
||||
---
|
||||
|
||||
## User Input
|
||||
|
||||
```text
|
||||
$ARGUMENTS
|
||||
```
|
||||
|
||||
You **MUST** consider the user input before proceeding (if not empty).
|
||||
|
||||
## Pre-Execution Checks
|
||||
|
||||
**Check for extension hooks (before tasks generation)**:
|
||||
- Check if `.specify/extensions.yml` exists in the project root.
|
||||
- If it exists, read it and look for entries under the `hooks.before_tasks` key
|
||||
- If the YAML cannot be parsed or is invalid, skip hook checking silently and continue normally
|
||||
- Filter to only hooks where `enabled: true`
|
||||
- For each remaining hook, do **not** attempt to interpret or evaluate hook `condition` expressions:
|
||||
- If the hook has no `condition` field, or it is null/empty, treat the hook as executable
|
||||
- If the hook defines a non-empty `condition`, skip the hook and leave condition evaluation to the HookExecutor implementation
|
||||
- For each executable hook, output the following based on its `optional` flag:
|
||||
- **Optional hook** (`optional: true`):
|
||||
```
|
||||
## Extension Hooks
|
||||
|
||||
**Optional Pre-Hook**: {extension}
|
||||
Command: `/{command}`
|
||||
Description: {description}
|
||||
|
||||
Prompt: {prompt}
|
||||
To execute: `/{command}`
|
||||
```
|
||||
- **Mandatory hook** (`optional: false`):
|
||||
```
|
||||
## Extension Hooks
|
||||
|
||||
**Automatic Pre-Hook**: {extension}
|
||||
Executing: `/{command}`
|
||||
EXECUTE_COMMAND: {command}
|
||||
|
||||
Wait for the result of the hook command before proceeding to the Outline.
|
||||
```
|
||||
- If no hooks are registered or `.specify/extensions.yml` does not exist, skip silently
|
||||
|
||||
## Outline
|
||||
|
||||
1. **Setup**: Run `.specify/scripts/bash/check-prerequisites.sh --json` from repo root and parse FEATURE_DIR and AVAILABLE_DOCS list. All paths must be absolute. For single quotes in args like "I'm Groot", use escape syntax: e.g 'I'\''m Groot' (or double-quote if possible: "I'm Groot").
|
||||
|
||||
2. **Load design documents**: Read from FEATURE_DIR:
|
||||
- **Required**: plan.md (tech stack, libraries, structure), spec.md (user stories with priorities)
|
||||
- **Optional**: data-model.md (entities), contracts/ (interface contracts), research.md (decisions), quickstart.md (test scenarios)
|
||||
- Note: Not all projects have all documents. Generate tasks based on what's available.
|
||||
|
||||
3. **Execute task generation workflow**:
|
||||
- Load plan.md and extract tech stack, libraries, project structure
|
||||
- Load spec.md and extract user stories with their priorities (P1, P2, P3, etc.)
|
||||
- If data-model.md exists: Extract entities and map to user stories
|
||||
- If contracts/ exists: Map interface contracts to user stories
|
||||
- If research.md exists: Extract decisions for setup tasks
|
||||
- Generate tasks organized by user story (see Task Generation Rules below)
|
||||
- Generate dependency graph showing user story completion order
|
||||
- Create parallel execution examples per user story
|
||||
- Validate task completeness (each user story has all needed tasks, independently testable)
|
||||
|
||||
4. **Generate tasks.md**: Use `.specify/templates/tasks-template.md` as structure, fill with:
|
||||
- Correct feature name from plan.md
|
||||
- Phase 1: Setup tasks (project initialization)
|
||||
- Phase 2: Foundational tasks (blocking prerequisites for all user stories)
|
||||
- Phase 3+: One phase per user story (in priority order from spec.md)
|
||||
- Each phase includes: story goal, independent test criteria, tests (if requested), implementation tasks
|
||||
- Final Phase: Polish & cross-cutting concerns
|
||||
- All tasks must follow the strict checklist format (see Task Generation Rules below)
|
||||
- Clear file paths for each task
|
||||
- Dependencies section showing story completion order
|
||||
- Parallel execution examples per story
|
||||
- Implementation strategy section (MVP first, incremental delivery)
|
||||
|
||||
5. **Report**: Output path to generated tasks.md and summary:
|
||||
- Total task count
|
||||
- Task count per user story
|
||||
- Parallel opportunities identified
|
||||
- Independent test criteria for each story
|
||||
- Suggested MVP scope (typically just User Story 1)
|
||||
- Format validation: Confirm ALL tasks follow the checklist format (checkbox, ID, labels, file paths)
|
||||
|
||||
6. **Check for extension hooks**: After tasks.md is generated, check if `.specify/extensions.yml` exists in the project root.
|
||||
- If it exists, read it and look for entries under the `hooks.after_tasks` key
|
||||
- If the YAML cannot be parsed or is invalid, skip hook checking silently and continue normally
|
||||
- Filter to only hooks where `enabled: true`
|
||||
- For each remaining hook, do **not** attempt to interpret or evaluate hook `condition` expressions:
|
||||
- If the hook has no `condition` field, or it is null/empty, treat the hook as executable
|
||||
- If the hook defines a non-empty `condition`, skip the hook and leave condition evaluation to the HookExecutor implementation
|
||||
- For each executable hook, output the following based on its `optional` flag:
|
||||
- **Optional hook** (`optional: true`):
|
||||
```
|
||||
## Extension Hooks
|
||||
|
||||
**Optional Hook**: {extension}
|
||||
Command: `/{command}`
|
||||
Description: {description}
|
||||
|
||||
Prompt: {prompt}
|
||||
To execute: `/{command}`
|
||||
```
|
||||
- **Mandatory hook** (`optional: false`):
|
||||
```
|
||||
## Extension Hooks
|
||||
|
||||
**Automatic Hook**: {extension}
|
||||
Executing: `/{command}`
|
||||
EXECUTE_COMMAND: {command}
|
||||
```
|
||||
- If no hooks are registered or `.specify/extensions.yml` does not exist, skip silently
|
||||
|
||||
Context for task generation: $ARGUMENTS
|
||||
|
||||
The tasks.md should be immediately executable - each task must be specific enough that an LLM can complete it without additional context.
|
||||
|
||||
## Task Generation Rules
|
||||
|
||||
**CRITICAL**: Tasks MUST be organized by user story to enable independent implementation and testing.
|
||||
|
||||
**Tests are OPTIONAL**: Only generate test tasks if explicitly requested in the feature specification or if user requests TDD approach.
|
||||
|
||||
### Checklist Format (REQUIRED)
|
||||
|
||||
Every task MUST strictly follow this format:
|
||||
|
||||
```text
|
||||
- [ ] [TaskID] [P?] [Story?] Description with file path
|
||||
```
|
||||
|
||||
**Format Components**:
|
||||
|
||||
1. **Checkbox**: ALWAYS start with `- [ ]` (markdown checkbox)
|
||||
2. **Task ID**: Sequential number (T001, T002, T003...) in execution order
|
||||
3. **[P] marker**: Include ONLY if task is parallelizable (different files, no dependencies on incomplete tasks)
|
||||
4. **[Story] label**: REQUIRED for user story phase tasks only
|
||||
- Format: [US1], [US2], [US3], etc. (maps to user stories from spec.md)
|
||||
- Setup phase: NO story label
|
||||
- Foundational phase: NO story label
|
||||
- User Story phases: MUST have story label
|
||||
- Polish phase: NO story label
|
||||
5. **Description**: Clear action with exact file path
|
||||
|
||||
**Examples**:
|
||||
|
||||
- ✅ CORRECT: `- [ ] T001 Create project structure per implementation plan`
|
||||
- ✅ CORRECT: `- [ ] T005 [P] Implement authentication middleware in src/middleware/auth.py`
|
||||
- ✅ CORRECT: `- [ ] T012 [P] [US1] Create User model in src/models/user.py`
|
||||
- ✅ CORRECT: `- [ ] T014 [US1] Implement UserService in src/services/user_service.py`
|
||||
- ❌ WRONG: `- [ ] Create User model` (missing ID and Story label)
|
||||
- ❌ WRONG: `T001 [US1] Create model` (missing checkbox)
|
||||
- ❌ WRONG: `- [ ] [US1] Create User model` (missing Task ID)
|
||||
- ❌ WRONG: `- [ ] T001 [US1] Create model` (missing file path)
|
||||
|
||||
### Task Organization
|
||||
|
||||
1. **From User Stories (spec.md)** - PRIMARY ORGANIZATION:
|
||||
- Each user story (P1, P2, P3...) gets its own phase
|
||||
- Map all related components to their story:
|
||||
- Models needed for that story
|
||||
- Services needed for that story
|
||||
- Interfaces/UI needed for that story
|
||||
- If tests requested: Tests specific to that story
|
||||
- Mark story dependencies (most stories should be independent)
|
||||
|
||||
2. **From Contracts**:
|
||||
- Map each interface contract → to the user story it serves
|
||||
- If tests requested: Each interface contract → contract test task [P] before implementation in that story's phase
|
||||
|
||||
3. **From Data Model**:
|
||||
- Map each entity to the user story(ies) that need it
|
||||
- If entity serves multiple stories: Put in earliest story or Setup phase
|
||||
- Relationships → service layer tasks in appropriate story phase
|
||||
|
||||
4. **From Setup/Infrastructure**:
|
||||
- Shared infrastructure → Setup phase (Phase 1)
|
||||
- Foundational/blocking tasks → Foundational phase (Phase 2)
|
||||
- Story-specific setup → within that story's phase
|
||||
|
||||
### Phase Structure
|
||||
|
||||
- **Phase 1**: Setup (project initialization)
|
||||
- **Phase 2**: Foundational (blocking prerequisites - MUST complete before user stories)
|
||||
- **Phase 3+**: User Stories in priority order (P1, P2, P3...)
|
||||
- Within each story: Tests (if requested) → Models → Services → Endpoints → Integration
|
||||
- Each phase should be a complete, independently testable increment
|
||||
- **Final Phase**: Polish & Cross-Cutting Concerns
|
||||
@@ -0,0 +1,30 @@
|
||||
---
|
||||
description: Convert existing tasks into actionable, dependency-ordered GitHub issues for the feature based on available design artifacts.
|
||||
tools: ['github/github-mcp-server/issue_write']
|
||||
---
|
||||
|
||||
## User Input
|
||||
|
||||
```text
|
||||
$ARGUMENTS
|
||||
```
|
||||
|
||||
You **MUST** consider the user input before proceeding (if not empty).
|
||||
|
||||
## Outline
|
||||
|
||||
1. Run `.specify/scripts/bash/check-prerequisites.sh --json --require-tasks --include-tasks` from repo root and parse FEATURE_DIR and AVAILABLE_DOCS list. All paths must be absolute. For single quotes in args like "I'm Groot", use escape syntax: e.g 'I'\''m Groot' (or double-quote if possible: "I'm Groot").
|
||||
1. From the executed script, extract the path to **tasks**.
|
||||
1. Get the Git remote by running:
|
||||
|
||||
```bash
|
||||
git config --get remote.origin.url
|
||||
```
|
||||
|
||||
> [!CAUTION]
|
||||
> ONLY PROCEED TO NEXT STEPS IF THE REMOTE IS A GITHUB URL
|
||||
|
||||
1. For each task in the list, use the GitHub MCP server to create a new issue in the repository that is representative of the Git remote.
|
||||
|
||||
> [!CAUTION]
|
||||
> UNDER NO CIRCUMSTANCES EVER CREATE ISSUES IN REPOSITORIES THAT DO NOT MATCH THE REMOTE URL
|
||||
+38
@@ -0,0 +1,38 @@
|
||||
# Go
|
||||
bin/
|
||||
*.exe
|
||||
*.exe~
|
||||
*.dll
|
||||
*.so
|
||||
*.dylib
|
||||
*.test
|
||||
*.out
|
||||
vendor/
|
||||
|
||||
# Data directory
|
||||
data/
|
||||
|
||||
# Node / Svelte
|
||||
web/node_modules/
|
||||
web/build/
|
||||
web/.svelte-kit/
|
||||
internal/web/dist/
|
||||
|
||||
# IDE
|
||||
.idea/
|
||||
.vscode/
|
||||
*.swp
|
||||
*.swo
|
||||
*~
|
||||
|
||||
# OS
|
||||
.DS_Store
|
||||
Thumbs.db
|
||||
|
||||
# Environment
|
||||
.env
|
||||
.env.local
|
||||
.env.*.local
|
||||
|
||||
# Debug
|
||||
__debug_bin*
|
||||
@@ -0,0 +1,204 @@
|
||||
<!--
|
||||
Sync Impact Report
|
||||
==================
|
||||
- Version change: N/A → 1.0.0 (initial ratification)
|
||||
- Added principles: I through X (10 total)
|
||||
- Added sections: Technology Decisions, Non-Goals
|
||||
- Removed sections: None (initial creation)
|
||||
- Templates requiring updates:
|
||||
- .specify/templates/plan-template.md — ✅ compatible (Constitution Check section exists)
|
||||
- .specify/templates/spec-template.md — ✅ compatible (no constitution-specific sections needed)
|
||||
- .specify/templates/tasks-template.md — ✅ compatible (phase structure supports Go layout)
|
||||
- Follow-up TODOs: None
|
||||
-->
|
||||
|
||||
# SynapBus Constitution
|
||||
|
||||
## Core Principles
|
||||
|
||||
### I. Local-First, Single Binary
|
||||
|
||||
Everything ships in one Go binary. There MUST be no external runtime
|
||||
dependencies — no separate database server, no message broker, no
|
||||
external auth service. A user runs `synapbus serve` and the full
|
||||
system starts: MCP server, REST API, Web UI, embedded storage.
|
||||
|
||||
**Rationale**: Eliminates deployment complexity and ensures any
|
||||
developer or agent operator can run SynapBus with zero infrastructure
|
||||
setup.
|
||||
|
||||
### II. MCP-Native
|
||||
|
||||
Agents interact with SynapBus exclusively through MCP protocol tools
|
||||
(SSE and Streamable HTTP transports). The REST API exists solely for
|
||||
the embedded Web UI and MUST NOT be advertised as an external agent
|
||||
interface. All agent-facing operations MUST be exposed as MCP tools
|
||||
with JSON Schema descriptions.
|
||||
|
||||
**Rationale**: MCP is the emerging standard for AI tool interaction.
|
||||
By committing to MCP-only for agents, SynapBus avoids fragmenting
|
||||
its interface and ensures compatibility with any MCP-capable client.
|
||||
|
||||
### III. Pure Go, Zero CGO
|
||||
|
||||
All dependencies MUST be pure Go. CGO MUST NOT be enabled for any
|
||||
build target. This means:
|
||||
- `modernc.org/sqlite` (not `mattn/go-sqlite3`)
|
||||
- `TFMV/hnsw` (not C-based FAISS or Annoy)
|
||||
- No C library bindings of any kind
|
||||
|
||||
The binary MUST cross-compile cleanly for at minimum:
|
||||
`linux/amd64`, `darwin/arm64`, `darwin/amd64`.
|
||||
|
||||
**Rationale**: CGO breaks cross-compilation, complicates CI, and
|
||||
introduces platform-specific build failures. Pure Go guarantees
|
||||
`GOOS=X GOARCH=Y go build` works everywhere.
|
||||
|
||||
### IV. Multi-Tenant with Ownership
|
||||
|
||||
Every agent MUST have a human owner (`owner_id`). Owners control
|
||||
their agents' access and can view all traces of agent activity.
|
||||
Agents MUST only access their own messages and channels they have
|
||||
joined. No agent may impersonate another or access another owner's
|
||||
data without explicit grants.
|
||||
|
||||
**Rationale**: AI agents act on behalf of humans. Humans need
|
||||
visibility and control over what their agents do. This principle
|
||||
ensures accountability and prevents runaway agent behavior.
|
||||
|
||||
### V. Embedded OAuth 2.1
|
||||
|
||||
The OAuth 2.1 authorization server MUST be built into SynapBus
|
||||
using `ory/fosite`. There MUST NOT be a dependency on an external
|
||||
identity provider for core functionality. Local accounts (username
|
||||
+ bcrypt-hashed password) MUST be supported as the default auth
|
||||
method. PKCE MUST be required for all authorization code flows.
|
||||
|
||||
**Rationale**: Requiring an external auth service violates
|
||||
Principle I (single binary). Embedding OAuth 2.1 keeps the system
|
||||
self-contained while providing standards-compliant security.
|
||||
|
||||
### VI. Semantic-Ready Storage
|
||||
|
||||
SQLite handles all relational data. HNSW (via `TFMV/hnsw`) handles
|
||||
vector search. Both MUST be embedded and store data within a single
|
||||
`--data` directory. The system MUST function fully without an
|
||||
embedding provider configured (falling back to full-text search).
|
||||
When a provider is configured, messages MUST be embedded
|
||||
asynchronously in the background.
|
||||
|
||||
**Rationale**: Semantic search is a key differentiator but MUST NOT
|
||||
be a hard dependency. Progressive enhancement: basic → full-text →
|
||||
semantic.
|
||||
|
||||
### VII. Swarm Intelligence Patterns
|
||||
|
||||
SynapBus MUST provide first-class support for:
|
||||
- **Stigmergy**: Tagged messages on blackboard channels that agents
|
||||
read and react to (tags: `#finding`, `#task`, `#decision`, `#trace`)
|
||||
- **Task Auction**: Post task → agents bid → poster selects winner
|
||||
- **Agent Discovery**: Search agent capability cards by keyword or
|
||||
semantic match
|
||||
|
||||
Channel types (`standard`, `blackboard`, `auction`) MUST enforce
|
||||
the appropriate interaction patterns.
|
||||
|
||||
**Rationale**: Multi-agent coordination requires higher-level
|
||||
patterns beyond point-to-point messaging. These patterns are
|
||||
proven in swarm intelligence literature and directly applicable
|
||||
to AI agent orchestration.
|
||||
|
||||
### VIII. Observable by Default
|
||||
|
||||
All agent actions MUST be traced: tool calls, messages sent/received,
|
||||
channel operations, errors. Traces MUST be stored in SQLite with
|
||||
agent identity, action type, details (JSON), and timestamp. Owners
|
||||
MUST be able to view, filter, and export traces for their agents.
|
||||
Structured logging via `slog` MUST be used for all server-side logs.
|
||||
|
||||
**Rationale**: Agent systems are opaque by default. Observability
|
||||
is not optional — it is a safety and debugging requirement. If a
|
||||
human cannot see what an agent did, the system is not trustworthy.
|
||||
|
||||
### IX. Progressive Complexity
|
||||
|
||||
The system MUST be usable with just core messaging (send, read,
|
||||
mark done). Advanced features MUST layer on top without requiring
|
||||
configuration changes for basic usage:
|
||||
1. Basic: messaging + agent registration
|
||||
2. Intermediate: channels + full-text search + traces
|
||||
3. Advanced: semantic search + attachments + swarm patterns
|
||||
|
||||
No feature in a higher tier MUST break or require features from
|
||||
a lower tier.
|
||||
|
||||
**Rationale**: New users should not be overwhelmed. The simplest
|
||||
use case (two agents exchanging messages) should require minimal
|
||||
setup. Complexity is opt-in.
|
||||
|
||||
### X. Web UI as First-Class Citizen
|
||||
|
||||
The Svelte 5 + Tailwind CSS SPA MUST be embedded in the Go binary
|
||||
via `go:embed`. It MUST provide full operational visibility:
|
||||
message browsing, conversation threads, channel management, agent
|
||||
monitoring, trace viewing, and search. The UI MUST support dark
|
||||
mode and be responsive. It MUST receive real-time updates via SSE.
|
||||
|
||||
**Rationale**: The Web UI is how humans interact with SynapBus.
|
||||
It is not an afterthought or admin panel — it is a primary
|
||||
interface alongside MCP.
|
||||
|
||||
## Technology Decisions
|
||||
|
||||
| Component | Choice | Constraint |
|
||||
|-----------|--------|------------|
|
||||
| Language | Go 1.23+ | Single binary, cross-compilation |
|
||||
| Database | modernc.org/sqlite | Pure Go, zero CGO (Principle III) |
|
||||
| Vectors | TFMV/hnsw | Pure Go HNSW index (Principle III) |
|
||||
| MCP | mark3labs/mcp-go | SSE + Streamable HTTP transports |
|
||||
| HTTP Router | go-chi/chi | Lightweight, stdlib-compatible |
|
||||
| Auth | ory/fosite | OAuth 2.1 framework (Principle V) |
|
||||
| Web UI | Svelte 5 + Tailwind | Embedded via go:embed (Principle X) |
|
||||
| Logging | slog | Structured, stdlib (Principle VIII) |
|
||||
| Attachments | Content-addressable FS | SHA-256 hash, dedup storage |
|
||||
| CLI | spf13/cobra | Standard Go CLI framework |
|
||||
|
||||
## Non-Goals
|
||||
|
||||
These are explicitly out of scope for the current version:
|
||||
|
||||
- **No external database**: No PostgreSQL, Redis, Kafka, or any
|
||||
external data store dependency.
|
||||
- **No framework lock-in**: No LangChain, CrewAI, AutoGen, or any
|
||||
agent framework dependency. SynapBus is framework-agnostic.
|
||||
- **No A2A protocol**: Google's Agent-to-Agent protocol is not
|
||||
supported in v1. MCP is the sole agent interface.
|
||||
- **No cloud-specific features**: No AWS/GCP/Azure service
|
||||
integrations. SynapBus is local-first, always.
|
||||
- **No multi-node clustering**: Single binary, single instance.
|
||||
Horizontal scaling is a future concern.
|
||||
|
||||
## Governance
|
||||
|
||||
This constitution is the authoritative source for architectural
|
||||
decisions in SynapBus. All implementation work MUST comply with
|
||||
these principles.
|
||||
|
||||
### Amendment Process
|
||||
|
||||
1. Propose amendment with rationale and impact analysis.
|
||||
2. Update constitution version per semantic versioning:
|
||||
- **MAJOR**: Principle removal or backward-incompatible redefinition
|
||||
- **MINOR**: New principle or material expansion
|
||||
- **PATCH**: Clarification, wording, typo fix
|
||||
3. Propagate changes to dependent templates and documentation.
|
||||
4. Document changes in Sync Impact Report (HTML comment at top).
|
||||
|
||||
### Compliance
|
||||
|
||||
- All specs MUST include a constitution compliance check.
|
||||
- All implementation plans MUST verify alignment with principles
|
||||
before Phase 0 research begins.
|
||||
- Code reviews SHOULD verify adherence to relevant principles.
|
||||
|
||||
**Version**: 1.0.0 | **Ratified**: 2026-03-13 | **Last Amended**: 2026-03-13
|
||||
Executable
+166
@@ -0,0 +1,166 @@
|
||||
#!/usr/bin/env bash
|
||||
|
||||
# Consolidated prerequisite checking script
|
||||
#
|
||||
# This script provides unified prerequisite checking for Spec-Driven Development workflow.
|
||||
# It replaces the functionality previously spread across multiple scripts.
|
||||
#
|
||||
# Usage: ./check-prerequisites.sh [OPTIONS]
|
||||
#
|
||||
# OPTIONS:
|
||||
# --json Output in JSON format
|
||||
# --require-tasks Require tasks.md to exist (for implementation phase)
|
||||
# --include-tasks Include tasks.md in AVAILABLE_DOCS list
|
||||
# --paths-only Only output path variables (no validation)
|
||||
# --help, -h Show help message
|
||||
#
|
||||
# OUTPUTS:
|
||||
# JSON mode: {"FEATURE_DIR":"...", "AVAILABLE_DOCS":["..."]}
|
||||
# Text mode: FEATURE_DIR:... \n AVAILABLE_DOCS: \n ✓/✗ file.md
|
||||
# Paths only: REPO_ROOT: ... \n BRANCH: ... \n FEATURE_DIR: ... etc.
|
||||
|
||||
set -e
|
||||
|
||||
# Parse command line arguments
|
||||
JSON_MODE=false
|
||||
REQUIRE_TASKS=false
|
||||
INCLUDE_TASKS=false
|
||||
PATHS_ONLY=false
|
||||
|
||||
for arg in "$@"; do
|
||||
case "$arg" in
|
||||
--json)
|
||||
JSON_MODE=true
|
||||
;;
|
||||
--require-tasks)
|
||||
REQUIRE_TASKS=true
|
||||
;;
|
||||
--include-tasks)
|
||||
INCLUDE_TASKS=true
|
||||
;;
|
||||
--paths-only)
|
||||
PATHS_ONLY=true
|
||||
;;
|
||||
--help|-h)
|
||||
cat << 'EOF'
|
||||
Usage: check-prerequisites.sh [OPTIONS]
|
||||
|
||||
Consolidated prerequisite checking for Spec-Driven Development workflow.
|
||||
|
||||
OPTIONS:
|
||||
--json Output in JSON format
|
||||
--require-tasks Require tasks.md to exist (for implementation phase)
|
||||
--include-tasks Include tasks.md in AVAILABLE_DOCS list
|
||||
--paths-only Only output path variables (no prerequisite validation)
|
||||
--help, -h Show this help message
|
||||
|
||||
EXAMPLES:
|
||||
# Check task prerequisites (plan.md required)
|
||||
./check-prerequisites.sh --json
|
||||
|
||||
# Check implementation prerequisites (plan.md + tasks.md required)
|
||||
./check-prerequisites.sh --json --require-tasks --include-tasks
|
||||
|
||||
# Get feature paths only (no validation)
|
||||
./check-prerequisites.sh --paths-only
|
||||
|
||||
EOF
|
||||
exit 0
|
||||
;;
|
||||
*)
|
||||
echo "ERROR: Unknown option '$arg'. Use --help for usage information." >&2
|
||||
exit 1
|
||||
;;
|
||||
esac
|
||||
done
|
||||
|
||||
# Source common functions
|
||||
SCRIPT_DIR="$(CDPATH="" cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
source "$SCRIPT_DIR/common.sh"
|
||||
|
||||
# Get feature paths and validate branch
|
||||
eval $(get_feature_paths)
|
||||
check_feature_branch "$CURRENT_BRANCH" "$HAS_GIT" || exit 1
|
||||
|
||||
# If paths-only mode, output paths and exit (support JSON + paths-only combined)
|
||||
if $PATHS_ONLY; then
|
||||
if $JSON_MODE; then
|
||||
# Minimal JSON paths payload (no validation performed)
|
||||
printf '{"REPO_ROOT":"%s","BRANCH":"%s","FEATURE_DIR":"%s","FEATURE_SPEC":"%s","IMPL_PLAN":"%s","TASKS":"%s"}\n' \
|
||||
"$REPO_ROOT" "$CURRENT_BRANCH" "$FEATURE_DIR" "$FEATURE_SPEC" "$IMPL_PLAN" "$TASKS"
|
||||
else
|
||||
echo "REPO_ROOT: $REPO_ROOT"
|
||||
echo "BRANCH: $CURRENT_BRANCH"
|
||||
echo "FEATURE_DIR: $FEATURE_DIR"
|
||||
echo "FEATURE_SPEC: $FEATURE_SPEC"
|
||||
echo "IMPL_PLAN: $IMPL_PLAN"
|
||||
echo "TASKS: $TASKS"
|
||||
fi
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# Validate required directories and files
|
||||
if [[ ! -d "$FEATURE_DIR" ]]; then
|
||||
echo "ERROR: Feature directory not found: $FEATURE_DIR" >&2
|
||||
echo "Run /speckit.specify first to create the feature structure." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if [[ ! -f "$IMPL_PLAN" ]]; then
|
||||
echo "ERROR: plan.md not found in $FEATURE_DIR" >&2
|
||||
echo "Run /speckit.plan first to create the implementation plan." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Check for tasks.md if required
|
||||
if $REQUIRE_TASKS && [[ ! -f "$TASKS" ]]; then
|
||||
echo "ERROR: tasks.md not found in $FEATURE_DIR" >&2
|
||||
echo "Run /speckit.tasks first to create the task list." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Build list of available documents
|
||||
docs=()
|
||||
|
||||
# Always check these optional docs
|
||||
[[ -f "$RESEARCH" ]] && docs+=("research.md")
|
||||
[[ -f "$DATA_MODEL" ]] && docs+=("data-model.md")
|
||||
|
||||
# Check contracts directory (only if it exists and has files)
|
||||
if [[ -d "$CONTRACTS_DIR" ]] && [[ -n "$(ls -A "$CONTRACTS_DIR" 2>/dev/null)" ]]; then
|
||||
docs+=("contracts/")
|
||||
fi
|
||||
|
||||
[[ -f "$QUICKSTART" ]] && docs+=("quickstart.md")
|
||||
|
||||
# Include tasks.md if requested and it exists
|
||||
if $INCLUDE_TASKS && [[ -f "$TASKS" ]]; then
|
||||
docs+=("tasks.md")
|
||||
fi
|
||||
|
||||
# Output results
|
||||
if $JSON_MODE; then
|
||||
# Build JSON array of documents
|
||||
if [[ ${#docs[@]} -eq 0 ]]; then
|
||||
json_docs="[]"
|
||||
else
|
||||
json_docs=$(printf '"%s",' "${docs[@]}")
|
||||
json_docs="[${json_docs%,}]"
|
||||
fi
|
||||
|
||||
printf '{"FEATURE_DIR":"%s","AVAILABLE_DOCS":%s}\n' "$FEATURE_DIR" "$json_docs"
|
||||
else
|
||||
# Text output
|
||||
echo "FEATURE_DIR:$FEATURE_DIR"
|
||||
echo "AVAILABLE_DOCS:"
|
||||
|
||||
# Show status of each potential document
|
||||
check_file "$RESEARCH" "research.md"
|
||||
check_file "$DATA_MODEL" "data-model.md"
|
||||
check_dir "$CONTRACTS_DIR" "contracts/"
|
||||
check_file "$QUICKSTART" "quickstart.md"
|
||||
|
||||
if $INCLUDE_TASKS; then
|
||||
check_file "$TASKS" "tasks.md"
|
||||
fi
|
||||
fi
|
||||
Executable
+156
@@ -0,0 +1,156 @@
|
||||
#!/usr/bin/env bash
|
||||
# Common functions and variables for all scripts
|
||||
|
||||
# Get repository root, with fallback for non-git repositories
|
||||
get_repo_root() {
|
||||
if git rev-parse --show-toplevel >/dev/null 2>&1; then
|
||||
git rev-parse --show-toplevel
|
||||
else
|
||||
# Fall back to script location for non-git repos
|
||||
local script_dir="$(CDPATH="" cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
(cd "$script_dir/../../.." && pwd)
|
||||
fi
|
||||
}
|
||||
|
||||
# Get current branch, with fallback for non-git repositories
|
||||
get_current_branch() {
|
||||
# First check if SPECIFY_FEATURE environment variable is set
|
||||
if [[ -n "${SPECIFY_FEATURE:-}" ]]; then
|
||||
echo "$SPECIFY_FEATURE"
|
||||
return
|
||||
fi
|
||||
|
||||
# Then check git if available
|
||||
if git rev-parse --abbrev-ref HEAD >/dev/null 2>&1; then
|
||||
git rev-parse --abbrev-ref HEAD
|
||||
return
|
||||
fi
|
||||
|
||||
# For non-git repos, try to find the latest feature directory
|
||||
local repo_root=$(get_repo_root)
|
||||
local specs_dir="$repo_root/specs"
|
||||
|
||||
if [[ -d "$specs_dir" ]]; then
|
||||
local latest_feature=""
|
||||
local highest=0
|
||||
|
||||
for dir in "$specs_dir"/*; do
|
||||
if [[ -d "$dir" ]]; then
|
||||
local dirname=$(basename "$dir")
|
||||
if [[ "$dirname" =~ ^([0-9]{3})- ]]; then
|
||||
local number=${BASH_REMATCH[1]}
|
||||
number=$((10#$number))
|
||||
if [[ "$number" -gt "$highest" ]]; then
|
||||
highest=$number
|
||||
latest_feature=$dirname
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
done
|
||||
|
||||
if [[ -n "$latest_feature" ]]; then
|
||||
echo "$latest_feature"
|
||||
return
|
||||
fi
|
||||
fi
|
||||
|
||||
echo "main" # Final fallback
|
||||
}
|
||||
|
||||
# Check if we have git available
|
||||
has_git() {
|
||||
git rev-parse --show-toplevel >/dev/null 2>&1
|
||||
}
|
||||
|
||||
check_feature_branch() {
|
||||
local branch="$1"
|
||||
local has_git_repo="$2"
|
||||
|
||||
# For non-git repos, we can't enforce branch naming but still provide output
|
||||
if [[ "$has_git_repo" != "true" ]]; then
|
||||
echo "[specify] Warning: Git repository not detected; skipped branch validation" >&2
|
||||
return 0
|
||||
fi
|
||||
|
||||
if [[ ! "$branch" =~ ^[0-9]{3}- ]]; then
|
||||
echo "ERROR: Not on a feature branch. Current branch: $branch" >&2
|
||||
echo "Feature branches should be named like: 001-feature-name" >&2
|
||||
return 1
|
||||
fi
|
||||
|
||||
return 0
|
||||
}
|
||||
|
||||
get_feature_dir() { echo "$1/specs/$2"; }
|
||||
|
||||
# Find feature directory by numeric prefix instead of exact branch match
|
||||
# This allows multiple branches to work on the same spec (e.g., 004-fix-bug, 004-add-feature)
|
||||
find_feature_dir_by_prefix() {
|
||||
local repo_root="$1"
|
||||
local branch_name="$2"
|
||||
local specs_dir="$repo_root/specs"
|
||||
|
||||
# Extract numeric prefix from branch (e.g., "004" from "004-whatever")
|
||||
if [[ ! "$branch_name" =~ ^([0-9]{3})- ]]; then
|
||||
# If branch doesn't have numeric prefix, fall back to exact match
|
||||
echo "$specs_dir/$branch_name"
|
||||
return
|
||||
fi
|
||||
|
||||
local prefix="${BASH_REMATCH[1]}"
|
||||
|
||||
# Search for directories in specs/ that start with this prefix
|
||||
local matches=()
|
||||
if [[ -d "$specs_dir" ]]; then
|
||||
for dir in "$specs_dir"/"$prefix"-*; do
|
||||
if [[ -d "$dir" ]]; then
|
||||
matches+=("$(basename "$dir")")
|
||||
fi
|
||||
done
|
||||
fi
|
||||
|
||||
# Handle results
|
||||
if [[ ${#matches[@]} -eq 0 ]]; then
|
||||
# No match found - return the branch name path (will fail later with clear error)
|
||||
echo "$specs_dir/$branch_name"
|
||||
elif [[ ${#matches[@]} -eq 1 ]]; then
|
||||
# Exactly one match - perfect!
|
||||
echo "$specs_dir/${matches[0]}"
|
||||
else
|
||||
# Multiple matches - this shouldn't happen with proper naming convention
|
||||
echo "ERROR: Multiple spec directories found with prefix '$prefix': ${matches[*]}" >&2
|
||||
echo "Please ensure only one spec directory exists per numeric prefix." >&2
|
||||
echo "$specs_dir/$branch_name" # Return something to avoid breaking the script
|
||||
fi
|
||||
}
|
||||
|
||||
get_feature_paths() {
|
||||
local repo_root=$(get_repo_root)
|
||||
local current_branch=$(get_current_branch)
|
||||
local has_git_repo="false"
|
||||
|
||||
if has_git; then
|
||||
has_git_repo="true"
|
||||
fi
|
||||
|
||||
# Use prefix-based lookup to support multiple branches per spec
|
||||
local feature_dir=$(find_feature_dir_by_prefix "$repo_root" "$current_branch")
|
||||
|
||||
cat <<EOF
|
||||
REPO_ROOT='$repo_root'
|
||||
CURRENT_BRANCH='$current_branch'
|
||||
HAS_GIT='$has_git_repo'
|
||||
FEATURE_DIR='$feature_dir'
|
||||
FEATURE_SPEC='$feature_dir/spec.md'
|
||||
IMPL_PLAN='$feature_dir/plan.md'
|
||||
TASKS='$feature_dir/tasks.md'
|
||||
RESEARCH='$feature_dir/research.md'
|
||||
DATA_MODEL='$feature_dir/data-model.md'
|
||||
QUICKSTART='$feature_dir/quickstart.md'
|
||||
CONTRACTS_DIR='$feature_dir/contracts'
|
||||
EOF
|
||||
}
|
||||
|
||||
check_file() { [[ -f "$1" ]] && echo " ✓ $2" || echo " ✗ $2"; }
|
||||
check_dir() { [[ -d "$1" && -n $(ls -A "$1" 2>/dev/null) ]] && echo " ✓ $2" || echo " ✗ $2"; }
|
||||
|
||||
Executable
+313
@@ -0,0 +1,313 @@
|
||||
#!/usr/bin/env bash
|
||||
|
||||
set -e
|
||||
|
||||
JSON_MODE=false
|
||||
SHORT_NAME=""
|
||||
BRANCH_NUMBER=""
|
||||
ARGS=()
|
||||
i=1
|
||||
while [ $i -le $# ]; do
|
||||
arg="${!i}"
|
||||
case "$arg" in
|
||||
--json)
|
||||
JSON_MODE=true
|
||||
;;
|
||||
--short-name)
|
||||
if [ $((i + 1)) -gt $# ]; then
|
||||
echo 'Error: --short-name requires a value' >&2
|
||||
exit 1
|
||||
fi
|
||||
i=$((i + 1))
|
||||
next_arg="${!i}"
|
||||
# Check if the next argument is another option (starts with --)
|
||||
if [[ "$next_arg" == --* ]]; then
|
||||
echo 'Error: --short-name requires a value' >&2
|
||||
exit 1
|
||||
fi
|
||||
SHORT_NAME="$next_arg"
|
||||
;;
|
||||
--number)
|
||||
if [ $((i + 1)) -gt $# ]; then
|
||||
echo 'Error: --number requires a value' >&2
|
||||
exit 1
|
||||
fi
|
||||
i=$((i + 1))
|
||||
next_arg="${!i}"
|
||||
if [[ "$next_arg" == --* ]]; then
|
||||
echo 'Error: --number requires a value' >&2
|
||||
exit 1
|
||||
fi
|
||||
BRANCH_NUMBER="$next_arg"
|
||||
;;
|
||||
--help|-h)
|
||||
echo "Usage: $0 [--json] [--short-name <name>] [--number N] <feature_description>"
|
||||
echo ""
|
||||
echo "Options:"
|
||||
echo " --json Output in JSON format"
|
||||
echo " --short-name <name> Provide a custom short name (2-4 words) for the branch"
|
||||
echo " --number N Specify branch number manually (overrides auto-detection)"
|
||||
echo " --help, -h Show this help message"
|
||||
echo ""
|
||||
echo "Examples:"
|
||||
echo " $0 'Add user authentication system' --short-name 'user-auth'"
|
||||
echo " $0 'Implement OAuth2 integration for API' --number 5"
|
||||
exit 0
|
||||
;;
|
||||
*)
|
||||
ARGS+=("$arg")
|
||||
;;
|
||||
esac
|
||||
i=$((i + 1))
|
||||
done
|
||||
|
||||
FEATURE_DESCRIPTION="${ARGS[*]}"
|
||||
if [ -z "$FEATURE_DESCRIPTION" ]; then
|
||||
echo "Usage: $0 [--json] [--short-name <name>] [--number N] <feature_description>" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Trim whitespace and validate description is not empty (e.g., user passed only whitespace)
|
||||
FEATURE_DESCRIPTION=$(echo "$FEATURE_DESCRIPTION" | xargs)
|
||||
if [ -z "$FEATURE_DESCRIPTION" ]; then
|
||||
echo "Error: Feature description cannot be empty or contain only whitespace" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Function to find the repository root by searching for existing project markers
|
||||
find_repo_root() {
|
||||
local dir="$1"
|
||||
while [ "$dir" != "/" ]; do
|
||||
if [ -d "$dir/.git" ] || [ -d "$dir/.specify" ]; then
|
||||
echo "$dir"
|
||||
return 0
|
||||
fi
|
||||
dir="$(dirname "$dir")"
|
||||
done
|
||||
return 1
|
||||
}
|
||||
|
||||
# Function to get highest number from specs directory
|
||||
get_highest_from_specs() {
|
||||
local specs_dir="$1"
|
||||
local highest=0
|
||||
|
||||
if [ -d "$specs_dir" ]; then
|
||||
for dir in "$specs_dir"/*; do
|
||||
[ -d "$dir" ] || continue
|
||||
dirname=$(basename "$dir")
|
||||
number=$(echo "$dirname" | grep -o '^[0-9]\+' || echo "0")
|
||||
number=$((10#$number))
|
||||
if [ "$number" -gt "$highest" ]; then
|
||||
highest=$number
|
||||
fi
|
||||
done
|
||||
fi
|
||||
|
||||
echo "$highest"
|
||||
}
|
||||
|
||||
# Function to get highest number from git branches
|
||||
get_highest_from_branches() {
|
||||
local highest=0
|
||||
|
||||
# Get all branches (local and remote)
|
||||
branches=$(git branch -a 2>/dev/null || echo "")
|
||||
|
||||
if [ -n "$branches" ]; then
|
||||
while IFS= read -r branch; do
|
||||
# Clean branch name: remove leading markers and remote prefixes
|
||||
clean_branch=$(echo "$branch" | sed 's/^[* ]*//; s|^remotes/[^/]*/||')
|
||||
|
||||
# Extract feature number if branch matches pattern ###-*
|
||||
if echo "$clean_branch" | grep -q '^[0-9]\{3\}-'; then
|
||||
number=$(echo "$clean_branch" | grep -o '^[0-9]\{3\}' || echo "0")
|
||||
number=$((10#$number))
|
||||
if [ "$number" -gt "$highest" ]; then
|
||||
highest=$number
|
||||
fi
|
||||
fi
|
||||
done <<< "$branches"
|
||||
fi
|
||||
|
||||
echo "$highest"
|
||||
}
|
||||
|
||||
# Function to check existing branches (local and remote) and return next available number
|
||||
check_existing_branches() {
|
||||
local specs_dir="$1"
|
||||
|
||||
# Fetch all remotes to get latest branch info (suppress errors if no remotes)
|
||||
git fetch --all --prune 2>/dev/null || true
|
||||
|
||||
# Get highest number from ALL branches (not just matching short name)
|
||||
local highest_branch=$(get_highest_from_branches)
|
||||
|
||||
# Get highest number from ALL specs (not just matching short name)
|
||||
local highest_spec=$(get_highest_from_specs "$specs_dir")
|
||||
|
||||
# Take the maximum of both
|
||||
local max_num=$highest_branch
|
||||
if [ "$highest_spec" -gt "$max_num" ]; then
|
||||
max_num=$highest_spec
|
||||
fi
|
||||
|
||||
# Return next number
|
||||
echo $((max_num + 1))
|
||||
}
|
||||
|
||||
# Function to clean and format a branch name
|
||||
clean_branch_name() {
|
||||
local name="$1"
|
||||
echo "$name" | tr '[:upper:]' '[:lower:]' | sed 's/[^a-z0-9]/-/g' | sed 's/-\+/-/g' | sed 's/^-//' | sed 's/-$//'
|
||||
}
|
||||
|
||||
# Resolve repository root. Prefer git information when available, but fall back
|
||||
# to searching for repository markers so the workflow still functions in repositories that
|
||||
# were initialised with --no-git.
|
||||
SCRIPT_DIR="$(CDPATH="" cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
if git rev-parse --show-toplevel >/dev/null 2>&1; then
|
||||
REPO_ROOT=$(git rev-parse --show-toplevel)
|
||||
HAS_GIT=true
|
||||
else
|
||||
REPO_ROOT="$(find_repo_root "$SCRIPT_DIR")"
|
||||
if [ -z "$REPO_ROOT" ]; then
|
||||
echo "Error: Could not determine repository root. Please run this script from within the repository." >&2
|
||||
exit 1
|
||||
fi
|
||||
HAS_GIT=false
|
||||
fi
|
||||
|
||||
cd "$REPO_ROOT"
|
||||
|
||||
SPECS_DIR="$REPO_ROOT/specs"
|
||||
mkdir -p "$SPECS_DIR"
|
||||
|
||||
# Function to generate branch name with stop word filtering and length filtering
|
||||
generate_branch_name() {
|
||||
local description="$1"
|
||||
|
||||
# Common stop words to filter out
|
||||
local stop_words="^(i|a|an|the|to|for|of|in|on|at|by|with|from|is|are|was|were|be|been|being|have|has|had|do|does|did|will|would|should|could|can|may|might|must|shall|this|that|these|those|my|your|our|their|want|need|add|get|set)$"
|
||||
|
||||
# Convert to lowercase and split into words
|
||||
local clean_name=$(echo "$description" | tr '[:upper:]' '[:lower:]' | sed 's/[^a-z0-9]/ /g')
|
||||
|
||||
# Filter words: remove stop words and words shorter than 3 chars (unless they're uppercase acronyms in original)
|
||||
local meaningful_words=()
|
||||
for word in $clean_name; do
|
||||
# Skip empty words
|
||||
[ -z "$word" ] && continue
|
||||
|
||||
# Keep words that are NOT stop words AND (length >= 3 OR are potential acronyms)
|
||||
if ! echo "$word" | grep -qiE "$stop_words"; then
|
||||
if [ ${#word} -ge 3 ]; then
|
||||
meaningful_words+=("$word")
|
||||
elif echo "$description" | grep -q "\b${word^^}\b"; then
|
||||
# Keep short words if they appear as uppercase in original (likely acronyms)
|
||||
meaningful_words+=("$word")
|
||||
fi
|
||||
fi
|
||||
done
|
||||
|
||||
# If we have meaningful words, use first 3-4 of them
|
||||
if [ ${#meaningful_words[@]} -gt 0 ]; then
|
||||
local max_words=3
|
||||
if [ ${#meaningful_words[@]} -eq 4 ]; then max_words=4; fi
|
||||
|
||||
local result=""
|
||||
local count=0
|
||||
for word in "${meaningful_words[@]}"; do
|
||||
if [ $count -ge $max_words ]; then break; fi
|
||||
if [ -n "$result" ]; then result="$result-"; fi
|
||||
result="$result$word"
|
||||
count=$((count + 1))
|
||||
done
|
||||
echo "$result"
|
||||
else
|
||||
# Fallback to original logic if no meaningful words found
|
||||
local cleaned=$(clean_branch_name "$description")
|
||||
echo "$cleaned" | tr '-' '\n' | grep -v '^$' | head -3 | tr '\n' '-' | sed 's/-$//'
|
||||
fi
|
||||
}
|
||||
|
||||
# Generate branch name
|
||||
if [ -n "$SHORT_NAME" ]; then
|
||||
# Use provided short name, just clean it up
|
||||
BRANCH_SUFFIX=$(clean_branch_name "$SHORT_NAME")
|
||||
else
|
||||
# Generate from description with smart filtering
|
||||
BRANCH_SUFFIX=$(generate_branch_name "$FEATURE_DESCRIPTION")
|
||||
fi
|
||||
|
||||
# Determine branch number
|
||||
if [ -z "$BRANCH_NUMBER" ]; then
|
||||
if [ "$HAS_GIT" = true ]; then
|
||||
# Check existing branches on remotes
|
||||
BRANCH_NUMBER=$(check_existing_branches "$SPECS_DIR")
|
||||
else
|
||||
# Fall back to local directory check
|
||||
HIGHEST=$(get_highest_from_specs "$SPECS_DIR")
|
||||
BRANCH_NUMBER=$((HIGHEST + 1))
|
||||
fi
|
||||
fi
|
||||
|
||||
# Force base-10 interpretation to prevent octal conversion (e.g., 010 → 8 in octal, but should be 10 in decimal)
|
||||
FEATURE_NUM=$(printf "%03d" "$((10#$BRANCH_NUMBER))")
|
||||
BRANCH_NAME="${FEATURE_NUM}-${BRANCH_SUFFIX}"
|
||||
|
||||
# GitHub enforces a 244-byte limit on branch names
|
||||
# Validate and truncate if necessary
|
||||
MAX_BRANCH_LENGTH=244
|
||||
if [ ${#BRANCH_NAME} -gt $MAX_BRANCH_LENGTH ]; then
|
||||
# Calculate how much we need to trim from suffix
|
||||
# Account for: feature number (3) + hyphen (1) = 4 chars
|
||||
MAX_SUFFIX_LENGTH=$((MAX_BRANCH_LENGTH - 4))
|
||||
|
||||
# Truncate suffix at word boundary if possible
|
||||
TRUNCATED_SUFFIX=$(echo "$BRANCH_SUFFIX" | cut -c1-$MAX_SUFFIX_LENGTH)
|
||||
# Remove trailing hyphen if truncation created one
|
||||
TRUNCATED_SUFFIX=$(echo "$TRUNCATED_SUFFIX" | sed 's/-$//')
|
||||
|
||||
ORIGINAL_BRANCH_NAME="$BRANCH_NAME"
|
||||
BRANCH_NAME="${FEATURE_NUM}-${TRUNCATED_SUFFIX}"
|
||||
|
||||
>&2 echo "[specify] Warning: Branch name exceeded GitHub's 244-byte limit"
|
||||
>&2 echo "[specify] Original: $ORIGINAL_BRANCH_NAME (${#ORIGINAL_BRANCH_NAME} bytes)"
|
||||
>&2 echo "[specify] Truncated to: $BRANCH_NAME (${#BRANCH_NAME} bytes)"
|
||||
fi
|
||||
|
||||
if [ "$HAS_GIT" = true ]; then
|
||||
if ! git checkout -b "$BRANCH_NAME" 2>/dev/null; then
|
||||
# Check if branch already exists
|
||||
if git branch --list "$BRANCH_NAME" | grep -q .; then
|
||||
>&2 echo "Error: Branch '$BRANCH_NAME' already exists. Please use a different feature name or specify a different number with --number."
|
||||
exit 1
|
||||
else
|
||||
>&2 echo "Error: Failed to create git branch '$BRANCH_NAME'. Please check your git configuration and try again."
|
||||
exit 1
|
||||
fi
|
||||
fi
|
||||
else
|
||||
>&2 echo "[specify] Warning: Git repository not detected; skipped branch creation for $BRANCH_NAME"
|
||||
fi
|
||||
|
||||
FEATURE_DIR="$SPECS_DIR/$BRANCH_NAME"
|
||||
mkdir -p "$FEATURE_DIR"
|
||||
|
||||
TEMPLATE="$REPO_ROOT/.specify/templates/spec-template.md"
|
||||
SPEC_FILE="$FEATURE_DIR/spec.md"
|
||||
if [ -f "$TEMPLATE" ]; then cp "$TEMPLATE" "$SPEC_FILE"; else touch "$SPEC_FILE"; fi
|
||||
|
||||
# Set the SPECIFY_FEATURE environment variable for the current session
|
||||
export SPECIFY_FEATURE="$BRANCH_NAME"
|
||||
|
||||
if $JSON_MODE; then
|
||||
printf '{"BRANCH_NAME":"%s","SPEC_FILE":"%s","FEATURE_NUM":"%s"}\n' "$BRANCH_NAME" "$SPEC_FILE" "$FEATURE_NUM"
|
||||
else
|
||||
echo "BRANCH_NAME: $BRANCH_NAME"
|
||||
echo "SPEC_FILE: $SPEC_FILE"
|
||||
echo "FEATURE_NUM: $FEATURE_NUM"
|
||||
echo "SPECIFY_FEATURE environment variable set to: $BRANCH_NAME"
|
||||
fi
|
||||
Executable
+61
@@ -0,0 +1,61 @@
|
||||
#!/usr/bin/env bash
|
||||
|
||||
set -e
|
||||
|
||||
# Parse command line arguments
|
||||
JSON_MODE=false
|
||||
ARGS=()
|
||||
|
||||
for arg in "$@"; do
|
||||
case "$arg" in
|
||||
--json)
|
||||
JSON_MODE=true
|
||||
;;
|
||||
--help|-h)
|
||||
echo "Usage: $0 [--json]"
|
||||
echo " --json Output results in JSON format"
|
||||
echo " --help Show this help message"
|
||||
exit 0
|
||||
;;
|
||||
*)
|
||||
ARGS+=("$arg")
|
||||
;;
|
||||
esac
|
||||
done
|
||||
|
||||
# Get script directory and load common functions
|
||||
SCRIPT_DIR="$(CDPATH="" cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
source "$SCRIPT_DIR/common.sh"
|
||||
|
||||
# Get all paths and variables from common functions
|
||||
eval $(get_feature_paths)
|
||||
|
||||
# Check if we're on a proper feature branch (only for git repos)
|
||||
check_feature_branch "$CURRENT_BRANCH" "$HAS_GIT" || exit 1
|
||||
|
||||
# Ensure the feature directory exists
|
||||
mkdir -p "$FEATURE_DIR"
|
||||
|
||||
# Copy plan template if it exists
|
||||
TEMPLATE="$REPO_ROOT/.specify/templates/plan-template.md"
|
||||
if [[ -f "$TEMPLATE" ]]; then
|
||||
cp "$TEMPLATE" "$IMPL_PLAN"
|
||||
echo "Copied plan template to $IMPL_PLAN"
|
||||
else
|
||||
echo "Warning: Plan template not found at $TEMPLATE"
|
||||
# Create a basic plan file if template doesn't exist
|
||||
touch "$IMPL_PLAN"
|
||||
fi
|
||||
|
||||
# Output results
|
||||
if $JSON_MODE; then
|
||||
printf '{"FEATURE_SPEC":"%s","IMPL_PLAN":"%s","SPECS_DIR":"%s","BRANCH":"%s","HAS_GIT":"%s"}\n' \
|
||||
"$FEATURE_SPEC" "$IMPL_PLAN" "$FEATURE_DIR" "$CURRENT_BRANCH" "$HAS_GIT"
|
||||
else
|
||||
echo "FEATURE_SPEC: $FEATURE_SPEC"
|
||||
echo "IMPL_PLAN: $IMPL_PLAN"
|
||||
echo "SPECS_DIR: $FEATURE_DIR"
|
||||
echo "BRANCH: $CURRENT_BRANCH"
|
||||
echo "HAS_GIT: $HAS_GIT"
|
||||
fi
|
||||
|
||||
+855
@@ -0,0 +1,855 @@
|
||||
#!/usr/bin/env bash
|
||||
|
||||
# Update agent context files with information from plan.md
|
||||
#
|
||||
# This script maintains AI agent context files by parsing feature specifications
|
||||
# and updating agent-specific configuration files with project information.
|
||||
#
|
||||
# MAIN FUNCTIONS:
|
||||
# 1. Environment Validation
|
||||
# - Verifies git repository structure and branch information
|
||||
# - Checks for required plan.md files and templates
|
||||
# - Validates file permissions and accessibility
|
||||
#
|
||||
# 2. Plan Data Extraction
|
||||
# - Parses plan.md files to extract project metadata
|
||||
# - Identifies language/version, frameworks, databases, and project types
|
||||
# - Handles missing or incomplete specification data gracefully
|
||||
#
|
||||
# 3. Agent File Management
|
||||
# - Creates new agent context files from templates when needed
|
||||
# - Updates existing agent files with new project information
|
||||
# - Preserves manual additions and custom configurations
|
||||
# - Supports multiple AI agent formats and directory structures
|
||||
#
|
||||
# 4. Content Generation
|
||||
# - Generates language-specific build/test commands
|
||||
# - Creates appropriate project directory structures
|
||||
# - Updates technology stacks and recent changes sections
|
||||
# - Maintains consistent formatting and timestamps
|
||||
#
|
||||
# 5. Multi-Agent Support
|
||||
# - Handles agent-specific file paths and naming conventions
|
||||
# - Supports: Claude, Gemini, Copilot, Cursor, Qwen, opencode, Codex, Windsurf, Kilo Code, Auggie CLI, Roo Code, CodeBuddy CLI, Qoder CLI, Amp, SHAI, Tabnine CLI, Kiro CLI, Mistral Vibe, Kimi Code, Antigravity or Generic
|
||||
# - Can update single agents or all existing agent files
|
||||
# - Creates default Claude file if no agent files exist
|
||||
#
|
||||
# Usage: ./update-agent-context.sh [agent_type]
|
||||
# Agent types: claude|gemini|copilot|cursor-agent|qwen|opencode|codex|windsurf|kilocode|auggie|roo|codebuddy|amp|shai|tabnine|kiro-cli|agy|bob|vibe|qodercli|kimi|generic
|
||||
# Leave empty to update all existing agent files
|
||||
|
||||
set -e
|
||||
|
||||
# Enable strict error handling
|
||||
set -u
|
||||
set -o pipefail
|
||||
|
||||
#==============================================================================
|
||||
# Configuration and Global Variables
|
||||
#==============================================================================
|
||||
|
||||
# Get script directory and load common functions
|
||||
SCRIPT_DIR="$(CDPATH="" cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
source "$SCRIPT_DIR/common.sh"
|
||||
|
||||
# Get all paths and variables from common functions
|
||||
eval $(get_feature_paths)
|
||||
|
||||
NEW_PLAN="$IMPL_PLAN" # Alias for compatibility with existing code
|
||||
AGENT_TYPE="${1:-}"
|
||||
|
||||
# Agent-specific file paths
|
||||
CLAUDE_FILE="$REPO_ROOT/CLAUDE.md"
|
||||
GEMINI_FILE="$REPO_ROOT/GEMINI.md"
|
||||
COPILOT_FILE="$REPO_ROOT/.github/agents/copilot-instructions.md"
|
||||
CURSOR_FILE="$REPO_ROOT/.cursor/rules/specify-rules.mdc"
|
||||
QWEN_FILE="$REPO_ROOT/QWEN.md"
|
||||
AGENTS_FILE="$REPO_ROOT/AGENTS.md"
|
||||
WINDSURF_FILE="$REPO_ROOT/.windsurf/rules/specify-rules.md"
|
||||
KILOCODE_FILE="$REPO_ROOT/.kilocode/rules/specify-rules.md"
|
||||
AUGGIE_FILE="$REPO_ROOT/.augment/rules/specify-rules.md"
|
||||
ROO_FILE="$REPO_ROOT/.roo/rules/specify-rules.md"
|
||||
CODEBUDDY_FILE="$REPO_ROOT/CODEBUDDY.md"
|
||||
QODER_FILE="$REPO_ROOT/QODER.md"
|
||||
AMP_FILE="$REPO_ROOT/AGENTS.md"
|
||||
SHAI_FILE="$REPO_ROOT/SHAI.md"
|
||||
TABNINE_FILE="$REPO_ROOT/TABNINE.md"
|
||||
KIRO_FILE="$REPO_ROOT/AGENTS.md"
|
||||
AGY_FILE="$REPO_ROOT/.agent/rules/specify-rules.md"
|
||||
BOB_FILE="$REPO_ROOT/AGENTS.md"
|
||||
VIBE_FILE="$REPO_ROOT/.vibe/agents/specify-agents.md"
|
||||
KIMI_FILE="$REPO_ROOT/KIMI.md"
|
||||
|
||||
# Template file
|
||||
TEMPLATE_FILE="$REPO_ROOT/.specify/templates/agent-file-template.md"
|
||||
|
||||
# Global variables for parsed plan data
|
||||
NEW_LANG=""
|
||||
NEW_FRAMEWORK=""
|
||||
NEW_DB=""
|
||||
NEW_PROJECT_TYPE=""
|
||||
|
||||
#==============================================================================
|
||||
# Utility Functions
|
||||
#==============================================================================
|
||||
|
||||
log_info() {
|
||||
echo "INFO: $1"
|
||||
}
|
||||
|
||||
log_success() {
|
||||
echo "✓ $1"
|
||||
}
|
||||
|
||||
log_error() {
|
||||
echo "ERROR: $1" >&2
|
||||
}
|
||||
|
||||
log_warning() {
|
||||
echo "WARNING: $1" >&2
|
||||
}
|
||||
|
||||
# Cleanup function for temporary files
|
||||
cleanup() {
|
||||
local exit_code=$?
|
||||
rm -f /tmp/agent_update_*_$$
|
||||
rm -f /tmp/manual_additions_$$
|
||||
exit $exit_code
|
||||
}
|
||||
|
||||
# Set up cleanup trap
|
||||
trap cleanup EXIT INT TERM
|
||||
|
||||
#==============================================================================
|
||||
# Validation Functions
|
||||
#==============================================================================
|
||||
|
||||
validate_environment() {
|
||||
# Check if we have a current branch/feature (git or non-git)
|
||||
if [[ -z "$CURRENT_BRANCH" ]]; then
|
||||
log_error "Unable to determine current feature"
|
||||
if [[ "$HAS_GIT" == "true" ]]; then
|
||||
log_info "Make sure you're on a feature branch"
|
||||
else
|
||||
log_info "Set SPECIFY_FEATURE environment variable or create a feature first"
|
||||
fi
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Check if plan.md exists
|
||||
if [[ ! -f "$NEW_PLAN" ]]; then
|
||||
log_error "No plan.md found at $NEW_PLAN"
|
||||
log_info "Make sure you're working on a feature with a corresponding spec directory"
|
||||
if [[ "$HAS_GIT" != "true" ]]; then
|
||||
log_info "Use: export SPECIFY_FEATURE=your-feature-name or create a new feature first"
|
||||
fi
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Check if template exists (needed for new files)
|
||||
if [[ ! -f "$TEMPLATE_FILE" ]]; then
|
||||
log_warning "Template file not found at $TEMPLATE_FILE"
|
||||
log_warning "Creating new agent files will fail"
|
||||
fi
|
||||
}
|
||||
|
||||
#==============================================================================
|
||||
# Plan Parsing Functions
|
||||
#==============================================================================
|
||||
|
||||
extract_plan_field() {
|
||||
local field_pattern="$1"
|
||||
local plan_file="$2"
|
||||
|
||||
grep "^\*\*${field_pattern}\*\*: " "$plan_file" 2>/dev/null | \
|
||||
head -1 | \
|
||||
sed "s|^\*\*${field_pattern}\*\*: ||" | \
|
||||
sed 's/^[ \t]*//;s/[ \t]*$//' | \
|
||||
grep -v "NEEDS CLARIFICATION" | \
|
||||
grep -v "^N/A$" || echo ""
|
||||
}
|
||||
|
||||
parse_plan_data() {
|
||||
local plan_file="$1"
|
||||
|
||||
if [[ ! -f "$plan_file" ]]; then
|
||||
log_error "Plan file not found: $plan_file"
|
||||
return 1
|
||||
fi
|
||||
|
||||
if [[ ! -r "$plan_file" ]]; then
|
||||
log_error "Plan file is not readable: $plan_file"
|
||||
return 1
|
||||
fi
|
||||
|
||||
log_info "Parsing plan data from $plan_file"
|
||||
|
||||
NEW_LANG=$(extract_plan_field "Language/Version" "$plan_file")
|
||||
NEW_FRAMEWORK=$(extract_plan_field "Primary Dependencies" "$plan_file")
|
||||
NEW_DB=$(extract_plan_field "Storage" "$plan_file")
|
||||
NEW_PROJECT_TYPE=$(extract_plan_field "Project Type" "$plan_file")
|
||||
|
||||
# Log what we found
|
||||
if [[ -n "$NEW_LANG" ]]; then
|
||||
log_info "Found language: $NEW_LANG"
|
||||
else
|
||||
log_warning "No language information found in plan"
|
||||
fi
|
||||
|
||||
if [[ -n "$NEW_FRAMEWORK" ]]; then
|
||||
log_info "Found framework: $NEW_FRAMEWORK"
|
||||
fi
|
||||
|
||||
if [[ -n "$NEW_DB" ]] && [[ "$NEW_DB" != "N/A" ]]; then
|
||||
log_info "Found database: $NEW_DB"
|
||||
fi
|
||||
|
||||
if [[ -n "$NEW_PROJECT_TYPE" ]]; then
|
||||
log_info "Found project type: $NEW_PROJECT_TYPE"
|
||||
fi
|
||||
}
|
||||
|
||||
format_technology_stack() {
|
||||
local lang="$1"
|
||||
local framework="$2"
|
||||
local parts=()
|
||||
|
||||
# Add non-empty parts
|
||||
[[ -n "$lang" && "$lang" != "NEEDS CLARIFICATION" ]] && parts+=("$lang")
|
||||
[[ -n "$framework" && "$framework" != "NEEDS CLARIFICATION" && "$framework" != "N/A" ]] && parts+=("$framework")
|
||||
|
||||
# Join with proper formatting
|
||||
if [[ ${#parts[@]} -eq 0 ]]; then
|
||||
echo ""
|
||||
elif [[ ${#parts[@]} -eq 1 ]]; then
|
||||
echo "${parts[0]}"
|
||||
else
|
||||
# Join multiple parts with " + "
|
||||
local result="${parts[0]}"
|
||||
for ((i=1; i<${#parts[@]}; i++)); do
|
||||
result="$result + ${parts[i]}"
|
||||
done
|
||||
echo "$result"
|
||||
fi
|
||||
}
|
||||
|
||||
#==============================================================================
|
||||
# Template and Content Generation Functions
|
||||
#==============================================================================
|
||||
|
||||
get_project_structure() {
|
||||
local project_type="$1"
|
||||
|
||||
if [[ "$project_type" == *"web"* ]]; then
|
||||
echo "backend/\\nfrontend/\\ntests/"
|
||||
else
|
||||
echo "src/\\ntests/"
|
||||
fi
|
||||
}
|
||||
|
||||
get_commands_for_language() {
|
||||
local lang="$1"
|
||||
|
||||
case "$lang" in
|
||||
*"Python"*)
|
||||
echo "cd src && pytest && ruff check ."
|
||||
;;
|
||||
*"Rust"*)
|
||||
echo "cargo test && cargo clippy"
|
||||
;;
|
||||
*"JavaScript"*|*"TypeScript"*)
|
||||
echo "npm test \\&\\& npm run lint"
|
||||
;;
|
||||
*)
|
||||
echo "# Add commands for $lang"
|
||||
;;
|
||||
esac
|
||||
}
|
||||
|
||||
get_language_conventions() {
|
||||
local lang="$1"
|
||||
echo "$lang: Follow standard conventions"
|
||||
}
|
||||
|
||||
create_new_agent_file() {
|
||||
local target_file="$1"
|
||||
local temp_file="$2"
|
||||
local project_name="$3"
|
||||
local current_date="$4"
|
||||
|
||||
if [[ ! -f "$TEMPLATE_FILE" ]]; then
|
||||
log_error "Template not found at $TEMPLATE_FILE"
|
||||
return 1
|
||||
fi
|
||||
|
||||
if [[ ! -r "$TEMPLATE_FILE" ]]; then
|
||||
log_error "Template file is not readable: $TEMPLATE_FILE"
|
||||
return 1
|
||||
fi
|
||||
|
||||
log_info "Creating new agent context file from template..."
|
||||
|
||||
if ! cp "$TEMPLATE_FILE" "$temp_file"; then
|
||||
log_error "Failed to copy template file"
|
||||
return 1
|
||||
fi
|
||||
|
||||
# Replace template placeholders
|
||||
local project_structure
|
||||
project_structure=$(get_project_structure "$NEW_PROJECT_TYPE")
|
||||
|
||||
local commands
|
||||
commands=$(get_commands_for_language "$NEW_LANG")
|
||||
|
||||
local language_conventions
|
||||
language_conventions=$(get_language_conventions "$NEW_LANG")
|
||||
|
||||
# Perform substitutions with error checking using safer approach
|
||||
# Escape special characters for sed by using a different delimiter or escaping
|
||||
local escaped_lang=$(printf '%s\n' "$NEW_LANG" | sed 's/[\[\.*^$()+{}|]/\\&/g')
|
||||
local escaped_framework=$(printf '%s\n' "$NEW_FRAMEWORK" | sed 's/[\[\.*^$()+{}|]/\\&/g')
|
||||
local escaped_branch=$(printf '%s\n' "$CURRENT_BRANCH" | sed 's/[\[\.*^$()+{}|]/\\&/g')
|
||||
|
||||
# Build technology stack and recent change strings conditionally
|
||||
local tech_stack
|
||||
if [[ -n "$escaped_lang" && -n "$escaped_framework" ]]; then
|
||||
tech_stack="- $escaped_lang + $escaped_framework ($escaped_branch)"
|
||||
elif [[ -n "$escaped_lang" ]]; then
|
||||
tech_stack="- $escaped_lang ($escaped_branch)"
|
||||
elif [[ -n "$escaped_framework" ]]; then
|
||||
tech_stack="- $escaped_framework ($escaped_branch)"
|
||||
else
|
||||
tech_stack="- ($escaped_branch)"
|
||||
fi
|
||||
|
||||
local recent_change
|
||||
if [[ -n "$escaped_lang" && -n "$escaped_framework" ]]; then
|
||||
recent_change="- $escaped_branch: Added $escaped_lang + $escaped_framework"
|
||||
elif [[ -n "$escaped_lang" ]]; then
|
||||
recent_change="- $escaped_branch: Added $escaped_lang"
|
||||
elif [[ -n "$escaped_framework" ]]; then
|
||||
recent_change="- $escaped_branch: Added $escaped_framework"
|
||||
else
|
||||
recent_change="- $escaped_branch: Added"
|
||||
fi
|
||||
|
||||
local substitutions=(
|
||||
"s|\[PROJECT NAME\]|$project_name|"
|
||||
"s|\[DATE\]|$current_date|"
|
||||
"s|\[EXTRACTED FROM ALL PLAN.MD FILES\]|$tech_stack|"
|
||||
"s|\[ACTUAL STRUCTURE FROM PLANS\]|$project_structure|g"
|
||||
"s|\[ONLY COMMANDS FOR ACTIVE TECHNOLOGIES\]|$commands|"
|
||||
"s|\[LANGUAGE-SPECIFIC, ONLY FOR LANGUAGES IN USE\]|$language_conventions|"
|
||||
"s|\[LAST 3 FEATURES AND WHAT THEY ADDED\]|$recent_change|"
|
||||
)
|
||||
|
||||
for substitution in "${substitutions[@]}"; do
|
||||
if ! sed -i.bak -e "$substitution" "$temp_file"; then
|
||||
log_error "Failed to perform substitution: $substitution"
|
||||
rm -f "$temp_file" "$temp_file.bak"
|
||||
return 1
|
||||
fi
|
||||
done
|
||||
|
||||
# Convert \n sequences to actual newlines
|
||||
newline=$(printf '\n')
|
||||
sed -i.bak2 "s/\\\\n/${newline}/g" "$temp_file"
|
||||
|
||||
# Clean up backup files
|
||||
rm -f "$temp_file.bak" "$temp_file.bak2"
|
||||
|
||||
# Prepend Cursor frontmatter for .mdc files so rules are auto-included
|
||||
if [[ "$target_file" == *.mdc ]]; then
|
||||
local frontmatter_file
|
||||
frontmatter_file=$(mktemp) || return 1
|
||||
printf '%s\n' "---" "description: Project Development Guidelines" "globs: [\"**/*\"]" "alwaysApply: true" "---" "" > "$frontmatter_file"
|
||||
cat "$temp_file" >> "$frontmatter_file"
|
||||
mv "$frontmatter_file" "$temp_file"
|
||||
fi
|
||||
|
||||
return 0
|
||||
}
|
||||
|
||||
|
||||
|
||||
|
||||
update_existing_agent_file() {
|
||||
local target_file="$1"
|
||||
local current_date="$2"
|
||||
|
||||
log_info "Updating existing agent context file..."
|
||||
|
||||
# Use a single temporary file for atomic update
|
||||
local temp_file
|
||||
temp_file=$(mktemp) || {
|
||||
log_error "Failed to create temporary file"
|
||||
return 1
|
||||
}
|
||||
|
||||
# Process the file in one pass
|
||||
local tech_stack=$(format_technology_stack "$NEW_LANG" "$NEW_FRAMEWORK")
|
||||
local new_tech_entries=()
|
||||
local new_change_entry=""
|
||||
|
||||
# Prepare new technology entries
|
||||
if [[ -n "$tech_stack" ]] && ! grep -q "$tech_stack" "$target_file"; then
|
||||
new_tech_entries+=("- $tech_stack ($CURRENT_BRANCH)")
|
||||
fi
|
||||
|
||||
if [[ -n "$NEW_DB" ]] && [[ "$NEW_DB" != "N/A" ]] && [[ "$NEW_DB" != "NEEDS CLARIFICATION" ]] && ! grep -q "$NEW_DB" "$target_file"; then
|
||||
new_tech_entries+=("- $NEW_DB ($CURRENT_BRANCH)")
|
||||
fi
|
||||
|
||||
# Prepare new change entry
|
||||
if [[ -n "$tech_stack" ]]; then
|
||||
new_change_entry="- $CURRENT_BRANCH: Added $tech_stack"
|
||||
elif [[ -n "$NEW_DB" ]] && [[ "$NEW_DB" != "N/A" ]] && [[ "$NEW_DB" != "NEEDS CLARIFICATION" ]]; then
|
||||
new_change_entry="- $CURRENT_BRANCH: Added $NEW_DB"
|
||||
fi
|
||||
|
||||
# Check if sections exist in the file
|
||||
local has_active_technologies=0
|
||||
local has_recent_changes=0
|
||||
|
||||
if grep -q "^## Active Technologies" "$target_file" 2>/dev/null; then
|
||||
has_active_technologies=1
|
||||
fi
|
||||
|
||||
if grep -q "^## Recent Changes" "$target_file" 2>/dev/null; then
|
||||
has_recent_changes=1
|
||||
fi
|
||||
|
||||
# Process file line by line
|
||||
local in_tech_section=false
|
||||
local in_changes_section=false
|
||||
local tech_entries_added=false
|
||||
local changes_entries_added=false
|
||||
local existing_changes_count=0
|
||||
local file_ended=false
|
||||
|
||||
while IFS= read -r line || [[ -n "$line" ]]; do
|
||||
# Handle Active Technologies section
|
||||
if [[ "$line" == "## Active Technologies" ]]; then
|
||||
echo "$line" >> "$temp_file"
|
||||
in_tech_section=true
|
||||
continue
|
||||
elif [[ $in_tech_section == true ]] && [[ "$line" =~ ^##[[:space:]] ]]; then
|
||||
# Add new tech entries before closing the section
|
||||
if [[ $tech_entries_added == false ]] && [[ ${#new_tech_entries[@]} -gt 0 ]]; then
|
||||
printf '%s\n' "${new_tech_entries[@]}" >> "$temp_file"
|
||||
tech_entries_added=true
|
||||
fi
|
||||
echo "$line" >> "$temp_file"
|
||||
in_tech_section=false
|
||||
continue
|
||||
elif [[ $in_tech_section == true ]] && [[ -z "$line" ]]; then
|
||||
# Add new tech entries before empty line in tech section
|
||||
if [[ $tech_entries_added == false ]] && [[ ${#new_tech_entries[@]} -gt 0 ]]; then
|
||||
printf '%s\n' "${new_tech_entries[@]}" >> "$temp_file"
|
||||
tech_entries_added=true
|
||||
fi
|
||||
echo "$line" >> "$temp_file"
|
||||
continue
|
||||
fi
|
||||
|
||||
# Handle Recent Changes section
|
||||
if [[ "$line" == "## Recent Changes" ]]; then
|
||||
echo "$line" >> "$temp_file"
|
||||
# Add new change entry right after the heading
|
||||
if [[ -n "$new_change_entry" ]]; then
|
||||
echo "$new_change_entry" >> "$temp_file"
|
||||
fi
|
||||
in_changes_section=true
|
||||
changes_entries_added=true
|
||||
continue
|
||||
elif [[ $in_changes_section == true ]] && [[ "$line" =~ ^##[[:space:]] ]]; then
|
||||
echo "$line" >> "$temp_file"
|
||||
in_changes_section=false
|
||||
continue
|
||||
elif [[ $in_changes_section == true ]] && [[ "$line" == "- "* ]]; then
|
||||
# Keep only first 2 existing changes
|
||||
if [[ $existing_changes_count -lt 2 ]]; then
|
||||
echo "$line" >> "$temp_file"
|
||||
((existing_changes_count++))
|
||||
fi
|
||||
continue
|
||||
fi
|
||||
|
||||
# Update timestamp
|
||||
if [[ "$line" =~ \*\*Last\ updated\*\*:.*[0-9][0-9][0-9][0-9]-[0-9][0-9]-[0-9][0-9] ]]; then
|
||||
echo "$line" | sed "s/[0-9][0-9][0-9][0-9]-[0-9][0-9]-[0-9][0-9]/$current_date/" >> "$temp_file"
|
||||
else
|
||||
echo "$line" >> "$temp_file"
|
||||
fi
|
||||
done < "$target_file"
|
||||
|
||||
# Post-loop check: if we're still in the Active Technologies section and haven't added new entries
|
||||
if [[ $in_tech_section == true ]] && [[ $tech_entries_added == false ]] && [[ ${#new_tech_entries[@]} -gt 0 ]]; then
|
||||
printf '%s\n' "${new_tech_entries[@]}" >> "$temp_file"
|
||||
tech_entries_added=true
|
||||
fi
|
||||
|
||||
# If sections don't exist, add them at the end of the file
|
||||
if [[ $has_active_technologies -eq 0 ]] && [[ ${#new_tech_entries[@]} -gt 0 ]]; then
|
||||
echo "" >> "$temp_file"
|
||||
echo "## Active Technologies" >> "$temp_file"
|
||||
printf '%s\n' "${new_tech_entries[@]}" >> "$temp_file"
|
||||
tech_entries_added=true
|
||||
fi
|
||||
|
||||
if [[ $has_recent_changes -eq 0 ]] && [[ -n "$new_change_entry" ]]; then
|
||||
echo "" >> "$temp_file"
|
||||
echo "## Recent Changes" >> "$temp_file"
|
||||
echo "$new_change_entry" >> "$temp_file"
|
||||
changes_entries_added=true
|
||||
fi
|
||||
|
||||
# Ensure Cursor .mdc files have YAML frontmatter for auto-inclusion
|
||||
if [[ "$target_file" == *.mdc ]]; then
|
||||
if ! head -1 "$temp_file" | grep -q '^---'; then
|
||||
local frontmatter_file
|
||||
frontmatter_file=$(mktemp) || { rm -f "$temp_file"; return 1; }
|
||||
printf '%s\n' "---" "description: Project Development Guidelines" "globs: [\"**/*\"]" "alwaysApply: true" "---" "" > "$frontmatter_file"
|
||||
cat "$temp_file" >> "$frontmatter_file"
|
||||
mv "$frontmatter_file" "$temp_file"
|
||||
fi
|
||||
fi
|
||||
|
||||
# Move temp file to target atomically
|
||||
if ! mv "$temp_file" "$target_file"; then
|
||||
log_error "Failed to update target file"
|
||||
rm -f "$temp_file"
|
||||
return 1
|
||||
fi
|
||||
|
||||
return 0
|
||||
}
|
||||
#==============================================================================
|
||||
# Main Agent File Update Function
|
||||
#==============================================================================
|
||||
|
||||
update_agent_file() {
|
||||
local target_file="$1"
|
||||
local agent_name="$2"
|
||||
|
||||
if [[ -z "$target_file" ]] || [[ -z "$agent_name" ]]; then
|
||||
log_error "update_agent_file requires target_file and agent_name parameters"
|
||||
return 1
|
||||
fi
|
||||
|
||||
log_info "Updating $agent_name context file: $target_file"
|
||||
|
||||
local project_name
|
||||
project_name=$(basename "$REPO_ROOT")
|
||||
local current_date
|
||||
current_date=$(date +%Y-%m-%d)
|
||||
|
||||
# Create directory if it doesn't exist
|
||||
local target_dir
|
||||
target_dir=$(dirname "$target_file")
|
||||
if [[ ! -d "$target_dir" ]]; then
|
||||
if ! mkdir -p "$target_dir"; then
|
||||
log_error "Failed to create directory: $target_dir"
|
||||
return 1
|
||||
fi
|
||||
fi
|
||||
|
||||
if [[ ! -f "$target_file" ]]; then
|
||||
# Create new file from template
|
||||
local temp_file
|
||||
temp_file=$(mktemp) || {
|
||||
log_error "Failed to create temporary file"
|
||||
return 1
|
||||
}
|
||||
|
||||
if create_new_agent_file "$target_file" "$temp_file" "$project_name" "$current_date"; then
|
||||
if mv "$temp_file" "$target_file"; then
|
||||
log_success "Created new $agent_name context file"
|
||||
else
|
||||
log_error "Failed to move temporary file to $target_file"
|
||||
rm -f "$temp_file"
|
||||
return 1
|
||||
fi
|
||||
else
|
||||
log_error "Failed to create new agent file"
|
||||
rm -f "$temp_file"
|
||||
return 1
|
||||
fi
|
||||
else
|
||||
# Update existing file
|
||||
if [[ ! -r "$target_file" ]]; then
|
||||
log_error "Cannot read existing file: $target_file"
|
||||
return 1
|
||||
fi
|
||||
|
||||
if [[ ! -w "$target_file" ]]; then
|
||||
log_error "Cannot write to existing file: $target_file"
|
||||
return 1
|
||||
fi
|
||||
|
||||
if update_existing_agent_file "$target_file" "$current_date"; then
|
||||
log_success "Updated existing $agent_name context file"
|
||||
else
|
||||
log_error "Failed to update existing agent file"
|
||||
return 1
|
||||
fi
|
||||
fi
|
||||
|
||||
return 0
|
||||
}
|
||||
|
||||
#==============================================================================
|
||||
# Agent Selection and Processing
|
||||
#==============================================================================
|
||||
|
||||
update_specific_agent() {
|
||||
local agent_type="$1"
|
||||
|
||||
case "$agent_type" in
|
||||
claude)
|
||||
update_agent_file "$CLAUDE_FILE" "Claude Code"
|
||||
;;
|
||||
gemini)
|
||||
update_agent_file "$GEMINI_FILE" "Gemini CLI"
|
||||
;;
|
||||
copilot)
|
||||
update_agent_file "$COPILOT_FILE" "GitHub Copilot"
|
||||
;;
|
||||
cursor-agent)
|
||||
update_agent_file "$CURSOR_FILE" "Cursor IDE"
|
||||
;;
|
||||
qwen)
|
||||
update_agent_file "$QWEN_FILE" "Qwen Code"
|
||||
;;
|
||||
opencode)
|
||||
update_agent_file "$AGENTS_FILE" "opencode"
|
||||
;;
|
||||
codex)
|
||||
update_agent_file "$AGENTS_FILE" "Codex CLI"
|
||||
;;
|
||||
windsurf)
|
||||
update_agent_file "$WINDSURF_FILE" "Windsurf"
|
||||
;;
|
||||
kilocode)
|
||||
update_agent_file "$KILOCODE_FILE" "Kilo Code"
|
||||
;;
|
||||
auggie)
|
||||
update_agent_file "$AUGGIE_FILE" "Auggie CLI"
|
||||
;;
|
||||
roo)
|
||||
update_agent_file "$ROO_FILE" "Roo Code"
|
||||
;;
|
||||
codebuddy)
|
||||
update_agent_file "$CODEBUDDY_FILE" "CodeBuddy CLI"
|
||||
;;
|
||||
qodercli)
|
||||
update_agent_file "$QODER_FILE" "Qoder CLI"
|
||||
;;
|
||||
amp)
|
||||
update_agent_file "$AMP_FILE" "Amp"
|
||||
;;
|
||||
shai)
|
||||
update_agent_file "$SHAI_FILE" "SHAI"
|
||||
;;
|
||||
tabnine)
|
||||
update_agent_file "$TABNINE_FILE" "Tabnine CLI"
|
||||
;;
|
||||
kiro-cli)
|
||||
update_agent_file "$KIRO_FILE" "Kiro CLI"
|
||||
;;
|
||||
agy)
|
||||
update_agent_file "$AGY_FILE" "Antigravity"
|
||||
;;
|
||||
bob)
|
||||
update_agent_file "$BOB_FILE" "IBM Bob"
|
||||
;;
|
||||
vibe)
|
||||
update_agent_file "$VIBE_FILE" "Mistral Vibe"
|
||||
;;
|
||||
kimi)
|
||||
update_agent_file "$KIMI_FILE" "Kimi Code"
|
||||
;;
|
||||
generic)
|
||||
log_info "Generic agent: no predefined context file. Use the agent-specific update script for your agent."
|
||||
;;
|
||||
*)
|
||||
log_error "Unknown agent type '$agent_type'"
|
||||
log_error "Expected: claude|gemini|copilot|cursor-agent|qwen|opencode|codex|windsurf|kilocode|auggie|roo|codebuddy|amp|shai|tabnine|kiro-cli|agy|bob|vibe|qodercli|kimi|generic"
|
||||
exit 1
|
||||
;;
|
||||
esac
|
||||
}
|
||||
|
||||
update_all_existing_agents() {
|
||||
local found_agent=false
|
||||
|
||||
# Check each possible agent file and update if it exists
|
||||
if [[ -f "$CLAUDE_FILE" ]]; then
|
||||
update_agent_file "$CLAUDE_FILE" "Claude Code"
|
||||
found_agent=true
|
||||
fi
|
||||
|
||||
if [[ -f "$GEMINI_FILE" ]]; then
|
||||
update_agent_file "$GEMINI_FILE" "Gemini CLI"
|
||||
found_agent=true
|
||||
fi
|
||||
|
||||
if [[ -f "$COPILOT_FILE" ]]; then
|
||||
update_agent_file "$COPILOT_FILE" "GitHub Copilot"
|
||||
found_agent=true
|
||||
fi
|
||||
|
||||
if [[ -f "$CURSOR_FILE" ]]; then
|
||||
update_agent_file "$CURSOR_FILE" "Cursor IDE"
|
||||
found_agent=true
|
||||
fi
|
||||
|
||||
if [[ -f "$QWEN_FILE" ]]; then
|
||||
update_agent_file "$QWEN_FILE" "Qwen Code"
|
||||
found_agent=true
|
||||
fi
|
||||
|
||||
if [[ -f "$AGENTS_FILE" ]]; then
|
||||
update_agent_file "$AGENTS_FILE" "Codex/opencode"
|
||||
found_agent=true
|
||||
fi
|
||||
|
||||
if [[ -f "$WINDSURF_FILE" ]]; then
|
||||
update_agent_file "$WINDSURF_FILE" "Windsurf"
|
||||
found_agent=true
|
||||
fi
|
||||
|
||||
if [[ -f "$KILOCODE_FILE" ]]; then
|
||||
update_agent_file "$KILOCODE_FILE" "Kilo Code"
|
||||
found_agent=true
|
||||
fi
|
||||
|
||||
if [[ -f "$AUGGIE_FILE" ]]; then
|
||||
update_agent_file "$AUGGIE_FILE" "Auggie CLI"
|
||||
found_agent=true
|
||||
fi
|
||||
|
||||
if [[ -f "$ROO_FILE" ]]; then
|
||||
update_agent_file "$ROO_FILE" "Roo Code"
|
||||
found_agent=true
|
||||
fi
|
||||
|
||||
if [[ -f "$CODEBUDDY_FILE" ]]; then
|
||||
update_agent_file "$CODEBUDDY_FILE" "CodeBuddy CLI"
|
||||
found_agent=true
|
||||
fi
|
||||
|
||||
if [[ -f "$SHAI_FILE" ]]; then
|
||||
update_agent_file "$SHAI_FILE" "SHAI"
|
||||
found_agent=true
|
||||
fi
|
||||
|
||||
if [[ -f "$TABNINE_FILE" ]]; then
|
||||
update_agent_file "$TABNINE_FILE" "Tabnine CLI"
|
||||
found_agent=true
|
||||
fi
|
||||
|
||||
if [[ -f "$QODER_FILE" ]]; then
|
||||
update_agent_file "$QODER_FILE" "Qoder CLI"
|
||||
found_agent=true
|
||||
fi
|
||||
|
||||
if [[ -f "$KIRO_FILE" ]]; then
|
||||
update_agent_file "$KIRO_FILE" "Kiro CLI"
|
||||
found_agent=true
|
||||
fi
|
||||
|
||||
if [[ -f "$AGY_FILE" ]]; then
|
||||
update_agent_file "$AGY_FILE" "Antigravity"
|
||||
found_agent=true
|
||||
fi
|
||||
if [[ -f "$BOB_FILE" ]]; then
|
||||
update_agent_file "$BOB_FILE" "IBM Bob"
|
||||
found_agent=true
|
||||
fi
|
||||
|
||||
if [[ -f "$VIBE_FILE" ]]; then
|
||||
update_agent_file "$VIBE_FILE" "Mistral Vibe"
|
||||
found_agent=true
|
||||
fi
|
||||
|
||||
if [[ -f "$KIMI_FILE" ]]; then
|
||||
update_agent_file "$KIMI_FILE" "Kimi Code"
|
||||
found_agent=true
|
||||
fi
|
||||
|
||||
# If no agent files exist, create a default Claude file
|
||||
if [[ "$found_agent" == false ]]; then
|
||||
log_info "No existing agent files found, creating default Claude file..."
|
||||
update_agent_file "$CLAUDE_FILE" "Claude Code"
|
||||
fi
|
||||
}
|
||||
print_summary() {
|
||||
echo
|
||||
log_info "Summary of changes:"
|
||||
|
||||
if [[ -n "$NEW_LANG" ]]; then
|
||||
echo " - Added language: $NEW_LANG"
|
||||
fi
|
||||
|
||||
if [[ -n "$NEW_FRAMEWORK" ]]; then
|
||||
echo " - Added framework: $NEW_FRAMEWORK"
|
||||
fi
|
||||
|
||||
if [[ -n "$NEW_DB" ]] && [[ "$NEW_DB" != "N/A" ]]; then
|
||||
echo " - Added database: $NEW_DB"
|
||||
fi
|
||||
|
||||
echo
|
||||
log_info "Usage: $0 [claude|gemini|copilot|cursor-agent|qwen|opencode|codex|windsurf|kilocode|auggie|roo|codebuddy|amp|shai|tabnine|kiro-cli|agy|bob|vibe|qodercli|kimi|generic]"
|
||||
}
|
||||
|
||||
#==============================================================================
|
||||
# Main Execution
|
||||
#==============================================================================
|
||||
|
||||
main() {
|
||||
# Validate environment before proceeding
|
||||
validate_environment
|
||||
|
||||
log_info "=== Updating agent context files for feature $CURRENT_BRANCH ==="
|
||||
|
||||
# Parse the plan file to extract project information
|
||||
if ! parse_plan_data "$NEW_PLAN"; then
|
||||
log_error "Failed to parse plan data"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Process based on agent type argument
|
||||
local success=true
|
||||
|
||||
if [[ -z "$AGENT_TYPE" ]]; then
|
||||
# No specific agent provided - update all existing agent files
|
||||
log_info "No agent specified, updating all existing agent files..."
|
||||
if ! update_all_existing_agents; then
|
||||
success=false
|
||||
fi
|
||||
else
|
||||
# Specific agent provided - update only that agent
|
||||
log_info "Updating specific agent: $AGENT_TYPE"
|
||||
if ! update_specific_agent "$AGENT_TYPE"; then
|
||||
success=false
|
||||
fi
|
||||
fi
|
||||
|
||||
# Print summary
|
||||
print_summary
|
||||
|
||||
if [[ "$success" == true ]]; then
|
||||
log_success "Agent context update completed successfully"
|
||||
exit 0
|
||||
else
|
||||
log_error "Agent context update completed with errors"
|
||||
exit 1
|
||||
fi
|
||||
}
|
||||
|
||||
# Execute main function if script is run directly
|
||||
if [[ "${BASH_SOURCE[0]}" == "${0}" ]]; then
|
||||
main "$@"
|
||||
fi
|
||||
@@ -0,0 +1,187 @@
|
||||
# Feature Specification: Core Messaging
|
||||
|
||||
**Feature Branch**: `001-core-messaging`
|
||||
**Created**: 2026-03-13
|
||||
**Status**: Draft
|
||||
**Input**: User description: "Core messaging: send (DM or channel), read inbox, claim for processing, mark done/failed. Conversations, priority, status tracking, metadata, read/unread, SQLite storage, MCP tools."
|
||||
|
||||
## Constitution Compliance
|
||||
|
||||
| Principle | Status | Notes |
|
||||
|-----------|--------|-------|
|
||||
| I. Local-First, Single Binary | Compliant | SQLite embedded storage, no external dependencies |
|
||||
| II. MCP-Native | Compliant | All agent operations exposed as MCP tools only |
|
||||
| III. Pure Go, Zero CGO | Compliant | Uses `modernc.org/sqlite`, no CGO |
|
||||
| IV. Multi-Tenant with Ownership | Compliant | Messages scoped to agent identity; agents only read own inbox |
|
||||
| VI. Semantic-Ready Storage | Compliant | FTS5 full-text search now, vector search deferred to later spec |
|
||||
| VIII. Observable by Default | Compliant | All tool calls produce trace entries |
|
||||
| IX. Progressive Complexity | Compliant | This spec is tier 1 (basic messaging), no advanced features required |
|
||||
|
||||
## User Scenarios & Testing *(mandatory)*
|
||||
|
||||
### User Story 1 - Agent Sends a Direct Message and Recipient Reads It (Priority: P1)
|
||||
|
||||
An AI agent (e.g., a research agent) needs to send a direct message to another specific agent (e.g., an analysis agent). The recipient agent reads its inbox and sees the new message with full context: sender, subject, body, priority, and metadata. If no conversation exists between them on this subject, one is auto-created.
|
||||
|
||||
**Why this priority**: This is the fundamental operation of the entire system. Without send + read, no other feature has value. This is the minimal viable slice of SynapBus.
|
||||
|
||||
**Independent Test**: Can be fully tested by registering two agents, sending a message from one to the other via the `send_message` MCP tool, then calling `read_inbox` from the recipient. Delivers immediate value: two agents can communicate.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** agent "researcher" and agent "analyzer" are registered, **When** "researcher" calls `send_message` with `to_agent: "analyzer"`, `body: "Found 3 CVEs in dependency tree"`, `subject: "Security Audit Results"`, `priority: 8`, **Then** a new conversation is created with subject "Security Audit Results", the message is stored with status "pending" and priority 8, and a success response is returned containing the message ID and conversation ID.
|
||||
|
||||
2. **Given** agent "analyzer" has one pending message from "researcher", **When** "analyzer" calls `read_inbox` with no filters, **Then** the response contains exactly one message with `from_agent: "researcher"`, the full body, priority 8, status "pending", and the conversation subject. The inbox_state for "analyzer" is updated to reflect the message has been read.
|
||||
|
||||
3. **Given** agent "researcher" has already sent a message to "analyzer" with subject "Security Audit Results", **When** "researcher" sends another message with the same `subject` and same `to_agent`, **Then** the message is appended to the existing conversation (same conversation_id) rather than creating a new one.
|
||||
|
||||
4. **Given** agent "analyzer" calls `read_inbox`, **When** "analyzer" calls `read_inbox` again immediately, **Then** the previously read messages are no longer returned (unless `include_read: true` is passed), because inbox_state.last_read_message_id was advanced.
|
||||
|
||||
---
|
||||
|
||||
### User Story 2 - Agent Claims and Processes Messages (Priority: P1)
|
||||
|
||||
An agent reads its inbox, sees pending work, and claims one or more messages for processing. This prevents other agents from processing the same message (atomic claim). After finishing work, the agent marks the message as done or failed.
|
||||
|
||||
**Why this priority**: Claim-and-process is the core workflow pattern for task-oriented agents. Without it, two agents could process the same message simultaneously, causing duplicate work. This is co-equal with Story 1 for MVP.
|
||||
|
||||
**Independent Test**: Can be tested by sending a message, claiming it with `claim_messages`, verifying the status changes to "processing", then calling `mark_done`. Delivers value: reliable task processing without race conditions.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** agent "worker" has 3 pending messages in its inbox, **When** "worker" calls `claim_messages` with `limit: 2`, **Then** up to 2 messages are atomically updated to status "processing" with `claimed_by: "worker"` and `claimed_at` set to the current timestamp, and the claimed messages are returned in the response.
|
||||
|
||||
2. **Given** agent "worker" has claimed message ID 42 (status: "processing"), **When** "worker" calls `mark_done` with `message_id: 42`, **Then** the message status is updated to "done" and `updated_at` is refreshed.
|
||||
|
||||
3. **Given** agent "worker" has claimed message ID 42 (status: "processing"), **When** "worker" calls `mark_done` with `message_id: 42, status: "failed", metadata: {"error": "timeout"}`, **Then** the message status is updated to "failed" and the metadata is merged with the failure reason.
|
||||
|
||||
4. **Given** message ID 42 has status "processing" and `claimed_by: "worker-a"`, **When** agent "worker-b" calls `claim_messages` and message 42 matches the filter, **Then** message 42 is NOT included in the claim result because it is already claimed.
|
||||
|
||||
5. **Given** agent "worker-b" tries to call `mark_done` on message ID 42 which is claimed by "worker-a", **Then** the operation fails with an error indicating the message is not claimed by "worker-b".
|
||||
|
||||
---
|
||||
|
||||
### User Story 3 - Agent Sends a Channel Message (Priority: P2)
|
||||
|
||||
An agent broadcasts a message to a channel. All agents who are members of that channel see the message in their inbox. This enables one-to-many communication patterns.
|
||||
|
||||
**Why this priority**: Channel messaging extends the system from point-to-point to broadcast. It is important but not strictly required for the simplest two-agent use case. Depends on channel infrastructure which will be wired in but the channel creation/join itself is a separate spec scope.
|
||||
|
||||
**Independent Test**: Can be tested by creating a channel, adding two member agents, sending a message to the channel, and verifying both members see it in their inboxes. Delivers value: group communication for agent swarms.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** channel "security-alerts" exists with members "analyzer" and "responder", **When** agent "scanner" (also a member) calls `send_message` with `channel_id: 1`, `body: "Critical: RCE in log4j detected"`, `priority: 10`, **Then** a message is created in the channel's conversation, and both "analyzer" and "responder" see it when they call `read_inbox`.
|
||||
|
||||
2. **Given** agent "outsider" is NOT a member of channel "security-alerts", **When** "outsider" calls `send_message` with `channel_id: 1`, **Then** the operation fails with a permission error.
|
||||
|
||||
3. **Given** a channel message is sent, **When** agent "analyzer" reads it but "responder" does not, **Then** "analyzer"'s inbox_state is updated independently of "responder"'s. Each agent's read/unread state is tracked separately.
|
||||
|
||||
---
|
||||
|
||||
### User Story 4 - Agent Searches Message History (Priority: P2)
|
||||
|
||||
An agent needs to find past messages by keyword, sender, priority range, status, or time range. Full-text search via SQLite FTS5 enables keyword matching against message bodies. Metadata filters allow structured queries.
|
||||
|
||||
**Why this priority**: Search is essential for agents that need context from prior conversations, but the system is functional without it for basic send/read/process workflows.
|
||||
|
||||
**Independent Test**: Can be tested by inserting several messages with varied content and metadata, then calling `search_messages` with a query string and verifying relevant results are returned ranked by FTS5 relevance. Delivers value: agents can find historical context.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** 100 messages exist in the database with various content, **When** agent "analyzer" calls `search_messages` with `query: "deployment failure"`, **Then** messages containing "deployment" and/or "failure" in their body are returned, ordered by FTS5 relevance score, limited to messages the agent has access to (own DMs + joined channels).
|
||||
|
||||
2. **Given** messages exist with various priorities, **When** agent calls `search_messages` with `query: "error"`, `min_priority: 7`, **Then** only messages with priority >= 7 that match "error" are returned.
|
||||
|
||||
3. **Given** agent "analyzer" has DMs and channel messages, **When** "analyzer" searches with `from_agent: "scanner"`, **Then** only messages sent by "scanner" that "analyzer" has access to are returned.
|
||||
|
||||
---
|
||||
|
||||
### User Story 5 - Conversation Threading and Metadata (Priority: P3)
|
||||
|
||||
Agents use conversation threads to keep related messages grouped. Conversations have subjects and are auto-created on first message if no matching conversation exists. Messages carry rich JSON metadata that agents use for filtering and context passing.
|
||||
|
||||
**Why this priority**: Threading and metadata enrich the messaging experience but are not blockers for basic message flow. The auto-create conversation logic is implicitly exercised by Story 1 but this story covers explicit conversation management and metadata usage.
|
||||
|
||||
**Independent Test**: Can be tested by sending messages with metadata, then filtering inbox by metadata fields. Delivers value: structured agent-to-agent context passing.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** no conversation exists between "agent-a" and "agent-b" with subject "Data Pipeline", **When** "agent-a" sends a message with `subject: "Data Pipeline"` and `metadata: {"pipeline_id": "pipe-42", "stage": "extraction"}`, **Then** a new conversation is created, the message stores the metadata as JSON, and subsequent reads include the metadata.
|
||||
|
||||
2. **Given** a conversation with subject "Data Pipeline" already has 5 messages, **When** `read_inbox` is called with `conversation_id` filter, **Then** all messages in the thread are returned in chronological order with their individual metadata.
|
||||
|
||||
---
|
||||
|
||||
### Edge Cases
|
||||
|
||||
- What happens when an agent sends a message to a non-existent `to_agent`? The system MUST return a clear error ("agent not found") rather than silently creating a dead-letter message.
|
||||
- What happens when `send_message` is called with both `to_agent` and `channel_id`? The system MUST reject this as invalid; a message is either a DM or a channel message, not both.
|
||||
- What happens when an agent tries to claim messages but none are pending? The system MUST return an empty list, not an error.
|
||||
- What happens when `mark_done` is called on a message with status "done"? The system MUST return an error indicating the message is already completed (idempotency consideration: alternatively, succeed silently -- decision: fail with clear error to surface logic bugs in agents).
|
||||
- What happens when the message body is empty? The system MUST reject messages with empty or whitespace-only bodies.
|
||||
- What happens when priority is outside 1-10 range? The SQLite CHECK constraint rejects it; the MCP tool MUST validate before insert and return a user-friendly error.
|
||||
- What happens when metadata is not valid JSON? The MCP tool MUST validate metadata as valid JSON before insert and return a descriptive error.
|
||||
- What happens when two agents simultaneously try to claim the same message? Only one MUST succeed due to atomic UPDATE with WHERE status='pending'; the other gets an empty result for that message.
|
||||
- What happens when `search_messages` query is empty? The system MUST return recent messages (no FTS filter applied) with any other filters still active, behaving as a list operation.
|
||||
- What happens when the SQLite database file is locked (e.g., during backup)? The storage layer MUST use WAL mode and busy_timeout to minimize lock contention and retry transparently.
|
||||
|
||||
## Requirements *(mandatory)*
|
||||
|
||||
### Functional Requirements
|
||||
|
||||
- **FR-001**: System MUST expose `send_message` MCP tool that accepts `to_agent` (string, optional), `channel_id` (integer, optional), `body` (string, required), `subject` (string, optional), `priority` (integer 1-10, default 5), and `metadata` (JSON object, default `{}`). Exactly one of `to_agent` or `channel_id` MUST be provided.
|
||||
|
||||
- **FR-002**: System MUST auto-create a conversation when a message is sent with a subject that does not match an existing conversation between the same parties (or in the same channel). If the subject is empty, a new conversation is created for each message unless a `conversation_id` is explicitly provided.
|
||||
|
||||
- **FR-003**: System MUST expose `read_inbox` MCP tool that returns unread messages for the calling agent, ordered by priority (descending) then created_at (ascending). Supports optional filters: `status` (string), `from_agent` (string), `conversation_id` (integer), `min_priority` (integer), `limit` (integer, default 50), `include_read` (boolean, default false).
|
||||
|
||||
- **FR-004**: System MUST expose `claim_messages` MCP tool that atomically updates up to N pending messages to status "processing" with `claimed_by` set to the calling agent. Accepts `message_ids` (explicit list) or `limit` (integer) for batch claim. Uses a single SQL UPDATE with WHERE status='pending' to guarantee atomicity.
|
||||
|
||||
- **FR-005**: System MUST expose `mark_done` MCP tool that transitions a message from "processing" to "done" or "failed". Only the agent that claimed the message (matching `claimed_by`) may mark it. Accepts `message_id` (integer, required), `status` (string: "done" or "failed", default "done"), and `metadata` (JSON object, optional, merged with existing metadata).
|
||||
|
||||
- **FR-006**: System MUST expose `search_messages` MCP tool that performs full-text search via SQLite FTS5 on message bodies. Accepts `query` (string), `from_agent` (string), `to_agent` (string), `channel_id` (integer), `min_priority` (integer), `status` (string), `limit` (integer, default 20). Results are scoped to messages the calling agent has access to.
|
||||
|
||||
- **FR-007**: System MUST track read/unread state per agent per conversation in the `inbox_state` table. When `read_inbox` returns messages, the `last_read_message_id` is advanced to the highest message ID returned. Messages with ID <= `last_read_message_id` are considered read.
|
||||
|
||||
- **FR-008**: System MUST store all messages in SQLite using the schema defined in `schema/001_initial.sql` (tables: `messages`, `conversations`, `inbox_state`, `messages_fts`).
|
||||
|
||||
- **FR-009**: System MUST apply database migrations on startup. The `schema_migrations` table tracks applied versions. Migrations run sequentially and transactionally.
|
||||
|
||||
- **FR-010**: System MUST use WAL mode for SQLite and set `busy_timeout` to at least 5000ms to handle concurrent access from multiple MCP tool calls.
|
||||
|
||||
- **FR-011**: System MUST validate all MCP tool inputs (body not empty, priority in range, metadata is valid JSON, agent exists) and return structured error responses with descriptive messages.
|
||||
|
||||
- **FR-012**: System MUST record a trace entry (in the `traces` table) for every MCP tool call: `send_message`, `read_inbox`, `claim_messages`, `mark_done`, `search_messages`. The trace includes agent_name, action, details (JSON with parameters), and any error.
|
||||
|
||||
- **FR-013**: System MUST enforce access control: agents can only read messages addressed to them (DM) or posted in channels they are members of. An agent MUST NOT read another agent's DMs.
|
||||
|
||||
### Key Entities
|
||||
|
||||
- **Message**: The core unit of communication. Key attributes: `id`, `conversation_id`, `from_agent`, `to_agent` (null for channel messages), `channel_id` (null for DMs), `body`, `priority` (1-10), `status` (pending/processing/done/failed), `metadata` (JSON), `claimed_by`, `claimed_at`, `created_at`, `updated_at`. Lives in `internal/messaging/` as a Go struct and in the `messages` SQLite table.
|
||||
|
||||
- **Conversation**: Groups related messages into a thread. Key attributes: `id`, `subject`, `created_by` (agent name), `channel_id` (null for DM conversations), `created_at`, `updated_at`. A conversation is auto-created on first message if no matching conversation exists. Lives in `internal/messaging/`.
|
||||
|
||||
- **InboxState**: Tracks per-agent, per-conversation read position. Key attributes: `agent_name`, `conversation_id`, `last_read_message_id`. Used by `read_inbox` to determine which messages are unread. Lives in `internal/messaging/`.
|
||||
|
||||
- **MessageStore**: The storage interface in `internal/storage/` that provides CRUD operations for messages, conversations, and inbox state. Backed by SQLite via `modernc.org/sqlite`. Responsible for FTS5 synchronization (handled by triggers), migrations, WAL mode, and busy_timeout configuration.
|
||||
|
||||
## Success Criteria *(mandatory)*
|
||||
|
||||
### Measurable Outcomes
|
||||
|
||||
- **SC-001**: An agent can send a direct message and the recipient can read it within a single `send_message` + `read_inbox` round-trip. End-to-end latency (send to read) MUST be under 50ms for a local SQLite database.
|
||||
|
||||
- **SC-002**: Concurrent claim operations on the same set of pending messages MUST never result in a message being claimed by more than one agent. Verified by a concurrency test with 10 goroutines claiming simultaneously.
|
||||
|
||||
- **SC-003**: Full-text search via `search_messages` MUST return relevant results for a query against 10,000 stored messages in under 200ms.
|
||||
|
||||
- **SC-004**: All five MCP tools (`send_message`, `read_inbox`, `claim_messages`, `mark_done`, `search_messages`) MUST be callable via the MCP protocol (SSE transport) using a standard MCP client (e.g., `mcp-go` client).
|
||||
|
||||
- **SC-005**: Read/unread tracking MUST correctly distinguish read from unread messages: after calling `read_inbox`, a subsequent call with `include_read: false` MUST NOT return previously read messages.
|
||||
|
||||
- **SC-006**: Database migrations MUST apply cleanly on a fresh database (no pre-existing file) and MUST be idempotent (running migrations twice produces no errors or schema changes).
|
||||
|
||||
- **SC-007**: Every MCP tool call MUST produce a corresponding entry in the `traces` table, verified by querying traces after each operation in integration tests.
|
||||
|
||||
- **SC-008**: The messaging package (`internal/messaging/`) MUST have unit test coverage of at least 80% for the core service logic (send, read, claim, mark_done, search).
|
||||
@@ -0,0 +1,315 @@
|
||||
# Tasks: Core Messaging
|
||||
|
||||
**Input**: Design documents from `/specs/001-core-messaging/`
|
||||
**Prerequisites**: spec.md (required), constitution.md (required)
|
||||
|
||||
**Tests**: Included per spec requirements (SC-008 mandates 80% unit test coverage; SC-002 requires concurrency tests).
|
||||
|
||||
**Organization**: Tasks are grouped by user story to enable independent implementation and testing of each story.
|
||||
|
||||
## Format: `[ID] [P?] [Story] Description`
|
||||
|
||||
- **[P]**: Can run in parallel (different files, no dependencies)
|
||||
- **[Story]**: Which user story this task belongs to (e.g., US1, US2, US3)
|
||||
- Include exact file paths in descriptions
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Setup (Shared Infrastructure)
|
||||
|
||||
**Purpose**: Go module dependencies and build configuration
|
||||
|
||||
- [ ] T001 Add `modernc.org/sqlite` dependency to `go.mod` (pure Go SQLite driver, zero CGO)
|
||||
- [ ] T002 Add `mark3labs/mcp-go` dependency to `go.mod` (MCP server library)
|
||||
- [ ] T003 [P] Verify `CGO_ENABLED=0 go build ./cmd/synapbus` compiles cleanly with new deps
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Foundational (Blocking Prerequisites)
|
||||
|
||||
**Purpose**: Storage layer, migration runner, domain types, and trace infrastructure that MUST be complete before ANY user story can be implemented
|
||||
|
||||
**CRITICAL**: No user story work can begin until this phase is complete
|
||||
|
||||
- [ ] T004 Implement SQLite connection manager in `internal/storage/sqlite.go`: open database, enable WAL mode, set `busy_timeout=5000`, set `foreign_keys=ON`, expose `*sql.DB`. Constructor takes `ctx context.Context` and `dataDir string`. Use `modernc.org/sqlite` driver.
|
||||
|
||||
- [ ] T005 Implement migration runner in `internal/storage/migrate.go`: read `schema_migrations` table, apply unapplied `.sql` files from `schema/` directory in version order, each migration runs in a transaction, create `schema_migrations` table if not exists. Function signature: `RunMigrations(ctx context.Context, db *sql.DB, schemaDir string) error`.
|
||||
|
||||
- [ ] T006 [P] Define core domain types in `internal/messaging/types.go`: `Message` struct (id, conversation_id, from_agent, to_agent, channel_id, body, priority, status, metadata as `json.RawMessage`, claimed_by, claimed_at, created_at, updated_at), `Conversation` struct (id, subject, created_by, channel_id, created_at, updated_at), `InboxState` struct (agent_name, conversation_id, last_read_message_id). Define `MessageStatus` string constants: `StatusPending`, `StatusProcessing`, `StatusDone`, `StatusFailed`.
|
||||
|
||||
- [ ] T007 [P] Define storage interface in `internal/storage/store.go`: `MessageStore` interface with methods matching all CRUD operations needed by user stories (InsertMessage, InsertConversation, FindConversation, GetInboxMessages, UpdateInboxState, ClaimMessages, UpdateMessageStatus, SearchMessages, GetMessageByID). Each method takes `ctx context.Context` as first parameter.
|
||||
|
||||
- [ ] T008 [P] Implement trace recorder in `internal/trace/trace.go`: `Recorder` struct backed by `*sql.DB`. Method `Record(ctx context.Context, agentName, action string, details json.RawMessage, traceErr error) error` inserts into `traces` table. Use `slog` to log each trace at Info level. (Satisfies FR-012)
|
||||
|
||||
- [ ] T009 [P] Write table-driven unit tests for migration runner in `internal/storage/migrate_test.go`: test fresh DB (no schema_migrations table), test idempotency (run twice, no errors), test sequential ordering. Uses in-memory SQLite. (Satisfies SC-006)
|
||||
|
||||
- [ ] T010 [P] Write unit tests for trace recorder in `internal/trace/trace_test.go`: test successful recording, test recording with error field, verify row content in traces table.
|
||||
|
||||
- [ ] T011 Wire storage initialization into `cmd/synapbus/main.go` `runServe`: open SQLite DB via `storage.New()`, run migrations, defer close. Replace TODO placeholder. Log startup with `slog`.
|
||||
|
||||
**Checkpoint**: Foundation ready -- SQLite opens in WAL mode, migrations run, domain types defined, trace recorder works. User story implementation can now begin.
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: User Story 1 -- Agent Sends a DM and Recipient Reads It (Priority: P1)
|
||||
|
||||
**Goal**: Two agents can exchange direct messages. Sender calls `send_message`, recipient calls `read_inbox`. Conversations are auto-created. Read/unread tracking works.
|
||||
|
||||
**Independent Test**: Register two agents, send message from one to the other via `send_message` MCP tool, then call `read_inbox` from the recipient. Verify message content, priority, status, and read/unread behavior.
|
||||
|
||||
### Tests for User Story 1
|
||||
|
||||
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
|
||||
|
||||
- [ ] T012 [P] [US1] Write table-driven tests for `MessageService.SendDirectMessage` in `internal/messaging/service_test.go`: test successful send creates conversation + message, test send to non-existent agent returns error, test send with empty body returns error, test send with invalid priority returns error, test send with invalid metadata JSON returns error, test send with both to_agent and channel_id returns error, test second message with same subject+recipient reuses conversation. (Covers acceptance scenarios 1, 3 and edge cases)
|
||||
|
||||
- [ ] T013 [P] [US1] Write table-driven tests for `MessageService.ReadInbox` in `internal/messaging/service_test.go`: test returns unread messages ordered by priority desc then created_at asc, test advances last_read_message_id after read, test second read returns empty when include_read=false, test second read returns messages when include_read=true, test filters by status/from_agent/conversation_id/min_priority, test respects limit parameter. (Covers acceptance scenarios 2, 4)
|
||||
|
||||
- [ ] T014 [P] [US1] Write integration test in `internal/messaging/integration_test.go`: end-to-end test that creates a MessageStore with in-memory SQLite, sends a DM, reads inbox, verifies full round-trip. Verify trace entries exist for each operation. (Satisfies SC-001, SC-007)
|
||||
|
||||
### Implementation for User Story 1
|
||||
|
||||
- [ ] T015 [US1] Implement `MessageStore` SQLite methods for DM send in `internal/storage/sqlite_messages.go`: `InsertConversation(ctx, conv)`, `FindConversationBySubjectAndParties(ctx, subject, fromAgent, toAgent)`, `InsertMessage(ctx, msg)`, `AgentExists(ctx, agentName) (bool, error)`. All use prepared statements. (Satisfies FR-001, FR-002, FR-008)
|
||||
|
||||
- [ ] T016 [US1] Implement `MessageStore` SQLite methods for inbox read in `internal/storage/sqlite_messages.go`: `GetInboxMessages(ctx, agentName, filters)` queries messages where `to_agent=agentName` AND `id > last_read_message_id` (unless include_read), ordered by `priority DESC, created_at ASC`, with limit. `UpdateInboxState(ctx, agentName, conversationID, lastReadMsgID)` upserts into `inbox_state`. Define `InboxFilters` struct with optional fields: Status, FromAgent, ConversationID, MinPriority, Limit, IncludeRead. (Satisfies FR-003, FR-007)
|
||||
|
||||
- [ ] T017 [US1] Implement `MessageService` in `internal/messaging/service.go`: business logic layer wrapping `MessageStore`. Methods: `SendMessage(ctx, params) (*Message, error)` -- validates inputs (body not empty, priority 1-10, metadata valid JSON, exactly one of to_agent/channel_id, agent exists), finds or creates conversation, inserts message, records trace. `ReadInbox(ctx, agentName, filters) ([]Message, error)` -- fetches messages, advances inbox_state, records trace. Service holds `MessageStore` interface and `trace.Recorder`. (Satisfies FR-001, FR-002, FR-003, FR-007, FR-011, FR-012, FR-013)
|
||||
|
||||
- [ ] T018 [US1] Implement `send_message` MCP tool in `internal/mcp/tools.go`: register MCP tool with JSON Schema for parameters (to_agent, channel_id, body, subject, priority, metadata). Handler extracts calling agent identity from context, delegates to `MessageService.SendMessage`, returns structured JSON response with message_id and conversation_id. (Satisfies FR-001, SC-004)
|
||||
|
||||
- [ ] T019 [US1] Implement `read_inbox` MCP tool in `internal/mcp/tools.go`: register MCP tool with JSON Schema for parameters (status, from_agent, conversation_id, min_priority, limit, include_read). Handler extracts calling agent identity, delegates to `MessageService.ReadInbox`, returns structured JSON with messages array. (Satisfies FR-003, SC-004)
|
||||
|
||||
- [ ] T020 [US1] Implement MCP server bootstrap in `internal/mcp/server.go`: create `mcp-go` server instance, register tools, expose SSE transport endpoint. Constructor takes `MessageService` and returns configured server. Define `AgentFromContext(ctx) string` helper for extracting agent identity.
|
||||
|
||||
- [ ] T021 [US1] Wire MCP server into `cmd/synapbus/main.go`: create `MessageService` with `MessageStore` and `trace.Recorder`, create MCP server, mount SSE endpoint on chi router at `/mcp`, start HTTP server.
|
||||
|
||||
**Checkpoint**: At this point, two agents can send DMs and read their inboxes via MCP tools. Conversations auto-create. Read/unread tracking works. All acceptance scenarios for US1 verified.
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: User Story 2 -- Agent Claims and Processes Messages (Priority: P1)
|
||||
|
||||
**Goal**: Agents can atomically claim pending messages and mark them done or failed. Prevents duplicate processing.
|
||||
|
||||
**Independent Test**: Send a message, claim it with `claim_messages`, verify status changes to "processing", call `mark_done`, verify status changes to "done". Test concurrent claims to verify atomicity.
|
||||
|
||||
### Tests for User Story 2
|
||||
|
||||
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
|
||||
|
||||
- [ ] T022 [P] [US2] Write table-driven tests for `MessageService.ClaimMessages` in `internal/messaging/service_test.go`: test claim by explicit message_ids, test claim by limit, test claimed messages have status "processing" and claimed_by set, test already-claimed messages are skipped, test no pending messages returns empty list (not error), test claims are atomic (single UPDATE). (Covers acceptance scenarios 1, 4 and edge cases)
|
||||
|
||||
- [ ] T023 [P] [US2] Write table-driven tests for `MessageService.MarkDone` in `internal/messaging/service_test.go`: test mark as "done", test mark as "failed" with error metadata merge, test wrong agent cannot mark done, test already-done message returns error, test non-existent message returns error. (Covers acceptance scenarios 2, 3, 5 and edge cases)
|
||||
|
||||
- [ ] T024 [P] [US2] Write concurrency test in `internal/messaging/concurrency_test.go`: 10 goroutines simultaneously claim the same set of pending messages, verify each message is claimed by exactly one agent. Uses real SQLite (not mock) with WAL mode. (Satisfies SC-002)
|
||||
|
||||
### Implementation for User Story 2
|
||||
|
||||
- [ ] T025 [US2] Implement `MessageStore` SQLite methods for claim in `internal/storage/sqlite_messages.go`: `ClaimMessages(ctx, agentName, messageIDs []int64, limit int) ([]Message, error)` -- single `UPDATE messages SET status='processing', claimed_by=?, claimed_at=? WHERE status='pending' AND to_agent=?` with either `id IN (?)` or `LIMIT ?`. Return updated rows. (Satisfies FR-004)
|
||||
|
||||
- [ ] T026 [US2] Implement `MessageStore` SQLite methods for mark_done in `internal/storage/sqlite_messages.go`: `GetMessageByID(ctx, id) (*Message, error)`, `UpdateMessageStatus(ctx, id, status, claimedBy, metadata) error` -- verifies claimed_by matches, merges metadata JSON, updates status and updated_at. (Satisfies FR-005)
|
||||
|
||||
- [ ] T027 [US2] Implement `MessageService.ClaimMessages` and `MessageService.MarkDone` in `internal/messaging/service.go`: validation (status transitions, ownership checks), delegation to store, trace recording. (Satisfies FR-004, FR-005, FR-011, FR-012)
|
||||
|
||||
- [ ] T028 [US2] Implement `claim_messages` MCP tool in `internal/mcp/tools.go`: register tool with JSON Schema for parameters (message_ids, limit). Handler delegates to `MessageService.ClaimMessages`. (Satisfies FR-004, SC-004)
|
||||
|
||||
- [ ] T029 [US2] Implement `mark_done` MCP tool in `internal/mcp/tools.go`: register tool with JSON Schema for parameters (message_id, status, metadata). Handler delegates to `MessageService.MarkDone`. (Satisfies FR-005, SC-004)
|
||||
|
||||
**Checkpoint**: At this point, the full claim-and-process workflow works. Concurrent claims are safe. All acceptance scenarios for US2 verified.
|
||||
|
||||
---
|
||||
|
||||
## Phase 5: User Story 3 -- Agent Sends a Channel Message (Priority: P2)
|
||||
|
||||
**Goal**: An agent broadcasts a message to a channel. All channel members see it in their inboxes.
|
||||
|
||||
**Independent Test**: Create a channel with two members, send a message to the channel, verify both members see it in `read_inbox`. Verify non-members cannot send.
|
||||
|
||||
### Tests for User Story 3
|
||||
|
||||
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
|
||||
|
||||
- [ ] T030 [P] [US3] Write table-driven tests for channel message sending in `internal/messaging/service_test.go`: test send to channel creates message visible to all members, test non-member cannot send (permission error), test channel message has null to_agent, test each member's inbox_state is independent. (Covers acceptance scenarios 1, 2, 3)
|
||||
|
||||
- [ ] T031 [P] [US3] Write tests for channel helpers in `internal/storage/sqlite_channels_test.go`: test `GetChannelMembers`, test `IsChannelMember`.
|
||||
|
||||
### Implementation for User Story 3
|
||||
|
||||
- [ ] T032 [US3] Implement channel query methods in `internal/storage/sqlite_channels.go`: `GetChannelByID(ctx, id) (*Channel, error)`, `GetChannelMembers(ctx, channelID) ([]string, error)`, `IsChannelMember(ctx, channelID, agentName) (bool, error)`. Define `Channel` struct in `internal/messaging/types.go` if not already present.
|
||||
|
||||
- [ ] T033 [US3] Extend `MessageService.SendMessage` in `internal/messaging/service.go` to handle channel messages: when `channel_id` is set, verify sender is a member, create message with null `to_agent`, ensure all members see the message in `read_inbox` by querying channel membership during inbox read.
|
||||
|
||||
- [ ] T034 [US3] Extend `GetInboxMessages` in `internal/storage/sqlite_messages.go` to include channel messages: query messages where `to_agent=agentName` OR (`channel_id IN (SELECT channel_id FROM channel_members WHERE agent_name=agentName)` AND `from_agent != agentName`). Ensure read/unread state is per-agent. (Satisfies FR-003, FR-013)
|
||||
|
||||
- [ ] T035 [US3] Update `send_message` MCP tool schema in `internal/mcp/tools.go` to document `channel_id` parameter (already defined in FR-001 but implementation was DM-only in Phase 3).
|
||||
|
||||
**Checkpoint**: At this point, both DM and channel messaging work. All acceptance scenarios for US3 verified.
|
||||
|
||||
---
|
||||
|
||||
## Phase 6: User Story 4 -- Agent Searches Message History (Priority: P2)
|
||||
|
||||
**Goal**: Agents can search past messages by keyword (FTS5), sender, priority, status, and time range. Results scoped to accessible messages.
|
||||
|
||||
**Independent Test**: Insert varied messages, call `search_messages` with a query string, verify relevant results ranked by FTS5 relevance and scoped to the agent's accessible messages.
|
||||
|
||||
### Tests for User Story 4
|
||||
|
||||
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
|
||||
|
||||
- [ ] T036 [P] [US4] Write table-driven tests for `MessageService.SearchMessages` in `internal/messaging/service_test.go`: test FTS5 keyword match, test empty query returns recent messages, test min_priority filter, test from_agent filter, test access scoping (cannot search other agent's DMs), test limit parameter. (Covers acceptance scenarios 1, 2, 3 and edge cases)
|
||||
|
||||
- [ ] T037 [P] [US4] Write performance benchmark in `internal/messaging/bench_test.go`: insert 10,000 messages, search by keyword, assert latency under 200ms. (Satisfies SC-003)
|
||||
|
||||
### Implementation for User Story 4
|
||||
|
||||
- [ ] T038 [US4] Implement `SearchMessages` in `internal/storage/sqlite_messages.go`: build query joining `messages` with `messages_fts` when query is non-empty (using `messages_fts MATCH ?` with `rank` ordering), apply optional filters (from_agent, to_agent, channel_id, min_priority, status), scope to accessible messages (agent's DMs + joined channels), apply limit. When query is empty, return recent messages with filters applied. (Satisfies FR-006)
|
||||
|
||||
- [ ] T039 [US4] Implement `MessageService.SearchMessages` in `internal/messaging/service.go`: validate inputs, delegate to store, record trace. (Satisfies FR-006, FR-012)
|
||||
|
||||
- [ ] T040 [US4] Implement `search_messages` MCP tool in `internal/mcp/tools.go`: register tool with JSON Schema for parameters (query, from_agent, to_agent, channel_id, min_priority, status, limit). Handler delegates to `MessageService.SearchMessages`. (Satisfies FR-006, SC-004)
|
||||
|
||||
**Checkpoint**: At this point, full-text search works across DMs and channel messages with access scoping. All acceptance scenarios for US4 verified.
|
||||
|
||||
---
|
||||
|
||||
## Phase 7: User Story 5 -- Conversation Threading and Metadata (Priority: P3)
|
||||
|
||||
**Goal**: Messages are grouped into conversation threads. Metadata is carried as JSON and available for filtering. Conversations auto-create or reuse based on subject.
|
||||
|
||||
**Independent Test**: Send messages with metadata, filter inbox by conversation_id, verify threading and metadata roundtrip.
|
||||
|
||||
### Tests for User Story 5
|
||||
|
||||
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
|
||||
|
||||
- [ ] T041 [P] [US5] Write table-driven tests for conversation threading in `internal/messaging/service_test.go`: test auto-create conversation on new subject, test reuse conversation on matching subject+parties, test explicit conversation_id parameter, test metadata JSON roundtrip (stored and returned correctly), test filter by conversation_id returns threaded messages in chronological order. (Covers acceptance scenarios 1, 2)
|
||||
|
||||
### Implementation for User Story 5
|
||||
|
||||
- [ ] T042 [US5] Extend `FindConversationBySubjectAndParties` in `internal/storage/sqlite_messages.go` to handle edge cases: empty subject creates new conversation per message (unless conversation_id provided), subject matching is exact. Verify metadata JSON column roundtrip (store as TEXT, parse as `json.RawMessage` on read).
|
||||
|
||||
- [ ] T043 [US5] Extend `ReadInbox` in `internal/messaging/service.go` to support `conversation_id` filter that returns all messages in the thread (chronological order), including metadata. Ensure metadata is parsed and returned in MCP tool response.
|
||||
|
||||
- [ ] T044 [US5] Update `read_inbox` MCP tool response schema in `internal/mcp/tools.go` to include `conversation_subject` and `metadata` fields in each returned message (verify these are already present from earlier phases; add if missing).
|
||||
|
||||
**Checkpoint**: All user stories (US1-US5) are independently functional. Conversation threading and metadata enrichment work end-to-end.
|
||||
|
||||
---
|
||||
|
||||
## Phase 8: Polish & Cross-Cutting Concerns
|
||||
|
||||
**Purpose**: Improvements that affect multiple user stories
|
||||
|
||||
- [ ] T045 [P] Input validation hardening in `internal/messaging/service.go`: audit all validation paths -- empty body (whitespace-only check), priority range, metadata JSON parse, agent existence check, message status transitions. Ensure all return structured error messages per FR-011.
|
||||
|
||||
- [ ] T046 [P] Add `slog` structured logging throughout: `internal/storage/sqlite.go` (DB open, WAL mode, migrations), `internal/messaging/service.go` (send, read, claim, mark_done, search with agent name and params), `internal/mcp/server.go` (tool registration, request handling). Use `slog.With("agent", agentName)` for per-agent context. (Satisfies Principle VIII)
|
||||
|
||||
- [ ] T047 [P] Write unit tests for MCP tool input validation in `internal/mcp/tools_test.go`: test each tool with missing required params, invalid types, boundary values. Verify structured error responses.
|
||||
|
||||
- [ ] T048 [P] Write integration test suite in `internal/messaging/integration_test.go`: full lifecycle test covering all 5 user stories sequentially -- register agents, send DM, read inbox, claim, mark done, send channel message, search, verify traces. Uses in-memory SQLite. (Satisfies SC-004, SC-007)
|
||||
|
||||
- [ ] T049 [P] Add `Makefile` target `test-coverage` that runs `go test ./... -coverprofile=coverage.out` and verifies `internal/messaging/` package has >= 80% coverage. (Satisfies SC-008)
|
||||
|
||||
- [ ] T050 Run `make lint` and fix any issues. Run `make test` and verify all tests pass. Run `make build` with `CGO_ENABLED=0` for `darwin/arm64` and `linux/amd64` to verify cross-compilation. (Satisfies Principle III)
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & Execution Order
|
||||
|
||||
### Phase Dependencies
|
||||
|
||||
- **Setup (Phase 1)**: No dependencies -- can start immediately
|
||||
- **Foundational (Phase 2)**: Depends on Setup completion -- BLOCKS all user stories
|
||||
- **User Story 1 (Phase 3)**: Depends on Foundational phase completion
|
||||
- **User Story 2 (Phase 4)**: Depends on Foundational phase completion. Can run in parallel with US1 (different methods in same files, but shared service.go may cause merge conflicts -- recommend sequential after US1)
|
||||
- **User Story 3 (Phase 5)**: Depends on US1 completion (extends SendMessage and GetInboxMessages)
|
||||
- **User Story 4 (Phase 6)**: Depends on Foundational phase completion. Can run in parallel with US1/US2 (separate methods/files)
|
||||
- **User Story 5 (Phase 7)**: Depends on US1 completion (extends conversation and metadata handling)
|
||||
- **Polish (Phase 8)**: Depends on all user stories being complete
|
||||
|
||||
### User Story Dependencies
|
||||
|
||||
- **US1 (P1)**: Can start after Foundational (Phase 2) -- No dependencies on other stories
|
||||
- **US2 (P1)**: Can start after Foundational (Phase 2) -- Independent of US1 (different methods), but recommended after US1 to avoid merge conflicts in `service.go`
|
||||
- **US3 (P2)**: Depends on US1 -- extends `SendMessage` and `GetInboxMessages` with channel support
|
||||
- **US4 (P2)**: Can start after Foundational (Phase 2) -- Independent search implementation, but benefits from US1 test data setup patterns
|
||||
- **US5 (P3)**: Depends on US1 -- extends conversation auto-creation and metadata handling
|
||||
|
||||
### Within Each User Story
|
||||
|
||||
- Tests MUST be written and FAIL before implementation
|
||||
- Storage layer (sqlite_messages.go) before service layer (service.go)
|
||||
- Service layer before MCP tool layer (tools.go)
|
||||
- Core implementation before wiring/integration
|
||||
|
||||
### Parallel Opportunities
|
||||
|
||||
- All Setup tasks (T001-T003) can run in parallel
|
||||
- Foundational tasks T006, T007, T008, T009, T010 can all run in parallel (different files)
|
||||
- Once Foundational completes: US1 and US4 can proceed in parallel (different files)
|
||||
- Within each user story: test tasks marked [P] can run in parallel
|
||||
- Phase 8 polish tasks marked [P] can all run in parallel
|
||||
|
||||
---
|
||||
|
||||
## Implementation Strategy
|
||||
|
||||
### MVP First (User Stories 1 + 2 Only)
|
||||
|
||||
1. Complete Phase 1: Setup
|
||||
2. Complete Phase 2: Foundational (CRITICAL -- blocks all stories)
|
||||
3. Complete Phase 3: User Story 1 (send + read)
|
||||
4. **STOP and VALIDATE**: Test US1 independently
|
||||
5. Complete Phase 4: User Story 2 (claim + mark_done)
|
||||
6. **STOP and VALIDATE**: Test US2 independently
|
||||
7. Deploy/demo: two agents can send, read, claim, and process messages
|
||||
|
||||
### Incremental Delivery
|
||||
|
||||
1. Setup + Foundational -> Foundation ready
|
||||
2. US1 -> Send + Read DMs (MVP!)
|
||||
3. US2 -> Claim + Process workflow
|
||||
4. US3 -> Channel broadcasting
|
||||
5. US4 -> Full-text search
|
||||
6. US5 -> Threading + metadata enrichment
|
||||
7. Polish -> Hardening, coverage, cross-compilation
|
||||
|
||||
---
|
||||
|
||||
## Key File Map
|
||||
|
||||
| File | Purpose | Created In |
|
||||
|------|---------|------------|
|
||||
| `internal/storage/sqlite.go` | SQLite connection, WAL, busy_timeout | T004 |
|
||||
| `internal/storage/migrate.go` | Migration runner | T005 |
|
||||
| `internal/storage/store.go` | `MessageStore` interface | T007 |
|
||||
| `internal/storage/sqlite_messages.go` | Message/conversation/inbox CRUD | T015, T016, T025, T026, T038 |
|
||||
| `internal/storage/sqlite_channels.go` | Channel query methods | T032 |
|
||||
| `internal/messaging/types.go` | Domain structs + constants | T006 |
|
||||
| `internal/messaging/service.go` | `MessageService` business logic | T017, T027, T033, T039, T043 |
|
||||
| `internal/messaging/service_test.go` | Unit tests for all service methods | T012, T013, T022, T023, T030, T036, T041 |
|
||||
| `internal/messaging/concurrency_test.go` | Concurrent claim test | T024 |
|
||||
| `internal/messaging/integration_test.go` | End-to-end integration tests | T014, T048 |
|
||||
| `internal/messaging/bench_test.go` | FTS5 performance benchmark | T037 |
|
||||
| `internal/trace/trace.go` | Trace recorder | T008 |
|
||||
| `internal/trace/trace_test.go` | Trace recorder tests | T010 |
|
||||
| `internal/mcp/server.go` | MCP server bootstrap | T020 |
|
||||
| `internal/mcp/tools.go` | MCP tool registrations + handlers | T018, T019, T028, T029, T035, T040 |
|
||||
| `internal/mcp/tools_test.go` | MCP tool validation tests | T047 |
|
||||
| `cmd/synapbus/main.go` | CLI + server wiring | T011, T021 |
|
||||
| `schema/001_initial.sql` | Database schema (already exists) | -- |
|
||||
|
||||
---
|
||||
|
||||
## Notes
|
||||
|
||||
- [P] tasks = different files, no dependencies
|
||||
- [Story] label maps task to specific user story for traceability
|
||||
- Each user story should be independently completable and testable
|
||||
- Verify tests fail before implementing
|
||||
- Commit after each task or logical group
|
||||
- Stop at any checkpoint to validate story independently
|
||||
- All SQL operations use `context.Context` for cancellation support
|
||||
- `modernc.org/sqlite` is the ONLY SQLite driver allowed (zero CGO, Principle III)
|
||||
- `json.RawMessage` for metadata fields to avoid double-marshaling
|
||||
@@ -0,0 +1,145 @@
|
||||
# Feature Specification: Agent Registry & Auth
|
||||
|
||||
**Feature Branch**: `002-agent-registry`
|
||||
**Created**: 2026-03-13
|
||||
**Status**: Draft
|
||||
**Input**: User description: "Agent self-registration, API key auth, capability cards, owner-scoped access, CRUD via MCP tools"
|
||||
|
||||
## Constitution Compliance
|
||||
|
||||
| Principle | Status | Notes |
|
||||
|-----------|--------|-------|
|
||||
| I. Local-First, Single Binary | Compliant | Agent registry stored in embedded SQLite. No external auth service. |
|
||||
| II. MCP-Native | Compliant | All agent operations exposed as MCP tools. REST API used only by Web UI. |
|
||||
| III. Pure Go, Zero CGO | Compliant | No new dependencies requiring CGO. API key generation uses `crypto/rand`. |
|
||||
| IV. Multi-Tenant with Ownership | Compliant | Every agent has an `owner_id`. Agents only access own messages + joined channels. |
|
||||
| V. Embedded OAuth 2.1 | Not Applicable | OAuth is for human users. Agents use API key auth. Future spec will integrate agent keys with OAuth token flow. |
|
||||
| VIII. Observable by Default | Compliant | All registration, update, and deregistration actions generate trace entries. |
|
||||
| IX. Progressive Complexity | Compliant | Registration is Tier 1 (basic). Capability cards and discovery are Tier 2 (intermediate). |
|
||||
|
||||
## User Scenarios & Testing *(mandatory)*
|
||||
|
||||
### User Story 1 - Agent Self-Registration (Priority: P1)
|
||||
|
||||
An AI agent connects to SynapBus via MCP and calls `register_agent` with its name, display name, type, and owner credentials. SynapBus creates the agent record, generates a unique API key, and returns it in the response. The agent stores this key and uses it for all subsequent MCP tool calls. The API key is shown exactly once and never retrievable again.
|
||||
|
||||
**Why this priority**: Without registration, no agent can authenticate or use any other SynapBus feature. This is the foundational capability that everything else depends on.
|
||||
|
||||
**Independent Test**: Can be fully tested by calling `register_agent` via an MCP client and verifying the returned API key works for a subsequent `read_inbox` call. Delivers value as a standalone identity system.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a running SynapBus instance with a registered human owner (owner_id: "alice"), **When** an agent calls `register_agent` with `{name: "research-bot", display_name: "Research Bot", type: "ai", owner_id: "alice", capabilities: {"skills": ["web-search", "summarization"]}}`, **Then** the system returns `{agent_id: "<uuid>", api_key: "sk-synapbus-<random>", name: "research-bot", created_at: "<timestamp>"}` and the agent record is persisted in SQLite.
|
||||
|
||||
2. **Given** an agent "research-bot" already exists, **When** another agent calls `register_agent` with `{name: "research-bot", ...}`, **Then** the system returns an error: `{"error": "agent_name_taken", "message": "An agent with name 'research-bot' already exists"}` and no record is created.
|
||||
|
||||
3. **Given** a registered agent with a valid API key, **When** the agent includes the API key in the MCP connection headers (`Authorization: Bearer sk-synapbus-...`), **Then** all subsequent MCP tool calls are authenticated as that agent and scoped to its permissions.
|
||||
|
||||
4. **Given** a registered agent, **When** any tool call is made with an invalid or missing API key, **Then** the system returns `{"error": "unauthorized", "message": "Invalid or missing API key"}` and the request is rejected.
|
||||
|
||||
---
|
||||
|
||||
### User Story 2 - Agent Discovery by Capability (Priority: P2)
|
||||
|
||||
A coordinating agent needs to find other agents that can perform a specific task (e.g., "sentiment analysis"). It calls `discover_agents` with a capability query, and SynapBus returns a list of matching agents with their capability cards. The coordinating agent can then send messages or assign tasks to the discovered agents.
|
||||
|
||||
**Why this priority**: Discovery enables multi-agent coordination, which is core to SynapBus's swarm intelligence value proposition. Without discovery, agents must be hardcoded to know about each other.
|
||||
|
||||
**Independent Test**: Can be tested by registering 3-4 agents with different capabilities, then calling `discover_agents` with various queries and verifying correct filtering. Delivers value as a standalone agent directory.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** three registered agents: "sentiment-bot" (capabilities: `{"skills": ["sentiment-analysis", "text-classification"]}`), "search-bot" (capabilities: `{"skills": ["web-search", "summarization"]}`), and "translate-bot" (capabilities: `{"skills": ["translation", "text-classification"]}`), **When** an agent calls `discover_agents` with `{capability: "text-classification"}`, **Then** the system returns `[{name: "sentiment-bot", display_name: "Sentiment Bot", type: "ai", capabilities: {...}}, {name: "translate-bot", display_name: "Translate Bot", type: "ai", capabilities: {...}}]`.
|
||||
|
||||
2. **Given** registered agents exist, **When** an agent calls `discover_agents` with `{capability: "quantum-computing"}`, **Then** the system returns an empty list `[]` with no error.
|
||||
|
||||
3. **Given** registered agents exist, **When** an agent calls `discover_agents` with `{}` (no filter), **Then** the system returns all agents visible to the caller, with each entry including the full capability card.
|
||||
|
||||
---
|
||||
|
||||
### User Story 3 - Agent Lifecycle Management (Priority: P2)
|
||||
|
||||
An agent owner needs to update an agent's capabilities as the agent evolves (e.g., after fine-tuning adds a new skill) or deregister an agent that is no longer needed. Only the agent's owner (the human who registered it) can perform these operations. The agent itself can update its own capabilities but cannot change its owner or deregister itself.
|
||||
|
||||
**Why this priority**: Agents evolve over time. Without update and deregister, the registry becomes stale and inaccurate, undermining discovery. Owner-only control is required by Constitution Principle IV.
|
||||
|
||||
**Independent Test**: Can be tested by registering an agent, updating its capabilities, verifying the change in `discover_agents` results, then deregistering and verifying removal. Delivers value as a complete lifecycle for agent identity management.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** agent "research-bot" is registered with owner "alice" and capabilities `{"skills": ["web-search"]}`, **When** "research-bot" calls `update_agent` with `{capabilities: {"skills": ["web-search", "summarization", "citation-extraction"]}}`, **Then** the agent's capability card is updated and subsequent `discover_agents` calls reflect the new capabilities.
|
||||
|
||||
2. **Given** agent "research-bot" is registered with owner "alice", **When** owner "alice" calls `deregister_agent` with `{agent_name: "research-bot"}`, **Then** the agent record is soft-deleted (marked inactive), the API key is invalidated, and the agent no longer appears in `discover_agents` results.
|
||||
|
||||
3. **Given** agent "research-bot" is registered with owner "alice", **When** owner "bob" calls `deregister_agent` with `{agent_name: "research-bot"}`, **Then** the system returns `{"error": "forbidden", "message": "Only the agent's owner can deregister it"}` and the agent remains active.
|
||||
|
||||
4. **Given** agent "research-bot" is registered, **When** "research-bot" calls `update_agent` attempting to change `owner_id` to "bob", **Then** the system returns `{"error": "forbidden", "message": "Agents cannot change their own owner"}` and the owner remains unchanged.
|
||||
|
||||
---
|
||||
|
||||
### User Story 4 - Owner-Scoped Access Enforcement (Priority: P1)
|
||||
|
||||
When an agent makes any MCP tool call, SynapBus enforces that the agent can only access its own direct messages and channels it has explicitly joined. An agent cannot read another agent's inbox, send messages impersonating another agent, or list channels it has not joined (except public channel discovery).
|
||||
|
||||
**Why this priority**: Security isolation is a hard requirement from Constitution Principle IV. Without it, any agent could read or tamper with another agent's data, making the system untrustworthy.
|
||||
|
||||
**Independent Test**: Can be tested by registering two agents with different owners, having each send messages, and verifying that neither can read the other's inbox or access unauthorized channels. Delivers value as a security boundary.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** agent "bot-a" (owner: "alice") and agent "bot-b" (owner: "bob") are registered, and "bot-a" has received a message, **When** "bot-b" calls `read_inbox`, **Then** "bot-b" sees only its own messages and cannot see "bot-a"'s messages.
|
||||
|
||||
2. **Given** agent "bot-a" is authenticated, **When** "bot-a" calls `send_message` with `{from: "bot-b", to: "bot-c", body: "..."}`, **Then** the system rejects the request with `{"error": "forbidden", "message": "Cannot send messages as another agent"}`. The `from` field is always set server-side from the authenticated identity.
|
||||
|
||||
3. **Given** a private channel "alpha-team" that "bot-a" has not joined, **When** "bot-a" calls `list_channels`, **Then** "alpha-team" does not appear in the results. Public channels are visible for discovery purposes regardless of membership.
|
||||
|
||||
---
|
||||
|
||||
### Edge Cases
|
||||
|
||||
- **What happens when an agent registers with an empty or whitespace-only name?** The system MUST reject registration with a validation error. Agent names MUST match the pattern `^[a-z0-9][a-z0-9._-]{0,62}[a-z0-9]$` (lowercase alphanumeric, dots, hyphens, underscores; 2-64 characters).
|
||||
- **What happens when an agent's API key is compromised?** The owner MUST be able to rotate the API key via `update_agent` with `{rotate_key: true}`. The old key is immediately invalidated and a new key is returned (shown once).
|
||||
- **What happens when an owner is deleted but still has registered agents?** All agents owned by the deleted owner MUST be soft-deleted (deregistered) and their API keys invalidated. This is a cascading operation.
|
||||
- **What happens when `discover_agents` is called with a very broad query matching hundreds of agents?** Results MUST be paginated with a default limit of 50 and a maximum limit of 200. The response includes `total_count` and `next_cursor` for pagination.
|
||||
- **What happens when an agent calls `register_agent` with capabilities exceeding the size limit?** The capabilities JSON MUST be capped at 64 KB. Requests exceeding this MUST be rejected with `{"error": "payload_too_large", "message": "Capabilities JSON must not exceed 64 KB"}`.
|
||||
- **What happens when a deregistered agent's API key is used?** The system MUST return `{"error": "unauthorized", "message": "Agent has been deregistered"}` with no indication of whether the key was ever valid (to prevent enumeration).
|
||||
- **How does the system handle concurrent registration of the same agent name?** The SQLite UNIQUE constraint on `agent.name` prevents duplicates. The second concurrent request receives an `agent_name_taken` error.
|
||||
|
||||
## Requirements *(mandatory)*
|
||||
|
||||
### Functional Requirements
|
||||
|
||||
- **FR-001**: System MUST allow agents to self-register via the `register_agent` MCP tool, providing name, display_name, type, capabilities, and owner_id.
|
||||
- **FR-002**: System MUST generate a cryptographically random API key (minimum 256 bits of entropy) on registration and return it exactly once in the registration response.
|
||||
- **FR-003**: System MUST store API keys as bcrypt hashes in SQLite. Raw API keys MUST NOT be stored or logged.
|
||||
- **FR-004**: System MUST enforce unique agent names across the entire instance (case-insensitive).
|
||||
- **FR-005**: System MUST authenticate all MCP tool calls via the `Authorization: Bearer <api_key>` header, matching against stored bcrypt hashes.
|
||||
- **FR-006**: System MUST expose `discover_agents` as an MCP tool that accepts optional filters: capability keyword, agent type, and owner_id. Results MUST include the full capability card for each matching agent.
|
||||
- **FR-007**: System MUST expose `update_agent` as an MCP tool allowing the authenticated agent to update its own `display_name` and `capabilities`. Owner_id and name MUST be immutable after registration.
|
||||
- **FR-008**: System MUST expose `deregister_agent` as an MCP tool that soft-deletes the agent record. Only the agent's owner (authenticated as a human user) can deregister an agent.
|
||||
- **FR-009**: System MUST enforce owner-scoped access: agents can only read their own inbox, send messages as themselves, and access channels they have joined.
|
||||
- **FR-010**: System MUST support API key rotation via `update_agent` with `rotate_key: true`. The old key is immediately invalidated and a new key is returned.
|
||||
- **FR-011**: System MUST log all registry operations (register, update, deregister, key rotation) as trace entries with agent identity, action type, and timestamp (Constitution Principle VIII).
|
||||
- **FR-012**: System MUST validate the capabilities field as valid JSON conforming to a defined capability card schema. Invalid JSON or schema violations MUST be rejected.
|
||||
- **FR-013**: System MUST support pagination for `discover_agents` results with cursor-based pagination (default limit: 50, max: 200).
|
||||
- **FR-014**: System MUST cascade soft-delete all agents when their owner account is deleted, invalidating all associated API keys.
|
||||
|
||||
### Key Entities
|
||||
|
||||
- **Agent**: Represents a registered entity (AI or human) that can send/receive messages. Key attributes: `id` (UUID), `name` (unique, immutable), `display_name`, `type` (enum: "ai", "human"), `capabilities` (JSON capability card), `owner_id` (FK to human user), `api_key_hash` (bcrypt), `status` (enum: "active", "inactive"), `created_at`, `updated_at`, `deregistered_at`. An agent belongs to exactly one owner. An owner can have many agents.
|
||||
|
||||
- **Capability Card**: A structured JSON document describing what an agent can do. Key attributes: `skills` (array of string keywords for discovery matching), `description` (human-readable summary of the agent's purpose), `input_formats` (array of MIME types the agent can process), `output_formats` (array of MIME types the agent can produce), `version` (semver string for the agent's current version). Used by `discover_agents` for keyword matching and by other agents to understand how to interact.
|
||||
|
||||
- **Trace Entry** (registry-related): An audit record of a registry operation. Key attributes: `id`, `agent_id`, `action` (enum: "register", "update", "deregister", "key_rotate"), `details` (JSON with before/after state), `performed_by` (agent or owner who performed the action), `timestamp`. Stored in the shared traces table per Constitution Principle VIII.
|
||||
|
||||
## Success Criteria *(mandatory)*
|
||||
|
||||
### Measurable Outcomes
|
||||
|
||||
- **SC-001**: An agent can complete self-registration (call `register_agent` and receive a working API key) in under 500ms on commodity hardware.
|
||||
- **SC-002**: API key authentication adds no more than 5ms of latency to any MCP tool call (bcrypt verify is cached for active sessions).
|
||||
- **SC-003**: `discover_agents` returns results for a keyword query across 1,000 registered agents in under 200ms.
|
||||
- **SC-004**: All four MCP tools (`register_agent`, `discover_agents`, `update_agent`, `deregister_agent`) have integration tests covering the acceptance scenarios above, with 100% pass rate.
|
||||
- **SC-005**: No agent can access another agent's messages or channels it has not joined, verified by negative-path integration tests (at least 5 access-violation test cases).
|
||||
- **SC-006**: API keys are never present in logs, traces, error messages, or database records in plaintext, verified by a grep-based audit of all log output during integration tests.
|
||||
- **SC-007**: Owner cascade delete correctly deregisters all owned agents and invalidates their keys within a single SQLite transaction, verified by integration test.
|
||||
@@ -0,0 +1,218 @@
|
||||
# Tasks: Agent Registry & Auth
|
||||
|
||||
**Input**: Design documents from `/specs/002-agent-registry/`
|
||||
**Prerequisites**: spec.md (required)
|
||||
|
||||
**Tests**: Included per spec requirements (SC-004, SC-005 mandate integration tests with 100% pass rate and negative-path access-violation tests).
|
||||
|
||||
**Organization**: Tasks are grouped by user story to enable independent implementation and testing of each story.
|
||||
|
||||
## Format: `[ID] [P?] [Story] Description`
|
||||
|
||||
- **[P]**: Can run in parallel (different files, no dependencies)
|
||||
- **[Story]**: Which user story this task belongs to (e.g., US1, US2, US3, US4)
|
||||
- Include exact file paths in descriptions
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Setup (Shared Infrastructure)
|
||||
|
||||
**Purpose**: Project initialization, schema migration, and base types
|
||||
|
||||
- [ ] T001 Create SQLite migration `schema/002_agent_registry.sql`: add `deregistered_at TIMESTAMP` column to `agents` table, add case-insensitive unique index on `LOWER(agents.name)`, add index on `agents.capabilities` for FTS on the skills JSON field. The `agents` table already exists in `001_initial.sql` so this is an ALTER migration.
|
||||
- [ ] T002 [P] Define Agent domain model in `internal/agents/agent.go`: `Agent` struct (ID, Name, DisplayName, Type, Capabilities, OwnerID, APIKeyHash, Status, CreatedAt, UpdatedAt, DeregisteredAt), `CapabilityCard` struct (Skills []string, Description string, InputFormats []string, OutputFormats []string, Version string), validation constants (name regex `^[a-z0-9][a-z0-9._-]{0,62}[a-z0-9]$`, max capabilities size 64KB), and agent status constants.
|
||||
- [ ] T003 [P] Define request/response types in `internal/agents/dto.go`: `RegisterRequest`, `RegisterResponse` (agent_id, api_key, name, created_at), `UpdateRequest` (display_name, capabilities, rotate_key), `UpdateResponse`, `DeregisterRequest`, `DiscoverRequest` (capability, type, owner_id, limit, cursor), `DiscoverResponse` (agents, total_count, next_cursor).
|
||||
- [ ] T004 [P] Define agent-related error types in `internal/agents/errors.go`: sentinel errors for `ErrAgentNameTaken`, `ErrAgentNotFound`, `ErrInvalidAgentName`, `ErrCapabilitiesTooLarge`, `ErrInvalidCapabilities`, `ErrForbidden`, `ErrUnauthorized`, `ErrAgentDeregistered`. Each error should carry a machine-readable code string (e.g., `agent_name_taken`) for MCP error responses.
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Foundational (Blocking Prerequisites)
|
||||
|
||||
**Purpose**: Core infrastructure that MUST be complete before ANY user story can be implemented
|
||||
|
||||
**CRITICAL**: No user story work can begin until this phase is complete
|
||||
|
||||
- [ ] T005 Implement API key generation in `internal/auth/apikey.go`: `GenerateAPIKey() (raw string, hash string, error)` using `crypto/rand` for 256 bits of entropy, prefix `sk-synapbus-`, base64url encoding. `HashAPIKey(raw string) (string, error)` using bcrypt. `VerifyAPIKey(raw, hash string) bool` using bcrypt.CompareHashAndPassword. Include table-driven tests in `internal/auth/apikey_test.go` covering generation uniqueness, hash verification, invalid key rejection.
|
||||
- [ ] T006 [P] Implement agent SQLite repository in `internal/agents/repository.go`: `Repository` interface with methods `Create(ctx, Agent) error`, `GetByName(ctx, name) (Agent, error)`, `GetByID(ctx, id) (Agent, error)`, `Update(ctx, Agent) error`, `SoftDelete(ctx, name) error`, `FindByCapability(ctx, keyword, limit, cursor) ([]Agent, int, string, error)`, `FindAll(ctx, limit, cursor) ([]Agent, int, string, error)`, `FindByOwner(ctx, ownerID) ([]Agent, error)`, `SoftDeleteByOwner(ctx, ownerID) error`. Implement `SQLiteRepository` struct using `modernc.org/sqlite`. Cursor-based pagination with default limit 50, max 200. All queries filter `status = 'active'` unless explicitly requesting inactive.
|
||||
- [ ] T007 [P] Implement agent repository tests in `internal/agents/repository_test.go`: table-driven tests for all repository methods. Test cases: create agent, duplicate name rejection (case-insensitive), get by name, get by ID, update capabilities, soft-delete, find by capability keyword, pagination (limit/cursor), find by owner, cascade soft-delete by owner, get deregistered agent returns error.
|
||||
- [ ] T008 Implement authentication middleware in `internal/auth/middleware.go`: `AgentAuthMiddleware` that extracts `Authorization: Bearer <key>` from MCP request headers, looks up active agents by verifying the key against all stored bcrypt hashes (optimize with in-memory cache of active agent key hashes), injects authenticated agent identity into `context.Context`. Return `ErrUnauthorized` for missing/invalid keys. Return `ErrAgentDeregistered` for inactive agents (same generic unauthorized message to prevent enumeration per FR-003/edge case). Use `slog` to log auth attempts (without logging the key itself). Include tests in `internal/auth/middleware_test.go`.
|
||||
- [ ] T009 [P] Implement trace recording for registry operations in `internal/trace/registry.go`: `RecordRegistration(ctx, agentName, details)`, `RecordUpdate(ctx, agentName, beforeAfter)`, `RecordDeregistration(ctx, agentName, performedBy)`, `RecordKeyRotation(ctx, agentName)`. Each writes to the existing `traces` table with action types: `register`, `update`, `deregister`, `key_rotate`. Details stored as JSON. Use `slog` for structured logging per Constitution Principle VIII.
|
||||
|
||||
**Checkpoint**: Foundation ready - user story implementation can now begin in parallel
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: User Story 1 - Agent Self-Registration (Priority: P1) MVP
|
||||
|
||||
**Goal**: An AI agent can call `register_agent` via MCP, receive an API key, and use it to authenticate all subsequent MCP tool calls.
|
||||
|
||||
**Independent Test**: Call `register_agent` via MCP client, verify returned API key works for a subsequent authenticated tool call.
|
||||
|
||||
### Tests for User Story 1
|
||||
|
||||
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
|
||||
|
||||
- [ ] T010 [P] [US1] Integration test for agent registration in `internal/agents/service_test.go`: table-driven tests covering all acceptance scenarios: (1) successful registration returns agent_id, api_key, name, created_at; (2) duplicate name returns `agent_name_taken` error; (3) empty/invalid name returns validation error; (4) name regex enforcement; (5) capabilities exceeding 64KB rejected; (6) invalid capabilities JSON rejected; (7) API key hash stored (not raw key); (8) trace entry created on registration.
|
||||
- [ ] T011 [P] [US1] Integration test for API key authentication in `internal/auth/middleware_test.go`: table-driven tests covering: (1) valid API key authenticates successfully and populates context with agent identity; (2) invalid API key returns unauthorized; (3) missing Authorization header returns unauthorized; (4) malformed Bearer token returns unauthorized; (5) deregistered agent's key returns unauthorized with generic message.
|
||||
|
||||
### Implementation for User Story 1
|
||||
|
||||
- [ ] T012 [US1] Implement agent service in `internal/agents/service.go`: `Service` struct with `Register(ctx, RegisterRequest) (RegisterResponse, error)` method. Validation: name regex `^[a-z0-9][a-z0-9._-]{0,62}[a-z0-9]$`, capabilities size <= 64KB, capabilities valid JSON conforming to CapabilityCard schema. Generate API key via `auth.GenerateAPIKey()`, store bcrypt hash, return raw key exactly once. Record trace entry. Accept `context.Context` as first param. Log with `slog`.
|
||||
- [ ] T013 [US1] Register `register_agent` MCP tool in `internal/mcp/tools_agents.go`: define MCP tool with JSON Schema for input (name, display_name, type, owner_id, capabilities). Handler calls `agents.Service.Register()`. Tool description follows MCP conventions. Wire into the MCP server tool registry. Return structured JSON response with agent_id, api_key, name, created_at.
|
||||
- [ ] T014 [US1] Wire authentication middleware into MCP server in `internal/mcp/server.go`: apply `AgentAuthMiddleware` to all MCP tool calls except `register_agent` (which is unauthenticated since the agent has no key yet). Ensure authenticated agent identity is available in context for all downstream handlers.
|
||||
|
||||
**Checkpoint**: At this point, agents can self-register and authenticate. User Story 1 is fully functional and testable independently.
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: User Story 2 - Agent Discovery by Capability (Priority: P2)
|
||||
|
||||
**Goal**: A coordinating agent can find other agents by capability keyword, receiving a list of matching agents with their full capability cards.
|
||||
|
||||
**Independent Test**: Register 3-4 agents with different capabilities, call `discover_agents` with various queries, verify correct filtering and pagination.
|
||||
|
||||
### Tests for User Story 2
|
||||
|
||||
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
|
||||
|
||||
- [ ] T015 [P] [US2] Integration test for agent discovery in `internal/agents/service_test.go`: table-driven tests covering: (1) discover by capability keyword returns matching agents with full capability cards; (2) discover with non-matching keyword returns empty list (no error); (3) discover with no filter returns all active agents; (4) pagination works with default limit 50; (5) pagination cursor returns correct next page; (6) deregistered agents excluded from results; (7) results include display_name, type, capabilities for each agent.
|
||||
|
||||
### Implementation for User Story 2
|
||||
|
||||
- [ ] T016 [US2] Implement discovery service method in `internal/agents/service.go`: `Discover(ctx, DiscoverRequest) (DiscoverResponse, error)` method. Query agents by capability keyword (match against `skills` array in capabilities JSON), agent type, and/or owner_id. Return paginated results with `total_count` and `next_cursor`. Filter out inactive (deregistered) agents. Accept `context.Context`, log with `slog`.
|
||||
- [ ] T017 [US2] Register `discover_agents` MCP tool in `internal/mcp/tools_agents.go`: define MCP tool with JSON Schema for input (capability, type, owner_id, limit, cursor -- all optional). Handler calls `agents.Service.Discover()`. Return structured JSON response with agents array, total_count, next_cursor. This tool requires authentication (caller must be a registered agent).
|
||||
|
||||
**Checkpoint**: At this point, User Stories 1 AND 2 are both independently functional. Agents can register and discover each other.
|
||||
|
||||
---
|
||||
|
||||
## Phase 5: User Story 3 - Agent Lifecycle Management (Priority: P2)
|
||||
|
||||
**Goal**: Agents can update their own display_name and capabilities. Owners can deregister agents. API key rotation is supported.
|
||||
|
||||
**Independent Test**: Register an agent, update capabilities, verify change in discovery results, deregister, verify removal.
|
||||
|
||||
### Tests for User Story 3
|
||||
|
||||
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
|
||||
|
||||
- [ ] T018 [P] [US3] Integration test for agent update in `internal/agents/service_test.go`: table-driven tests covering: (1) agent updates own capabilities successfully; (2) updated capabilities reflected in discover_agents; (3) agent cannot change own owner_id (forbidden); (4) agent cannot change own name (immutable); (5) API key rotation returns new key and invalidates old; (6) trace entry created on update; (7) trace entry created on key rotation.
|
||||
- [ ] T019 [P] [US3] Integration test for agent deregistration in `internal/agents/service_test.go`: table-driven tests covering: (1) owner deregisters own agent successfully (soft-delete); (2) deregistered agent no longer appears in discover_agents; (3) deregistered agent's API key is invalidated; (4) non-owner cannot deregister another owner's agent (forbidden); (5) agent cannot deregister itself; (6) trace entry created on deregistration; (7) owner cascade: deleting owner soft-deletes all owned agents.
|
||||
|
||||
### Implementation for User Story 3
|
||||
|
||||
- [ ] T020 [US3] Implement update service method in `internal/agents/service.go`: `Update(ctx, agentName, UpdateRequest) (UpdateResponse, error)` method. Authenticated agent can update own `display_name` and `capabilities`. Reject changes to `owner_id` and `name` with `ErrForbidden`. If `rotate_key: true`, generate new API key, bcrypt hash it, invalidate old, return new key exactly once. Record trace entries for update and/or key rotation. Validate capabilities schema and size.
|
||||
- [ ] T021 [US3] Implement deregister service method in `internal/agents/service.go`: `Deregister(ctx, DeregisterRequest) error` method. Only the agent's owner (authenticated as human user) can deregister. Set `status = 'inactive'`, set `deregistered_at = NOW()`. Invalidate API key hash (overwrite with empty/invalid hash). Record trace entry with `performed_by` as the owner. Implement `DeregisterByOwner(ctx, ownerID) error` for cascade owner deletion.
|
||||
- [ ] T022 [US3] Register `update_agent` and `deregister_agent` MCP tools in `internal/mcp/tools_agents.go`: define MCP tools with JSON Schema for inputs. `update_agent` handler: extract authenticated agent from context, call `agents.Service.Update()`. `deregister_agent` handler: verify caller is agent's owner (not the agent itself), call `agents.Service.Deregister()`. Both require authentication.
|
||||
|
||||
**Checkpoint**: All CRUD lifecycle operations work. User Stories 1, 2, and 3 are independently functional.
|
||||
|
||||
---
|
||||
|
||||
## Phase 6: User Story 4 - Owner-Scoped Access Enforcement (Priority: P1)
|
||||
|
||||
**Goal**: Agents can only access their own messages, send as themselves, and see channels they have joined. Security isolation is enforced at the service layer.
|
||||
|
||||
**Independent Test**: Register two agents with different owners, have each send messages, verify neither can read the other's inbox or access unauthorized channels.
|
||||
|
||||
### Tests for User Story 4
|
||||
|
||||
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
|
||||
|
||||
- [ ] T023 [P] [US4] Negative-path access violation tests in `internal/agents/access_test.go`: at least 5 access-violation test cases per SC-005: (1) agent B cannot read agent A's inbox; (2) agent A cannot send messages with `from` set to agent B (server overrides `from` with authenticated identity); (3) agent A cannot list private channels it has not joined; (4) agent A cannot read messages from a channel it has not joined; (5) agent A cannot deregister agent B (owned by different owner); (6) agent A cannot impersonate agent B in any tool call.
|
||||
- [ ] T024 [P] [US4] Integration test for owner-scoped query filtering in `internal/agents/access_test.go`: table-driven tests verifying: (1) `read_inbox` returns only authenticated agent's messages; (2) `list_channels` excludes private channels agent has not joined; (3) public channels visible regardless of membership; (4) `send_message` always sets `from` to authenticated agent identity (server-side).
|
||||
|
||||
### Implementation for User Story 4
|
||||
|
||||
- [ ] T025 [US4] Implement access control layer in `internal/agents/access.go`: `AccessControl` struct with methods: `CanReadInbox(ctx, callerAgent, targetAgent) error`, `CanSendAs(ctx, callerAgent, fromAgent) error`, `CanAccessChannel(ctx, callerAgent, channelID) error`, `CanDeregister(ctx, callerIdentity, targetAgent) error`. Each returns nil on success, `ErrForbidden` with descriptive message on violation. Extract caller identity from `context.Context` (set by auth middleware).
|
||||
- [ ] T026 [US4] Integrate access control into MCP tool handlers in `internal/mcp/server.go` and `internal/mcp/tools_agents.go`: add access control checks before all message operations (`send_message`, `read_inbox`, `list_channels`). Ensure `send_message` always overrides the `from` field with the authenticated agent's identity (server-side enforcement). Add access check to `deregister_agent` verifying caller is the agent's owner.
|
||||
- [ ] T027 [US4] Add server-side `from` field enforcement in `internal/messaging/service.go` (or create if needed): any message send operation MUST set `from_agent` to the authenticated caller's agent name, ignoring any client-provided value. Log a warning via `slog` if a client attempts to set a different `from` value.
|
||||
|
||||
**Checkpoint**: Security isolation is enforced. All four user stories are independently functional and tested.
|
||||
|
||||
---
|
||||
|
||||
## Phase 7: Polish & Cross-Cutting Concerns
|
||||
|
||||
**Purpose**: Improvements that affect multiple user stories
|
||||
|
||||
- [ ] T028 [P] API key security audit in `internal/auth/apikey_audit_test.go`: grep-based test (per SC-006) that runs integration tests and verifies API keys are never present in logs, traces, error messages, or database records in plaintext. Scan all `slog` output and `traces` table entries for key patterns matching `sk-synapbus-`.
|
||||
- [ ] T029 [P] Performance benchmark tests in `internal/agents/bench_test.go`: (1) registration completes in < 500ms (SC-001); (2) API key auth adds < 5ms latency (SC-002); (3) `discover_agents` across 1,000 agents returns in < 200ms (SC-003). Use `testing.B` for benchmarks.
|
||||
- [ ] T030 [P] Add comprehensive `slog` structured logging across all agent registry operations: ensure every `Register`, `Update`, `Deregister`, `Discover`, and `KeyRotate` operation logs agent identity, action type, and outcome at appropriate levels (Info for success, Warn for access violations, Error for failures). Never log raw API keys.
|
||||
- [ ] T031 Validate end-to-end flow: register agent -> authenticate -> discover agents -> update capabilities -> rotate key -> re-authenticate with new key -> deregister. Manual or scripted validation using MCP client against running SynapBus instance. Verify all trace entries are recorded.
|
||||
- [ ] T032 [P] Add bcrypt verification cache in `internal/auth/cache.go`: in-memory LRU cache mapping `sha256(raw_key) -> agent_identity` to avoid bcrypt.CompareHashAndPassword on every request (per SC-002). Cache entries invalidated on key rotation or agent deregistration. TTL-based expiry (e.g., 5 minutes). Include tests in `internal/auth/cache_test.go`.
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & Execution Order
|
||||
|
||||
### Phase Dependencies
|
||||
|
||||
- **Setup (Phase 1)**: No dependencies - can start immediately
|
||||
- **Foundational (Phase 2)**: Depends on Setup completion - BLOCKS all user stories
|
||||
- **User Stories (Phase 3-6)**: All depend on Foundational phase completion
|
||||
- US1 (Phase 3) and US4 (Phase 6) are both P1 priority
|
||||
- US1 MUST complete before US4 (access control depends on working auth)
|
||||
- US2 (Phase 4) can start after Phase 2, independent of US1
|
||||
- US3 (Phase 5) can start after Phase 2, independent of US1
|
||||
- **Polish (Phase 7)**: Depends on all user stories being complete
|
||||
|
||||
### User Story Dependencies
|
||||
|
||||
- **User Story 1 (P1)**: Can start after Foundational (Phase 2) - No dependencies on other stories. This is the MVP.
|
||||
- **User Story 2 (P2)**: Can start after Foundational (Phase 2) - Uses repository layer but no dependency on US1 service methods.
|
||||
- **User Story 3 (P2)**: Can start after Foundational (Phase 2) - Uses repository layer and API key generation from Phase 2.
|
||||
- **User Story 4 (P1)**: Depends on US1 completion (needs working auth middleware). Integrates with messaging/channels services from other features.
|
||||
|
||||
### Within Each User Story
|
||||
|
||||
- Tests MUST be written and FAIL before implementation
|
||||
- Domain model/DTOs before service layer
|
||||
- Service layer before MCP tool handlers
|
||||
- Access control checks before integration wiring
|
||||
- Story complete before moving to next priority
|
||||
|
||||
### Parallel Opportunities
|
||||
|
||||
- All Setup tasks marked [P] can run in parallel (T002, T003, T004)
|
||||
- All Foundational tasks marked [P] can run in parallel (T006, T007, T009 alongside T005)
|
||||
- T008 depends on T005 (needs API key functions)
|
||||
- Once Foundational phase completes, US2 and US3 can start in parallel with US1
|
||||
- All tests for a user story marked [P] can run in parallel
|
||||
- All Polish tasks marked [P] can run in parallel
|
||||
|
||||
---
|
||||
|
||||
## Implementation Strategy
|
||||
|
||||
### MVP First (User Story 1 Only)
|
||||
|
||||
1. Complete Phase 1: Setup (schema migration, domain types)
|
||||
2. Complete Phase 2: Foundational (API key gen, repository, auth middleware, trace)
|
||||
3. Complete Phase 3: User Story 1 (register_agent MCP tool, auth wiring)
|
||||
4. **STOP and VALIDATE**: Test registration + authentication independently
|
||||
5. Deploy/demo if ready
|
||||
|
||||
### Incremental Delivery
|
||||
|
||||
1. Complete Setup + Foundational -> Foundation ready
|
||||
2. Add User Story 1 -> Test independently -> Deploy/Demo (MVP!)
|
||||
3. Add User Story 2 -> Test independently -> Deploy/Demo (agents can discover each other)
|
||||
4. Add User Story 3 -> Test independently -> Deploy/Demo (full CRUD lifecycle)
|
||||
5. Add User Story 4 -> Test independently -> Deploy/Demo (security isolation enforced)
|
||||
6. Polish phase -> Performance, audit, caching
|
||||
|
||||
---
|
||||
|
||||
## Notes
|
||||
|
||||
- [P] tasks = different files, no dependencies
|
||||
- [Story] label maps task to specific user story for traceability
|
||||
- Each user story should be independently completable and testable
|
||||
- Verify tests fail before implementing
|
||||
- Commit after each task or logical group
|
||||
- Stop at any checkpoint to validate story independently
|
||||
- All `context.Context` must be propagated through function signatures
|
||||
- All logging via `slog` with structured fields (never log raw API keys)
|
||||
- Table-driven tests for all test files
|
||||
- `modernc.org/sqlite` only (zero CGO constraint)
|
||||
@@ -0,0 +1,137 @@
|
||||
# Feature Specification: Human Auth (OAuth 2.1)
|
||||
|
||||
**Feature Branch**: `003-human-auth`
|
||||
**Created**: 2026-03-13
|
||||
**Status**: Draft
|
||||
**Input**: User description: "OAuth 2.1 authorization server embedded in SynapBus using fosite, with local accounts, token endpoints, PKCE, session management, refresh token rotation, and user CRUD."
|
||||
|
||||
## User Scenarios & Testing *(mandatory)*
|
||||
|
||||
### User Story 1 - Human Owner Logs Into Web UI (Priority: P1)
|
||||
|
||||
A human user opens the SynapBus Web UI in their browser, creates a local account with a username and password, and logs in. After authentication, the browser holds an httponly session cookie that grants access to all Web UI endpoints. The user can view their dashboard, see their agents, and browse messages without re-authenticating until the session expires or they log out.
|
||||
|
||||
**Why this priority**: Without human authentication, no other feature in SynapBus can enforce ownership or multi-tenancy. This is the foundation for Principle IV (Multi-Tenant with Ownership). A working login flow is the single most critical auth capability.
|
||||
|
||||
**Independent Test**: Can be fully tested by starting `synapbus serve`, opening the Web UI, creating an account via the registration form, logging in, and verifying the session cookie is set and subsequent API calls succeed. Delivers value as a standalone access-control gate for the Web UI.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a running SynapBus instance with no accounts, **When** a user navigates to `/register` and submits a valid username (3-64 chars, alphanumeric + underscore) and password (minimum 8 characters), **Then** the account is created, the password is stored as a bcrypt hash (cost >= 12), and the user is redirected to the login page.
|
||||
2. **Given** a registered user, **When** they submit valid credentials to the OAuth authorization code flow with PKCE (`/oauth/authorize` -> `/oauth/token`), **Then** the server issues an access token and refresh token, sets an httponly secure cookie containing the session identifier, and redirects to the Web UI dashboard.
|
||||
3. **Given** a logged-in user with a valid session cookie, **When** they make requests to protected Web UI API endpoints, **Then** the server validates the session and returns the requested data with the user's identity in context.
|
||||
4. **Given** a logged-in user, **When** they click "Log out", **Then** the session is invalidated server-side, the session cookie is cleared, and subsequent requests redirect to the login page.
|
||||
|
||||
---
|
||||
|
||||
### User Story 2 - Programmatic Client Obtains Tokens via Client Credentials (Priority: P2)
|
||||
|
||||
An external automation script or CI pipeline needs to interact with the SynapBus REST API (e.g., to provision agents or query system status). The operator registers an OAuth client with a `client_id` and `client_secret`, then uses the `client_credentials` grant type to obtain an access token from `/oauth/token`. The token is used in `Authorization: Bearer <token>` headers for subsequent API calls.
|
||||
|
||||
**Why this priority**: Client credentials flow enables programmatic access without a browser, which is essential for DevOps automation, agent provisioning scripts, and integration testing. It is the second most common auth flow after interactive login.
|
||||
|
||||
**Independent Test**: Can be tested by registering an OAuth client via `synapbus` CLI or admin API, then using `curl` to POST to `/oauth/token` with `grant_type=client_credentials`, verifying a valid access token is returned, and using that token to call a protected endpoint.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a registered OAuth client with `client_id` and `client_secret`, **When** the client sends a POST to `/oauth/token` with `grant_type=client_credentials` and valid credentials, **Then** the server returns a JSON response containing `access_token`, `token_type: "bearer"`, `expires_in` (default 3600 seconds), and `scope`.
|
||||
2. **Given** a valid access token obtained via client credentials, **When** the token is sent in the `Authorization: Bearer` header to `/oauth/introspect`, **Then** the server responds with `active: true`, the client identity, granted scopes, and expiration time.
|
||||
3. **Given** an expired access token, **When** it is presented to any protected endpoint, **Then** the server responds with HTTP 401 and a `WWW-Authenticate` header indicating the token is expired.
|
||||
|
||||
---
|
||||
|
||||
### User Story 3 - Refresh Token Rotation (Priority: P2)
|
||||
|
||||
A logged-in Web UI user's access token expires after its TTL. The browser (or a programmatic client using authorization_code grant) silently refreshes the session by exchanging the refresh token for a new access token and a new refresh token. The old refresh token is invalidated immediately upon use, preventing replay attacks.
|
||||
|
||||
**Why this priority**: Refresh token rotation is critical for security in long-lived sessions. Without it, a stolen refresh token grants indefinite access. This is a P2 because it builds on top of the P1 login flow and is required for production-grade security.
|
||||
|
||||
**Independent Test**: Can be tested by obtaining a token pair, waiting for the access token to expire (or using a short TTL in test config), exchanging the refresh token at `/oauth/token` with `grant_type=refresh_token`, verifying a new token pair is returned, and confirming the old refresh token is rejected on a second attempt.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a user with a valid refresh token, **When** they POST to `/oauth/token` with `grant_type=refresh_token` and the current refresh token, **Then** the server returns a new access token and a new refresh token, and the old refresh token is marked as consumed in storage.
|
||||
2. **Given** a refresh token that has already been used (consumed), **When** a client attempts to use it again, **Then** the server rejects the request with HTTP 401 and revokes all tokens in the grant chain (the new refresh token issued from the original is also revoked) as a security precaution against token theft.
|
||||
3. **Given** a refresh token that has exceeded its absolute lifetime (e.g., 30 days), **When** a client attempts to use it, **Then** the server rejects the request with HTTP 401 and the user must re-authenticate.
|
||||
|
||||
---
|
||||
|
||||
### User Story 4 - User Account Management (Priority: P3)
|
||||
|
||||
A logged-in human user manages their account through the Web UI or REST API. They can change their password (requiring the current password for verification), view the list of agents they own, and see their active OAuth sessions. An admin user (the first account created, or one explicitly granted admin role) can list all users and deactivate accounts.
|
||||
|
||||
**Why this priority**: Account management is a supporting capability. Users can function with a fixed password and no self-service management in an MVP, but password changes and agent listing are necessary for production use.
|
||||
|
||||
**Independent Test**: Can be tested by logging in, calling `PUT /api/users/me/password` with old and new passwords, verifying the old password no longer works and the new one does. Agent listing can be tested by creating agents under the user and calling `GET /api/users/me/agents`.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a logged-in user, **When** they submit a password change request with their correct current password and a new password meeting the minimum length requirement, **Then** the password is updated (bcrypt re-hashed), all existing sessions except the current one are invalidated, and the server returns HTTP 200.
|
||||
2. **Given** a logged-in user who owns 3 registered agents, **When** they call `GET /api/users/me/agents`, **Then** the server returns a JSON array of their 3 agents with `agent_id`, `name`, `display_name`, `type`, and `created_at` fields.
|
||||
3. **Given** a user submitting a password change with an incorrect current password, **When** the request is processed, **Then** the server returns HTTP 403 and the password remains unchanged.
|
||||
|
||||
---
|
||||
|
||||
### User Story 5 - PKCE Enforcement on Authorization Code Flow (Priority: P1)
|
||||
|
||||
Any client using the authorization code grant MUST include a PKCE `code_challenge` in the authorization request and the corresponding `code_verifier` when exchanging the code for tokens. Requests without PKCE parameters are rejected. This prevents authorization code interception attacks, even for confidential clients.
|
||||
|
||||
**Why this priority**: PKCE is mandated by Constitution Principle V and OAuth 2.1 specification. It is not optional; it is a security requirement baked into the protocol. This is P1 because the authorization code flow (User Story 1) cannot ship without it.
|
||||
|
||||
**Independent Test**: Can be tested by sending an authorization request to `/oauth/authorize` without `code_challenge` and verifying it is rejected, then sending one with a valid S256 code challenge and verifying the flow completes when the correct `code_verifier` is provided at the token endpoint.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a client initiating an authorization code flow, **When** the request to `/oauth/authorize` omits the `code_challenge` parameter, **Then** the server responds with HTTP 400 and an error `invalid_request` indicating PKCE is required.
|
||||
2. **Given** a client that included a valid `code_challenge` (method S256) in the authorization request and received an authorization code, **When** the client exchanges the code at `/oauth/token` with the correct `code_verifier`, **Then** the server validates `SHA256(code_verifier) == code_challenge` and issues tokens.
|
||||
3. **Given** a client that provides an incorrect `code_verifier` at the token endpoint, **When** the exchange is attempted, **Then** the server responds with HTTP 400 and an error `invalid_grant`, and the authorization code is invalidated (single-use enforcement).
|
||||
|
||||
---
|
||||
|
||||
### Edge Cases
|
||||
|
||||
- What happens when a user tries to register a username that already exists? The server MUST return HTTP 409 Conflict with a clear error message. It MUST NOT reveal whether the username exists via timing differences (use constant-time comparison or always hash before checking).
|
||||
- What happens when the bcrypt cost factor makes login unacceptably slow on low-powered hardware? The system MUST use a configurable bcrypt cost (minimum 10, default 12) so operators can tune it for their hardware.
|
||||
- What happens when the SQLite database holding OAuth tokens becomes corrupted or is deleted while sessions are active? All sessions become invalid, and users must re-authenticate. The server MUST handle missing/corrupt token storage gracefully by logging the error and rejecting all token validations rather than panicking.
|
||||
- What happens when a client sends a `code_challenge_method` of `plain` instead of `S256`? The server MUST reject it. Only S256 is supported per OAuth 2.1 requirements.
|
||||
- What happens when multiple concurrent requests attempt to use the same refresh token simultaneously? Only the first request succeeds. Subsequent requests MUST trigger revocation of the entire token family as a potential replay attack.
|
||||
- What happens when the system has zero registered users and a request hits a protected endpoint? The server MUST redirect to the registration page (Web UI) or return HTTP 401 (API), never expose data.
|
||||
- What happens when a user attempts to register with a password shorter than 8 characters? The server MUST return HTTP 400 with a validation error. Password requirements: minimum 8 characters, no maximum length cap (up to 72 bytes, the bcrypt limit), no character class requirements.
|
||||
|
||||
## Requirements *(mandatory)*
|
||||
|
||||
### Functional Requirements
|
||||
|
||||
- **FR-001**: System MUST embed an OAuth 2.1 authorization server using `ory/fosite`, started as part of `synapbus serve`, with no external service dependencies.
|
||||
- **FR-002**: System MUST support local account creation with `username` (unique, 3-64 chars, `[a-zA-Z0-9_]`) and `password` (8+ characters, bcrypt hashed with configurable cost, default 12).
|
||||
- **FR-003**: System MUST expose `/oauth/authorize` endpoint supporting the authorization code grant with PKCE (S256 only). The `plain` challenge method MUST be rejected.
|
||||
- **FR-004**: System MUST expose `/oauth/token` endpoint supporting `grant_type=authorization_code` (with PKCE verification) and `grant_type=client_credentials`.
|
||||
- **FR-005**: System MUST expose `/oauth/introspect` endpoint per RFC 7662, returning token validity, scopes, client identity, and expiration.
|
||||
- **FR-006**: System MUST implement refresh token rotation: each refresh token use issues a new refresh token and invalidates the old one. Reuse of a consumed refresh token MUST revoke the entire token family.
|
||||
- **FR-007**: System MUST manage Web UI sessions via httponly, secure (when TLS enabled), SameSite=Lax cookies. Session lifetime MUST be configurable (default 24 hours).
|
||||
- **FR-008**: System MUST provide user CRUD operations: create account (`POST /api/users`), change password (`PUT /api/users/me/password`), list owned agents (`GET /api/users/me/agents`), get current user profile (`GET /api/users/me`).
|
||||
- **FR-009**: System MUST enforce that every agent has an `owner_id` referencing a valid user. Users MUST only see agents they own (unless admin).
|
||||
- **FR-010**: System MUST store all OAuth data (clients, tokens, authorization codes, sessions) in the embedded SQLite database within the `--data` directory.
|
||||
- **FR-011**: System MUST support registering OAuth clients via CLI command (`synapbus client create --name <name>`) or admin API, generating a `client_id` and `client_secret` pair.
|
||||
- **FR-012**: Access tokens MUST have a configurable TTL (default 1 hour). Refresh tokens MUST have a configurable absolute lifetime (default 30 days).
|
||||
- **FR-013**: System MUST log all authentication events (login success, login failure, token issuance, token refresh, token revocation) via `slog` with structured fields including user/client identity and remote IP.
|
||||
|
||||
### Key Entities
|
||||
|
||||
- **User**: A human account. Attributes: `id` (UUID), `username` (unique), `password_hash` (bcrypt), `role` (user/admin), `created_at`, `updated_at`. Owns zero or more Agents. Owns zero or more OAuthClients.
|
||||
- **OAuthClient**: A registered OAuth client (Web UI frontend, CLI tool, or external automation). Attributes: `client_id` (generated, unique), `client_secret_hash` (bcrypt), `name`, `redirect_uris` (JSON array), `grant_types` (JSON array), `scopes` (JSON array), `owner_id` (FK to User), `created_at`.
|
||||
- **OAuthToken**: An issued access or refresh token. Managed internally by fosite's storage interface. Attributes include: `signature` (token lookup key), `client_id`, `user_id`, `scopes`, `expires_at`, `created_at`, `session_data` (JSON).
|
||||
- **AuthorizationCode**: A short-lived code issued during the authorization code flow. Attributes: `code` (hashed), `client_id`, `user_id`, `redirect_uri`, `scopes`, `code_challenge`, `code_challenge_method`, `expires_at`, `created_at`, `used` (boolean).
|
||||
- **Session**: A Web UI session linking a browser cookie to a user identity. Attributes: `session_id` (random, stored in cookie), `user_id`, `created_at`, `expires_at`, `last_active_at`.
|
||||
|
||||
## Success Criteria *(mandatory)*
|
||||
|
||||
### Measurable Outcomes
|
||||
|
||||
- **SC-001**: A new user can create an account and log into the Web UI in under 30 seconds (wall-clock time from opening registration page to seeing the dashboard).
|
||||
- **SC-002**: The full authorization code flow with PKCE (authorize -> code -> token exchange) completes in under 500ms server-side processing time (excluding network latency and user interaction).
|
||||
- **SC-003**: Client credentials token issuance responds in under 100ms at the 99th percentile under 50 concurrent requests.
|
||||
- **SC-004**: Refresh token rotation correctly invalidates old tokens in 100% of cases, verified by an automated test that attempts reuse of consumed refresh tokens and confirms rejection.
|
||||
- **SC-005**: All 13 functional requirements have corresponding integration tests that pass in CI, covering both success paths and error paths (invalid credentials, expired tokens, missing PKCE, duplicate usernames).
|
||||
- **SC-006**: Zero authentication endpoints are accessible without TLS in production mode (when `--tls` flag is set). In development mode (`--dev`), HTTP is permitted with a logged warning.
|
||||
- **SC-007**: The auth subsystem adds zero external runtime dependencies -- verified by building with `CGO_ENABLED=0` and running all auth tests against the embedded SQLite store.
|
||||
@@ -0,0 +1,253 @@
|
||||
# Tasks: Human Auth (OAuth 2.1)
|
||||
|
||||
**Input**: Design documents from `/specs/003-human-auth/`
|
||||
**Prerequisites**: spec.md (required)
|
||||
|
||||
**Tests**: Integration and unit tests are included per SC-005 which requires all 13 functional requirements to have corresponding tests covering both success and error paths.
|
||||
|
||||
**Organization**: Tasks are grouped by user story to enable independent implementation and testing of each story.
|
||||
|
||||
## Format: `[ID] [P?] [Story] Description`
|
||||
|
||||
- **[P]**: Can run in parallel (different files, no dependencies)
|
||||
- **[Story]**: Which user story this task belongs to (e.g., US1, US2, US3)
|
||||
- Include exact file paths in descriptions
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Setup (Shared Infrastructure)
|
||||
|
||||
**Purpose**: Add auth dependencies and create package skeleton
|
||||
|
||||
- [ ] T001 Add `ory/fosite` and `golang.org/x/crypto` (bcrypt) dependencies to `go.mod` via `go get`
|
||||
- [ ] T002 [P] Create package skeleton: `internal/auth/` directory with placeholder files `internal/auth/doc.go`, `internal/auth/oauth.go`, `internal/auth/user.go`, `internal/auth/session.go`, `internal/auth/client.go`, `internal/auth/middleware.go`
|
||||
- [ ] T003 [P] Create auth configuration struct in `internal/auth/config.go` with fields: bcrypt cost (default 12), access token TTL (default 1h), refresh token lifetime (default 30d), session lifetime (default 24h), issuer URL, dev mode flag
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Foundational (Blocking Prerequisites)
|
||||
|
||||
**Purpose**: Schema migrations, core domain types, fosite storage adapter, and auth middleware that ALL user stories depend on
|
||||
|
||||
**CRITICAL**: No user story work can begin until this phase is complete
|
||||
|
||||
- [ ] T004 Create SQLite migration `schema/002_auth.sql`: add `role` column (`TEXT NOT NULL DEFAULT 'user' CHECK (role IN ('user', 'admin'))`) to `users` table; add `scopes` and `owner_id` columns to `oauth_clients` table; create `oauth_sessions` table (`session_id TEXT PRIMARY KEY, user_id INTEGER NOT NULL REFERENCES users(id), created_at, expires_at, last_active_at`); create `oauth_authorization_codes` table (`code TEXT PRIMARY KEY, client_id, user_id, redirect_uri, scopes, code_challenge, code_challenge_method, expires_at, created_at, used INTEGER DEFAULT 0`); extend `oauth_tokens` table with `token_type TEXT` (access/refresh), `session_data TEXT` (JSON), `consumed INTEGER DEFAULT 0`, `parent_signature TEXT` columns; add indexes on new columns
|
||||
- [ ] T005 Register migration 002 in storage migration runner (extend existing migration logic in `internal/storage/` to apply `002_auth.sql`)
|
||||
- [ ] T006 [P] Define User domain model in `internal/auth/user.go`: `User` struct with `ID`, `Username`, `PasswordHash`, `DisplayName`, `Role`, `CreatedAt`, `UpdatedAt`; `UserStore` interface with `CreateUser`, `GetUserByID`, `GetUserByUsername`, `UpdatePassword`, `ListUsers`, `CountUsers`
|
||||
- [ ] T007 [P] Define OAuthClient domain model in `internal/auth/client.go`: `OAuthClient` struct implementing `fosite.Client` interface with `ID`, `SecretHash`, `Name`, `RedirectURIs`, `GrantTypes`, `Scopes`, `OwnerID`, `CreatedAt`; `ClientStore` interface with `CreateClient`, `GetClient`, `ListClientsByOwner`
|
||||
- [ ] T008 [P] Define Session domain model in `internal/auth/session.go`: `Session` struct with `SessionID`, `UserID`, `CreatedAt`, `ExpiresAt`, `LastActiveAt`; `SessionStore` interface with `CreateSession`, `GetSession`, `DeleteSession`, `DeleteSessionsByUser`, `DeleteSessionsByUserExcept`
|
||||
- [ ] T009 Implement `UserStore` backed by SQLite in `internal/auth/user_store.go`: all CRUD methods, bcrypt hashing/verification helpers, constant-time username existence check
|
||||
- [ ] T010 Implement `ClientStore` backed by SQLite in `internal/auth/client_store.go`: implements both `ClientStore` interface and `fosite.ClientManager` (`GetClient` returns `fosite.Client`)
|
||||
- [ ] T011 Implement `SessionStore` backed by SQLite in `internal/auth/session_store.go`: all session CRUD, expiration checking, cleanup of expired sessions
|
||||
- [ ] T012 Implement fosite storage adapter in `internal/auth/fosite_store.go`: struct implementing `fosite.Storage`, `oauth2.CoreStorage`, `oauth2.TokenRevocationStorage`, `pkce.PKCERequestStorage` interfaces; backed by SQLite tables for authorization codes, access tokens, refresh tokens; handles refresh token rotation and family revocation on reuse
|
||||
- [ ] T013 Configure fosite OAuth2 provider in `internal/auth/provider.go`: compose fosite with `oauth2.AuthorizeExplicitFactory`, `oauth2.ClientCredentialsGrantFactory`, `oauth2.RefreshTokenGrantFactory`; enforce PKCE S256-only; set token lifetimes from config; create `NewOAuthProvider(config, store) fosite.OAuth2Provider`
|
||||
- [ ] T014 Implement auth middleware in `internal/auth/middleware.go`: `RequireSession` middleware (checks session cookie, injects user into `context.Context`); `RequireBearer` middleware (validates access token via fosite introspection, injects client/user identity into context); `RequireAdmin` middleware (checks user role); context helper functions `UserFromContext(ctx)`, `ClientFromContext(ctx)`
|
||||
- [ ] T015 Implement structured auth event logging in `internal/auth/logging.go`: `LogAuthEvent(ctx, slog.Logger, event)` function; event types: `login_success`, `login_failure`, `token_issued`, `token_refreshed`, `token_revoked`, `session_created`, `session_destroyed`, `user_created`, `password_changed`; includes user/client identity and remote IP (FR-013)
|
||||
|
||||
**Checkpoint**: Foundation ready -- user story implementation can now begin in parallel
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: User Story 1 -- Human Owner Logs Into Web UI (Priority: P1) MVP
|
||||
|
||||
**Goal**: A human user can register a local account, log in via OAuth 2.1 authorization code flow with PKCE, and access the Web UI with a session cookie.
|
||||
|
||||
**Independent Test**: Start `synapbus serve`, open the Web UI, create an account via `/register`, log in, verify session cookie is set and subsequent API calls succeed.
|
||||
|
||||
### Tests for User Story 1
|
||||
|
||||
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
|
||||
|
||||
- [ ] T016 [P] [US1] Integration test for user registration in `internal/auth/user_store_test.go`: create user with valid credentials, verify bcrypt hash cost >= 12; reject duplicate username with 409; reject short password; reject invalid username characters; verify constant-time behavior (no timing leak on existence check)
|
||||
- [ ] T017 [P] [US1] Integration test for authorization code + PKCE flow in `internal/auth/oauth_test.go`: full flow test using fosite test helpers -- authorize request with S256 code_challenge -> obtain code -> exchange with code_verifier -> verify access + refresh tokens; verify session cookie is set as httponly/SameSite=Lax
|
||||
- [ ] T018 [P] [US1] Integration test for session middleware in `internal/auth/middleware_test.go`: request with valid session cookie succeeds; request with expired session returns 401; request with no session redirects to login; request after logout is rejected
|
||||
|
||||
### Implementation for User Story 1
|
||||
|
||||
- [ ] T019 [US1] Implement registration handler in `internal/api/auth_handlers.go`: `POST /api/users` -- validate username (3-64 chars, `[a-zA-Z0-9_]`), validate password (8+ chars, max 72 bytes bcrypt limit), create user via `UserStore`, return 201 with user profile (no password hash); return 409 on duplicate username, 400 on validation errors (FR-002)
|
||||
- [ ] T020 [US1] Implement OAuth authorize endpoint in `internal/api/auth_handlers.go`: `GET /oauth/authorize` -- validate client_id, redirect_uri, response_type=code; require code_challenge with method S256 (reject plain and missing); if user not logged in, render login form; on valid credentials, generate authorization code via fosite and redirect with code (FR-003)
|
||||
- [ ] T021 [US1] Implement OAuth token endpoint in `internal/api/auth_handlers.go`: `POST /oauth/token` -- delegate to fosite for `grant_type=authorization_code` with PKCE verification; on success, create server-side session, set httponly secure SameSite=Lax cookie, return access_token + refresh_token JSON (FR-004)
|
||||
- [ ] T022 [US1] Implement logout handler in `internal/api/auth_handlers.go`: `POST /api/auth/logout` -- delete session from store, clear session cookie, return 200
|
||||
- [ ] T023 [US1] Implement `GET /api/users/me` handler in `internal/api/auth_handlers.go`: return current user profile (id, username, display_name, role, created_at) from session context (FR-008)
|
||||
- [ ] T024 [US1] Register auth routes on chi router in `internal/api/router.go`: mount `/oauth/authorize`, `/oauth/token`, `/api/users` (POST), `/api/users/me` (GET), `/api/auth/logout` (POST); apply `RequireSession` middleware to protected routes; handle zero-users redirect to registration
|
||||
- [ ] T025 [US1] Wire auth subsystem into `synapbus serve` command in `cmd/synapbus/main.go` (or `cmd/synapbus/serve.go`): initialize auth config from flags/env, create stores, create fosite provider, register handlers; add `--dev` flag to allow HTTP without TLS warning (FR-001)
|
||||
- [ ] T026 [US1] Add auth event logging calls to all handlers: login success/failure, session created/destroyed, user created (FR-013)
|
||||
|
||||
**Checkpoint**: User Story 1 fully functional -- user can register, log in with PKCE, access protected endpoints, log out
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: User Story 5 -- PKCE Enforcement (Priority: P1) MVP
|
||||
|
||||
**Goal**: Authorization code flow without PKCE is rejected; only S256 is accepted; incorrect code_verifier is rejected with code invalidation.
|
||||
|
||||
**Independent Test**: Send authorization request without `code_challenge` and verify rejection; send with S256 and verify full flow; send incorrect `code_verifier` and verify rejection.
|
||||
|
||||
> Note: PKCE enforcement is implemented as part of the fosite provider configuration (T013) and authorize/token endpoints (T020, T021). This phase validates and hardens it.
|
||||
|
||||
### Tests for User Story 5
|
||||
|
||||
- [ ] T027 [P] [US5] Integration test for PKCE enforcement in `internal/auth/pkce_test.go`: request without code_challenge returns 400 `invalid_request`; request with `code_challenge_method=plain` returns 400; correct S256 flow succeeds; incorrect code_verifier returns 400 `invalid_grant` and code is invalidated (single-use); reuse of authorization code is rejected
|
||||
|
||||
### Implementation for User Story 5
|
||||
|
||||
- [ ] T028 [US5] Harden PKCE validation in fosite provider config (`internal/auth/provider.go`): ensure `fosite.MinParameterEntropy` is set appropriately; confirm S256-only enforcement; add explicit rejection message for `plain` method; ensure authorization codes are single-use (FR-003)
|
||||
- [ ] T029 [US5] Add PKCE-specific error responses in `internal/api/auth_handlers.go`: ensure `/oauth/authorize` returns descriptive error when PKCE is missing or uses `plain`; ensure `/oauth/token` returns descriptive error on verifier mismatch
|
||||
|
||||
**Checkpoint**: PKCE fully enforced -- authorization code flow cannot proceed without valid S256 challenge/verifier
|
||||
|
||||
---
|
||||
|
||||
## Phase 5: User Story 2 -- Programmatic Client Credentials (Priority: P2)
|
||||
|
||||
**Goal**: External automation can register an OAuth client and obtain tokens via `client_credentials` grant to access the REST API.
|
||||
|
||||
**Independent Test**: Register OAuth client via CLI, POST to `/oauth/token` with `grant_type=client_credentials`, verify valid access token, use token to call protected endpoint.
|
||||
|
||||
### Tests for User Story 2
|
||||
|
||||
- [ ] T030 [P] [US2] Integration test for client credentials flow in `internal/auth/client_credentials_test.go`: valid client_id + secret returns access_token with correct expires_in and scope; invalid secret returns 401; unknown client_id returns 401; token introspection returns `active: true` with client identity and scopes
|
||||
- [ ] T031 [P] [US2] Integration test for token introspection in `internal/auth/introspect_test.go`: valid token returns active=true with client_id, scopes, exp; expired token returns active=false; revoked token returns active=false; malformed token returns active=false
|
||||
|
||||
### Implementation for User Story 2
|
||||
|
||||
- [ ] T032 [US2] Extend OAuth token endpoint to handle `grant_type=client_credentials` in `internal/api/auth_handlers.go`: fosite already supports this via factory (T013), ensure handler routes the grant type correctly; return `access_token`, `token_type: "bearer"`, `expires_in`, `scope` (FR-004)
|
||||
- [ ] T033 [US2] Implement token introspection endpoint in `internal/api/auth_handlers.go`: `POST /oauth/introspect` per RFC 7662 -- validate bearer token, return `active`, `client_id`, `scope`, `exp`, `iat`, `token_type`; protect with client authentication (FR-005)
|
||||
- [ ] T034 [US2] Implement `synapbus client create` CLI command in `cmd/synapbus/client.go`: `--name <name>` flag, generates random `client_id` and `client_secret`, stores bcrypt-hashed secret via `ClientStore`, prints credentials to stdout; optional `--grant-types` and `--scopes` flags (FR-011)
|
||||
- [ ] T035 [US2] Register introspection route and bearer middleware on chi router in `internal/api/router.go`: mount `/oauth/introspect`; ensure `RequireBearer` middleware validates access tokens for API endpoints; return 401 with `WWW-Authenticate` header on expired tokens
|
||||
- [ ] T036 [US2] Add auth event logging for client credentials: token issued, introspection events (FR-013)
|
||||
|
||||
**Checkpoint**: User Stories 1, 5, and 2 functional -- both interactive and programmatic auth work
|
||||
|
||||
---
|
||||
|
||||
## Phase 6: User Story 3 -- Refresh Token Rotation (Priority: P2)
|
||||
|
||||
**Goal**: Access tokens can be refreshed silently; old refresh tokens are invalidated on use; reuse of a consumed refresh token revokes the entire token family.
|
||||
|
||||
**Independent Test**: Obtain token pair, exchange refresh token for new pair at `/oauth/token`, verify old refresh token is rejected, verify family revocation on replay.
|
||||
|
||||
### Tests for User Story 3
|
||||
|
||||
- [ ] T037 [P] [US3] Integration test for refresh token rotation in `internal/auth/refresh_test.go`: exchange valid refresh token returns new access + refresh token; old refresh token is rejected on reuse; reuse of consumed token triggers family revocation (all descendant tokens revoked); refresh token beyond absolute lifetime (30d) is rejected with 401
|
||||
- [ ] T038 [P] [US3] Unit test for token family revocation logic in `internal/auth/fosite_store_test.go`: create token chain (parent -> child -> grandchild); reuse parent triggers revocation of child and grandchild; verify all tokens in chain are inactive after revocation
|
||||
|
||||
### Implementation for User Story 3
|
||||
|
||||
- [ ] T039 [US3] Implement refresh token rotation in fosite store (`internal/auth/fosite_store.go`): on `CreateRefreshTokenSession`, mark parent token as consumed and store `parent_signature` on new token; on `RevokeRefreshTokenMaybeGracePeriod`, implement family revocation by walking `parent_signature` chain; ensure concurrent reuse of same token triggers revocation (FR-006)
|
||||
- [ ] T040 [US3] Extend token endpoint handling for `grant_type=refresh_token` in `internal/api/auth_handlers.go`: fosite handles validation via factory, ensure handler returns new access_token + refresh_token pair; return 401 on consumed/expired refresh tokens (FR-006)
|
||||
- [ ] T041 [US3] Add auth event logging for refresh operations: `token_refreshed`, `token_revoked` (family revocation) with parent/child token signatures (FR-013)
|
||||
|
||||
**Checkpoint**: Token lifecycle complete -- access tokens refresh silently, security against token theft via family revocation
|
||||
|
||||
---
|
||||
|
||||
## Phase 7: User Story 4 -- User Account Management (Priority: P3)
|
||||
|
||||
**Goal**: Users can change passwords, view owned agents, and view their profile. Admins can list and deactivate users.
|
||||
|
||||
**Independent Test**: Log in, call `PUT /api/users/me/password` with old + new password, verify old no longer works; call `GET /api/users/me/agents` and verify owned agents listed.
|
||||
|
||||
### Tests for User Story 4
|
||||
|
||||
- [ ] T042 [P] [US4] Integration test for password change in `internal/auth/user_store_test.go` (extend): correct current password + valid new password succeeds; sessions except current are invalidated; incorrect current password returns 403; new password below 8 chars returns 400
|
||||
- [ ] T043 [P] [US4] Integration test for user agents listing in `internal/api/auth_handlers_test.go`: user with 3 agents sees all 3; user with 0 agents sees empty array; user does not see agents owned by other users
|
||||
- [ ] T044 [P] [US4] Integration test for admin operations in `internal/api/auth_handlers_test.go`: admin can list all users; admin can deactivate account; non-admin gets 403 on admin endpoints
|
||||
|
||||
### Implementation for User Story 4
|
||||
|
||||
- [ ] T045 [US4] Implement password change handler in `internal/api/auth_handlers.go`: `PUT /api/users/me/password` -- require `current_password` and `new_password` in body; verify current password with bcrypt; re-hash new password; invalidate all sessions except current via `SessionStore.DeleteSessionsByUserExcept`; return 200 on success, 403 on wrong current password, 400 on invalid new password (FR-008)
|
||||
- [ ] T046 [US4] Implement user agents listing handler in `internal/api/auth_handlers.go`: `GET /api/users/me/agents` -- query agents table by `owner_id` from session context; return JSON array with `id`, `name`, `display_name`, `type`, `created_at` (FR-008, FR-009)
|
||||
- [ ] T047 [US4] Implement admin user management handlers in `internal/api/auth_handlers.go`: `GET /api/admin/users` -- list all users (admin only); `PUT /api/admin/users/{id}/deactivate` -- deactivate user account (admin only); first user created is automatically admin (FR-008)
|
||||
- [ ] T048 [US4] Register user management routes in `internal/api/router.go`: mount `/api/users/me/password` (PUT), `/api/users/me/agents` (GET), `/api/admin/users` (GET), `/api/admin/users/{id}/deactivate` (PUT); apply `RequireSession` + `RequireAdmin` where appropriate
|
||||
- [ ] T049 [US4] Add auth event logging for account management: `password_changed`, admin operations (FR-013)
|
||||
|
||||
**Checkpoint**: All user stories functional -- full auth lifecycle from registration to account management
|
||||
|
||||
---
|
||||
|
||||
## Phase 8: Polish & Cross-Cutting Concerns
|
||||
|
||||
**Purpose**: Edge cases, security hardening, and integration quality
|
||||
|
||||
- [ ] T050 [P] Handle zero-users edge case in middleware (`internal/auth/middleware.go`): when `UserStore.CountUsers() == 0`, redirect Web UI requests to registration; return 401 for API requests; never expose data
|
||||
- [ ] T051 [P] Handle corrupt/missing token storage gracefully in `internal/auth/fosite_store.go`: catch SQLite errors, log via slog, reject all token validations without panicking; return appropriate OAuth error responses
|
||||
- [ ] T052 [P] Implement TLS enforcement in `cmd/synapbus/main.go` (or serve command): when `--tls` is set, refuse to start OAuth endpoints on plain HTTP; in `--dev` mode, allow HTTP with logged `slog.Warn` (SC-006)
|
||||
- [ ] T053 [P] Add configurable bcrypt cost in `internal/auth/config.go` and `internal/auth/user_store.go`: minimum 10, default 12, configurable via `--bcrypt-cost` flag or `SYNAPBUS_BCRYPT_COST` env var; validate range at startup
|
||||
- [ ] T054 [P] Add configurable token TTLs via flags/env: `--access-token-ttl` (default 1h), `--refresh-token-lifetime` (default 30d), `--session-lifetime` (default 24h); wire into auth config (FR-012)
|
||||
- [ ] T055 [P] Verify zero-CGO build compatibility: add build tag test or CI step that runs `CGO_ENABLED=0 go build ./...` and `CGO_ENABLED=0 go test ./internal/auth/...` to confirm no C dependencies (SC-007)
|
||||
- [ ] T056 End-to-end integration test in `internal/auth/e2e_test.go`: full lifecycle -- register user -> log in with PKCE -> access protected endpoint -> refresh token -> change password -> re-login -> register OAuth client -> client_credentials token -> introspect -> logout; covers SC-001 through SC-005
|
||||
- [ ] T057 Code cleanup: ensure all exported types have godoc comments in `internal/auth/`; ensure error messages are consistent; ensure all SQL queries use parameterized statements (no SQL injection)
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & Execution Order
|
||||
|
||||
### Phase Dependencies
|
||||
|
||||
- **Setup (Phase 1)**: No dependencies -- can start immediately
|
||||
- **Foundational (Phase 2)**: Depends on Phase 1 completion -- BLOCKS all user stories
|
||||
- **User Story 1 (Phase 3)**: Depends on Phase 2 -- can start immediately after
|
||||
- **User Story 5 (Phase 4)**: Depends on Phase 3 (hardens PKCE already built in US1)
|
||||
- **User Story 2 (Phase 5)**: Depends on Phase 2 -- can run in parallel with Phase 3/4
|
||||
- **User Story 3 (Phase 6)**: Depends on Phase 2 -- can run in parallel with Phase 3/4/5 (fosite store built in Phase 2)
|
||||
- **User Story 4 (Phase 7)**: Depends on Phase 2 -- can run in parallel with others
|
||||
- **Polish (Phase 8)**: Depends on all user stories being complete
|
||||
|
||||
### User Story Dependencies
|
||||
|
||||
- **US1 (P1)**: Foundation only -- no cross-story dependencies
|
||||
- **US5 (P1)**: Validates PKCE enforcement built in US1 -- depends on US1 endpoints existing
|
||||
- **US2 (P2)**: Foundation only -- independent of US1 (different grant type)
|
||||
- **US3 (P2)**: Foundation only (fosite store) -- benefits from US1 being done for authorization_code refresh testing, but client_credentials refresh can test independently
|
||||
- **US4 (P3)**: Foundation + US1 (needs session infrastructure) -- depends on US1 for session context
|
||||
|
||||
### Within Each User Story
|
||||
|
||||
- Tests MUST be written and FAIL before implementation
|
||||
- Domain models/stores before handlers
|
||||
- Handlers before route registration
|
||||
- Core implementation before integration wiring
|
||||
- Story complete before moving to next priority
|
||||
|
||||
### Parallel Opportunities
|
||||
|
||||
- Phase 2: T006, T007, T008 (domain models) can run in parallel
|
||||
- Phase 2: T009, T010, T011 (store implementations) can run in parallel after their models
|
||||
- Phase 3+: All test tasks marked [P] within a story can run in parallel
|
||||
- US2, US3, US4 can theoretically start in parallel after Phase 2, but sequential P1->P2->P3 is recommended for a single developer
|
||||
|
||||
---
|
||||
|
||||
## Implementation Strategy
|
||||
|
||||
### MVP First (User Stories 1 + 5)
|
||||
|
||||
1. Complete Phase 1: Setup
|
||||
2. Complete Phase 2: Foundational (CRITICAL -- blocks all stories)
|
||||
3. Complete Phase 3: User Story 1 (registration + login + session)
|
||||
4. Complete Phase 4: User Story 5 (PKCE hardening)
|
||||
5. **STOP and VALIDATE**: Test full login flow independently
|
||||
6. Deploy/demo if ready -- Web UI is access-controlled
|
||||
|
||||
### Incremental Delivery
|
||||
|
||||
1. Setup + Foundational -> Foundation ready
|
||||
2. US1 + US5 -> Interactive login with PKCE -> Deploy (MVP!)
|
||||
3. US2 -> Programmatic access via client credentials -> Deploy
|
||||
4. US3 -> Refresh token rotation for long-lived sessions -> Deploy
|
||||
5. US4 -> Account management for self-service -> Deploy
|
||||
6. Polish -> Hardening and edge cases -> Production ready
|
||||
|
||||
---
|
||||
|
||||
## Notes
|
||||
|
||||
- [P] tasks = different files, no dependencies
|
||||
- [Story] label maps task to specific user story for traceability
|
||||
- All OAuth data stored in SQLite within `--data` directory per Constitution Principle I
|
||||
- `ory/fosite` is the OAuth 2.1 framework; storage adapter in `internal/auth/fosite_store.go` is the primary integration point
|
||||
- `modernc.org/sqlite` is already the project database (zero CGO, Principle III)
|
||||
- Existing schema (`001_initial.sql`) has `users`, `oauth_tokens`, `oauth_clients` tables; migration 002 extends them
|
||||
- Commit after each task or logical group
|
||||
- Stop at any checkpoint to validate story independently
|
||||
@@ -0,0 +1,138 @@
|
||||
# Feature Specification: Channels
|
||||
|
||||
**Feature Branch**: `004-channels`
|
||||
**Created**: 2026-03-13
|
||||
**Status**: Draft
|
||||
**Input**: User description: "Public and private channels with membership management, channel metadata, message broadcast, and MCP tool exposure."
|
||||
|
||||
## User Scenarios & Testing *(mandatory)*
|
||||
|
||||
### User Story 1 - Agent Creates and Broadcasts to a Public Channel (Priority: P1)
|
||||
|
||||
An AI agent (e.g., a monitoring agent) creates a public channel called `#alerts` so that any interested agent in the system can join and receive broadcast messages. The creating agent sends a message to the channel, and all members receive it. This is the fundamental channel workflow and the core value proposition of the feature.
|
||||
|
||||
**Why this priority**: Without the ability to create channels and broadcast messages to members, no other channel functionality has value. This is the minimum viable channel feature and maps directly to Constitution Principle IX (Progressive Complexity, Tier 2: channels).
|
||||
|
||||
**Independent Test**: Can be fully tested by registering two agents, having agent A call `create_channel` to create a public channel, having agent B call `join_channel`, then having agent A call `send_message` targeting the channel. Agent B reads its inbox and sees the broadcast message. Delivers group communication value independently of all other stories.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a registered agent A, **When** agent A calls `create_channel` with `name: "alerts"`, `type: "public"`, `description: "System alerts"`, **Then** the channel is created, agent A is automatically added as a member with role `owner`, and the tool returns the channel ID and metadata.
|
||||
2. **Given** a public channel `#alerts` exists, **When** registered agent B calls `join_channel` with the channel ID, **Then** agent B is added as a member with role `member` and receives a confirmation.
|
||||
3. **Given** agents A and B are both members of `#alerts`, **When** agent A calls `send_message` with `channel_id` set to the channel, **Then** agent B receives the message in its inbox, and the message `channel_id` field identifies the source channel.
|
||||
4. **Given** agent A is the only member of `#alerts`, **When** agent A sends a message to the channel, **Then** no delivery errors occur (the message is stored but delivered to zero other agents).
|
||||
|
||||
---
|
||||
|
||||
### User Story 2 - Agent Discovers and Lists Available Channels (Priority: P2)
|
||||
|
||||
An agent that has just been registered needs to discover what channels exist so it can decide which ones to join. The agent calls `list_channels` and receives a filtered list of channels it is eligible to join (all public channels, plus private channels it has been invited to). This supports the agent onboarding flow and is required before an agent can participate in any channel-based communication.
|
||||
|
||||
**Why this priority**: Discovery is the prerequisite for joining. Without listing channels, agents cannot find channels to join unless they already know the channel ID. This is essential for autonomous agent behavior per Constitution Principle VII (Swarm Intelligence Patterns).
|
||||
|
||||
**Independent Test**: Can be fully tested by creating several public and private channels, then having a new agent call `list_channels`. The response includes all public channels with their metadata (name, description, topic, member count) and excludes private channels the agent has not been invited to.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** three public channels and two private channels exist, **When** a registered agent calls `list_channels` with no filters, **Then** the response contains all three public channels with their `name`, `description`, `topic`, `created_by`, `member_count`, and `type` fields. The two private channels are not included.
|
||||
2. **Given** agent B has been invited to private channel `#core-team`, **When** agent B calls `list_channels`, **Then** the response includes `#core-team` along with all public channels.
|
||||
3. **Given** no channels exist, **When** an agent calls `list_channels`, **Then** the response is an empty list with no error.
|
||||
|
||||
---
|
||||
|
||||
### User Story 3 - Channel Owner Manages Private Channel Membership (Priority: P2)
|
||||
|
||||
A human owner creates a private channel `#core-team` for a select group of trusted agents. The owner invites specific agents, and only invited agents can join. The owner can also remove (kick) an agent that is no longer needed. Uninvited agents cannot see or join the channel. This enables controlled, secure communication groups.
|
||||
|
||||
**Why this priority**: Private channels with invite-only access are essential for multi-tenant security (Constitution Principle IV). Without this, all communication is visible to all agents, which is unacceptable for sensitive coordination tasks. Ranked P2 because public channels (P1) must work first.
|
||||
|
||||
**Independent Test**: Can be fully tested by creating a private channel, inviting agent B via `invite_to_channel`, confirming agent B can join, then verifying that uninvited agent C gets an authorization error when attempting to join.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** agent A creates a channel with `type: "private"` and `name: "core-team"`, **When** uninvited agent C calls `join_channel` with that channel ID, **Then** the system returns an error indicating the agent is not authorized to join the private channel.
|
||||
2. **Given** agent A owns private channel `#core-team`, **When** agent A calls `invite_to_channel` with `channel_id` and `agent_id` for agent B, **Then** agent B is added to the channel's invite list and can now call `join_channel` successfully.
|
||||
3. **Given** agent B is a member of `#core-team` and agent A is the owner, **When** agent A calls `kick_from_channel` with agent B's ID, **Then** agent B is removed from the channel and no longer receives messages broadcast to `#core-team`.
|
||||
4. **Given** agent B is a member (not owner) of `#core-team`, **When** agent B calls `invite_to_channel` or `kick_from_channel`, **Then** the system returns an authorization error because only the channel owner can manage membership.
|
||||
|
||||
---
|
||||
|
||||
### User Story 4 - Agent Leaves a Channel (Priority: P3)
|
||||
|
||||
An agent that no longer needs to participate in a channel can voluntarily leave it. After leaving, the agent stops receiving messages broadcast to that channel. The agent can rejoin a public channel later, but would need a new invite for a private channel.
|
||||
|
||||
**Why this priority**: Leaving is important for resource hygiene and agent lifecycle management, but is not required for core channel functionality. Agents can function without this feature by simply ignoring channel messages.
|
||||
|
||||
**Independent Test**: Can be fully tested by having an agent join a public channel, calling `leave_channel`, then verifying that subsequent messages to the channel are not delivered to the departed agent.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** agent B is a member of public channel `#alerts`, **When** agent B calls `leave_channel` with the channel ID, **Then** agent B is removed from the membership list and subsequent channel messages are not delivered to agent B.
|
||||
2. **Given** agent B has left public channel `#alerts`, **When** agent B calls `join_channel` again, **Then** agent B is re-added as a member and begins receiving new messages.
|
||||
3. **Given** agent A is the owner and sole member of a channel, **When** agent A calls `leave_channel`, **Then** the system returns an error indicating the channel owner cannot leave without transferring ownership or deleting the channel.
|
||||
|
||||
---
|
||||
|
||||
### User Story 5 - Channel Metadata and Topic Management (Priority: P3)
|
||||
|
||||
A channel owner updates the channel's topic and description to reflect the current purpose of the channel. Other members can read the metadata but cannot modify it. This supports long-running channels whose purpose evolves over time.
|
||||
|
||||
**Why this priority**: Metadata updates are a quality-of-life feature. Channels are fully functional without topic changes. Ranked P3 because the initial metadata set at creation time is sufficient for MVP.
|
||||
|
||||
**Independent Test**: Can be fully tested by creating a channel with a topic, updating the topic via an update call, then verifying `list_channels` returns the updated metadata.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** agent A owns channel `#research` with topic "Q1 findings", **When** agent A calls `update_channel` with `topic: "Q2 planning"`, **Then** the channel's topic is updated and reflected in subsequent `list_channels` responses.
|
||||
2. **Given** agent B is a member (not owner) of `#research`, **When** agent B calls `update_channel`, **Then** the system returns an authorization error.
|
||||
|
||||
---
|
||||
|
||||
### Edge Cases
|
||||
|
||||
- What happens when an agent tries to create a channel with a name that already exists? The system MUST return a conflict error with a clear message. Channel names MUST be unique within the system.
|
||||
- What happens when a message is sent to a channel with zero other members (only the sender)? The message MUST be stored successfully. No delivery errors should occur.
|
||||
- What happens when the channel owner's agent is deregistered? The channel MUST remain accessible. Ownership SHOULD transfer to the agent's human owner or the channel becomes ownerless but still functional.
|
||||
- What happens when an agent is invited to a channel it is already a member of? The system MUST return a no-op success (idempotent) rather than an error.
|
||||
- What happens when `join_channel` is called with a non-existent channel ID? The system MUST return a not-found error.
|
||||
- What happens when an agent tries to send a message to a channel it has not joined? The system MUST return an authorization error. Only members can send messages to a channel.
|
||||
- What happens when a channel name contains special characters or exceeds a reasonable length? The system MUST validate channel names: alphanumeric plus hyphens and underscores, maximum 64 characters.
|
||||
|
||||
## Requirements *(mandatory)*
|
||||
|
||||
### Functional Requirements
|
||||
|
||||
- **FR-001**: System MUST support creating channels with a `type` of either `public` or `private`.
|
||||
- **FR-002**: System MUST store channel metadata: `id`, `name`, `description`, `topic`, `type`, `created_by` (agent ID), `created_at`, `updated_at`.
|
||||
- **FR-003**: Channel names MUST be unique, case-insensitive, and restricted to alphanumeric characters, hyphens, and underscores (max 64 characters).
|
||||
- **FR-004**: The agent that creates a channel MUST be automatically added as a member with role `owner`.
|
||||
- **FR-005**: Any registered agent MUST be able to join a `public` channel via the `join_channel` MCP tool.
|
||||
- **FR-006**: Only agents with a pending invite MUST be able to join a `private` channel.
|
||||
- **FR-007**: Only the channel owner MUST be able to invite agents to a private channel via `invite_to_channel`.
|
||||
- **FR-008**: Only the channel owner MUST be able to remove members from a channel (kick).
|
||||
- **FR-009**: Any member MUST be able to leave a channel voluntarily via `leave_channel`, except the owner (who must transfer ownership or delete the channel first).
|
||||
- **FR-010**: Messages sent to a channel MUST be delivered to all current members except the sender.
|
||||
- **FR-011**: The `list_channels` MCP tool MUST return all public channels and any private channels the calling agent has been invited to or is a member of.
|
||||
- **FR-012**: All channel operations (create, join, leave, invite, kick, message broadcast) MUST be logged as traces per Constitution Principle VIII.
|
||||
- **FR-013**: Channel operations MUST be exposed exclusively as MCP tools per Constitution Principle II. The REST API MUST only serve the Web UI for channel management.
|
||||
- **FR-014**: Channel data MUST be stored in embedded SQLite per Constitution Principle VI. No external database dependencies.
|
||||
- **FR-015**: Agents MUST only be able to send messages to channels they are a member of.
|
||||
- **FR-016**: The system MUST support the following MCP tools for channels: `create_channel`, `join_channel`, `leave_channel`, `list_channels`, `invite_to_channel`.
|
||||
|
||||
### Key Entities
|
||||
|
||||
- **Channel**: Represents a named group communication space. Key attributes: `id` (UUID), `name` (unique, case-insensitive), `description` (text), `topic` (text), `type` (public | private), `created_by` (references agent), `created_at`, `updated_at`. A channel has many members through the Membership entity.
|
||||
- **Membership**: Represents the relationship between an agent and a channel. Key attributes: `channel_id` (references channel), `agent_id` (references agent), `role` (owner | member), `joined_at`. Composite unique constraint on (`channel_id`, `agent_id`).
|
||||
- **Channel Invite**: Represents a pending invitation for an agent to join a private channel. Key attributes: `channel_id`, `agent_id` (the invitee), `invited_by` (the agent who sent the invite), `created_at`, `status` (pending | accepted | declined). Used to gate access to private channels.
|
||||
- **Channel Message**: Not a new entity; uses the existing Message entity with a non-null `channel_id` field. When `channel_id` is set, the message is a broadcast to all channel members rather than a direct message.
|
||||
|
||||
## Success Criteria *(mandatory)*
|
||||
|
||||
### Measurable Outcomes
|
||||
|
||||
- **SC-001**: An agent can create a channel, and a second agent can join and receive a broadcast message within a single MCP session, end to end in under 5 seconds on local deployment.
|
||||
- **SC-002**: The `list_channels` tool returns accurate results reflecting the calling agent's visibility permissions (public channels visible, uninvited private channels hidden) with 100% correctness.
|
||||
- **SC-003**: Private channel access control is enforced: unauthorized `join_channel` attempts on private channels are rejected 100% of the time.
|
||||
- **SC-004**: All five MCP tools (`create_channel`, `join_channel`, `leave_channel`, `list_channels`, `invite_to_channel`) are registered with complete JSON Schema descriptions and are discoverable by any MCP-capable client.
|
||||
- **SC-005**: Channel operations generate trace entries viewable by the agent's owner through the Web UI or trace query tools.
|
||||
- **SC-006**: The system handles at least 50 concurrent channel members receiving broadcast messages without message loss or delivery failure.
|
||||
@@ -0,0 +1,256 @@
|
||||
# Tasks: Channels
|
||||
|
||||
**Input**: Design documents from `/specs/004-channels/`
|
||||
**Prerequisites**: spec.md (required), constitution.md (required)
|
||||
|
||||
**Tests**: Included per Go project conventions (table-driven tests, context propagation, slog logging).
|
||||
|
||||
**Organization**: Tasks are grouped by user story to enable independent implementation and testing of each story.
|
||||
|
||||
## Format: `[ID] [P?] [Story] Description`
|
||||
|
||||
- **[P]**: Can run in parallel (different files, no dependencies)
|
||||
- **[Story]**: Which user story this task belongs to (e.g., US1, US2, US3)
|
||||
- Include exact file paths in descriptions
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Setup (Shared Infrastructure)
|
||||
|
||||
**Purpose**: Database migration and package scaffolding for channels feature
|
||||
|
||||
- [ ] T001 Create database migration `schema/002_channels.sql` adding `channel_invites` table (`channel_id`, `agent_name`, `invited_by`, `created_at`, `status` with CHECK IN ('pending', 'accepted', 'declined')). Add composite unique constraint on (`channel_id`, `agent_name`). Add index on `channel_invites(agent_name)`. Update `channels.type` CHECK constraint to include `'public'` and `'private'` values alongside existing `'standard'`, `'blackboard'`, `'auction'` — or add an `is_private` boolean if the existing schema already handles it (note: `is_private INTEGER` already exists in 001_initial.sql, so this migration adds only the `channel_invites` table).
|
||||
- [ ] T002 [P] Create channel domain types in `internal/channels/types.go`: `Channel` struct (ID, Name, Description, Topic, Type, IsPrivate, CreatedBy, CreatedAt, UpdatedAt), `Membership` struct (ID, ChannelID, AgentName, Role, JoinedAt), `ChannelInvite` struct (ID, ChannelID, AgentName, InvitedBy, CreatedAt, Status), `ChannelType` and `MemberRole` string constants, `CreateChannelRequest`, `JoinChannelRequest`, `InviteRequest`, `UpdateChannelRequest` input structs.
|
||||
- [ ] T003 [P] Create channel name validation in `internal/channels/validate.go`: alphanumeric plus hyphens and underscores, max 64 characters, case-insensitive normalization (lowercase). Export `ValidateChannelName(name string) error`.
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Foundational (Blocking Prerequisites)
|
||||
|
||||
**Purpose**: Repository layer and service skeleton that ALL user stories depend on
|
||||
|
||||
**CRITICAL**: No user story work can begin until this phase is complete
|
||||
|
||||
- [ ] T004 Implement channel repository in `internal/channels/repository.go`: define `Repository` interface with methods `CreateChannel(ctx, Channel) (Channel, error)`, `GetChannelByID(ctx, int64) (Channel, error)`, `GetChannelByName(ctx, string) (Channel, error)`, `ListChannels(ctx, agentName string) ([]Channel, error)`, `UpdateChannel(ctx, Channel) error`, `DeleteChannel(ctx, int64) error`.
|
||||
- [ ] T005 [P] Implement membership repository in `internal/channels/membership_repo.go`: define `MembershipRepository` interface with methods `AddMember(ctx, Membership) error`, `RemoveMember(ctx, channelID int64, agentName string) error`, `GetMember(ctx, channelID int64, agentName string) (Membership, error)`, `ListMembers(ctx, channelID int64) ([]Membership, error)`, `IsMember(ctx, channelID int64, agentName string) (bool, error)`, `CountMembers(ctx, channelID int64) (int, error)`.
|
||||
- [ ] T006 [P] Implement invite repository in `internal/channels/invite_repo.go`: define `InviteRepository` interface with methods `CreateInvite(ctx, ChannelInvite) error`, `GetInvite(ctx, channelID int64, agentName string) (ChannelInvite, error)`, `HasPendingInvite(ctx, channelID int64, agentName string) (bool, error)`, `AcceptInvite(ctx, channelID int64, agentName string) error`.
|
||||
- [ ] T007 Implement SQLite channel repository in `internal/channels/sqlite_repository.go`: implement the `Repository` interface backed by `*sql.DB`. `ListChannels` must return all public channels plus private channels where the agent is a member or has a pending invite. Use `COLLATE NOCASE` for channel name uniqueness checks. Include slog logging for all operations.
|
||||
- [ ] T008 [P] Implement SQLite membership repository in `internal/channels/sqlite_membership_repo.go`: implement the `MembershipRepository` interface backed by `*sql.DB`. Include slog logging.
|
||||
- [ ] T009 [P] Implement SQLite invite repository in `internal/channels/sqlite_invite_repo.go`: implement the `InviteRepository` interface backed by `*sql.DB`. Include slog logging.
|
||||
- [ ] T010 Create channel service skeleton in `internal/channels/service.go`: define `Service` struct taking `Repository`, `MembershipRepository`, `InviteRepository`, and a trace recorder dependency. Constructor `NewService(...)`. This is the entry point for all channel business logic. Use `context.Context` for all public methods and `slog` for structured logging.
|
||||
|
||||
### Tests for Phase 2
|
||||
|
||||
- [ ] T011 [P] Write table-driven tests for channel repository in `internal/channels/sqlite_repository_test.go`: test CreateChannel (success, duplicate name conflict, name validation), GetChannelByID (found, not found), ListChannels (returns public channels, hides uninvited private channels). Use an in-memory SQLite database for test isolation.
|
||||
- [ ] T012 [P] Write table-driven tests for membership repository in `internal/channels/sqlite_membership_repo_test.go`: test AddMember (success, duplicate), RemoveMember (success, not found), IsMember, CountMembers.
|
||||
- [ ] T013 [P] Write table-driven tests for invite repository in `internal/channels/sqlite_invite_repo_test.go`: test CreateInvite (success, duplicate idempotent), HasPendingInvite, AcceptInvite.
|
||||
|
||||
**Checkpoint**: Foundation ready — user story implementation can now begin in parallel
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: User Story 1 — Agent Creates and Broadcasts to a Public Channel (Priority: P1) MVP
|
||||
|
||||
**Goal**: An agent can create a public channel, another agent joins, and broadcast messages reach all members.
|
||||
|
||||
**Independent Test**: Register two agents, agent A calls `create_channel` (public), agent B calls `join_channel`, agent A calls `send_message` with `channel_id`. Agent B reads inbox and sees the broadcast message.
|
||||
|
||||
### Tests for User Story 1
|
||||
|
||||
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
|
||||
|
||||
- [ ] T014 [P] [US1] Write table-driven tests for `Service.CreateChannel` in `internal/channels/service_test.go`: test creating public channel sets creator as owner, duplicate name returns conflict error, invalid name returns validation error, trace is recorded.
|
||||
- [ ] T015 [P] [US1] Write table-driven tests for `Service.JoinChannel` in `internal/channels/service_test.go`: test joining public channel succeeds, joining non-existent channel returns not-found error, joining channel agent is already a member of is idempotent (no error), trace is recorded.
|
||||
- [ ] T016 [P] [US1] Write integration test for channel broadcast in `internal/channels/broadcast_test.go`: test that when a message is sent to a channel, all members except the sender receive it in their inbox. Test with 0 other members (no error), 1 member, and multiple members.
|
||||
|
||||
### Implementation for User Story 1
|
||||
|
||||
- [ ] T017 [US1] Implement `Service.CreateChannel(ctx, CreateChannelRequest) (Channel, error)` in `internal/channels/service.go`: validate name via `ValidateChannelName`, check uniqueness, insert channel, auto-add creator as `owner` member, record trace entry, return channel with ID.
|
||||
- [ ] T018 [US1] Implement `Service.JoinChannel(ctx, agentName string, channelID int64) error` in `internal/channels/service.go`: verify channel exists, verify channel is public (for US1; private join gated by invite will be added in US3), check if already a member (idempotent return), add member with role `member`, record trace entry.
|
||||
- [ ] T019 [US1] Implement channel broadcast logic in `internal/channels/broadcast.go`: export `BroadcastToChannel(ctx, channelID int64, senderAgent string, messageBody string, metadata map[string]any) error`. Query all channel members, exclude the sender, create a message for each recipient with the `channel_id` field set. This function calls into the messaging repository (depends on `internal/messaging` package — accept an interface `MessageCreator` to avoid tight coupling).
|
||||
- [ ] T020 [US1] Register channel MCP tools (`create_channel`, `join_channel`) in `internal/mcp/channel_tools.go`: define JSON Schema `inputSchema` for each tool, implement tool handler functions that call `channels.Service` methods, wire into MCP tool registry. `create_channel` accepts `name` (required), `type` ("public"|"private", default "public"), `description` (optional), `topic` (optional). `join_channel` accepts `channel_id` (required integer).
|
||||
- [ ] T021 [US1] Extend `send_message` MCP tool to support channel broadcast in `internal/mcp/message_tools.go`: when `channel_id` is provided (and `to_agent` is omitted), verify sender is a channel member (authorization), then delegate to `BroadcastToChannel`. Return error if agent is not a member (FR-015). Record trace.
|
||||
|
||||
**Checkpoint**: At this point, agents can create public channels, join them, and broadcast messages. User Story 1 is fully functional and testable independently.
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: User Story 2 — Agent Discovers and Lists Available Channels (Priority: P2)
|
||||
|
||||
**Goal**: An agent can call `list_channels` and see all public channels plus private channels it has been invited to, with full metadata.
|
||||
|
||||
**Independent Test**: Create several public and private channels, call `list_channels` as a new agent. Response includes all public channels with metadata, excludes uninvited private channels.
|
||||
|
||||
### Tests for User Story 2
|
||||
|
||||
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
|
||||
|
||||
- [ ] T022 [P] [US2] Write table-driven tests for `Service.ListChannels` in `internal/channels/service_test.go`: test returns all public channels, excludes uninvited private channels, includes private channels where agent is a member, includes private channels where agent has a pending invite, returns empty list when no channels exist. Verify response includes `name`, `description`, `topic`, `created_by`, `member_count`, `type` fields.
|
||||
|
||||
### Implementation for User Story 2
|
||||
|
||||
- [ ] T023 [US2] Implement `Service.ListChannels(ctx, agentName string) ([]ChannelWithCount, error)` in `internal/channels/service.go`: define `ChannelWithCount` struct (embeds `Channel` + `MemberCount int`). Delegate to repository `ListChannels` which already filters by visibility. Enrich results with member count via `MembershipRepository.CountMembers`. Record trace entry.
|
||||
- [ ] T024 [US2] Register `list_channels` MCP tool in `internal/mcp/channel_tools.go`: no required input parameters (agent identity comes from MCP connection context). Return JSON array of channel objects with fields: `id`, `name`, `description`, `topic`, `type`, `created_by`, `member_count`. Add JSON Schema for the tool.
|
||||
|
||||
**Checkpoint**: At this point, User Stories 1 AND 2 should both work independently. Agents can create, join, discover, and broadcast to channels.
|
||||
|
||||
---
|
||||
|
||||
## Phase 5: User Story 3 — Channel Owner Manages Private Channel Membership (Priority: P2)
|
||||
|
||||
**Goal**: A channel owner creates a private channel, invites specific agents, and can kick members. Uninvited agents cannot join.
|
||||
|
||||
**Independent Test**: Create a private channel, invite agent B, confirm agent B can join. Verify uninvited agent C gets an authorization error on `join_channel`. Owner kicks agent B, verify agent B no longer receives channel messages.
|
||||
|
||||
### Tests for User Story 3
|
||||
|
||||
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
|
||||
|
||||
- [ ] T025 [P] [US3] Write table-driven tests for `Service.InviteToChannel` in `internal/channels/service_test.go`: test owner can invite, non-owner gets authorization error, inviting an already-member agent is idempotent, inviting to a non-existent channel returns not-found.
|
||||
- [ ] T026 [P] [US3] Write table-driven tests for `Service.KickFromChannel` in `internal/channels/service_test.go`: test owner can kick a member, non-owner gets authorization error, kicking a non-member returns not-found, owner cannot kick themselves.
|
||||
- [ ] T027 [P] [US3] Write table-driven tests for private channel join gating in `internal/channels/service_test.go`: test uninvited agent gets authorization error on `join_channel` for private channel, invited agent can join, invite status changes to `accepted` after join.
|
||||
|
||||
### Implementation for User Story 3
|
||||
|
||||
- [ ] T028 [US3] Extend `Service.JoinChannel` in `internal/channels/service.go` to enforce private channel invite gating: if channel `is_private`, check `InviteRepository.HasPendingInvite`. If no pending invite, return authorization error. On successful join, call `InviteRepository.AcceptInvite`.
|
||||
- [ ] T029 [US3] Implement `Service.InviteToChannel(ctx, ownerAgent string, channelID int64, inviteeAgent string) error` in `internal/channels/service.go`: verify channel exists, verify caller is the channel owner (role check via `MembershipRepository.GetMember`), check if invitee is already a member (idempotent return), create invite via `InviteRepository.CreateInvite`, record trace entry.
|
||||
- [ ] T030 [US3] Implement `Service.KickFromChannel(ctx, ownerAgent string, channelID int64, targetAgent string) error` in `internal/channels/service.go`: verify channel exists, verify caller is the channel owner, verify target is a member (not-found if not), prevent owner from kicking themselves, remove member via `MembershipRepository.RemoveMember`, record trace entry.
|
||||
- [ ] T031 [US3] Register `invite_to_channel` MCP tool in `internal/mcp/channel_tools.go`: accepts `channel_id` (required integer), `agent_id` (required string — the invitee agent name). Only callable by the channel owner. Add JSON Schema.
|
||||
- [ ] T032 [US3] Register `kick_from_channel` MCP tool in `internal/mcp/channel_tools.go`: accepts `channel_id` (required integer), `agent_id` (required string — the target agent name). Only callable by the channel owner. Add JSON Schema.
|
||||
|
||||
**Checkpoint**: At this point, User Stories 1, 2, and 3 are all functional. Public and private channels work with full membership management.
|
||||
|
||||
---
|
||||
|
||||
## Phase 6: User Story 4 — Agent Leaves a Channel (Priority: P3)
|
||||
|
||||
**Goal**: An agent can voluntarily leave a channel and stop receiving messages. Owner cannot leave without transferring ownership or deleting the channel.
|
||||
|
||||
**Independent Test**: Agent joins a public channel, calls `leave_channel`, verify subsequent channel messages are not delivered. Agent can rejoin. Owner leaving returns an error.
|
||||
|
||||
### Tests for User Story 4
|
||||
|
||||
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
|
||||
|
||||
- [ ] T033 [P] [US4] Write table-driven tests for `Service.LeaveChannel` in `internal/channels/service_test.go`: test member can leave, non-member returns not-found, owner cannot leave (returns error with guidance to transfer ownership or delete), after leaving agent does not receive broadcast messages, agent can rejoin a public channel.
|
||||
|
||||
### Implementation for User Story 4
|
||||
|
||||
- [ ] T034 [US4] Implement `Service.LeaveChannel(ctx, agentName string, channelID int64) error` in `internal/channels/service.go`: verify channel exists, verify agent is a member, if agent role is `owner` return error with message "channel owner cannot leave; transfer ownership or delete the channel first", remove member via `MembershipRepository.RemoveMember`, record trace entry.
|
||||
- [ ] T035 [US4] Register `leave_channel` MCP tool in `internal/mcp/channel_tools.go`: accepts `channel_id` (required integer). Agent identity from MCP context. Add JSON Schema.
|
||||
|
||||
**Checkpoint**: All five MCP tools from FR-016 are now registered: `create_channel`, `join_channel`, `leave_channel`, `list_channels`, `invite_to_channel`.
|
||||
|
||||
---
|
||||
|
||||
## Phase 7: User Story 5 — Channel Metadata and Topic Management (Priority: P3)
|
||||
|
||||
**Goal**: Channel owner can update topic and description. Non-owners get an authorization error.
|
||||
|
||||
**Independent Test**: Create a channel with a topic, update it via `update_channel`, verify `list_channels` reflects the change. Non-owner update returns error.
|
||||
|
||||
### Tests for User Story 5
|
||||
|
||||
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
|
||||
|
||||
- [ ] T036 [P] [US5] Write table-driven tests for `Service.UpdateChannel` in `internal/channels/service_test.go`: test owner can update topic, owner can update description, non-owner gets authorization error, non-existent channel returns not-found, updated_at timestamp is refreshed.
|
||||
|
||||
### Implementation for User Story 5
|
||||
|
||||
- [ ] T037 [US5] Implement `Service.UpdateChannel(ctx, agentName string, channelID int64, req UpdateChannelRequest) (Channel, error)` in `internal/channels/service.go`: verify channel exists, verify caller is owner, apply updates (topic, description — only non-nil fields), update `updated_at`, persist via `Repository.UpdateChannel`, record trace entry, return updated channel.
|
||||
- [ ] T038 [US5] Register `update_channel` MCP tool in `internal/mcp/channel_tools.go`: accepts `channel_id` (required integer), `topic` (optional string), `description` (optional string). Only callable by the channel owner. Add JSON Schema.
|
||||
|
||||
**Checkpoint**: All user stories are now independently functional.
|
||||
|
||||
---
|
||||
|
||||
## Phase 8: Polish & Cross-Cutting Concerns
|
||||
|
||||
**Purpose**: Edge cases, performance, and integration hardening
|
||||
|
||||
- [ ] T039 [P] Handle edge case: channel name conflict returns structured error with clear message in `internal/channels/service.go`. Ensure error type is distinguishable (e.g., `ErrChannelNameConflict`) for MCP tool handlers to return appropriate error codes.
|
||||
- [ ] T040 [P] Handle edge case: deregistered channel owner — in `internal/channels/service.go`, ensure channel remains accessible when owner agent is deregistered. Document behavior: channel becomes ownerless but functional, messages can still be broadcast.
|
||||
- [ ] T041 [P] Define sentinel errors in `internal/channels/errors.go`: `ErrChannelNotFound`, `ErrChannelNameConflict`, `ErrNotChannelMember`, `ErrNotChannelOwner`, `ErrOwnerCannotLeave`, `ErrNotInvited`, `ErrInvalidChannelName`. MCP tool handlers in `internal/mcp/channel_tools.go` should map these to appropriate MCP error responses.
|
||||
- [ ] T042 [P] Add REST API endpoints for Web UI channel management in `internal/api/channels.go`: `GET /api/channels` (list), `GET /api/channels/:id` (detail with members), `POST /api/channels` (create), `PUT /api/channels/:id` (update). These are internal endpoints for the Web UI only (per Constitution Principle II, not for agents).
|
||||
- [ ] T043 Write concurrency test in `internal/channels/concurrency_test.go`: verify that 50 agents can join a channel and receive broadcast messages concurrently without message loss or race conditions (SC-006). Use `sync.WaitGroup` and `t.Parallel()`.
|
||||
- [ ] T044 [P] Verify all channel MCP tools have complete JSON Schema descriptions with field types, required markers, and descriptions. Ensure `tools/list` returns all five channel tools (SC-004). Write a test in `internal/mcp/channel_tools_test.go`.
|
||||
- [ ] T045 Review and verify trace entries for all channel operations in `internal/channels/service.go`: create, join, leave, invite, kick, update, broadcast. Each trace must include `agent_name`, `action` (e.g., `channel.create`, `channel.join`), and `details` JSON with relevant IDs (SC-005, FR-012).
|
||||
- [ ] T046 Run `make lint` and `make test` to verify all code passes linting and tests.
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & Execution Order
|
||||
|
||||
### Phase Dependencies
|
||||
|
||||
- **Setup (Phase 1)**: No dependencies — can start immediately
|
||||
- **Foundational (Phase 2)**: Depends on Phase 1 (migration and types must exist) — BLOCKS all user stories
|
||||
- **User Story 1 (Phase 3)**: Depends on Phase 2 completion
|
||||
- **User Story 2 (Phase 4)**: Depends on Phase 2 completion; can run in parallel with Phase 3
|
||||
- **User Story 3 (Phase 5)**: Depends on Phase 2 completion; can run in parallel with Phases 3-4
|
||||
- **User Story 4 (Phase 6)**: Depends on Phase 2 completion; can run in parallel with Phases 3-5
|
||||
- **User Story 5 (Phase 7)**: Depends on Phase 2 completion; can run in parallel with Phases 3-6
|
||||
- **Polish (Phase 8)**: Depends on all user story phases being complete
|
||||
|
||||
### Cross-Package Dependencies
|
||||
|
||||
- `internal/channels/` depends on `internal/storage/` (SQLite connection)
|
||||
- `internal/channels/broadcast.go` depends on `internal/messaging/` (MessageCreator interface for delivering broadcast messages to agent inboxes)
|
||||
- `internal/mcp/channel_tools.go` depends on `internal/channels/` (Service) and `internal/mcp/` (tool registry)
|
||||
- `internal/api/channels.go` depends on `internal/channels/` (Service)
|
||||
- `internal/channels/service.go` depends on `internal/trace/` (trace recorder for FR-012)
|
||||
|
||||
### Within Each User Story
|
||||
|
||||
- Tests MUST be written and FAIL before implementation
|
||||
- Repository layer before service layer
|
||||
- Service layer before MCP tool registration
|
||||
- Core implementation before integration
|
||||
- Story complete before moving to next priority
|
||||
|
||||
### Parallel Opportunities
|
||||
|
||||
- All Phase 1 tasks marked [P] can run in parallel (T002, T003)
|
||||
- All Phase 2 repository interfaces marked [P] can run in parallel (T005, T006)
|
||||
- All Phase 2 SQLite implementations marked [P] can run in parallel (T008, T009)
|
||||
- All Phase 2 tests marked [P] can run in parallel (T011, T012, T013)
|
||||
- Once Phase 2 completes, all user story phases (3-7) can start in parallel
|
||||
- All test tasks within a story marked [P] can run in parallel
|
||||
- Phase 8 polish tasks marked [P] can run in parallel
|
||||
|
||||
---
|
||||
|
||||
## Implementation Strategy
|
||||
|
||||
### MVP First (User Story 1 Only)
|
||||
|
||||
1. Complete Phase 1: Setup (migration, types, validation)
|
||||
2. Complete Phase 2: Foundational (repositories, service skeleton)
|
||||
3. Complete Phase 3: User Story 1 (create, join, broadcast)
|
||||
4. **STOP and VALIDATE**: Test User Story 1 independently — two agents communicate via a public channel
|
||||
5. Deploy/demo if ready
|
||||
|
||||
### Incremental Delivery
|
||||
|
||||
1. Setup + Foundational -> Foundation ready
|
||||
2. Add User Story 1 -> Test independently -> Deploy/Demo (MVP!)
|
||||
3. Add User Story 2 -> Test independently -> Deploy/Demo (discovery)
|
||||
4. Add User Story 3 -> Test independently -> Deploy/Demo (private channels)
|
||||
5. Add User Story 4 -> Test independently -> Deploy/Demo (leave)
|
||||
6. Add User Story 5 -> Test independently -> Deploy/Demo (metadata updates)
|
||||
7. Each story adds value without breaking previous stories
|
||||
|
||||
---
|
||||
|
||||
## Notes
|
||||
|
||||
- [P] tasks = different files, no dependencies
|
||||
- [Story] label maps task to specific user story for traceability
|
||||
- Each user story is independently completable and testable
|
||||
- Verify tests fail before implementing
|
||||
- Commit after each task or logical group
|
||||
- Stop at any checkpoint to validate story independently
|
||||
- Channel names are stored lowercase; all comparisons are case-insensitive
|
||||
- The existing schema already has `channels` and `channel_members` tables in `schema/001_initial.sql` — the migration in T001 only adds the `channel_invites` table
|
||||
- `kick_from_channel` is not in FR-016's explicit MCP tool list but is required by US3 acceptance scenario 3; it is registered as an additional MCP tool
|
||||
- The `update_channel` tool is not in FR-016's explicit list but is required by US5; it is registered as an additional MCP tool
|
||||
@@ -0,0 +1,195 @@
|
||||
# Feature Specification: Web UI
|
||||
|
||||
**Feature Branch**: `005-web-ui`
|
||||
**Created**: 2026-03-13
|
||||
**Status**: Draft
|
||||
**Input**: User description: "Svelte 5 + Tailwind CSS embedded SPA with login, dashboard, conversations, channels, agents, settings pages. Real-time SSE updates, compose, search, agent management, dark mode, responsive, embedded in Go binary via go:embed."
|
||||
|
||||
## Constitution Compliance
|
||||
|
||||
| Principle | Status | Notes |
|
||||
|-----------|--------|-------|
|
||||
| I. Local-First, Single Binary | Compliant | SPA built at compile time, embedded via `go:embed` into the Go binary. No external CDN or asset server. |
|
||||
| II. MCP-Native | Compliant | Web UI consumes the internal REST API only. No MCP tools are exposed for UI operations. |
|
||||
| III. Pure Go, Zero CGO | Compliant | Svelte build produces static assets; Go embedding uses stdlib `embed` package. No CGO involved. |
|
||||
| IV. Multi-Tenant with Ownership | Compliant | UI enforces owner-scoped views: users see only their own agents, traces, and authorized channels. |
|
||||
| V. Embedded OAuth 2.1 | Compliant | Login page authenticates against the embedded OAuth 2.1 server. Session managed via httponly cookies. |
|
||||
| VIII. Observable by Default | Compliant | Trace viewer page lets owners inspect all agent activity. |
|
||||
| IX. Progressive Complexity | Compliant | Dashboard shows basic messaging by default; channels, traces, and search are separate pages users navigate to when needed. |
|
||||
| X. Web UI as First-Class Citizen | Compliant | This spec directly implements Principle X. |
|
||||
|
||||
## User Scenarios & Testing *(mandatory)*
|
||||
|
||||
### User Story 1 - Dashboard and Conversation Browsing (Priority: P1)
|
||||
|
||||
A human owner logs in and lands on the Dashboard, which shows their most recent messages across all conversations. They click a conversation to open the thread view, read the full history, and see real-time updates as new messages arrive via SSE without refreshing the page.
|
||||
|
||||
**Why this priority**: The primary reason humans open the Web UI is to see what their agents are doing. Without message browsing and real-time updates, the UI provides no value. This is the core loop that makes SynapBus observable.
|
||||
|
||||
**Independent Test**: Can be fully tested by logging in, viewing the dashboard, clicking into a conversation, and confirming that a message sent via MCP by an agent appears in the thread within 2 seconds without a page refresh. Delivers the core value of human oversight over agent communication.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a logged-in user with 3 conversations containing messages, **When** they load the Dashboard, **Then** they see the 3 conversations sorted by most recent activity, each showing the last message preview, sender name, timestamp, and unread count.
|
||||
2. **Given** a user viewing a conversation thread with 10 messages, **When** an agent sends a new message to that conversation via MCP, **Then** the message appears at the bottom of the thread within 2 seconds via SSE, without a page refresh.
|
||||
3. **Given** a user on the Dashboard, **When** they click a conversation with 5 unread messages, **Then** the thread opens, all messages display in chronological order, and the unread count resets to 0.
|
||||
4. **Given** a user viewing a conversation, **When** they scroll up past the initial 50 messages, **Then** older messages load incrementally (pagination) without losing their scroll position.
|
||||
|
||||
---
|
||||
|
||||
### User Story 2 - Compose and Send Messages (Priority: P1)
|
||||
|
||||
A human owner composes and sends a direct message to a specific agent or posts a message to a channel. They select the recipient from a searchable dropdown, type the message body, optionally set a subject and priority, and send. The message appears in the conversation immediately.
|
||||
|
||||
**Why this priority**: Two-way communication is essential. Without compose, the UI is read-only and humans cannot participate in agent workflows. This is tied with Story 1 as the minimum viable product.
|
||||
|
||||
**Independent Test**: Can be fully tested by opening the compose form, selecting an agent recipient, typing a message, clicking send, and confirming the message appears in the conversation thread and is received by the target agent via MCP `read_inbox`.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a logged-in user on the Dashboard, **When** they click "New Message", **Then** a compose form opens with a searchable recipient dropdown (populated from registered agents and channels), a subject field, a message body textarea, and a priority selector (1-10, default 5).
|
||||
2. **Given** a user in the compose form who has typed "deploy" in the recipient field, **When** the dropdown filters, **Then** only agents and channels whose name or display_name contains "deploy" appear in the results.
|
||||
3. **Given** a user who has filled in recipient (agent "researcher-01"), subject ("Analysis request"), body ("Please analyze Q1 data"), and priority (7), **When** they click Send, **Then** the message is created via the REST API, the compose form closes, and the user is navigated to the conversation thread showing their sent message.
|
||||
4. **Given** a user composing a message, **When** they submit with an empty body, **Then** the form shows a validation error "Message body is required" and does not submit.
|
||||
|
||||
---
|
||||
|
||||
### User Story 3 - Agent Management and Trace Viewing (Priority: P2)
|
||||
|
||||
A human owner navigates to the Agents page to see all agents they own. They click an agent to view its details, activity traces, and API key status. They can regenerate or revoke an agent's API key from this page.
|
||||
|
||||
**Why this priority**: Agent oversight is a constitutional requirement (Principles IV and VIII). Without this, owners cannot monitor or control their agents. It is P2 because the system is still useful for messaging without it, but it is essential for production trust.
|
||||
|
||||
**Independent Test**: Can be fully tested by navigating to the Agents page, clicking an agent, viewing its traces (filtered by time range and action type), and revoking its API key, then confirming the agent can no longer authenticate via MCP.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a logged-in user who owns 3 agents, **When** they navigate to the Agents page, **Then** they see a list of their 3 agents showing name, display_name, type (ai/human), status (active/inactive), and the number of messages sent in the last 24 hours.
|
||||
2. **Given** a user viewing agent "researcher-01" details, **When** they click "Activity Traces", **Then** they see a paginated, reverse-chronological list of traces (tool calls, messages sent, channel joins, errors) with filters for action type and date range.
|
||||
3. **Given** a user viewing agent "researcher-01" details, **When** they click "Revoke API Key" and confirm the dialog, **Then** the API key is invalidated, the agent status changes to "inactive", and subsequent MCP requests from that agent return 401 Unauthorized.
|
||||
4. **Given** a user viewing agent "researcher-01" details, **When** they click "Regenerate API Key", **Then** a new API key is generated and displayed once in a copiable field with a warning that it will not be shown again.
|
||||
|
||||
---
|
||||
|
||||
### User Story 4 - Channel Management (Priority: P2)
|
||||
|
||||
A human owner browses available channels, views channel details and membership, and creates new channels. For channels they own, they can manage members (invite, remove) and update channel metadata.
|
||||
|
||||
**Why this priority**: Channels are a core messaging concept and the foundation for swarm patterns. Channel management through the UI is essential for humans to organize agent communication, but basic DM messaging (Stories 1-2) works without it.
|
||||
|
||||
**Independent Test**: Can be fully tested by navigating to the Channels page, creating a new public channel, viewing its member list, and confirming that an agent can join it via MCP `join_channel`.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a logged-in user, **When** they navigate to the Channels page, **Then** they see a list of all public channels and private channels they are a member of, showing channel name, description, member count, and last activity timestamp.
|
||||
2. **Given** a user on the Channels page, **When** they click "Create Channel" and fill in name ("research-findings"), description ("Shared research results"), type (public), **Then** the channel is created, appears in the list, and the user is its owner.
|
||||
3. **Given** a user viewing a channel they own, **When** they click "Members" and then "Invite", **Then** they see a searchable list of agents not already in the channel and can select one or more to invite.
|
||||
4. **Given** a user viewing a public channel they do not own, **When** they view the channel, **Then** they can see messages and members but cannot see "Invite" or "Remove" controls.
|
||||
|
||||
---
|
||||
|
||||
### User Story 5 - Full-Text Search (Priority: P2)
|
||||
|
||||
A human owner uses the search bar to find messages across all their conversations and channels. Results show message previews with highlighted matches, grouped by conversation, and link directly to the message in its thread context.
|
||||
|
||||
**Why this priority**: Search is critical for operational use once message volume grows. Without it, finding past agent communications becomes impossible. It is P2 because a small number of conversations can be browsed manually.
|
||||
|
||||
**Independent Test**: Can be fully tested by sending several messages with known content via MCP, then searching for a keyword in the Web UI and confirming matching messages appear with highlighted terms and correct conversation links.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a logged-in user with messages containing the word "deployment" in 3 different conversations, **When** they type "deployment" in the global search bar and press Enter, **Then** they see results grouped by conversation, each showing the matching message preview with "deployment" highlighted, sender name, and timestamp.
|
||||
2. **Given** search results displayed, **When** the user clicks a result, **Then** they are navigated to the conversation thread scrolled to the specific message, with the matched message visually highlighted.
|
||||
3. **Given** a search for "xyznonexistent", **When** results load, **Then** an empty state is shown with the message "No messages found for 'xyznonexistent'".
|
||||
4. **Given** a user who does not have access to a private channel, **When** they search for content that exists only in that channel, **Then** those messages do not appear in results (access control enforced server-side).
|
||||
|
||||
---
|
||||
|
||||
### User Story 6 - Login and Authentication (Priority: P1)
|
||||
|
||||
A user navigates to SynapBus in their browser. If not authenticated, they are redirected to the login page. They enter their username and password, submit, and are redirected to the Dashboard. Sessions persist across browser restarts via httponly cookies until explicitly logged out.
|
||||
|
||||
**Why this priority**: Authentication gates all other functionality. Without login, no other story can be tested. This is a foundational P1 alongside Stories 1 and 2.
|
||||
|
||||
**Independent Test**: Can be fully tested by opening the Web UI in a browser, being redirected to login, entering valid credentials, and confirming redirect to the Dashboard. Then close and reopen the browser, confirm the session persists. Then click logout and confirm redirect back to login.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** an unauthenticated user, **When** they navigate to any UI route (e.g., `/dashboard`), **Then** they are redirected to `/login` with the original URL preserved as a return parameter.
|
||||
2. **Given** a user on the login page, **When** they enter valid username and password and click "Sign In", **Then** they are authenticated, a session cookie is set (httponly, secure, SameSite=Strict), and they are redirected to the Dashboard (or the original requested URL).
|
||||
3. **Given** a user on the login page, **When** they enter an incorrect password, **Then** they see the error "Invalid username or password" without revealing which field is wrong. The form is not rate-limited visually but the server enforces rate limiting.
|
||||
4. **Given** an authenticated user, **When** they click "Logout" in the navigation, **Then** the session is invalidated server-side, the cookie is cleared, and they are redirected to `/login`.
|
||||
|
||||
---
|
||||
|
||||
### User Story 7 - Settings and Dark Mode (Priority: P3)
|
||||
|
||||
A user navigates to Settings to change their password, toggle dark/light mode, and configure notification preferences. Dark mode preference persists in local storage and applies immediately without a page reload.
|
||||
|
||||
**Why this priority**: Settings and dark mode are quality-of-life features. The system is fully functional without them. Dark mode is important for developer comfort but does not block any core workflow.
|
||||
|
||||
**Independent Test**: Can be fully tested by toggling dark mode in Settings and confirming all pages render correctly in both themes. Change password and confirm login works with the new password.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a logged-in user on the Settings page, **When** they toggle the "Dark Mode" switch, **Then** the entire UI switches to dark theme immediately (no reload), and the preference is saved to `localStorage`.
|
||||
2. **Given** a user who previously enabled dark mode, **When** they close and reopen the browser, **Then** the UI loads in dark mode directly (no flash of light theme).
|
||||
3. **Given** a user on the Settings page, **When** they enter their current password, a new password, and confirm the new password, then click "Change Password", **Then** the password is updated and they see a success message. Their session remains active.
|
||||
4. **Given** a user changing their password, **When** the new password is shorter than 8 characters, **Then** a validation error is shown: "Password must be at least 8 characters".
|
||||
|
||||
---
|
||||
|
||||
### Edge Cases
|
||||
|
||||
- What happens when the SSE connection drops (e.g., network interruption)? The UI MUST detect the disconnect, show a "Reconnecting..." banner, attempt automatic reconnection with exponential backoff (1s, 2s, 4s, max 30s), and reconcile missed messages on reconnect by fetching updates since the last received event timestamp.
|
||||
- What happens when a user has no conversations or agents yet? The Dashboard MUST show a meaningful empty state with guidance: "No conversations yet. Send your first message or register an agent to get started." with action buttons.
|
||||
- What happens when two browser tabs are open and one logs out? The other tab MUST detect the invalidated session on the next API call (401 response) and redirect to login without data loss (no partial state).
|
||||
- How does the UI handle very long messages (>10,000 characters)? Messages MUST be truncated to 500 characters with a "Show more" toggle that expands inline. Code blocks within messages MUST be syntax-highlighted.
|
||||
- What happens when the agent list is very large (>100 agents) in the compose dropdown? The dropdown MUST use virtualized rendering and support keyboard navigation. Results MUST be debounced (300ms) to avoid excessive API calls during typing.
|
||||
- What happens when a user tries to access an agent they do not own via direct URL manipulation (e.g., `/agents/other-owners-agent`)? The API MUST return 403 Forbidden and the UI MUST show "You do not have access to this agent."
|
||||
- How does the UI behave on slow connections? All API calls MUST show loading skeletons (not spinners) for content areas. Actions (send, revoke) MUST disable the button and show a loading indicator to prevent double submission.
|
||||
|
||||
## Requirements *(mandatory)*
|
||||
|
||||
### Functional Requirements
|
||||
|
||||
- **FR-001**: The Web UI MUST be a Svelte 5 single-page application styled with Tailwind CSS, built to static assets via `make web` and embedded in the Go binary using `go:embed`.
|
||||
- **FR-002**: The UI MUST include these pages: Login, Dashboard, Conversations (thread view), Channels, Agents, Settings, each accessible via client-side routing with shareable URLs.
|
||||
- **FR-003**: The Dashboard MUST display the user's recent conversations sorted by last activity, showing last message preview, sender, timestamp, and unread count.
|
||||
- **FR-004**: The Conversations page MUST render a threaded message view with chronological ordering, pagination (50 messages per page), and infinite scroll for older messages.
|
||||
- **FR-005**: The UI MUST receive real-time updates via Server-Sent Events (SSE) from the REST API. Events MUST include: new messages, message status changes, agent status changes, and channel membership changes.
|
||||
- **FR-006**: The compose form MUST allow users to send direct messages or channel messages, with a searchable recipient dropdown, subject field, body textarea, and priority selector (1-10).
|
||||
- **FR-007**: The search feature MUST perform full-text search across all messages the user has access to, returning results grouped by conversation with highlighted match terms.
|
||||
- **FR-008**: The Agents page MUST list all agents owned by the authenticated user, showing name, type, status, and recent activity summary.
|
||||
- **FR-009**: The agent detail view MUST display the agent's activity traces in a paginated, filterable list (by action type and date range).
|
||||
- **FR-010**: Users MUST be able to revoke and regenerate API keys for their agents from the agent detail view.
|
||||
- **FR-011**: The Channels page MUST list all public channels and private channels the user belongs to, with options to create new channels and manage membership for owned channels.
|
||||
- **FR-012**: The UI MUST support dark mode and light mode, togglable from Settings, with preference persisted in `localStorage` and applied without page reload.
|
||||
- **FR-013**: The UI MUST be responsive, functioning correctly on viewports from 375px (mobile) to 2560px (ultrawide) width.
|
||||
- **FR-014**: All authenticated routes MUST redirect to `/login` when the session is invalid or expired. The login page MUST authenticate against the embedded OAuth 2.1 server.
|
||||
- **FR-015**: The UI MUST show loading skeletons for content areas during data fetches and disable action buttons during pending requests to prevent double submission.
|
||||
- **FR-016**: The SSE connection MUST automatically reconnect on disconnect with exponential backoff (1s, 2s, 4s, max 30s) and reconcile missed events on reconnection.
|
||||
- **FR-017**: The Settings page MUST allow users to change their password with current password verification and minimum 8-character validation.
|
||||
- **FR-018**: The UI MUST enforce access control client-side (hiding unauthorized actions) and the REST API MUST enforce it server-side (returning 403 for unauthorized access).
|
||||
|
||||
### Key Entities *(include if feature involves data)*
|
||||
|
||||
- **Page/Route**: Each UI page maps to a client-side route (e.g., `/dashboard`, `/conversations/:id`, `/channels`, `/channels/:id`, `/agents`, `/agents/:id`, `/settings`, `/login`). Routes are managed by the Svelte router with history-mode navigation.
|
||||
- **SSE Event Stream**: A persistent server-to-client connection at `GET /api/v1/events` that emits typed events (`message.new`, `message.status`, `agent.status`, `channel.member`). Each event includes a monotonic sequence ID for reconnection reconciliation.
|
||||
- **Compose Payload**: The data structure submitted when sending a message: `{ recipient_type: "agent" | "channel", recipient_id: string, subject?: string, body: string, priority: number }`.
|
||||
- **Search Result**: A result object containing `{ conversation_id, message_id, body_preview (highlighted), sender_name, timestamp, conversation_subject }`.
|
||||
- **Theme Preference**: A `localStorage` key (`synapbus-theme`) with values `"light"` or `"dark"`, read on app initialization before first render to prevent flash of wrong theme.
|
||||
- **Trace Entry (UI representation)**: Displayed in the agent detail view: `{ id, agent_name, action, details_summary, timestamp }` with expandable JSON details.
|
||||
|
||||
## Success Criteria *(mandatory)*
|
||||
|
||||
### Measurable Outcomes
|
||||
|
||||
- **SC-001**: A new user can create an account, log in, send a message, and read a response within 3 minutes without documentation, using only the Web UI.
|
||||
- **SC-002**: Messages sent by agents via MCP appear in the Web UI within 2 seconds via SSE, without manual refresh.
|
||||
- **SC-003**: The `make web` build produces a complete SPA bundle under 500KB gzipped (excluding source maps), embedded in the Go binary with zero additional runtime dependencies.
|
||||
- **SC-004**: All UI pages render correctly on Chrome, Firefox, and Safari (latest 2 versions) at mobile (375px), tablet (768px), and desktop (1440px) viewport widths, in both light and dark modes.
|
||||
- **SC-005**: Full-text search returns results within 500ms for a corpus of 10,000 messages.
|
||||
- **SC-006**: The SSE connection automatically recovers from a network interruption within 30 seconds and displays all messages that arrived during the disconnection without duplicates.
|
||||
- **SC-007**: An owner can view an agent's traces, filter by action type and date range, and revoke its API key entirely through the Web UI without using CLI or API tools directly.
|
||||
- **SC-008**: The Lighthouse accessibility score for all pages MUST be 90 or above, ensuring the UI is usable with screen readers and keyboard navigation.
|
||||
@@ -0,0 +1,276 @@
|
||||
# Tasks: Web UI
|
||||
|
||||
**Input**: Design documents from `/specs/005-web-ui/`
|
||||
**Prerequisites**: spec.md (required), constitution.md (required for principles)
|
||||
|
||||
**Tests**: Not explicitly requested in the feature specification. Test tasks are omitted.
|
||||
|
||||
**Organization**: Tasks are grouped by user story to enable independent implementation and testing of each story. User stories are ordered by priority (P1 first).
|
||||
|
||||
## Format: `[ID] [P?] [Story] Description`
|
||||
|
||||
- **[P]**: Can run in parallel (different files, no dependencies)
|
||||
- **[Story]**: Which user story this task belongs to (e.g., US1, US6)
|
||||
- Include exact file paths in descriptions
|
||||
|
||||
## Path Conventions
|
||||
|
||||
- **Svelte SPA**: `web/src/` (source), `internal/web/dist/` (build output)
|
||||
- **Go embedding**: `internal/web/embed.go`
|
||||
- **Go API handlers**: `internal/api/`
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Setup (Shared Infrastructure)
|
||||
|
||||
**Purpose**: Initialize Svelte 5 project, Tailwind CSS, and Go embedding scaffold
|
||||
|
||||
- [ ] T001 Initialize Svelte 5 project with Vite in `web/` directory: `package.json`, `svelte.config.js`, `vite.config.ts`, `tsconfig.json`. Configure build output to `../internal/web/dist/`.
|
||||
- [ ] T002 Install and configure Tailwind CSS v4 in `web/tailwind.config.js` and `web/src/app.css` with base, components, and utilities layers. Include CSS custom properties for dark mode theme variables.
|
||||
- [ ] T003 [P] Install client-side router (`svelte-spa-router` or equivalent) and create `web/src/App.svelte` with route definitions for: `/login`, `/dashboard`, `/conversations/:id`, `/channels`, `/channels/:id`, `/agents`, `/agents/:id`, `/settings`.
|
||||
- [ ] T004 [P] Create Go embedding scaffold at `internal/web/embed.go`: use `go:embed dist/*` to embed the built SPA. Export an `http.FileServer` handler that serves `index.html` for all non-API routes (SPA fallback).
|
||||
- [ ] T005 [P] Update `Makefile` `web` target to ensure build output lands in `internal/web/dist/`. Add a `web-dev` target for Vite dev server with API proxy to `localhost:8080`.
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Foundational (Blocking Prerequisites)
|
||||
|
||||
**Purpose**: Core infrastructure that MUST be complete before ANY user story can be implemented
|
||||
|
||||
**CRITICAL**: No user story work can begin until this phase is complete
|
||||
|
||||
- [ ] T006 Create shared API client module at `web/src/lib/api.ts`: base URL configuration, fetch wrapper with JSON parsing, automatic 401 detection (redirect to `/login`), CSRF token handling, and request/response type definitions.
|
||||
- [ ] T007 [P] Create auth store at `web/src/lib/stores/auth.ts`: Svelte 5 reactive state (`$state`) for current user, login status, and session validation. Export `login()`, `logout()`, `checkSession()` functions that call the API client.
|
||||
- [ ] T008 [P] Create theme system at `web/src/lib/stores/theme.ts`: read `localStorage` key `synapbus-theme` before first render (inline script in `web/index.html` to prevent flash), export `toggleTheme()`, apply `dark` class to `<html>` element. Svelte 5 `$state` rune for reactive theme.
|
||||
- [ ] T009 [P] Create SSE client module at `web/src/lib/sse.ts`: connect to `GET /api/v1/events`, parse typed events (`message.new`, `message.status`, `agent.status`, `channel.member`), track last event sequence ID, implement auto-reconnect with exponential backoff (1s, 2s, 4s, max 30s), reconcile missed events on reconnect by fetching since last sequence ID. Export reactive connection status (`connected`, `reconnecting`, `disconnected`).
|
||||
- [ ] T010 [P] Create shared UI components at `web/src/lib/components/`: `LoadingSkeleton.svelte` (content placeholder), `EmptyState.svelte` (icon + message + action button), `ReconnectBanner.svelte` (SSE disconnect notification), `SubmitButton.svelte` (loading state, disabled during pending).
|
||||
- [ ] T011 [P] Create layout component at `web/src/lib/components/Layout.svelte`: sidebar navigation (Dashboard, Conversations, Channels, Agents, Settings, Logout), top bar with global search input, user avatar, responsive hamburger menu for mobile (<768px). Include `ReconnectBanner` at top.
|
||||
- [ ] T012 Create auth guard wrapper at `web/src/lib/components/AuthGuard.svelte`: check session on mount, redirect to `/login?return=<current_url>` if unauthenticated. Wrap all authenticated routes in `App.svelte`.
|
||||
|
||||
**Checkpoint**: Foundation ready - user story implementation can now begin in parallel
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: User Story 6 - Login and Authentication (Priority: P1)
|
||||
|
||||
**Goal**: Users can log in, maintain sessions via httponly cookies, and log out. All other routes are gated behind authentication.
|
||||
|
||||
**Independent Test**: Open the Web UI, get redirected to login, enter valid credentials, confirm redirect to Dashboard. Close and reopen browser, confirm session persists. Click logout, confirm redirect to login.
|
||||
|
||||
### Implementation for User Story 6
|
||||
|
||||
- [ ] T013 [US6] Create login page at `web/src/routes/Login.svelte`: username and password fields, "Sign In" button with loading state, error display area. On submit, call `POST /api/v1/auth/login` via API client. On success, redirect to return URL or `/dashboard`. On 401, show "Invalid username or password".
|
||||
- [ ] T014 [US6] Implement return URL handling in login flow: `AuthGuard` saves current path as `?return=` query param, `Login.svelte` reads it and redirects after successful auth.
|
||||
- [ ] T015 [US6] Add logout handler: `Layout.svelte` logout button calls `POST /api/v1/auth/logout` via API client, clears auth store, redirects to `/login`.
|
||||
- [ ] T016 [US6] Handle cross-tab session invalidation: API client 401 interceptor in `web/src/lib/api.ts` clears auth store and redirects to `/login` on any 401 response, preventing partial state in stale tabs.
|
||||
|
||||
**Checkpoint**: Login, session persistence, and logout are fully functional. All subsequent stories depend on authenticated access.
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: User Story 1 - Dashboard and Conversation Browsing (Priority: P1)
|
||||
|
||||
**Goal**: Logged-in users see their recent conversations on the Dashboard and can open a conversation to view the full threaded message history with real-time SSE updates.
|
||||
|
||||
**Independent Test**: Log in, view dashboard, click into a conversation, confirm a message sent via MCP by an agent appears in the thread within 2 seconds without page refresh.
|
||||
|
||||
### Implementation for User Story 1
|
||||
|
||||
- [ ] T017 [P] [US1] Create TypeScript types at `web/src/lib/types.ts`: `Conversation` (id, subject, last_message_preview, sender_name, timestamp, unread_count), `Message` (id, conversation_id, sender_name, sender_type, subject, body, priority, timestamp, status), `PaginatedResponse<T>` (items, total, page, per_page).
|
||||
- [ ] T018 [P] [US1] Create conversation API functions at `web/src/lib/api/conversations.ts`: `getConversations(page, per_page)`, `getConversation(id)`, `getMessages(conversation_id, page, per_page)`, `markAsRead(conversation_id)`.
|
||||
- [ ] T019 [US1] Create Dashboard page at `web/src/routes/Dashboard.svelte`: fetch conversations sorted by last activity, render list with `LoadingSkeleton` during load, each item shows last message preview, sender name, timestamp, and unread badge. Show `EmptyState` with "No conversations yet" guidance and "New Message" action when empty. Click navigates to `/conversations/:id`.
|
||||
- [ ] T020 [US1] Create Conversation thread page at `web/src/routes/Conversation.svelte`: fetch messages for conversation ID from route param, render in chronological order, show `LoadingSkeleton` during load. Each message shows sender name, avatar placeholder, timestamp, and body. Long messages (>500 chars) truncated with "Show more" toggle.
|
||||
- [ ] T021 [US1] Implement infinite scroll pagination in `Conversation.svelte`: load 50 messages initially, detect scroll to top, fetch older page, prepend messages while preserving scroll position.
|
||||
- [ ] T022 [US1] Implement SSE integration for real-time messages: subscribe to `message.new` events in `Conversation.svelte`, append new messages matching current conversation ID to the bottom of the thread. Update conversation list on Dashboard when `message.new` arrives (increment unread count, update preview).
|
||||
- [ ] T023 [US1] Call `markAsRead(conversation_id)` when opening a conversation thread, reset unread count to 0 in the conversation list.
|
||||
|
||||
**Checkpoint**: Dashboard shows conversations, clicking opens threaded view, real-time messages arrive via SSE, pagination works.
|
||||
|
||||
---
|
||||
|
||||
## Phase 5: User Story 2 - Compose and Send Messages (Priority: P1)
|
||||
|
||||
**Goal**: Users can compose and send direct messages to agents or post messages to channels via a compose form with searchable recipient selection.
|
||||
|
||||
**Independent Test**: Open compose form, select an agent recipient, type a message, click send, confirm the message appears in the conversation thread and is received by the target agent via MCP `read_inbox`.
|
||||
|
||||
### Implementation for User Story 2
|
||||
|
||||
- [ ] T024 [P] [US2] Create message API functions at `web/src/lib/api/messages.ts`: `sendMessage(payload: ComposePayload)`, `searchRecipients(query: string)` returning agents and channels matching the query.
|
||||
- [ ] T025 [P] [US2] Create `RecipientSearch.svelte` component at `web/src/lib/components/RecipientSearch.svelte`: searchable dropdown input with 300ms debounce, keyboard navigation (arrow keys + enter), virtualized rendering for large lists (>100 items), shows agent/channel name and type icon.
|
||||
- [ ] T026 [US2] Create compose form at `web/src/lib/components/ComposeForm.svelte`: `RecipientSearch` for recipient selection, subject text input (optional), body textarea (required, validate non-empty), priority selector (1-10 range input, default 5). Submit button with loading state. On success, navigate to the conversation thread. On validation error, show inline "Message body is required".
|
||||
- [ ] T027 [US2] Add "New Message" button to `Dashboard.svelte` and `Layout.svelte` that opens `ComposeForm` as a modal or navigates to `/compose` route. Wire up route in `App.svelte` if using a dedicated page.
|
||||
|
||||
**Checkpoint**: Users can compose messages to agents/channels, messages appear in conversation threads. Combined with US1 and US6, this is the MVP.
|
||||
|
||||
---
|
||||
|
||||
## Phase 6: User Story 3 - Agent Management and Trace Viewing (Priority: P2)
|
||||
|
||||
**Goal**: Owners can view their agents, inspect activity traces with filters, and manage API keys (revoke, regenerate).
|
||||
|
||||
**Independent Test**: Navigate to Agents page, click an agent, view traces filtered by action type and date range, revoke API key, confirm agent can no longer authenticate.
|
||||
|
||||
### Implementation for User Story 3
|
||||
|
||||
- [ ] T028 [P] [US3] Create agent TypeScript types in `web/src/lib/types.ts`: `Agent` (id, name, display_name, type, status, messages_24h), `TraceEntry` (id, agent_name, action, details_summary, details_json, timestamp).
|
||||
- [ ] T029 [P] [US3] Create agent API functions at `web/src/lib/api/agents.ts`: `getAgents()`, `getAgent(id)`, `getAgentTraces(id, filters: { action_type?, date_from?, date_to?, page, per_page })`, `revokeApiKey(agent_id)`, `regenerateApiKey(agent_id)`.
|
||||
- [ ] T030 [US3] Create Agents list page at `web/src/routes/Agents.svelte`: fetch and display owned agents in a card/list layout with `LoadingSkeleton` during load. Each agent shows name, display_name, type badge (ai/human), status indicator (active=green, inactive=gray), and messages sent in last 24h. Click navigates to `/agents/:id`.
|
||||
- [ ] T031 [US3] Create Agent detail page at `web/src/routes/AgentDetail.svelte`: display agent info header (name, display_name, type, status), tabbed sections for "Activity Traces" and "API Key Management".
|
||||
- [ ] T032 [US3] Implement trace viewer tab in `AgentDetail.svelte`: paginated reverse-chronological trace list, filter controls for action type (dropdown) and date range (date inputs). Each trace row shows action, details_summary, timestamp. Expandable row reveals full JSON details with syntax highlighting.
|
||||
- [ ] T033 [US3] Implement API key management tab in `AgentDetail.svelte`: "Revoke API Key" button with confirmation dialog ("Are you sure? This agent will no longer be able to authenticate."), "Regenerate API Key" button that shows the new key once in a copiable field with warning "This key will not be shown again."
|
||||
- [ ] T034 [US3] Subscribe to `agent.status` SSE events in `Agents.svelte` and `AgentDetail.svelte` to update agent status in real-time when keys are revoked or agents go offline.
|
||||
|
||||
**Checkpoint**: Agent management is fully functional - list, detail, traces, API key operations, real-time status updates.
|
||||
|
||||
---
|
||||
|
||||
## Phase 7: User Story 4 - Channel Management (Priority: P2)
|
||||
|
||||
**Goal**: Users can browse channels, view details and membership, create new channels, and manage members for owned channels.
|
||||
|
||||
**Independent Test**: Navigate to Channels page, create a new public channel, view member list, confirm an agent can join via MCP `join_channel`.
|
||||
|
||||
### Implementation for User Story 4
|
||||
|
||||
- [ ] T035 [P] [US4] Create channel TypeScript types in `web/src/lib/types.ts`: `Channel` (id, name, description, type, member_count, last_activity, is_owner), `ChannelMember` (agent_id, agent_name, joined_at).
|
||||
- [ ] T036 [P] [US4] Create channel API functions at `web/src/lib/api/channels.ts`: `getChannels()`, `getChannel(id)`, `getChannelMembers(id)`, `createChannel(name, description, type)`, `inviteToChannel(channel_id, agent_ids)`, `removeFromChannel(channel_id, agent_id)`.
|
||||
- [ ] T037 [US4] Create Channels list page at `web/src/routes/Channels.svelte`: display public channels and private channels user belongs to, each showing name, description, member count, and last activity. "Create Channel" button at top. `LoadingSkeleton` during load.
|
||||
- [ ] T038 [US4] Create "Create Channel" modal/form in `web/src/lib/components/CreateChannelForm.svelte`: name input (required), description textarea, type selector (public/private). On success, navigate to the new channel's detail page.
|
||||
- [ ] T039 [US4] Create Channel detail page at `web/src/routes/ChannelDetail.svelte`: channel info header, message thread (reuse conversation message rendering from `Conversation.svelte` or extract shared `MessageList.svelte` component), member list sidebar.
|
||||
- [ ] T040 [US4] Implement member management in `ChannelDetail.svelte`: for owned channels, show "Invite" button that opens a searchable agent list (reuse `RecipientSearch`), and "Remove" button next to each member. For non-owned channels, hide management controls (read-only view of members and messages).
|
||||
- [ ] T041 [US4] Subscribe to `channel.member` SSE events to update member lists in real-time when agents join or leave channels.
|
||||
|
||||
**Checkpoint**: Channel browsing, creation, member management, and real-time membership updates are functional.
|
||||
|
||||
---
|
||||
|
||||
## Phase 8: User Story 5 - Full-Text Search (Priority: P2)
|
||||
|
||||
**Goal**: Users can search messages across all conversations and channels, with highlighted results grouped by conversation and direct links to message context.
|
||||
|
||||
**Independent Test**: Send several messages with known content via MCP, search for a keyword in the Web UI, confirm matching messages appear with highlighted terms and correct conversation links.
|
||||
|
||||
### Implementation for User Story 5
|
||||
|
||||
- [ ] T042 [P] [US5] Create search TypeScript types in `web/src/lib/types.ts`: `SearchResult` (conversation_id, message_id, body_preview, sender_name, timestamp, conversation_subject), `SearchResponse` (results grouped by conversation_id, total count).
|
||||
- [ ] T043 [P] [US5] Create search API function at `web/src/lib/api/search.ts`: `searchMessages(query: string, page?, per_page?)` returning `SearchResponse`.
|
||||
- [ ] T044 [US5] Implement global search bar in `Layout.svelte`: text input in top bar, submit on Enter, navigate to `/search?q=<query>`. Debounce is NOT needed here (search triggers on explicit submit, not on keystroke).
|
||||
- [ ] T045 [US5] Create Search results page at `web/src/routes/Search.svelte`: read query from URL params, call search API, display results grouped by conversation. Each result shows message preview with highlighted match term (use `<mark>` tags), sender name, timestamp. Show `EmptyState` "No messages found for '<query>'" when no results.
|
||||
- [ ] T046 [US5] Implement click-to-context in search results: clicking a result navigates to `/conversations/:conversation_id?highlight=:message_id`, and `Conversation.svelte` scrolls to and visually highlights the target message.
|
||||
|
||||
**Checkpoint**: Full-text search works end-to-end with highlighted results, grouped display, and click-through to message context.
|
||||
|
||||
---
|
||||
|
||||
## Phase 9: User Story 7 - Settings and Dark Mode (Priority: P3)
|
||||
|
||||
**Goal**: Users can change password, toggle dark/light mode, and manage preferences from the Settings page.
|
||||
|
||||
**Independent Test**: Toggle dark mode in Settings, confirm all pages render correctly in both themes. Change password, confirm login works with new password.
|
||||
|
||||
### Implementation for User Story 7
|
||||
|
||||
- [ ] T047 [P] [US7] Create settings API functions at `web/src/lib/api/settings.ts`: `changePassword(current_password, new_password)`.
|
||||
- [ ] T048 [US7] Create Settings page at `web/src/routes/Settings.svelte`: sections for "Appearance" (dark mode toggle) and "Security" (change password form).
|
||||
- [ ] T049 [US7] Implement dark mode toggle in Settings page: switch component bound to theme store, toggles immediately without reload. Wire to `toggleTheme()` from `web/src/lib/stores/theme.ts`.
|
||||
- [ ] T050 [US7] Implement change password form in Settings page: current password, new password, confirm new password fields. Client-side validation: new password minimum 8 characters ("Password must be at least 8 characters"), passwords match. On success, show success message. Session remains active.
|
||||
- [ ] T051 [US7] Audit all components and pages for dark mode support: verify Tailwind `dark:` variant classes are applied to all backgrounds, text colors, borders, inputs, buttons, cards, and modals across every page.
|
||||
|
||||
**Checkpoint**: Settings page functional with dark mode toggle and password change. Both themes render correctly across all pages.
|
||||
|
||||
---
|
||||
|
||||
## Phase 10: Polish & Cross-Cutting Concerns
|
||||
|
||||
**Purpose**: Improvements that affect multiple user stories
|
||||
|
||||
- [ ] T052 [P] Responsive audit: test all pages at 375px, 768px, 1440px, and 2560px viewport widths. Fix layout issues with Tailwind responsive breakpoints (`sm:`, `md:`, `lg:`, `xl:`). Ensure sidebar collapses to hamburger menu on mobile.
|
||||
- [ ] T053 [P] Accessibility audit: add ARIA labels to all interactive elements, ensure keyboard navigation works for all dropdowns and modals, verify focus management on route changes, test with screen reader. Target Lighthouse accessibility score >= 90.
|
||||
- [ ] T054 [P] Extract shared `MessageList.svelte` component from `Conversation.svelte` for reuse in `ChannelDetail.svelte` — unified message rendering with syntax-highlighted code blocks (use a lightweight highlighter like Prism or Shiki).
|
||||
- [ ] T055 [P] Add loading skeletons to all remaining pages that don't have them: `AgentDetail.svelte` trace list, `ChannelDetail.svelte` member list, `Search.svelte` results.
|
||||
- [ ] T056 Bundle size optimization: verify `make web` produces a gzipped bundle under 500KB (SC-003). Configure Vite build with tree-shaking, code splitting per route, and minification. Add bundle analyzer script.
|
||||
- [ ] T057 Error boundary component at `web/src/lib/components/ErrorBoundary.svelte`: catch rendering errors, display user-friendly "Something went wrong" message with retry button instead of blank screen.
|
||||
- [ ] T058 Verify Go embedding works end-to-end: run `make web && make build`, start binary, confirm all SPA routes serve correctly, API proxy works, and `index.html` fallback handles client-side routing.
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & Execution Order
|
||||
|
||||
### Phase Dependencies
|
||||
|
||||
- **Setup (Phase 1)**: No dependencies - can start immediately
|
||||
- **Foundational (Phase 2)**: Depends on Setup completion - BLOCKS all user stories
|
||||
- **US6 Login (Phase 3)**: Depends on Foundational - BLOCKS all other user stories (auth gates everything)
|
||||
- **US1 Dashboard (Phase 4)**: Depends on US6 (requires authenticated session)
|
||||
- **US2 Compose (Phase 5)**: Depends on US6; benefits from US1 (conversation thread to view sent message) but compose API call is independently testable
|
||||
- **US3 Agents (Phase 6)**: Depends on US6; independent of US1/US2
|
||||
- **US4 Channels (Phase 7)**: Depends on US6; can reuse components from US1 (message list) and US2 (recipient search)
|
||||
- **US5 Search (Phase 8)**: Depends on US6 and US1 (search results link to conversation thread view)
|
||||
- **US7 Settings (Phase 9)**: Depends on US6 and Phase 2 theme system; independent of all other stories
|
||||
- **Polish (Phase 10)**: Depends on all user stories being complete
|
||||
|
||||
### User Story Dependencies
|
||||
|
||||
- **US6 (P1, Login)**: Can start after Foundational (Phase 2) - BLOCKS all other stories
|
||||
- **US1 (P1, Dashboard)**: Can start after US6 - No dependencies on other stories
|
||||
- **US2 (P1, Compose)**: Can start after US6 - Benefits from US1 for conversation view but independently testable
|
||||
- **US3 (P2, Agents)**: Can start after US6 - Independent of US1/US2
|
||||
- **US4 (P2, Channels)**: Can start after US6 - May reuse components from US1/US2
|
||||
- **US5 (P2, Search)**: Can start after US6 + US1 - Depends on conversation thread view for click-through
|
||||
- **US7 (P3, Settings)**: Can start after US6 - Independent of all other stories
|
||||
|
||||
### Within Each User Story
|
||||
|
||||
- Types and API functions (marked [P]) can be built in parallel
|
||||
- Pages depend on their API functions and types
|
||||
- SSE integration depends on the base page being rendered
|
||||
- Story complete before moving to next priority
|
||||
|
||||
### Parallel Opportunities
|
||||
|
||||
- All Setup tasks T003, T004, T005 can run in parallel (after T001/T002)
|
||||
- All Foundational tasks T007-T012 marked [P] can run in parallel (after T006)
|
||||
- After US6 is complete: US1, US2, US3, US4, US7 can start in parallel
|
||||
- US5 requires US1 to be complete (conversation thread view)
|
||||
- Within each story: types and API function tasks marked [P] can run in parallel
|
||||
- All Polish tasks marked [P] can run in parallel
|
||||
|
||||
---
|
||||
|
||||
## Implementation Strategy
|
||||
|
||||
### MVP First (US6 + US1 + US2)
|
||||
|
||||
1. Complete Phase 1: Setup
|
||||
2. Complete Phase 2: Foundational
|
||||
3. Complete Phase 3: US6 - Login (gates everything)
|
||||
4. Complete Phase 4: US1 - Dashboard & Conversations
|
||||
5. Complete Phase 5: US2 - Compose & Send
|
||||
6. **STOP and VALIDATE**: Test the core loop - login, browse conversations, send messages, see real-time updates
|
||||
7. Deploy/demo if ready
|
||||
|
||||
### Incremental Delivery
|
||||
|
||||
1. Setup + Foundational -> Foundation ready
|
||||
2. Add US6 Login -> Authenticated access works
|
||||
3. Add US1 Dashboard -> Browse conversations with real-time updates (MVP!)
|
||||
4. Add US2 Compose -> Two-way messaging (full MVP!)
|
||||
5. Add US3 Agents -> Agent oversight and API key management
|
||||
6. Add US4 Channels -> Channel organization
|
||||
7. Add US5 Search -> Message discovery at scale
|
||||
8. Add US7 Settings -> Dark mode and password management
|
||||
9. Polish -> Responsive, accessible, optimized
|
||||
|
||||
---
|
||||
|
||||
## Notes
|
||||
|
||||
- [P] tasks = different files, no dependencies
|
||||
- [Story] label maps task to specific user story for traceability
|
||||
- US6 (Login) is extracted to its own phase before other P1 stories because it gates all authenticated functionality
|
||||
- Svelte 5 runes (`$state`, `$derived`, `$effect`) should be used throughout instead of Svelte 4 stores
|
||||
- All API calls go through the shared client in `web/src/lib/api.ts` for consistent auth handling
|
||||
- Theme initialization happens in `web/index.html` via inline script (before Svelte mounts) to prevent flash
|
||||
- SSE reconnection logic must track sequence IDs to avoid duplicate or missed messages
|
||||
- `internal/web/embed.go` must handle SPA fallback: serve `index.html` for any path not matching a static asset
|
||||
@@ -0,0 +1,153 @@
|
||||
# Feature Specification: MCP Server
|
||||
|
||||
**Feature Branch**: `006-mcp-server`
|
||||
**Created**: 2026-03-13
|
||||
**Status**: Draft
|
||||
**Input**: User description: "MCP server using mark3labs/mcp-go with SSE and Streamable HTTP transports, exposing all messaging operations as MCP tools with API key auth, JSON schema tool listing, health check, and connection management."
|
||||
|
||||
## User Scenarios & Testing *(mandatory)*
|
||||
|
||||
### User Story 1 - Agent Connects and Calls Messaging Tools via SSE (Priority: P1)
|
||||
|
||||
An AI agent operator configures their MCP client (e.g., Claude Desktop, a custom LLM agent) to connect to SynapBus's MCP server over SSE transport. The agent authenticates using an API key passed in the `Authorization` header. Once connected, the agent lists available tools, sees JSON schema descriptions for each, and calls `send_message` to send a direct message to another agent. The agent then calls `read_inbox` to check for incoming messages.
|
||||
|
||||
**Why this priority**: This is the foundational interaction path. Without SSE transport and tool execution working end-to-end, no agent can use SynapBus at all. SSE is the primary transport per the constitution (Principle II).
|
||||
|
||||
**Independent Test**: Can be fully tested by starting `synapbus serve`, connecting an MCP client over SSE with a valid API key, listing tools, and sending/reading a message. Delivers the core value of MCP-native agent messaging.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** SynapBus is running and an agent has a valid API key, **When** the agent connects via SSE transport at `/mcp/sse`, **Then** the connection is established, the server sends an SSE endpoint event, and the agent can issue MCP `initialize` and `tools/list` requests.
|
||||
2. **Given** an agent is connected via SSE, **When** the agent calls `tools/list`, **Then** the server returns all available MCP tools with their names, descriptions, and JSON Schema `inputSchema` definitions.
|
||||
3. **Given** an agent is connected via SSE, **When** the agent calls `tools/call` with `send_message` and a valid payload (recipient, body), **Then** the message is persisted and the server returns a success result containing the message ID.
|
||||
4. **Given** an agent is connected via SSE, **When** the agent calls `tools/call` with `read_inbox`, **Then** the server returns the agent's pending messages as structured JSON content.
|
||||
|
||||
---
|
||||
|
||||
### User Story 2 - Agent Authenticates with API Key in MCP Headers (Priority: P1)
|
||||
|
||||
An AI agent attempts to connect to the MCP server. The server validates the API key provided in the HTTP `Authorization` header (as `Bearer <api-key>`) during the initial SSE or Streamable HTTP connection handshake. If the key is valid, the server associates the connection with the corresponding agent identity and allows tool calls scoped to that agent. If the key is missing or invalid, the server rejects the connection immediately.
|
||||
|
||||
**Why this priority**: Authentication is a hard requirement before any tool execution. Without it, any client could impersonate any agent or access arbitrary messages, violating Principle IV (multi-tenant with ownership).
|
||||
|
||||
**Independent Test**: Can be tested by attempting connections with valid, invalid, and missing API keys and verifying the server accepts or rejects appropriately.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** an agent has a valid API key, **When** the agent connects to `/mcp/sse` with `Authorization: Bearer <valid-key>`, **Then** the connection is accepted and subsequent tool calls are scoped to that agent's identity.
|
||||
2. **Given** no API key is provided, **When** a client connects to `/mcp/sse` without an `Authorization` header, **Then** the server responds with HTTP 401 Unauthorized and closes the connection.
|
||||
3. **Given** an invalid API key is provided, **When** a client connects with `Authorization: Bearer <invalid-key>`, **Then** the server responds with HTTP 401 Unauthorized and the connection is not established.
|
||||
4. **Given** an agent is authenticated, **When** the agent calls `send_message` specifying itself as the sender, **Then** the server accepts the call. **When** the agent attempts to call `read_inbox` for a different agent, **Then** the server returns an MCP error result indicating access denied.
|
||||
|
||||
---
|
||||
|
||||
### User Story 3 - Agent Connects via Streamable HTTP Transport (Priority: P2)
|
||||
|
||||
An agent operator whose environment does not support long-lived SSE connections (e.g., serverless functions, firewalled networks) connects to SynapBus using the Streamable HTTP transport. The agent sends MCP requests as HTTP POST to `/mcp` and receives responses (including streaming results) over the same HTTP connection. All the same tools and authentication mechanisms work identically to the SSE transport.
|
||||
|
||||
**Why this priority**: Streamable HTTP is the secondary transport. Some deployment environments cannot maintain persistent SSE connections, so this transport broadens compatibility. However, SSE covers the majority of use cases, making this P2.
|
||||
|
||||
**Independent Test**: Can be tested by sending MCP JSON-RPC requests via HTTP POST to `/mcp` with a valid API key and verifying tool responses match SSE behavior.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** SynapBus is running, **When** an agent sends an MCP `initialize` request as HTTP POST to `/mcp` with `Authorization: Bearer <valid-key>`, **Then** the server responds with a valid MCP initialize result containing server capabilities.
|
||||
2. **Given** an agent is using Streamable HTTP, **When** the agent calls `tools/list` via POST, **Then** the response contains the same tool set with the same JSON schemas as the SSE transport.
|
||||
3. **Given** an agent is using Streamable HTTP, **When** the agent calls `tools/call` with `send_message`, **Then** the message is persisted identically to an SSE-originated call and the response format matches.
|
||||
|
||||
---
|
||||
|
||||
### User Story 4 - Operator Monitors Server Health and Connected Agents (Priority: P2)
|
||||
|
||||
A system operator or monitoring tool checks the MCP server's health endpoint to verify the service is running and responsive. The operator can also query the connection management subsystem to see how many agents are currently connected, their identities, transport type, and connection duration. This enables operational monitoring and capacity planning.
|
||||
|
||||
**Why this priority**: Health checks are essential for production deployments (load balancers, Kubernetes liveness probes, Docker health checks). Connection tracking supports observability (Principle VIII). However, these are operational concerns, not core messaging, making them P2.
|
||||
|
||||
**Independent Test**: Can be tested by calling the health endpoint and verifying the response, then connecting multiple agents and querying the connection list.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** SynapBus is running, **When** an HTTP GET request is sent to `/health`, **Then** the server responds with HTTP 200 and a JSON body containing at minimum `{"status": "ok"}` and the server version.
|
||||
2. **Given** SynapBus is running but the database is unreachable, **When** an HTTP GET request is sent to `/health`, **Then** the server responds with HTTP 503 and a JSON body indicating the unhealthy component.
|
||||
3. **Given** three agents are connected via SSE and one via Streamable HTTP, **When** an authenticated operator queries connected agents (via an internal MCP tool or REST endpoint), **Then** the response lists all four connections with agent ID, transport type (`sse` or `streamable-http`), connected-at timestamp, and last-activity timestamp.
|
||||
4. **Given** an agent disconnects (SSE connection drops), **When** the operator queries connected agents, **Then** the disconnected agent is no longer listed.
|
||||
|
||||
---
|
||||
|
||||
### User Story 5 - Agent Discovers Available Tools with Full JSON Schemas (Priority: P3)
|
||||
|
||||
A newly developed agent connects to SynapBus for the first time and needs to understand what operations are available. The agent calls `tools/list` and receives a comprehensive list of all MCP tools with human-readable descriptions and full JSON Schema definitions for each tool's input parameters. The schemas include property types, required fields, enums for constrained values, and description strings for each parameter.
|
||||
|
||||
**Why this priority**: Tool discovery is handled by the MCP protocol's built-in `tools/list` method. While essential for agent usability, the JSON schema quality is an incremental improvement over having the tools work at all (covered in P1). This story focuses on schema completeness and documentation quality.
|
||||
|
||||
**Independent Test**: Can be tested by connecting and calling `tools/list`, then validating each returned tool's `inputSchema` is a valid JSON Schema with descriptions on all parameters.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** an agent is connected, **When** the agent calls `tools/list`, **Then** every tool in the response has a non-empty `description` string explaining what the tool does in plain language.
|
||||
2. **Given** an agent is connected, **When** the agent calls `tools/list`, **Then** every tool's `inputSchema` is a valid JSON Schema object with `type`, `properties`, and `required` fields defined.
|
||||
3. **Given** an agent examines the `send_message` tool schema, **Then** the schema defines `to` (string, required), `body` (string, required), `subject` (string, optional), `priority` (integer, optional, minimum 1, maximum 10), `channel_id` (string, optional), and `metadata` (object, optional) with descriptions for each property.
|
||||
|
||||
---
|
||||
|
||||
### Edge Cases
|
||||
|
||||
- What happens when an agent's API key is revoked while the agent has an active SSE connection? The server MUST terminate the connection within a reasonable timeframe (e.g., on the next tool call or within 60 seconds via a background sweep).
|
||||
- What happens when an SSE connection drops unexpectedly (network failure, client crash)? The server MUST detect the closed connection, clean up the connection tracking entry, and release any resources associated with that connection.
|
||||
- What happens when the maximum number of concurrent connections is reached? The server MUST reject new connections with HTTP 503 Service Unavailable and a descriptive error message, rather than silently dropping or hanging.
|
||||
- What happens when an agent sends a malformed MCP JSON-RPC request? The server MUST respond with a standard JSON-RPC error (`-32700` parse error or `-32600` invalid request) rather than crashing or returning an HTTP error.
|
||||
- What happens when a tool call references an entity that does not exist (e.g., sending a message to a non-existent agent)? The server MUST return an MCP tool error result with a clear error message, not a protocol-level error.
|
||||
- What happens when two agents connect with the same API key simultaneously? The server MUST either reject the second connection or support multiple concurrent connections per agent, with connection tracking reflecting all active sessions.
|
||||
- What happens when the health check endpoint is called during server startup before the database is initialized? The server MUST return HTTP 503 with a status indicating the service is starting up.
|
||||
|
||||
## Requirements *(mandatory)*
|
||||
|
||||
### Functional Requirements
|
||||
|
||||
- **FR-001**: System MUST implement an MCP server using `mark3labs/mcp-go` that handles the full MCP protocol lifecycle: `initialize`, `initialized`, `tools/list`, `tools/call`, and `ping`.
|
||||
- **FR-002**: System MUST support SSE transport, accepting connections on a configurable path (default `/mcp/sse`) with standard MCP SSE semantics (server-sent events for server-to-client, HTTP POST for client-to-server).
|
||||
- **FR-003**: System MUST support Streamable HTTP transport, accepting MCP JSON-RPC requests via HTTP POST on a configurable path (default `/mcp`).
|
||||
- **FR-004**: System MUST authenticate MCP connections using API keys passed in the HTTP `Authorization` header as `Bearer <api-key>`. The API key MUST be validated on the initial connection (SSE) or on each request (Streamable HTTP).
|
||||
- **FR-005**: System MUST reject unauthenticated or invalidly authenticated connections with HTTP 401 Unauthorized before any MCP protocol messages are exchanged.
|
||||
- **FR-006**: System MUST expose all messaging operations as MCP tools: `send_message`, `read_inbox`, `claim_messages`, `mark_done`, `search_messages`, `create_channel`, `join_channel`, `list_channels`, `register_agent`, `discover_agents`.
|
||||
- **FR-007**: Each MCP tool MUST have a complete JSON Schema `inputSchema` with `type`, `properties`, `required` array, and human-readable `description` strings on both the tool itself and each input property.
|
||||
- **FR-008**: System MUST scope all tool calls to the authenticated agent's identity. An agent MUST NOT be able to read another agent's inbox, send messages impersonating another agent, or access channels it has not joined.
|
||||
- **FR-009**: System MUST provide a health check endpoint at `/health` (HTTP GET, no authentication required) that returns the server's health status, version, and component states (database, MCP server).
|
||||
- **FR-010**: System MUST track all active MCP connections, recording: agent ID, transport type (SSE or Streamable HTTP), connection timestamp, and last activity timestamp.
|
||||
- **FR-011**: System MUST clean up connection tracking entries when SSE connections are closed (client disconnect, server shutdown, or error).
|
||||
- **FR-012**: System MUST log all MCP tool calls via `slog` structured logging, including agent ID, tool name, request duration, and success/failure status (Principle VIII).
|
||||
- **FR-013**: System MUST return standard MCP error responses for tool failures (not HTTP errors), using the `isError` field in tool results with descriptive error messages.
|
||||
- **FR-014**: System MUST handle graceful shutdown: on SIGTERM/SIGINT, the server MUST stop accepting new connections, allow in-flight requests to complete (with a configurable timeout, default 30 seconds), and close all SSE connections.
|
||||
|
||||
### Key Entities
|
||||
|
||||
- **MCPServer**: The top-level server component that registers tools, manages transports, and dispatches tool calls to the core engine. Wraps `mark3labs/mcp-go` server instance. Configured with server name, version, and transport options.
|
||||
- **MCPConnection**: Represents an active agent connection. Attributes: connection ID (UUID), agent ID (resolved from API key), transport type (SSE or Streamable HTTP), connected-at timestamp, last-activity timestamp, remote address.
|
||||
- **MCPTool**: A registered MCP tool definition. Attributes: name, description, input JSON schema, handler function reference. Each tool maps to a core engine operation.
|
||||
- **ToolCallContext**: Per-request context created for each `tools/call` invocation. Contains: authenticated agent identity, request ID, tool name, raw arguments. Passed to the core engine handler for authorization and execution.
|
||||
- **HealthStatus**: Response object for the `/health` endpoint. Attributes: overall status (ok/degraded/unhealthy), server version, uptime, component statuses (database, mcp_server), active connection count.
|
||||
|
||||
## Success Criteria *(mandatory)*
|
||||
|
||||
### Measurable Outcomes
|
||||
|
||||
- **SC-001**: An MCP client (e.g., `npx @modelcontextprotocol/inspector`) can connect via SSE, list tools, and execute a `send_message` / `read_inbox` round-trip within 5 seconds on localhost.
|
||||
- **SC-002**: All MCP tools return valid JSON Schema `inputSchema` definitions that pass JSON Schema Draft 2020-12 validation.
|
||||
- **SC-003**: Unauthenticated connection attempts are rejected with HTTP 401 within 100ms, with zero tool calls executed.
|
||||
- **SC-004**: The health check endpoint responds within 200ms under normal conditions and correctly reports degraded status when the database is unavailable.
|
||||
- **SC-005**: The server handles at least 50 concurrent SSE connections without connection drops or degraded tool call latency (p99 < 500ms for simple tool calls).
|
||||
- **SC-006**: When an SSE client disconnects, the connection tracking entry is removed within 5 seconds.
|
||||
- **SC-007**: All tool calls are logged with agent ID, tool name, duration, and outcome, verifiable by inspecting structured log output.
|
||||
- **SC-008**: Both SSE and Streamable HTTP transports produce identical tool results for the same inputs, verified by running the same test suite against both transports.
|
||||
|
||||
## Constitution Compliance
|
||||
|
||||
| Principle | Compliance |
|
||||
|-----------|------------|
|
||||
| I. Single Binary | MCP server is embedded in the main binary; no external MCP broker or proxy |
|
||||
| II. MCP-Native | This spec implements the core MCP interface; all agent operations are MCP tools |
|
||||
| III. Pure Go, Zero CGO | `mark3labs/mcp-go` is pure Go; no CGO dependencies introduced |
|
||||
| IV. Multi-Tenant | API key authentication scopes every tool call to the authenticated agent |
|
||||
| V. Embedded OAuth 2.1 | API keys serve as the agent authentication mechanism; OAuth applies to human users (separate spec) |
|
||||
| VIII. Observable | All tool calls logged via `slog`; connection tracking enables monitoring |
|
||||
| IX. Progressive Complexity | Basic tools (send, read, mark done) work without channels or search configured |
|
||||
@@ -0,0 +1,206 @@
|
||||
# Tasks: MCP Server
|
||||
|
||||
**Input**: Design documents from `/specs/006-mcp-server/`
|
||||
**Prerequisites**: spec.md (required), constitution.md (required)
|
||||
|
||||
**Tests**: Not explicitly requested in the feature specification. Test tasks are omitted.
|
||||
|
||||
**Organization**: Tasks are grouped by user story to enable independent implementation and testing of each story.
|
||||
|
||||
## Format: `[ID] [P?] [Story] Description`
|
||||
|
||||
- **[P]**: Can run in parallel (different files, no dependencies)
|
||||
- **[Story]**: Which user story this task belongs to (e.g., US1, US2, US3)
|
||||
- Include exact file paths in descriptions
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Setup (Shared Infrastructure)
|
||||
|
||||
**Purpose**: Add `mark3labs/mcp-go` dependency and create the base package structure for the MCP server.
|
||||
|
||||
- [ ] T001 Add `mark3labs/mcp-go` and `go-chi/chi` dependencies to `go.mod` via `go get`
|
||||
- [ ] T002 [P] Create package scaffold `internal/mcp/doc.go` with package-level doc comment describing the MCP server subsystem
|
||||
- [ ] T003 [P] Create `internal/mcp/config.go` with `Config` struct: server name, version, SSE path (default `/mcp/sse`), Streamable HTTP path (default `/mcp`), health path (default `/health`), max connections (default 100), shutdown timeout (default 30s)
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Foundational (Blocking Prerequisites)
|
||||
|
||||
**Purpose**: Core types, interfaces, and the engine bridge that ALL user stories depend on. No MCP tool or transport can work without these.
|
||||
|
||||
**CRITICAL**: No user story work can begin until this phase is complete.
|
||||
|
||||
- [ ] T004 Define `ToolCallContext` struct in `internal/mcp/context.go` containing: agent ID (resolved from API key), request ID, tool name, raw arguments map, and a `context.Context` field for propagation
|
||||
- [ ] T005 [P] Define `MCPConnection` struct in `internal/mcp/connection.go` with: connection ID (UUID), agent ID, transport type enum (`sse` | `streamable-http`), connected-at timestamp, last-activity timestamp, remote address
|
||||
- [ ] T006 [P] Define `HealthStatus` struct in `internal/mcp/health.go` with: overall status (`ok` | `degraded` | `unhealthy`), server version, uptime, component statuses map (database, mcp_server), active connection count
|
||||
- [ ] T007 Define `Engine` interface in `internal/mcp/engine.go` that abstracts the core operations the MCP tools will call: `SendMessage`, `ReadInbox`, `ClaimMessages`, `MarkDone`, `SearchMessages`, `CreateChannel`, `JoinChannel`, `ListChannels`, `RegisterAgent`, `DiscoverAgents`, `ValidateAPIKey(ctx, key) (agentID, error)`, `HealthCheck(ctx) HealthStatus`
|
||||
- [ ] T008 Implement `ConnectionManager` in `internal/mcp/connmgr.go`: thread-safe map of active `MCPConnection` entries with `Add`, `Remove`, `Get`, `List`, `UpdateActivity`, and `Count` methods using `sync.RWMutex`
|
||||
- [ ] T009 [P] Add structured `slog` logging helpers in `internal/mcp/logging.go`: `LogToolCall(logger, agentID, toolName, duration, err)` and `LogConnection(logger, agentID, transport, event)` that emit structured key-value log entries per FR-012
|
||||
|
||||
**Checkpoint**: Foundation ready - all types, interfaces, and connection tracking in place. User story implementation can now begin.
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: User Story 1 - Agent Connects and Calls Messaging Tools via SSE (Priority: P1) MVP
|
||||
|
||||
**Goal**: An MCP client can connect over SSE, list all tools with JSON schemas, and execute `send_message` / `read_inbox` round-trips.
|
||||
|
||||
**Independent Test**: Start `synapbus serve`, connect an MCP client (e.g., `@modelcontextprotocol/inspector`) over SSE at `/mcp/sse`, list tools, send a message, read inbox.
|
||||
|
||||
### Implementation for User Story 1
|
||||
|
||||
- [ ] T010 [US1] Implement tool definitions in `internal/mcp/tools.go`: register all 10 MCP tools (`send_message`, `read_inbox`, `claim_messages`, `mark_done`, `search_messages`, `create_channel`, `join_channel`, `list_channels`, `register_agent`, `discover_agents`) with `mcp.NewTool()` from `mark3labs/mcp-go`, each with name, description string, and JSON Schema `inputSchema` defining `type`, `properties`, `required`, and per-property `description` fields (FR-006, FR-007)
|
||||
- [ ] T011 [US1] Implement tool handler dispatch in `internal/mcp/handlers.go`: a `HandleToolCall(ctx, toolCallContext) (*mcp.CallToolResult, error)` function that switches on tool name, extracts typed arguments from the raw map, calls the corresponding `Engine` interface method, and wraps the result as MCP tool content (text JSON). Return `isError: true` with descriptive messages on failures (FR-013)
|
||||
- [ ] T012 [US1] Implement `MCPServer` in `internal/mcp/server.go`: constructor `NewMCPServer(cfg Config, engine Engine, logger *slog.Logger)` that creates a `mark3labs/mcp-go` server instance via `server.NewMCPServer()`, registers all tools from T010, wires the tool call handler from T011, and stores references to the `ConnectionManager` and `Engine`
|
||||
- [ ] T013 [US1] Implement SSE transport setup in `internal/mcp/transport_sse.go`: function `NewSSEHandler(mcpServer, connMgr, logger) http.Handler` that creates a `mark3labs/mcp-go` SSE transport handler, wraps it to track connections in `ConnectionManager` on connect/disconnect, and updates last-activity on each message (FR-002, FR-010, FR-011)
|
||||
- [ ] T014 [US1] Implement `Mount(router chi.Router)` method on `MCPServer` in `internal/mcp/server.go` that mounts the SSE handler at the configured SSE path (default `/mcp/sse`) and the health endpoint at `/health` on the provided chi router
|
||||
- [ ] T015 [US1] Wire MCP server into `cmd/synapbus/main.go`: in `runServe`, create a chi router, instantiate `MCPServer` with config and a stub `Engine` implementation, call `Mount`, and start `http.Server` with graceful shutdown on SIGTERM/SIGINT (FR-014)
|
||||
- [ ] T016 [US1] Implement graceful shutdown in `internal/mcp/server.go`: `Shutdown(ctx context.Context) error` method that stops accepting new connections, waits for in-flight requests up to the configured timeout, closes all tracked SSE connections, and logs shutdown progress (FR-014)
|
||||
- [ ] T017 [US1] Add `slog` logging to all tool call paths in `internal/mcp/handlers.go`: wrap each tool call with timing, log agent ID, tool name, duration, and success/failure using the helpers from T009 (FR-012)
|
||||
|
||||
**Checkpoint**: At this point, an MCP client can connect via SSE, list all 10 tools with full JSON schemas, and call any tool (dispatched to the Engine interface). This is the MVP.
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: User Story 2 - Agent Authenticates with API Key in MCP Headers (Priority: P1)
|
||||
|
||||
**Goal**: MCP connections are authenticated via `Authorization: Bearer <api-key>` header. Unauthenticated requests are rejected with HTTP 401. All tool calls are scoped to the authenticated agent's identity.
|
||||
|
||||
**Independent Test**: Attempt connections with valid, invalid, and missing API keys. Verify 401 rejection for bad keys. Verify tool calls are scoped to the authenticated agent.
|
||||
|
||||
### Implementation for User Story 2
|
||||
|
||||
- [ ] T018 [US2] Implement auth middleware in `internal/mcp/auth.go`: `AuthMiddleware(engine Engine, logger *slog.Logger) func(http.Handler) http.Handler` that extracts `Authorization: Bearer <key>` from the request header, calls `engine.ValidateAPIKey()`, and either injects the agent ID into the request context or responds with HTTP 401 Unauthorized (FR-004, FR-005)
|
||||
- [ ] T019 [US2] Define context key and helpers in `internal/mcp/auth.go`: `AgentIDFromContext(ctx) (string, bool)` to retrieve the authenticated agent ID from context, used by tool handlers to scope operations
|
||||
- [ ] T020 [US2] Update `HandleToolCall` in `internal/mcp/handlers.go` to extract agent ID from context via `AgentIDFromContext`, populate `ToolCallContext.AgentID`, and pass it to all `Engine` method calls. Reject calls where agent ID is missing with an MCP error result
|
||||
- [ ] T021 [US2] Add authorization enforcement in `internal/mcp/handlers.go`: for `read_inbox`, verify the requested agent matches the authenticated agent; for `send_message`, enforce the sender is the authenticated agent. Return MCP error result with "access denied" for violations (FR-008)
|
||||
- [ ] T022 [US2] Wire auth middleware into `Mount` in `internal/mcp/server.go`: apply `AuthMiddleware` to the SSE route group so that authentication happens before the MCP protocol handshake. Unauthenticated clients never reach the SSE handler
|
||||
|
||||
**Checkpoint**: At this point, User Stories 1 AND 2 are complete. SSE connections require valid API keys, tool calls are identity-scoped, and unauthorized access is rejected.
|
||||
|
||||
---
|
||||
|
||||
## Phase 5: User Story 3 - Agent Connects via Streamable HTTP Transport (Priority: P2)
|
||||
|
||||
**Goal**: Agents can use Streamable HTTP (POST to `/mcp`) as an alternative to SSE. Same tools, same auth, same results.
|
||||
|
||||
**Independent Test**: Send MCP JSON-RPC requests via HTTP POST to `/mcp` with a valid API key. Verify responses match SSE behavior for the same tool calls.
|
||||
|
||||
### Implementation for User Story 3
|
||||
|
||||
- [ ] T023 [US3] Implement Streamable HTTP transport in `internal/mcp/transport_streamhttp.go`: function `NewStreamableHTTPHandler(mcpServer, connMgr, logger) http.Handler` that creates a `mark3labs/mcp-go` Streamable HTTP transport handler, tracks per-request logical connections in `ConnectionManager` (FR-003, FR-010)
|
||||
- [ ] T024 [US3] Update `Mount` in `internal/mcp/server.go` to mount the Streamable HTTP handler at the configured path (default `/mcp`) with the same `AuthMiddleware` applied. Validate API key on each POST request (FR-004)
|
||||
- [ ] T025 [US3] Update `ConnectionManager` in `internal/mcp/connmgr.go` to handle Streamable HTTP logical connections: create connection entry on request start, remove on request completion, set transport type to `streamable-http`
|
||||
|
||||
**Checkpoint**: Both SSE and Streamable HTTP transports are functional with identical tool behavior and authentication. SC-008 is achievable.
|
||||
|
||||
---
|
||||
|
||||
## Phase 6: User Story 4 - Operator Monitors Server Health and Connected Agents (Priority: P2)
|
||||
|
||||
**Goal**: Health endpoint at `/health` reports server status, version, and component health. Operators can query active connections.
|
||||
|
||||
**Independent Test**: Call `/health` and verify JSON response. Connect multiple agents, query connection list, disconnect one, verify it disappears.
|
||||
|
||||
### Implementation for User Story 4
|
||||
|
||||
- [ ] T026 [US4] Implement health check handler in `internal/mcp/health.go`: `NewHealthHandler(engine Engine, connMgr *ConnectionManager, version string, startTime time.Time, logger *slog.Logger) http.HandlerFunc` that calls `engine.HealthCheck()`, adds connection count from `ConnectionManager`, computes uptime, and returns JSON `HealthStatus` with HTTP 200 (ok) or HTTP 503 (degraded/unhealthy) (FR-009)
|
||||
- [ ] T027 [US4] Handle startup and database-unavailable states in `internal/mcp/health.go`: if the engine returns a database error, set component status to unhealthy and overall status to `degraded` or `unhealthy`. During startup (before engine is ready), return HTTP 503 with `{"status": "starting"}`
|
||||
- [ ] T028 [US4] Add `ListConnections` endpoint or MCP tool in `internal/mcp/connmgr.go`: return all active connections with agent ID, transport type, connected-at, and last-activity timestamps. Expose via the health handler as an optional `?connections=true` query parameter (FR-010)
|
||||
- [ ] T029 [US4] Wire health endpoint in `Mount` in `internal/mcp/server.go`: mount the health handler at `/health` WITHOUT auth middleware (unauthenticated per FR-009)
|
||||
|
||||
**Checkpoint**: Operators can monitor server health, see connected agents, and integrate with load balancers / Kubernetes probes.
|
||||
|
||||
---
|
||||
|
||||
## Phase 7: User Story 5 - Agent Discovers Available Tools with Full JSON Schemas (Priority: P3)
|
||||
|
||||
**Goal**: Every MCP tool has comprehensive, human-readable JSON Schema definitions with descriptions on all parameters, correct types, required fields, and constraints.
|
||||
|
||||
**Independent Test**: Connect and call `tools/list`. Validate every tool's `inputSchema` is a valid JSON Schema with descriptions on all properties.
|
||||
|
||||
### Implementation for User Story 5
|
||||
|
||||
- [ ] T030 [US5] Enhance `send_message` tool schema in `internal/mcp/tools.go`: define `to` (string, required, description), `body` (string, required, description), `subject` (string, optional, description), `priority` (integer, optional, minimum 1, maximum 10, description), `channel_id` (string, optional, description), `metadata` (object, optional, description) per acceptance scenario 3
|
||||
- [ ] T031 [US5] Enhance all remaining tool schemas in `internal/mcp/tools.go`: for each of the 10 tools, ensure every input property has a `description` string, correct `type`, and that `required` arrays are accurate. Add enum constraints where applicable (e.g., channel type in `create_channel`)
|
||||
- [ ] T032 [US5] Add tool-level descriptions in `internal/mcp/tools.go`: ensure every tool's top-level `description` is a clear, plain-language sentence explaining what the tool does, suitable for LLM consumption
|
||||
|
||||
**Checkpoint**: All tools have production-quality JSON schemas. SC-002 is achievable.
|
||||
|
||||
---
|
||||
|
||||
## Phase 8: Edge Cases & Robustness
|
||||
|
||||
**Purpose**: Handle the edge cases enumerated in the spec: revoked keys, dropped connections, max connections, malformed requests, non-existent entities, duplicate connections.
|
||||
|
||||
- [ ] T033 [P] Implement connection limit enforcement in `internal/mcp/connmgr.go`: reject new connections with HTTP 503 when `Count()` reaches `Config.MaxConnections`, return descriptive error message
|
||||
- [ ] T034 [P] Implement SSE disconnect detection in `internal/mcp/transport_sse.go`: detect closed client connections (context cancellation, write errors), clean up `ConnectionManager` entry within 5 seconds (SC-006)
|
||||
- [ ] T035 [P] Implement revoked API key handling in `internal/mcp/auth.go`: on each tool call (not just connection), re-validate the API key via `Engine.ValidateAPIKey`. If revoked, return MCP error result and close the SSE connection
|
||||
- [ ] T036 [P] Implement malformed request handling in `internal/mcp/transport_sse.go` and `internal/mcp/transport_streamhttp.go`: ensure `mark3labs/mcp-go` returns standard JSON-RPC errors (`-32700` parse error, `-32600` invalid request) for malformed input rather than crashing
|
||||
- [ ] T037 Implement non-existent entity errors in `internal/mcp/handlers.go`: when `Engine` methods return "not found" errors (e.g., sending to a non-existent agent), wrap them as MCP tool error results (`isError: true`) with clear messages, not protocol-level errors
|
||||
- [ ] T038 [P] Handle concurrent connections from same API key in `internal/mcp/connmgr.go`: support multiple simultaneous connections per agent, track each with a unique connection ID, list all sessions for the same agent
|
||||
|
||||
---
|
||||
|
||||
## Phase 9: Polish & Cross-Cutting Concerns
|
||||
|
||||
**Purpose**: Improvements that affect multiple user stories.
|
||||
|
||||
- [ ] T039 [P] Review and verify all `slog` structured log fields across `internal/mcp/` are consistent: agent_id, tool_name, transport, duration_ms, connection_id, error (SC-007)
|
||||
- [ ] T040 [P] Verify graceful shutdown sequence in `internal/mcp/server.go`: SIGTERM stops new connections, drains in-flight requests within timeout, closes all SSE connections, logs each step (FR-014)
|
||||
- [ ] T041 Add server version injection: pass build version (via `-ldflags`) from `cmd/synapbus/main.go` through to `MCPServer` config and health endpoint response
|
||||
- [ ] T042 Verify both transports produce identical results: manually or via script, run the same tool call sequence against SSE and Streamable HTTP and confirm matching output (SC-008)
|
||||
- [ ] T043 Run `make lint` and fix any linting issues in `internal/mcp/` package
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & Execution Order
|
||||
|
||||
### Phase Dependencies
|
||||
|
||||
- **Setup (Phase 1)**: No dependencies - can start immediately
|
||||
- **Foundational (Phase 2)**: Depends on Phase 1 completion - BLOCKS all user stories
|
||||
- **User Story 1 (Phase 3)**: Depends on Phase 2 - delivers MVP
|
||||
- **User Story 2 (Phase 4)**: Depends on Phase 2, integrates with Phase 3 (adds auth to existing SSE)
|
||||
- **User Story 3 (Phase 5)**: Depends on Phase 2, integrates with Phase 3 and Phase 4 (adds second transport)
|
||||
- **User Story 4 (Phase 6)**: Depends on Phase 2, uses ConnectionManager from Phase 3
|
||||
- **User Story 5 (Phase 7)**: Depends on Phase 3 (tool definitions must exist to enhance)
|
||||
- **Edge Cases (Phase 8)**: Depends on Phases 3-6 being complete
|
||||
- **Polish (Phase 9)**: Depends on all prior phases
|
||||
|
||||
### User Story Dependencies
|
||||
|
||||
- **US1 (P1)**: Can start after Phase 2 - no dependencies on other stories
|
||||
- **US2 (P1)**: Can start after Phase 2 - integrates with US1's SSE transport but is independently testable
|
||||
- **US3 (P2)**: Can start after Phase 2 - reuses tools from US1, auth from US2, but adds an independent transport
|
||||
- **US4 (P2)**: Can start after Phase 2 - uses ConnectionManager but is otherwise independent
|
||||
- **US5 (P3)**: Depends on US1 tool definitions existing - enhances schema quality
|
||||
|
||||
### Recommended Execution Order (Single Developer)
|
||||
|
||||
1. Phase 1 (Setup) + Phase 2 (Foundational)
|
||||
2. Phase 3 (US1 - SSE + Tools) - **MVP checkpoint**
|
||||
3. Phase 4 (US2 - Auth) - secures the MVP
|
||||
4. Phase 5 (US3 - Streamable HTTP) + Phase 6 (US4 - Health) in parallel
|
||||
5. Phase 7 (US5 - Schema polish)
|
||||
6. Phase 8 (Edge cases)
|
||||
7. Phase 9 (Polish)
|
||||
|
||||
### Parallel Opportunities
|
||||
|
||||
- T002 and T003 (Phase 1) can run in parallel
|
||||
- T005, T006, and T009 (Phase 2) can run in parallel
|
||||
- T033, T034, T035, T036, and T038 (Phase 8) can run in parallel
|
||||
- T039 and T040 (Phase 9) can run in parallel
|
||||
- Phase 5 and Phase 6 can run in parallel after Phase 4
|
||||
|
||||
---
|
||||
|
||||
## Notes
|
||||
|
||||
- All file paths are under `internal/mcp/` per the project's directory structure in CLAUDE.md
|
||||
- The `Engine` interface (T007) is the critical abstraction: it decouples MCP tools from the core messaging/channels/agents implementations in other `internal/` packages
|
||||
- A stub `Engine` implementation is needed for US1 to be testable before core messaging is built (T015)
|
||||
- `mark3labs/mcp-go` handles MCP protocol details (JSON-RPC, SSE framing); our code handles auth, connection tracking, tool dispatch, and engine bridging
|
||||
- Zero CGO constraint is satisfied: `mark3labs/mcp-go` and `go-chi/chi` are pure Go (Principle III)
|
||||
@@ -0,0 +1,136 @@
|
||||
# Feature Specification: Trace Logging & Observability
|
||||
|
||||
**Feature Branch**: `007-trace-logging`
|
||||
**Created**: 2026-03-13
|
||||
**Status**: Draft
|
||||
**Input**: User description: "All agent actions logged: tool calls, messages sent/received, channel joins, errors. Traces stored in SQLite with agent_name, action, details, timestamp. Owner can view traces for their agents via Web UI. Filterable by agent, action type, time range. Exportable as JSON/CSV. Optional Prometheus metrics endpoint (/metrics). Structured logging (slog) to stdout."
|
||||
|
||||
## User Scenarios & Testing *(mandatory)*
|
||||
|
||||
### User Story 1 - Owner Inspects Agent Activity Traces (Priority: P1)
|
||||
|
||||
An owner logs into the SynapBus Web UI and navigates to the "Traces" view to see everything their agents have done. They see a reverse-chronological log of all actions (tool calls, messages sent, messages received, channel joins/leaves, errors) across all their agents. They click on an individual trace entry to expand the full JSON details. This is the foundational use case: a human gaining visibility into what their agents are doing.
|
||||
|
||||
**Why this priority**: Without trace storage and a basic viewing interface, none of the other stories (filtering, exporting, metrics) have anything to build on. This directly implements Constitution Principle VIII (Observable by Default) and Principle IV (owners control and view agent activity).
|
||||
|
||||
**Independent Test**: Can be fully tested by registering an agent, performing several MCP tool calls (send_message, join_channel, read_inbox), then logging into the Web UI as the agent's owner and verifying each action appears in the trace view with correct agent name, action type, timestamp, and JSON details.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** an agent "research-bot" owned by user "alice" has sent 3 messages and joined 1 channel, **When** alice opens the Traces view in the Web UI, **Then** she sees 4 trace entries in reverse chronological order, each showing agent_name="research-bot", the action type, and a human-readable timestamp.
|
||||
2. **Given** alice is viewing the Traces list, **When** she clicks on a trace entry for a `send_message` action, **Then** she sees the full JSON details including the recipient, channel (if any), message body preview, and message ID.
|
||||
3. **Given** agent "research-bot" encounters an error (e.g., sending a message to a non-existent agent), **When** alice views the Traces, **Then** she sees a trace entry with action="error", and the details JSON includes the error message and the originating tool call.
|
||||
4. **Given** user "bob" also has agents, **When** bob opens the Traces view, **Then** he sees only traces for his own agents and never sees traces belonging to alice's agents.
|
||||
|
||||
---
|
||||
|
||||
### User Story 2 - Owner Filters and Searches Traces (Priority: P2)
|
||||
|
||||
An owner has accumulated hundreds or thousands of trace entries across multiple agents. They need to narrow down to specific activity: a particular agent, a specific action type (e.g., only errors), or a specific time window. The Web UI provides filter controls for agent name, action type, and time range. Filters can be combined. Results update in real time as filters change.
|
||||
|
||||
**Why this priority**: Trace viewing (P1) becomes unwieldy at scale without filtering. This story makes traces operationally useful for debugging and monitoring. It depends on P1's trace storage being in place.
|
||||
|
||||
**Independent Test**: Can be tested by generating 50+ trace entries across 2 agents with mixed action types, then applying each filter individually and in combination, verifying the result set matches expectations.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** alice has two agents ("research-bot" and "writer-bot") with 100 combined trace entries, **When** she selects agent="research-bot" in the filter, **Then** only traces for "research-bot" are displayed and the count updates accordingly.
|
||||
2. **Given** alice is viewing traces with no filters, **When** she selects action_type="error" from the action type dropdown, **Then** only error traces are shown.
|
||||
3. **Given** alice selects a time range of "last 1 hour", **When** the filter is applied, **Then** only traces with timestamps within the last 60 minutes appear, and older entries are excluded.
|
||||
4. **Given** alice has set agent="research-bot" AND action_type="send_message" AND time range="today", **When** results are displayed, **Then** only traces matching all three criteria are shown.
|
||||
|
||||
---
|
||||
|
||||
### User Story 3 - Owner Exports Traces as JSON or CSV (Priority: P3)
|
||||
|
||||
An owner needs to share agent activity logs with a colleague, feed them into an external analysis tool, or archive them for compliance. They apply filters (or leave them unfiltered) and click an export button. They choose JSON or CSV format. The file downloads to their browser containing all matching trace entries with full details.
|
||||
|
||||
**Why this priority**: Export is a value-add on top of viewing and filtering. It enables external workflows (compliance, analytics, debugging outside the UI) but is not required for core observability.
|
||||
|
||||
**Independent Test**: Can be tested by generating trace entries, optionally applying filters, then exporting as JSON and CSV separately, and verifying both files parse correctly and contain the expected number of entries with all fields.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** alice is viewing unfiltered traces (50 entries), **When** she clicks "Export as JSON", **Then** a file `traces-YYYY-MM-DD.json` downloads containing a JSON array of 50 trace objects, each with fields: id, agent_name, action, details (object), timestamp.
|
||||
2. **Given** alice has filtered traces to agent="research-bot" (20 entries), **When** she clicks "Export as CSV", **Then** a file `traces-YYYY-MM-DD.csv` downloads with 20 data rows plus a header row. Columns: id, agent_name, action, details (JSON-encoded string), timestamp.
|
||||
3. **Given** alice exports traces with no matching results (empty filter), **When** the export completes, **Then** the JSON file contains an empty array `[]` and the CSV file contains only the header row.
|
||||
|
||||
---
|
||||
|
||||
### User Story 4 - Operator Enables Prometheus Metrics (Priority: P3)
|
||||
|
||||
A SynapBus operator running the service on a team server wants to monitor system health using their existing Prometheus + Grafana stack. They start SynapBus with `--metrics` (or set `SYNAPBUS_METRICS=true`). A `/metrics` endpoint becomes available on the HTTP server, exposing counters and histograms for trace actions, message throughput, active agents, and error rates. This endpoint is unauthenticated (standard for Prometheus scrape targets behind a firewall).
|
||||
|
||||
**Why this priority**: Prometheus metrics are explicitly optional and serve operators with existing monitoring infrastructure. Core observability (traces in SQLite + Web UI) works without this.
|
||||
|
||||
**Independent Test**: Can be tested by starting SynapBus with `--metrics`, performing several agent actions, then curling `/metrics` and verifying Prometheus-formatted output includes expected metric names and non-zero counters.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** SynapBus is started with `--metrics`, **When** an operator sends `GET /metrics`, **Then** the response is `text/plain` in Prometheus exposition format and includes at least: `synapbus_traces_total` (counter), `synapbus_traces_by_action` (counter vec with action label), `synapbus_active_agents` (gauge).
|
||||
2. **Given** SynapBus is started without `--metrics`, **When** an operator sends `GET /metrics`, **Then** the server returns 404 Not Found.
|
||||
3. **Given** SynapBus is running with `--metrics` and agent "research-bot" has sent 5 messages, **When** an operator scrapes `/metrics`, **Then** `synapbus_traces_by_action{action="send_message"}` reports a value of at least 5.
|
||||
|
||||
---
|
||||
|
||||
### User Story 5 - Structured Logging to stdout (Priority: P2)
|
||||
|
||||
A SynapBus operator wants structured, machine-parseable logs on stdout for integration with log aggregation systems (Loki, CloudWatch, ELK). All server-side log output uses Go's `slog` package with JSON format. Each log line includes timestamp, level, message, and relevant context fields (agent_name, action, request_id, error). Log level is configurable via `--log-level` flag (debug, info, warn, error).
|
||||
|
||||
**Why this priority**: Structured logging is foundational infrastructure that benefits both development and production. It is required by Constitution Principle VIII and improves debuggability from day one, independent of the Web UI trace viewer.
|
||||
|
||||
**Independent Test**: Can be tested by starting SynapBus with `--log-level=debug`, performing agent actions, and piping stdout through `jq` to verify each line is valid JSON with the expected fields.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** SynapBus is started with `--log-level=info`, **When** an agent sends a message via MCP, **Then** stdout emits a JSON log line with keys: `time`, `level` ("INFO"), `msg`, `agent_name`, `action` ("send_message"), and `request_id`.
|
||||
2. **Given** SynapBus is started with `--log-level=error`, **When** an agent sends a message successfully, **Then** no log line is emitted for that action (since it is info-level). Only errors appear on stdout.
|
||||
3. **Given** SynapBus is started with `--log-level=debug`, **When** any MCP tool call is received, **Then** a debug-level log line is emitted containing the full tool call parameters as a JSON field.
|
||||
|
||||
---
|
||||
|
||||
### Edge Cases
|
||||
|
||||
- What happens when the traces table grows very large (millions of rows)? Trace queries MUST use indexed columns (agent_name, action, timestamp) and paginate results. The system SHOULD support a configurable retention period (`--trace-retention=30d`) that auto-deletes traces older than the threshold.
|
||||
- How does the system handle a burst of concurrent agent actions generating traces? Trace inserts MUST NOT block the MCP tool call response. Traces SHOULD be buffered in-memory and flushed to SQLite in batches to avoid write contention.
|
||||
- What happens if the SQLite write fails during trace insertion (e.g., disk full)? The agent's tool call MUST still succeed. The failure MUST be logged to stderr/slog at error level. The trace is lost but the agent operation is not impacted.
|
||||
- What happens when an owner has zero agents or zero traces? The Web UI MUST display an empty state with a helpful message ("No traces found. Agent activity will appear here once your agents start performing actions.").
|
||||
- What happens if an agent is deleted but its traces remain? Traces MUST be retained even after agent deletion. The agent_name field in trace records is a denormalized string, not a foreign key, so traces survive agent removal. A filter for deleted agents SHOULD still work.
|
||||
- How does export handle very large result sets (100k+ traces)? Export MUST stream the response rather than buffering the entire result in memory. The HTTP response SHOULD use `Transfer-Encoding: chunked` for large exports.
|
||||
- What happens when Prometheus metrics are enabled but no scraper connects? No impact. The `/metrics` endpoint is passive; metrics are maintained in-memory regardless, with negligible overhead.
|
||||
|
||||
## Requirements *(mandatory)*
|
||||
|
||||
### Functional Requirements
|
||||
|
||||
- **FR-001**: System MUST record a trace entry for every MCP tool call received, including: agent_name, action (tool name), details (JSON object with call parameters and result summary), and timestamp (UTC).
|
||||
- **FR-002**: System MUST record trace entries for the following action types at minimum: `send_message`, `read_inbox`, `claim_messages`, `mark_done`, `search_messages`, `create_channel`, `join_channel`, `list_channels`, `register_agent`, `discover_agents`, `post_task`, `bid_task`, `upload_attachment`, `read_attachment`, and `error`.
|
||||
- **FR-003**: Trace entries MUST include an `owner_id` field derived from the agent's owner, enabling owner-scoped queries without joining the agents table.
|
||||
- **FR-004**: System MUST enforce owner isolation: REST API trace endpoints and Web UI MUST only return traces belonging to the authenticated owner's agents. No cross-owner trace access is permitted.
|
||||
- **FR-005**: System MUST provide a REST API endpoint (`GET /api/traces`) that returns paginated traces, accepting query parameters: `agent_name`, `action`, `since` (ISO 8601 timestamp), `until` (ISO 8601 timestamp), `page`, and `page_size` (default 50, max 200).
|
||||
- **FR-006**: System MUST provide a REST API endpoint (`GET /api/traces/export`) that streams traces in the format specified by the `Accept` header or `format` query parameter (`json` or `csv`).
|
||||
- **FR-007**: The Web UI MUST display a Traces view accessible from the main navigation, showing a paginated, reverse-chronological list of trace entries with columns: timestamp, agent name, action, and a details preview.
|
||||
- **FR-008**: The Web UI Traces view MUST provide filter controls for agent name (dropdown of owner's agents), action type (dropdown), and time range (date/time pickers or preset ranges: last hour, last 24h, last 7d, custom).
|
||||
- **FR-009**: System MUST use Go's `slog` package for all server-side logging, with JSON output format on stdout.
|
||||
- **FR-010**: Log level MUST be configurable via `--log-level` CLI flag and `SYNAPBUS_LOG_LEVEL` environment variable, supporting values: `debug`, `info`, `warn`, `error`. Default: `info`.
|
||||
- **FR-011**: When `--metrics` is enabled, the system MUST expose a Prometheus-compatible `/metrics` HTTP endpoint with at least: `synapbus_traces_total` (counter), `synapbus_traces_by_action` (counter vec, label: action), `synapbus_active_agents` (gauge), `synapbus_errors_total` (counter).
|
||||
- **FR-012**: When `--metrics` is not enabled, the `/metrics` endpoint MUST NOT be registered (404 response).
|
||||
- **FR-013**: Trace insertion MUST NOT block or slow down the MCP tool call that triggered it. Traces SHOULD be written asynchronously.
|
||||
- **FR-014**: System SHOULD support a `--trace-retention` flag (e.g., `30d`, `90d`, `0` for unlimited) that triggers periodic cleanup of traces older than the specified duration.
|
||||
|
||||
### Key Entities
|
||||
|
||||
- **Trace**: Represents a single recorded agent action. Attributes: `id` (integer, auto-increment), `owner_id` (string, denormalized from agent), `agent_name` (string, denormalized), `action` (string, e.g. "send_message", "error"), `details` (JSON text, contains tool call parameters, result summary, error info), `timestamp` (UTC datetime). Indexed on: `(owner_id, timestamp)`, `(owner_id, agent_name, timestamp)`, `(owner_id, action, timestamp)`.
|
||||
- **Metric**: In-memory Prometheus metric (counter, gauge, or histogram). Not persisted to SQLite. Registered conditionally when `--metrics` is enabled. Managed via `prometheus/client_golang` or a pure-Go metrics library.
|
||||
|
||||
## Success Criteria *(mandatory)*
|
||||
|
||||
### Measurable Outcomes
|
||||
|
||||
- **SC-001**: Every MCP tool call results in a corresponding trace entry in SQLite within 1 second, with zero tool calls lost under normal operation (no disk-full or crash conditions).
|
||||
- **SC-002**: An owner viewing traces in the Web UI can filter by agent, action type, and time range, and see results update within 500ms for datasets up to 100,000 trace entries.
|
||||
- **SC-003**: Trace export (JSON/CSV) completes and initiates download for 10,000 entries in under 5 seconds.
|
||||
- **SC-004**: Trace insertion adds less than 5ms of latency to MCP tool call response times (measured as p99).
|
||||
- **SC-005**: All server log output on stdout is valid JSON parseable by `jq`, with no unstructured log lines emitted under any code path.
|
||||
- **SC-006**: When `--metrics` is enabled, `/metrics` returns valid Prometheus exposition format that can be scraped by a standard Prometheus server without errors.
|
||||
- **SC-007**: Owner isolation is enforced: no REST API call or Web UI interaction allows an owner to access traces belonging to another owner's agents, verified by automated tests with multi-owner scenarios.
|
||||
@@ -0,0 +1,198 @@
|
||||
# Tasks: Trace Logging & Observability
|
||||
|
||||
**Input**: Design documents from `/specs/007-trace-logging/`
|
||||
**Prerequisites**: spec.md (required)
|
||||
|
||||
**Tests**: Tests are included where specified by the feature specification acceptance scenarios.
|
||||
|
||||
**Organization**: Tasks are grouped by user story to enable independent implementation and testing of each story.
|
||||
|
||||
## Format: `[ID] [P?] [Story] Description`
|
||||
|
||||
- **[P]**: Can run in parallel (different files, no dependencies)
|
||||
- **[Story]**: Which user story this task belongs to (e.g., US1, US2, US3)
|
||||
- Include exact file paths in descriptions
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Setup (Shared Infrastructure)
|
||||
|
||||
**Purpose**: Project initialization, dependencies, and schema changes for trace logging
|
||||
|
||||
- [ ] T001 Add `--log-level` flag and `SYNAPBUS_LOG_LEVEL` env var to `cmd/synapbus/main.go`. Accept values: `debug`, `info`, `warn`, `error`. Default: `info`.
|
||||
- [ ] T002 Add `--metrics` flag and `SYNAPBUS_METRICS` env var to `cmd/synapbus/main.go`. Boolean, default false.
|
||||
- [ ] T003 Add `--trace-retention` flag and `SYNAPBUS_TRACE_RETENTION` env var to `cmd/synapbus/main.go`. Accept duration strings like `30d`, `90d`, `0` (unlimited). Default: `0`.
|
||||
- [ ] T004 [P] Add `prometheus/client_golang` dependency to `go.mod` for optional metrics support.
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Foundational (Blocking Prerequisites)
|
||||
|
||||
**Purpose**: Core infrastructure that MUST be complete before ANY user story can be implemented
|
||||
|
||||
**CRITICAL**: No user story work can begin until this phase is complete
|
||||
|
||||
- [ ] T005 Create SQLite migration `schema/002_trace_logging.sql`: add `owner_id TEXT NOT NULL DEFAULT ''` column to existing `traces` table, add composite indexes `(owner_id, timestamp)`, `(owner_id, agent_name, timestamp)`, `(owner_id, action, timestamp)`. Update `schema_migrations`.
|
||||
- [ ] T006 Configure slog JSON handler as the global logger in `cmd/synapbus/main.go`. Initialize `slog.New(slog.NewJSONHandler(os.Stdout, &slog.HandlerOptions{Level: parsedLevel}))` and set as default. Wire log level from the `--log-level` flag.
|
||||
- [ ] T007 [P] Define the `Trace` domain struct in `internal/trace/model.go` with fields: `ID int64`, `OwnerID string`, `AgentName string`, `Action string`, `Details json.RawMessage`, `Timestamp time.Time`.
|
||||
- [ ] T008 [P] Define the `TraceFilter` struct in `internal/trace/model.go` with fields: `OwnerID string`, `AgentName string`, `Action string`, `Since *time.Time`, `Until *time.Time`, `Page int`, `PageSize int`.
|
||||
- [ ] T009 Define the `TraceStore` interface in `internal/trace/store.go` with methods: `Insert(ctx context.Context, t *Trace) error`, `Query(ctx context.Context, f TraceFilter) ([]Trace, int, error)` (returns traces + total count), `DeleteOlderThan(ctx context.Context, before time.Time) (int64, error)`.
|
||||
- [ ] T010 Implement `SQLiteTraceStore` in `internal/trace/sqlite_store.go` satisfying the `TraceStore` interface. Use `modernc.org/sqlite` via the existing storage layer. Insert must be fast (single row insert). Query must use indexed columns, enforce `PageSize` max 200, default 50. `DeleteOlderThan` for retention cleanup.
|
||||
- [ ] T011 Implement the async `Tracer` service in `internal/trace/tracer.go`. Accepts trace entries via a buffered channel, batches writes to SQLite in a background goroutine (flush every 100ms or when buffer reaches 64 entries). Expose `Record(ctx context.Context, ownerID, agentName, action string, details any)` method that serializes details to JSON and enqueues without blocking. On SQLite write failure, log error via slog but do not propagate to caller. Provide `Close()` for graceful shutdown (flush remaining).
|
||||
- [ ] T012 Write table-driven tests for `SQLiteTraceStore` in `internal/trace/sqlite_store_test.go`: test insert, query with each filter combination, pagination, `DeleteOlderThan`, owner isolation (query for owner A must not return owner B traces).
|
||||
- [ ] T013 Write tests for `Tracer` in `internal/trace/tracer_test.go`: test async recording (Record returns immediately), batch flush behavior, graceful shutdown flushes pending traces.
|
||||
|
||||
**Checkpoint**: Foundation ready — trace storage, async tracer, slog configured, schema migrated. User story implementation can now begin in parallel.
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: User Story 1 — Owner Inspects Agent Activity Traces (Priority: P1) MVP
|
||||
|
||||
**Goal**: Owners can view a reverse-chronological list of all their agents' traced actions via REST API and Web UI, with expandable JSON details.
|
||||
|
||||
**Independent Test**: Register an agent, perform several MCP tool calls, then query `GET /api/traces` as the owner and verify each action appears with correct agent_name, action, timestamp, and details.
|
||||
|
||||
### Implementation for User Story 1
|
||||
|
||||
- [ ] T014 [US1] Instrument MCP tool handlers to call `Tracer.Record()` after every tool call in `internal/mcp/`. For each MCP tool (send_message, read_inbox, claim_messages, mark_done, search_messages, create_channel, join_channel, list_channels, register_agent, discover_agents, post_task, bid_task, upload_attachment, read_attachment), add a `tracer.Record(ctx, ownerID, agentName, toolName, detailsMap)` call capturing input params and result summary. On tool error, record a separate trace with action="error" including the error message and originating tool name.
|
||||
- [ ] T015 [US1] Implement `GET /api/traces` handler in `internal/api/traces_handler.go`. Accept query params: `agent_name`, `action`, `since`, `until`, `page`, `page_size`. Extract `owner_id` from authenticated session. Call `TraceStore.Query()` with owner-scoped filter. Return JSON response: `{ "traces": [...], "total": N, "page": N, "page_size": N }`. Return 200 with empty array if no results.
|
||||
- [ ] T016 [US1] Register the `/api/traces` route in `internal/api/router.go` (or equivalent). Apply authentication middleware so only authenticated owners can access. Wire the `TraceStore` dependency.
|
||||
- [ ] T017 [P] [US1] Create the Svelte Traces list page in `web/src/routes/traces/+page.svelte`. Display a paginated, reverse-chronological table with columns: timestamp (human-readable), agent name, action, details preview (truncated to 80 chars). Clicking a row expands an inline panel showing the full JSON details formatted with syntax highlighting or `<pre>` block.
|
||||
- [ ] T018 [P] [US1] Add "Traces" navigation link to the Web UI sidebar/nav in `web/src/lib/components/Nav.svelte` (or equivalent layout component).
|
||||
- [ ] T019 [US1] Handle empty state in the Traces view: when zero traces exist, show a helpful message: "No traces found. Agent activity will appear here once your agents start performing actions."
|
||||
- [ ] T020 [US1] Write integration test in `internal/api/traces_handler_test.go`: create two owners with agents, generate traces for both, verify `GET /api/traces` returns only the authenticated owner's traces (owner isolation). Verify response structure matches expected JSON format.
|
||||
|
||||
**Checkpoint**: User Story 1 complete. An owner can log into the Web UI, navigate to Traces, and see all their agents' actions in reverse-chronological order with expandable details. Owner isolation is enforced.
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: User Story 5 — Structured Logging to stdout (Priority: P2)
|
||||
|
||||
**Goal**: All server-side log output uses slog JSON format with structured fields. Log level is configurable.
|
||||
|
||||
**Independent Test**: Start SynapBus with `--log-level=debug`, perform agent actions, pipe stdout through `jq` to verify every line is valid JSON with expected fields.
|
||||
|
||||
### Implementation for User Story 5
|
||||
|
||||
- [ ] T021 [US5] Create a slog middleware for chi in `internal/api/middleware_logging.go`. Log every HTTP request with fields: `method`, `path`, `status`, `duration_ms`, `request_id`. Generate `request_id` (UUID) per request and store in context.
|
||||
- [ ] T022 [US5] Add structured slog calls to MCP tool handlers in `internal/mcp/`. Each tool call logs at `info` level with fields: `agent_name`, `action` (tool name), `request_id`. At `debug` level, include full tool call parameters as a JSON field. Errors log at `error` level with the error message.
|
||||
- [ ] T023 [US5] Audit all existing `fmt.Printf` / `fmt.Println` calls in `cmd/synapbus/main.go` and any other files. Replace with `slog.Info()`, `slog.Debug()`, or `slog.Error()` calls with appropriate structured fields. Ensure zero unstructured log lines are emitted.
|
||||
- [ ] T024 [US5] Write test in `internal/api/middleware_logging_test.go`: capture stdout, make HTTP requests at various log levels, parse each line as JSON, verify required fields are present. Verify `--log-level=error` suppresses info-level output.
|
||||
|
||||
**Checkpoint**: User Story 5 complete. All stdout output is valid JSON parseable by `jq`. Log level is configurable. No unstructured log lines.
|
||||
|
||||
---
|
||||
|
||||
## Phase 5: User Story 2 — Owner Filters and Searches Traces (Priority: P2)
|
||||
|
||||
**Goal**: Owners can filter traces by agent name, action type, and time range. Filters combine with AND logic. Results update as filters change.
|
||||
|
||||
**Independent Test**: Generate 50+ traces across 2 agents with mixed action types. Apply each filter individually and in combination. Verify result sets match expectations.
|
||||
|
||||
### Implementation for User Story 2
|
||||
|
||||
- [ ] T025 [US2] Add filter controls to the Traces Svelte page in `web/src/routes/traces/+page.svelte`: agent name dropdown (populated from owner's agents via `GET /api/agents`), action type dropdown (hardcoded list of known action types), time range selector (presets: "last hour", "last 24h", "last 7d", "custom" with date/time pickers). Filters update query params and re-fetch traces on change.
|
||||
- [ ] T026 [US2] Implement `GET /api/agents` handler (if not already present) in `internal/api/agents_handler.go` to return the authenticated owner's agents. Used by the filter dropdown.
|
||||
- [ ] T027 [US2] Write integration test in `internal/api/traces_handler_test.go`: insert 50+ traces across 2 agents with mixed actions and timestamps. Test each filter individually (agent_name, action, since/until) and combined filters. Verify correct counts and that no unmatched traces leak through.
|
||||
|
||||
**Checkpoint**: User Story 2 complete. Owners can narrow traces by agent, action type, and time range. All filters combine correctly.
|
||||
|
||||
---
|
||||
|
||||
## Phase 6: User Story 3 — Owner Exports Traces as JSON or CSV (Priority: P3)
|
||||
|
||||
**Goal**: Owners can export filtered or unfiltered traces as a JSON or CSV file download. Export streams results to avoid memory exhaustion on large datasets.
|
||||
|
||||
**Independent Test**: Generate traces, apply filters, export as JSON and CSV. Verify files parse correctly and contain expected entries.
|
||||
|
||||
### Implementation for User Story 3
|
||||
|
||||
- [ ] T028 [US3] Implement `GET /api/traces/export` handler in `internal/api/traces_export_handler.go`. Accept same filter params as `GET /api/traces` plus `format` query param (`json` or `csv`; also respect `Accept` header). Stream results using `Transfer-Encoding: chunked`. For JSON: open with `[`, stream each trace object comma-separated, close with `]`. For CSV: write header row (`id,agent_name,action,details,timestamp`), then stream each row with `details` as a JSON-encoded string. Set `Content-Disposition: attachment; filename="traces-YYYY-MM-DD.{json|csv}"`.
|
||||
- [ ] T029 [US3] Add a streaming query method `QueryStream(ctx context.Context, f TraceFilter, fn func(Trace) error) error` to `TraceStore` interface and `SQLiteTraceStore` in `internal/trace/store.go` and `internal/trace/sqlite_store.go`. Iterates rows without loading all into memory. Calls `fn` for each row.
|
||||
- [ ] T030 [US3] Register `/api/traces/export` route in `internal/api/router.go` with authentication middleware.
|
||||
- [ ] T031 [P] [US3] Add export buttons ("Export JSON", "Export CSV") to the Traces Svelte page in `web/src/routes/traces/+page.svelte`. Buttons construct the export URL with current filter params and trigger browser download.
|
||||
- [ ] T032 [US3] Write integration test in `internal/api/traces_export_handler_test.go`: export as JSON, parse the response body as `[]Trace`, verify count. Export as CSV, parse rows, verify header and data row count. Test with filters applied. Test empty result (empty array / header-only CSV).
|
||||
|
||||
**Checkpoint**: User Story 3 complete. Owners can export traces as JSON or CSV with current filters applied. Large exports stream without memory issues.
|
||||
|
||||
---
|
||||
|
||||
## Phase 7: User Story 4 — Operator Enables Prometheus Metrics (Priority: P3)
|
||||
|
||||
**Goal**: When `--metrics` is enabled, a `/metrics` endpoint exposes Prometheus-formatted counters and gauges. When disabled, `/metrics` returns 404.
|
||||
|
||||
**Independent Test**: Start with `--metrics`, perform agent actions, curl `/metrics`, verify Prometheus-formatted output with expected metric names.
|
||||
|
||||
### Implementation for User Story 4
|
||||
|
||||
- [ ] T033 [P] [US4] Create `internal/trace/metrics.go`. Define Prometheus metrics using `prometheus/client_golang`: `synapbus_traces_total` (counter), `synapbus_traces_by_action` (counter vec, label: `action`), `synapbus_active_agents` (gauge), `synapbus_errors_total` (counter). Provide a `Metrics` struct with methods `IncTrace(action string)`, `IncError()`, `SetActiveAgents(n int)`, and a no-op `NullMetrics` implementation for when metrics are disabled.
|
||||
- [ ] T034 [US4] Conditionally register `/metrics` route in `internal/api/router.go`. When `--metrics` is enabled, register `promhttp.Handler()` at `/metrics`. When disabled, do not register the route (chi will 404 by default).
|
||||
- [ ] T035 [US4] Wire metrics into the `Tracer` in `internal/trace/tracer.go`. After each successful trace batch write, call `metrics.IncTrace(action)` for each trace in the batch. On error traces, also call `metrics.IncError()`. Periodically update `metrics.SetActiveAgents()` by querying distinct active agent count.
|
||||
- [ ] T036 [US4] Write test in `internal/trace/metrics_test.go`: verify counter increments, verify `NullMetrics` does not panic. Write integration test: start server with `--metrics`, perform actions, scrape `/metrics`, verify output contains expected metric names with correct values. Verify server without `--metrics` returns 404 on `/metrics`.
|
||||
|
||||
**Checkpoint**: User Story 4 complete. Operators with existing Prometheus+Grafana stacks can scrape SynapBus metrics. The endpoint is passive when no scraper connects.
|
||||
|
||||
---
|
||||
|
||||
## Phase 8: Polish & Cross-Cutting Concerns
|
||||
|
||||
**Purpose**: Improvements that span multiple user stories
|
||||
|
||||
- [ ] T037 [P] Implement trace retention cleanup in `internal/trace/retention.go`. Start a background goroutine that runs every hour (configurable). If `--trace-retention` is set to a non-zero duration, call `TraceStore.DeleteOlderThan()` with the computed cutoff time. Log deletions at info level via slog.
|
||||
- [ ] T038 [P] Wire retention cleanup into the server startup in `cmd/synapbus/main.go`. Parse the `--trace-retention` flag, initialize the retention goroutine if duration > 0, ensure graceful shutdown cancels it.
|
||||
- [ ] T039 Add context-propagated `request_id` to trace details in `internal/trace/tracer.go`. Extract `request_id` from context (set by logging middleware) and include it in the trace `details` JSON for cross-referencing logs and traces.
|
||||
- [ ] T040 [P] Verify owner isolation end-to-end: write a multi-owner integration test in `internal/api/traces_handler_test.go` that creates 3 owners, generates traces for each, and verifies no API call or export leaks traces across owners. Covers SC-007.
|
||||
- [ ] T041 Performance validation: write a benchmark test in `internal/trace/sqlite_store_test.go` using `testing.B`. Verify trace insertion adds < 5ms p99 latency (SC-004). Verify query on 100k traces with filters returns within 500ms (SC-002).
|
||||
- [ ] T042 [P] Run `make lint` and fix any linting issues across all new files.
|
||||
- [ ] T043 Graceful shutdown: ensure `Tracer.Close()` is called on server shutdown in `cmd/synapbus/main.go` so all buffered traces are flushed to SQLite before exit.
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & Execution Order
|
||||
|
||||
### Phase Dependencies
|
||||
|
||||
- **Setup (Phase 1)**: No dependencies — can start immediately
|
||||
- **Foundational (Phase 2)**: Depends on Setup completion — BLOCKS all user stories
|
||||
- **User Stories (Phases 3–7)**: All depend on Foundational phase completion
|
||||
- US1 (Phase 3, P1) should be completed first as MVP
|
||||
- US5 (Phase 4, P2) and US2 (Phase 5, P2) can proceed in parallel after US1, or sequentially
|
||||
- US3 (Phase 6, P3) depends on US2 filter infrastructure in the API (already built in Phase 2 foundation)
|
||||
- US4 (Phase 7, P3) is fully independent of other user stories
|
||||
- **Polish (Phase 8)**: Depends on all user stories being complete
|
||||
|
||||
### User Story Dependencies
|
||||
|
||||
- **User Story 1 (P1)**: Can start after Foundational (Phase 2) — No dependencies on other stories. This is the MVP.
|
||||
- **User Story 5 (P2)**: Can start after Foundational (Phase 2) — Independent. Enhances logging infrastructure.
|
||||
- **User Story 2 (P2)**: Can start after Foundational (Phase 2) — Independent from US5. Uses same API/store built in foundation. Builds on US1's Svelte page.
|
||||
- **User Story 3 (P3)**: Can start after Foundational (Phase 2) — Adds export to the API. Adds streaming query to store. UI builds on US2's filter controls.
|
||||
- **User Story 4 (P3)**: Can start after Foundational (Phase 2) — Fully independent. Only needs the `Tracer` from foundation.
|
||||
|
||||
### Within Each User Story
|
||||
|
||||
- Models/interfaces before service implementations
|
||||
- Services before API handlers
|
||||
- API handlers before Svelte UI components
|
||||
- Core implementation before integration tests
|
||||
- Story complete before moving to next priority
|
||||
|
||||
### Parallel Opportunities
|
||||
|
||||
- Phase 1: T001/T002/T003 touch the same file (`main.go`) — do sequentially. T004 is independent [P].
|
||||
- Phase 2: T007/T008 are parallel [P] (both in `model.go` but logically grouped). T009 depends on T007/T008. T010 depends on T009. T011 depends on T010. T012/T013 are parallel [P] after their respective implementations.
|
||||
- Phase 3: T017 and T018 are parallel [P] (different Svelte files). T014–T016 are sequential (instrument → handler → route).
|
||||
- Phase 6: T031 is parallel [P] with T028–T030 (Svelte vs Go).
|
||||
- Phase 7: T033 is parallel [P] with other Go work (new file, no dependencies).
|
||||
- Phase 8: T037, T038, T040, T041, T042 are marked [P] where applicable.
|
||||
|
||||
---
|
||||
|
||||
## Notes
|
||||
|
||||
- All storage uses `modernc.org/sqlite` (pure Go, zero CGO per Constitution Principle III)
|
||||
- Trace insertion is async via buffered channel — must not block MCP tool calls (FR-013, SC-004)
|
||||
- Owner isolation is enforced at every layer: store queries always include `owner_id`, API handlers extract owner from session, tests verify no cross-owner leakage (FR-004, SC-007)
|
||||
- The existing `traces` table in `schema/001_initial.sql` lacks `owner_id` — migration `002_trace_logging.sql` adds it
|
||||
- Prometheus metrics use `prometheus/client_golang` with a `NullMetrics` no-op for when `--metrics` is disabled
|
||||
- Structured logging via `slog` replaces all `fmt.Print*` calls — every stdout line must be valid JSON (SC-005)
|
||||
@@ -0,0 +1,144 @@
|
||||
# Feature Specification: Semantic Search
|
||||
|
||||
**Feature Branch**: `008-semantic-search`
|
||||
**Created**: 2026-03-13
|
||||
**Status**: Draft
|
||||
**Input**: User description: "Message embedding on ingest, configurable providers (OpenAI/Gemini/Ollama), HNSW vector index, combined search with filters, MCP tool search_messages, incremental background indexing, full-text fallback"
|
||||
|
||||
## User Scenarios & Testing *(mandatory)*
|
||||
|
||||
### User Story 1 - Agent Searches Messages by Meaning (Priority: P1)
|
||||
|
||||
An AI agent working on a code review task needs to find previous discussions about "deployment failures in the staging environment." The agent calls the `search_messages` MCP tool with a natural language query. SynapBus embeds the query using the configured provider, performs ANN search against the HNSW index, and returns the top-K most semantically relevant messages ranked by similarity score. The agent receives messages that discuss staging deployment issues even if they never used the exact phrase "deployment failures."
|
||||
|
||||
**Why this priority**: Semantic search is the core differentiator of this feature. Without it, agents are limited to exact-match or keyword search, which fails when different terminology is used for the same concept. This is the fundamental value proposition.
|
||||
|
||||
**Independent Test**: Can be fully tested by configuring an embedding provider, sending several messages with varied vocabulary about related topics, then issuing a `search_messages` call with a query that uses different wording. Delivers value by returning semantically relevant results that keyword search would miss.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** an embedding provider is configured and 50 messages exist across multiple channels, **When** an agent calls `search_messages` with `query: "database connection timeouts"`, **Then** the system returns messages discussing DB connection issues, pool exhaustion, and query latency — ranked by cosine similarity — even if none contain the exact phrase "database connection timeouts."
|
||||
2. **Given** an agent has access to channels A and B but not channel C, **When** the agent calls `search_messages` with a query, **Then** results MUST only include messages from channels A and B, never from channel C, regardless of similarity score.
|
||||
3. **Given** a `search_messages` call with `query: "memory leak"` and `limit: 5`, **When** there are 20 semantically relevant messages, **Then** the system returns exactly 5 results ordered by descending similarity score.
|
||||
|
||||
---
|
||||
|
||||
### User Story 2 - Combined Search with Metadata Filters (Priority: P2)
|
||||
|
||||
An agent needs to find messages about a specific topic but only within a certain channel, from a specific sender, or within a time range. The agent calls `search_messages` with both a semantic query and structured filters (channel_id, sender_id, priority range, tags, date range). SynapBus first applies the metadata filters via SQLite, then ranks the filtered results by vector similarity, returning a precise intersection of structural and semantic relevance.
|
||||
|
||||
**Why this priority**: Pure semantic search is often too broad. Agents operating in multi-channel environments need to scope searches to specific contexts. This builds on P1 by adding the filtering layer that makes semantic search practically useful in real workflows.
|
||||
|
||||
**Independent Test**: Can be tested by sending messages about the same topic across multiple channels and from different agents, then issuing filtered searches and verifying that results respect all filter constraints while still being ranked by semantic relevance.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** messages about "performance optimization" exist in channels `#backend` and `#frontend`, **When** an agent calls `search_messages` with `query: "performance optimization"` and `filters: { channel_id: "#backend" }`, **Then** only messages from `#backend` are returned, ranked by similarity.
|
||||
2. **Given** messages from agent-A and agent-B both discuss "API rate limiting," **When** an agent calls `search_messages` with `query: "rate limiting"` and `filters: { sender_id: "agent-A", after: "2026-03-01" }`, **Then** only agent-A's messages sent after March 1 are returned.
|
||||
3. **Given** a search with `query: "error handling"` and `filters: { priority_min: 7 }`, **When** matching messages exist at priorities 3, 5, 8, and 10, **Then** only messages with priority 8 and 10 are returned, ranked by similarity.
|
||||
4. **Given** a search with `query: "deployment"` and `filters: { tags: ["#finding", "#trace"] }`, **When** matching messages exist with various tags, **Then** only messages tagged with at least one of the specified tags are returned.
|
||||
|
||||
---
|
||||
|
||||
### User Story 3 - Graceful Fallback to Full-Text Search (Priority: P2)
|
||||
|
||||
A SynapBus operator runs the system without configuring any embedding provider (fully local, no API keys). When an agent calls `search_messages`, the system transparently falls back to SQLite FTS5 full-text search. The agent receives keyword-matched results without errors or degraded API contracts. The response includes a field indicating the search mode used (`semantic` vs `fulltext`).
|
||||
|
||||
**Why this priority**: Per Constitution Principle IX (Progressive Complexity), SynapBus MUST function fully without an embedding provider. This ensures the search MCP tool is always available regardless of deployment configuration, preventing agents from encountering broken tools.
|
||||
|
||||
**Independent Test**: Can be tested by starting SynapBus with no embedding provider configured, sending messages, and calling `search_messages`. Verify that results are returned via full-text matching and the response indicates `search_mode: "fulltext"`.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** no embedding provider is configured, **When** an agent calls `search_messages` with `query: "deployment"`, **Then** the system returns messages containing the word "deployment" (or stemmed variants) via FTS5, and the response includes `search_mode: "fulltext"`.
|
||||
2. **Given** an embedding provider was configured but becomes unreachable (API key revoked, network failure), **When** an agent calls `search_messages`, **Then** the system falls back to full-text search, returns results, and includes `search_mode: "fulltext"` with a `warning: "embedding provider unavailable, using full-text fallback"`.
|
||||
3. **Given** the system is running in full-text mode, **When** an operator configures an embedding provider and restarts, **Then** existing messages are backfilled with embeddings in the background, and subsequent searches use semantic mode.
|
||||
|
||||
---
|
||||
|
||||
### User Story 4 - Background Embedding on Message Ingest (Priority: P1)
|
||||
|
||||
When a new message is sent via `send_message`, SynapBus asynchronously generates an embedding for the message body and indexes it in the HNSW vector index. The `send_message` call returns immediately without waiting for embedding generation. The embedding pipeline processes messages in the background with configurable concurrency, and the HNSW index is updated incrementally.
|
||||
|
||||
**Why this priority**: This is the data pipeline that powers P1 (semantic search). Without background embedding, there is nothing to search against. It is equally critical as search itself because it determines data freshness and system responsiveness.
|
||||
|
||||
**Independent Test**: Can be tested by sending a message and immediately checking that the `send_message` response is fast (< 100ms overhead), then polling the embedding status endpoint or waiting briefly and confirming the message appears in semantic search results.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** an embedding provider is configured, **When** an agent sends a message via `send_message`, **Then** the message is persisted and returned to the sender within normal latency (no embedding delay), and within 5 seconds the message becomes searchable via semantic search.
|
||||
2. **Given** 100 messages are sent in rapid succession, **When** the embedding pipeline is processing, **Then** messages are embedded in FIFO order with configurable concurrency (default: 4 workers), and no messages are dropped or lost.
|
||||
3. **Given** the embedding provider returns a transient error (rate limit, timeout), **When** embedding fails for a message, **Then** the system retries with exponential backoff (max 3 retries) and logs the failure. The message remains searchable via full-text search in the meantime.
|
||||
|
||||
---
|
||||
|
||||
### User Story 5 - Configurable Embedding Providers (Priority: P3)
|
||||
|
||||
An operator chooses their embedding provider based on their deployment constraints: OpenAI `text-embedding-3-small` for cloud deployments with API access, Google Gemini embedding for GCP-adjacent setups, or Ollama for fully air-gapped local deployments. The provider is configured via SynapBus config file or CLI flags. Switching providers triggers a backfill of existing message embeddings using the new provider.
|
||||
|
||||
**Why this priority**: The system can ship with a single provider initially. Multiple providers add deployment flexibility but are not required for core functionality. OpenAI support alone covers the majority of use cases.
|
||||
|
||||
**Independent Test**: Can be tested by configuring each provider independently, sending messages, and verifying that embeddings are generated and searchable. Provider switching can be tested by changing config and verifying backfill completes.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** config `embedding.provider: "openai"` and `embedding.api_key: "sk-..."`, **When** a message is sent, **Then** the system calls OpenAI's `text-embedding-3-small` endpoint and stores the resulting 1536-dimensional vector.
|
||||
2. **Given** config `embedding.provider: "ollama"` and `embedding.endpoint: "http://localhost:11434"`, **When** a message is sent, **Then** the system calls the local Ollama API with the configured model and stores the embedding vector.
|
||||
3. **Given** an operator switches from `openai` to `ollama` and restarts, **When** the system starts, **Then** it detects the provider change, marks all existing embeddings as stale, and re-embeds messages in the background using the new provider.
|
||||
|
||||
---
|
||||
|
||||
### Edge Cases
|
||||
|
||||
- What happens when a message body is empty or contains only whitespace? The system MUST skip embedding generation and exclude the message from vector search (it remains findable via metadata filters only).
|
||||
- What happens when a message body exceeds the embedding provider's token limit (e.g., 8191 tokens for OpenAI)? The system MUST truncate the text to the provider's limit before embedding, log a warning with the message ID, and store the truncated embedding.
|
||||
- What happens when the HNSW index file is corrupted or missing on startup? The system MUST detect the corruption, rebuild the index from stored embeddings in SQLite, and log the rebuild event.
|
||||
- What happens when two messages have identical bodies? Both MUST receive their own embedding entries and be independently searchable (no deduplication of embeddings).
|
||||
- What happens when a message is deleted? The corresponding embedding MUST be removed from the HNSW index and the SQLite embedding record.
|
||||
- What happens during a provider switch when the old and new providers produce different vector dimensions? The system MUST clear the entire HNSW index, rebuild it with the new dimension, and re-embed all messages.
|
||||
- What happens when the system has thousands of unembedded messages at startup (backfill scenario)? The backfill MUST be rate-limited to avoid overwhelming the embedding provider, with configurable batch size and delay between batches.
|
||||
- How does the system handle concurrent reads from the HNSW index while it is being updated? The index MUST support concurrent read access during incremental writes without returning corrupted results or panicking.
|
||||
|
||||
## Requirements *(mandatory)*
|
||||
|
||||
### Functional Requirements
|
||||
|
||||
- **FR-001**: System MUST embed message bodies asynchronously upon ingest when an embedding provider is configured, without blocking the `send_message` response.
|
||||
- **FR-002**: System MUST support three embedding providers: OpenAI (`text-embedding-3-small`), Google Gemini embedding, and Ollama (local). Provider selection MUST be configurable via the SynapBus configuration file.
|
||||
- **FR-003**: System MUST maintain an HNSW vector index (via `TFMV/hnsw`, pure Go, zero CGO) stored within the `--data` directory alongside SQLite data.
|
||||
- **FR-004**: System MUST expose an MCP tool `search_messages` accepting: `query` (string, required), `filters` (object, optional: `channel_id`, `sender_id`, `priority_min`, `priority_max`, `tags`, `after`, `before`), `limit` (integer, optional, default 10, max 100), and `search_mode` (string, optional: `"auto"`, `"semantic"`, `"fulltext"`, default `"auto"`).
|
||||
- **FR-005**: The `search_messages` tool MUST enforce agent access control: results MUST only include messages from channels the calling agent has joined or direct messages addressed to/from the calling agent.
|
||||
- **FR-006**: When `search_mode` is `"auto"`, the system MUST use semantic search if an embedding provider is configured and healthy, otherwise fall back to FTS5 full-text search.
|
||||
- **FR-007**: System MUST fall back to SQLite FTS5 full-text search when no embedding provider is configured. The FTS5 index MUST be maintained on the `messages.body` column.
|
||||
- **FR-008**: Search responses MUST include: `results` (array of message objects with `similarity_score` for semantic or `relevance_score` for full-text), `search_mode` (string: `"semantic"` or `"fulltext"`), `total_results` (integer), and optionally `warning` (string, for degraded mode).
|
||||
- **FR-009**: System MUST persist embeddings in SQLite (message_id, provider, model, vector BLOB, created_at) so the HNSW index can be rebuilt from stored data if corrupted or after a provider switch.
|
||||
- **FR-010**: System MUST process embedding backlog in FIFO order with configurable concurrency (`embedding.workers`, default 4) and retry failed embeddings with exponential backoff (max 3 retries, base delay 1s).
|
||||
- **FR-011**: When the embedding provider changes, the system MUST invalidate all existing embeddings and re-embed messages in the background using the new provider.
|
||||
- **FR-012**: The HNSW index MUST support concurrent read access during incremental write operations without data corruption.
|
||||
- **FR-013**: System MUST truncate message bodies exceeding the provider's token limit before embedding, and log a warning with the affected message ID.
|
||||
- **FR-014**: System MUST remove embeddings and HNSW index entries when the corresponding message is deleted.
|
||||
- **FR-015**: The `search_messages` MCP tool MUST include a JSON Schema description compliant with Constitution Principle II (MCP-Native).
|
||||
|
||||
### Key Entities *(include if feature involves data)*
|
||||
|
||||
- **Embedding**: Represents a vector embedding for a single message. Key attributes: `id`, `message_id` (FK to messages), `provider` (string: "openai", "gemini", "ollama"), `model` (string: provider-specific model identifier), `vector` (BLOB: serialized float32 array), `dimensions` (integer), `created_at` (timestamp). One-to-one relationship with Message. Stored in SQLite for durability; loaded into HNSW index for search.
|
||||
|
||||
- **EmbeddingProvider**: A configured embedding service. Key attributes: `type` (enum: openai, gemini, ollama), `api_key` (string, optional for ollama), `endpoint` (string, custom URL or default), `model` (string, provider-specific model name), `dimensions` (integer, determined by model). Configured via SynapBus config file, not persisted in database.
|
||||
|
||||
- **EmbeddingQueue**: Tracks messages pending embedding. Key attributes: `message_id` (FK to messages), `status` (enum: pending, processing, completed, failed), `attempts` (integer), `last_error` (string, nullable), `created_at` (timestamp), `completed_at` (timestamp, nullable). Stored in SQLite. Drives the background embedding pipeline.
|
||||
|
||||
- **HNSWIndex**: In-memory ANN index backed by `TFMV/hnsw`. Key attributes: configured `dimensions`, `ef_construction` (index build quality, default 200), `M` (max connections per node, default 16), `ef_search` (query-time quality, default 100). Persisted to disk in the `--data` directory. Rebuilt from SQLite embeddings table on startup or corruption.
|
||||
|
||||
- **SearchResult**: Returned by `search_messages`. Key attributes: `message` (full message object), `similarity_score` (float64, 0.0-1.0 for semantic) or `relevance_score` (float64 for full-text), `search_mode` (string). Not persisted.
|
||||
|
||||
## Success Criteria *(mandatory)*
|
||||
|
||||
### Measurable Outcomes
|
||||
|
||||
- **SC-001**: Semantic search returns relevant results for queries using different terminology than the indexed messages, with the top-3 results containing at least one genuinely relevant message in 90%+ of test cases (measured against a curated test set of 50 query-message pairs).
|
||||
- **SC-002**: The `send_message` MCP tool adds no more than 10ms of latency due to embedding pipeline overhead (embedding itself happens asynchronously).
|
||||
- **SC-003**: A newly sent message becomes searchable via semantic search within 5 seconds under normal load (< 100 pending embeddings in queue).
|
||||
- **SC-004**: The HNSW index handles at least 100,000 vectors with search latency under 50ms (p99) for top-10 queries.
|
||||
- **SC-005**: When no embedding provider is configured, `search_messages` returns full-text results with zero errors and no configuration changes required by the operator.
|
||||
- **SC-006**: Provider switching (e.g., OpenAI to Ollama) completes backfill of 10,000 messages within 30 minutes without blocking ongoing search operations (full-text fallback available during backfill).
|
||||
- **SC-007**: The system correctly enforces access control on 100% of search results — no message from an inaccessible channel or unrelated DM thread is ever returned, regardless of similarity score.
|
||||
@@ -0,0 +1,216 @@
|
||||
# Tasks: Semantic Search
|
||||
|
||||
**Input**: Design documents from `/specs/008-semantic-search/`
|
||||
**Prerequisites**: spec.md (required)
|
||||
|
||||
**Organization**: Tasks are grouped by user story to enable independent implementation and testing of each story. US1 (Semantic Search) and US4 (Background Embedding) are both P1 and deeply coupled — they are split across Phases 3 and 4 but share foundational infrastructure from Phase 2.
|
||||
|
||||
## Format: `[ID] [P?] [Story] Description`
|
||||
|
||||
- **[P]**: Can run in parallel (different files, no dependencies)
|
||||
- **[Story]**: Which user story this task belongs to (e.g., US1, US2, US3, US4, US5)
|
||||
- Include exact file paths in descriptions
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Setup (Shared Infrastructure)
|
||||
|
||||
**Purpose**: Add dependencies and create package scaffolding for semantic search
|
||||
|
||||
- [ ] T001 Add `TFMV/hnsw` dependency to `go.mod` — run `go get github.com/TFMV/hnsw`
|
||||
- [ ] T002 [P] Create `internal/search/` package structure with placeholder files: `internal/search/search.go` (package declaration and doc comment), `internal/search/types.go` (shared types)
|
||||
- [ ] T003 [P] Create `internal/search/embedding/` sub-package structure: `internal/search/embedding/embedding.go` (package declaration)
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Foundational (Blocking Prerequisites)
|
||||
|
||||
**Purpose**: Core infrastructure that MUST be complete before ANY user story can be implemented — embedding provider interface, SQLite schema for embeddings/queue, HNSW index wrapper, and config plumbing
|
||||
|
||||
**CRITICAL**: No user story work can begin until this phase is complete
|
||||
|
||||
- [ ] T004 Create SQLite migration `schema/002_semantic_search.sql` — add `embeddings` table (`id`, `message_id` FK, `provider`, `model`, `vector` BLOB, `dimensions` INTEGER, `created_at`), `embedding_queue` table (`id`, `message_id` FK, `status` enum pending/processing/completed/failed, `attempts` INTEGER DEFAULT 0, `last_error` TEXT, `created_at`, `completed_at`), add `tags` TEXT column to `messages` table (JSON array, DEFAULT '[]'), and relevant indexes (`idx_embeddings_message` UNIQUE, `idx_embedding_queue_status`, `idx_messages_tags`)
|
||||
- [ ] T005 [P] Define embedding provider interface in `internal/search/embedding/provider.go` — `EmbeddingProvider` interface with methods: `Embed(ctx context.Context, text string) ([]float32, error)`, `EmbedBatch(ctx context.Context, texts []string) ([][]float32, error)`, `Dimensions() int`, `Name() string`, `Model() string`, `MaxTokens() int`. Also define `ProviderConfig` struct (Type, APIKey, Endpoint, Model string)
|
||||
- [ ] T006 [P] Define search domain types in `internal/search/types.go` — `SearchRequest` struct (Query string, Filters *SearchFilters, Limit int, SearchMode string), `SearchFilters` struct (ChannelID *int, SenderID *string, PriorityMin/PriorityMax *int, Tags []string, After/Before *time.Time), `SearchResult` struct (Message, SimilarityScore/RelevanceScore float64, SearchMode string), `SearchResponse` struct (Results []SearchResult, SearchMode string, TotalResults int, Warning string)
|
||||
- [ ] T007 [P] Implement HNSW index wrapper in `internal/search/hnsw.go` — `HNSWIndex` struct wrapping `TFMV/hnsw`, methods: `NewHNSWIndex(dimensions int, dataDir string) (*HNSWIndex, error)`, `Add(id uint64, vector []float32) error`, `Remove(id uint64) error`, `Search(query []float32, k int) ([]HNSWResult, error)` returning (id, distance) pairs, `Save() error`, `Load() error`, `Rebuild(vectors map[uint64][]float32) error`. Use `sync.RWMutex` for concurrent read safety (FR-012). Config: efConstruction=200, M=16, efSearch=100
|
||||
- [ ] T008 [P] Implement embedding repository in `internal/search/repository.go` — `EmbeddingRepository` struct with SQLite `*sql.DB`, methods: `SaveEmbedding(ctx, messageID int64, provider, model string, vector []float32, dimensions int) error`, `GetEmbedding(ctx, messageID int64) (*Embedding, error)`, `GetAllEmbeddings(ctx) ([]Embedding, error)`, `DeleteEmbedding(ctx, messageID int64) error`, `DeleteAllEmbeddings(ctx) error`, `GetEmbeddingCount(ctx) (int64, error)`. Serialize float32 vectors to/from BLOB using `encoding/binary`
|
||||
- [ ] T009 [P] Implement embedding queue repository in `internal/search/queue.go` — `QueueRepository` struct with SQLite `*sql.DB`, methods: `Enqueue(ctx, messageID int64) error`, `Dequeue(ctx, batchSize int) ([]QueueItem, error)` (atomically sets status=processing), `MarkCompleted(ctx, messageID int64) error`, `MarkFailed(ctx, messageID int64, errMsg string) error`, `RetryFailed(ctx, maxAttempts int) (int64, error)` (re-queues items below max attempts), `PendingCount(ctx) (int64, error)`, `GetStaleItems(ctx, olderThan time.Duration) ([]QueueItem, error)`
|
||||
- [ ] T010 Implement search configuration in `internal/search/config.go` — `Config` struct (Provider ProviderConfig, Workers int default 4, BatchSize int default 50, RetryMaxAttempts int default 3, RetryBaseDelay time.Duration default 1s, HNSWEfConstruction int default 200, HNSWM int default 16, HNSWEfSearch int default 100). Parse from environment variables (`SYNAPBUS_EMBEDDING_PROVIDER`, `SYNAPBUS_EMBEDDING_API_KEY`, `SYNAPBUS_OLLAMA_URL`) and config file. Method `IsEnabled() bool` returns true if provider is configured
|
||||
|
||||
**Checkpoint**: Foundation ready — embedding provider interface defined, HNSW wrapper built, SQLite schema for embeddings and queue created, configuration plumbed. User story implementation can now begin in parallel.
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: User Story 4 — Background Embedding on Message Ingest (Priority: P1)
|
||||
|
||||
**Goal**: When a message is sent, asynchronously generate an embedding and index it in HNSW. The `send_message` call returns immediately without blocking on embedding generation.
|
||||
|
||||
**Independent Test**: Send a message, verify `send_message` responds in < 100ms overhead, then wait briefly and confirm the message appears in the HNSW index with a valid embedding stored in SQLite.
|
||||
|
||||
**Rationale for Phase 3 (before US1)**: US1 (search) requires data in the index. The embedding pipeline must exist first so there is something to search against.
|
||||
|
||||
### Implementation for User Story 4
|
||||
|
||||
- [ ] T011 [P] [US4] Implement OpenAI embedding provider in `internal/search/embedding/openai.go` — `OpenAIProvider` struct implementing `EmbeddingProvider`, calls `https://api.openai.com/v1/embeddings` with model `text-embedding-3-small` (1536 dimensions). Handle HTTP POST with JSON body, parse response, extract float32 vectors. Support batch embedding (max 2048 inputs per batch per OpenAI limits). Truncate text to MaxTokens (8191) before sending. Return descriptive errors for 401/429/500 status codes
|
||||
- [ ] T012 [P] [US4] Implement Ollama embedding provider in `internal/search/embedding/ollama.go` — `OllamaProvider` struct implementing `EmbeddingProvider`, calls `POST {endpoint}/api/embeddings` with configurable model (default `nomic-embed-text`, 768 dimensions). Support custom endpoint URL. Implement `EmbedBatch` by sequential calls (Ollama does not support batch). Handle connection refused and timeout errors gracefully
|
||||
- [ ] T013 [P] [US4] Implement Gemini embedding provider in `internal/search/embedding/gemini.go` — `GeminiProvider` struct implementing `EmbeddingProvider`, calls Google Gemini embedding API (`POST https://generativelanguage.googleapis.com/v1beta/models/{model}:embedContent`) with model `text-embedding-004` (768 dimensions). API key passed as query param. Handle batch via `batchEmbedContents` endpoint
|
||||
- [ ] T014 [US4] Implement provider factory in `internal/search/embedding/factory.go` — `NewProvider(cfg ProviderConfig) (EmbeddingProvider, error)` function that returns the appropriate provider based on `cfg.Type` ("openai", "gemini", "ollama"). Return clear error for unknown provider type. Validate required fields (API key for openai/gemini, endpoint for ollama)
|
||||
- [ ] T015 [US4] Implement embedding pipeline worker in `internal/search/pipeline.go` — `Pipeline` struct with dependencies: `EmbeddingProvider`, `EmbeddingRepository`, `QueueRepository`, `*HNSWIndex`, `*slog.Logger`. Method `Start(ctx context.Context)` launches N worker goroutines (from config.Workers). Each worker loops: dequeue batch from queue, call `provider.EmbedBatch`, save embeddings to SQLite via repository, add vectors to HNSW index. Implement exponential backoff retry (base 1s, max 3 attempts per FR-010). Handle empty/whitespace message bodies by skipping (mark completed). Handle text truncation to provider's MaxTokens with slog warning (FR-013). Graceful shutdown via context cancellation
|
||||
- [ ] T016 [US4] Implement message ingest hook in `internal/search/ingest.go` — `IngestHook` struct that receives new message events and enqueues them for embedding. Method `OnMessageCreated(ctx, messageID int64, body string)` — if body is empty/whitespace, skip; otherwise insert into embedding_queue with status=pending. Method `OnMessageDeleted(ctx, messageID int64)` — remove from embedding_queue if pending, delete embedding from repository, remove from HNSW index (FR-014). This must be non-blocking: use a buffered channel or direct SQLite insert (queue table)
|
||||
- [ ] T017 [US4] Implement HNSW index initialization and recovery in `internal/search/index_manager.go` — `IndexManager` struct, method `Initialize(ctx) error`: load HNSW index from disk if file exists, validate dimensions match configured provider, if dimensions mismatch or file missing/corrupted then rebuild from SQLite embeddings table. Method `DetectProviderChange(ctx, currentProvider string) (bool, error)`: compare stored provider/model in embeddings table against current config, return true if changed. Method `TriggerBackfill(ctx) error`: mark all existing embeddings as stale (delete from embeddings, re-enqueue all message IDs to embedding_queue)
|
||||
|
||||
**Checkpoint**: At this point, messages sent via `send_message` are asynchronously embedded and indexed. The HNSW index contains searchable vectors. The pipeline handles errors, retries, and edge cases.
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: User Story 1 — Agent Searches Messages by Meaning (Priority: P1) MVP
|
||||
|
||||
**Goal**: An agent calls the `search_messages` MCP tool with a natural language query and receives semantically relevant messages ranked by cosine similarity, with access control enforced.
|
||||
|
||||
**Independent Test**: Configure an embedding provider, send several messages with varied vocabulary about related topics, then issue a `search_messages` call with different wording. Verify the top results are semantically relevant even without exact keyword matches.
|
||||
|
||||
### Implementation for User Story 1
|
||||
|
||||
- [ ] T018 [US1] Implement semantic search engine in `internal/search/semantic.go` — `SemanticEngine` struct with dependencies: `EmbeddingProvider`, `*HNSWIndex`, `EmbeddingRepository`, `*sql.DB`. Method `Search(ctx, query string, limit int) ([]SearchResult, error)`: embed the query text using the provider, call `HNSWIndex.Search(queryVec, limit * 3)` for over-fetch (to account for access control filtering), convert HNSW distances to cosine similarity scores (1.0 - distance), look up message objects by ID from SQLite, return results sorted by descending similarity. Handle provider errors by returning error (caller decides fallback)
|
||||
- [ ] T019 [US1] Implement full-text search engine in `internal/search/fulltext.go` — `FullTextEngine` struct with `*sql.DB`. Method `Search(ctx, query string, limit int) ([]SearchResult, error)`: query `messages_fts` using FTS5 `MATCH` with `bm25()` ranking function, join to `messages` table for full message data, return results with `relevance_score` (BM25 score normalized). Handle FTS5 syntax errors gracefully (e.g., special characters in query)
|
||||
- [ ] T020 [US1] Implement unified search service in `internal/search/service.go` — `Service` struct orchestrating semantic and full-text engines. Method `Search(ctx, req SearchRequest, callerAgent string) (*SearchResponse, error)`: determine search mode (auto/semantic/fulltext per FR-006), if semantic: try `SemanticEngine.Search()`, on error fall back to full-text with warning. Apply access control filter: query `channel_members` to get channels the callerAgent has joined, filter results to only include messages from those channels or DMs to/from callerAgent (FR-005). Enforce limit (default 10, max 100). Build `SearchResponse` with `search_mode` and optional `warning` fields (FR-008)
|
||||
- [ ] T021 [US1] Implement `search_messages` MCP tool in `internal/mcp/search_tool.go` — register MCP tool `search_messages` with JSON Schema per FR-015: input schema with `query` (string, required), `filters` (object, optional), `limit` (integer, optional, default 10, max 100), `search_mode` (string, optional, enum: auto/semantic/fulltext, default auto). Handler: extract caller agent identity from MCP session context, build `SearchRequest` from tool input, call `search.Service.Search()`, marshal `SearchResponse` to MCP tool result. Include field descriptions in JSON Schema for agent discoverability
|
||||
- [ ] T022 [US1] Wire search service into server startup in `cmd/synapbus/main.go` — in `runServe()`: parse search config from env/config, if embedding provider configured: create provider via factory, initialize HNSW index via IndexManager (with recovery), create repositories, start Pipeline workers, create SemanticEngine. Always create FullTextEngine. Create search Service with available engines. Register `search_messages` MCP tool. Handle graceful shutdown of pipeline workers
|
||||
|
||||
**Checkpoint**: At this point, the `search_messages` MCP tool is functional. Agents can search by meaning with results ranked by cosine similarity. Access control is enforced. Full system is wired end-to-end.
|
||||
|
||||
---
|
||||
|
||||
## Phase 5: User Story 2 — Combined Search with Metadata Filters (Priority: P2)
|
||||
|
||||
**Goal**: An agent calls `search_messages` with both a semantic query and structured filters (channel_id, sender_id, priority range, tags, date range). Results respect all filters while being ranked by semantic relevance.
|
||||
|
||||
**Independent Test**: Send messages about the same topic across multiple channels and from different agents. Issue filtered searches and verify results respect all filter constraints while maintaining semantic ranking.
|
||||
|
||||
### Implementation for User Story 2
|
||||
|
||||
- [ ] T023 [US2] Implement filter query builder in `internal/search/filters.go` — `FilterBuilder` struct, method `BuildWhereClause(filters *SearchFilters) (clause string, args []interface{})`: generate SQL WHERE conditions for channel_id, from_agent (sender_id), priority BETWEEN min/max, created_at after/before, tags containment (JSON `json_each` for SQLite JSON array matching). Return composable clause and parameterized args. Handle nil filters (no-op). Handle tags filter using `EXISTS (SELECT 1 FROM json_each(messages.metadata, '$.tags') WHERE json_each.value IN (...))` or use the `tags` column added in T004
|
||||
- [ ] T024 [US2] Integrate filters into semantic search engine in `internal/search/semantic.go` — modify `SemanticEngine.Search()` to accept `*SearchFilters`, apply pre-filtering strategy: first query SQLite with filters to get candidate message IDs, then intersect with HNSW ANN results. If filter is very selective (< 1000 candidates), compute cosine similarity directly against candidate vectors instead of using HNSW (brute-force on small set). This avoids HNSW returning results that are later all filtered out
|
||||
- [ ] T025 [US2] Integrate filters into full-text search engine in `internal/search/fulltext.go` — modify `FullTextEngine.Search()` to accept `*SearchFilters`, append filter WHERE clauses to FTS5 query JOIN, ensuring filters and FTS ranking work together
|
||||
- [ ] T026 [US2] Update search service and MCP tool for filter support in `internal/search/service.go` and `internal/mcp/search_tool.go` — parse `filters` object from MCP tool input into `SearchFilters` struct (channel_id string, sender_id string, priority_min/priority_max int, tags []string, after/before RFC3339 strings parsed to time.Time). Pass filters through to search engines. Update MCP tool JSON Schema to document filter fields with types and descriptions
|
||||
|
||||
**Checkpoint**: At this point, `search_messages` supports full metadata filtering combined with semantic or full-text ranking. US1 and US2 are both functional and independently testable.
|
||||
|
||||
---
|
||||
|
||||
## Phase 6: User Story 3 — Graceful Fallback to Full-Text Search (Priority: P2)
|
||||
|
||||
**Goal**: When no embedding provider is configured or the provider becomes unreachable, `search_messages` transparently falls back to FTS5 full-text search. The response indicates the search mode used.
|
||||
|
||||
**Independent Test**: Start SynapBus with no embedding provider configured. Send messages and call `search_messages`. Verify results come from FTS5 and the response includes `search_mode: "fulltext"`.
|
||||
|
||||
### Implementation for User Story 3
|
||||
|
||||
- [ ] T027 [US3] Implement provider health checker in `internal/search/embedding/health.go` — `HealthChecker` struct wrapping `EmbeddingProvider`, method `IsHealthy(ctx) bool`: attempt a lightweight embed call (e.g., embed "health check" string) with short timeout (2s), cache result for 30s to avoid hammering provider. Method `LastError() error`. Used by search service to determine if semantic search is available at query time
|
||||
- [ ] T028 [US3] Implement fallback logic in search service `internal/search/service.go` — modify `Search()` method: when `search_mode` is "auto", check `HealthChecker.IsHealthy()` — if provider unavailable, use full-text with warning "embedding provider unavailable, using full-text fallback" (FR-006, acceptance scenario 2). When `search_mode` is "semantic" but provider is down, return error explaining provider is unavailable. When no provider is configured at all, always use full-text mode with no warning (it is the expected mode)
|
||||
- [ ] T029 [US3] Implement startup backfill trigger in `internal/search/index_manager.go` — extend `Initialize()`: after loading index, if provider was just configured (embeddings table empty but messages exist), enqueue all existing message IDs to embedding_queue for background processing. Use configurable batch size and delay between batches to avoid overwhelming the provider (edge case: thousands of unembedded messages). Log progress: "backfill: enqueued N messages for embedding"
|
||||
|
||||
**Checkpoint**: SynapBus works fully without an embedding provider (FTS5 fallback). Runtime provider failures are handled gracefully with automatic fallback and warning messages. Provider addition triggers backfill.
|
||||
|
||||
---
|
||||
|
||||
## Phase 7: User Story 5 — Configurable Embedding Providers (Priority: P3)
|
||||
|
||||
**Goal**: Operators can choose between OpenAI, Gemini, and Ollama embedding providers. Switching providers triggers a full re-embedding backfill.
|
||||
|
||||
**Independent Test**: Configure each provider independently, send messages, verify embeddings are generated. Switch provider in config and restart — verify old embeddings are invalidated and messages are re-embedded with the new provider.
|
||||
|
||||
### Implementation for User Story 5
|
||||
|
||||
- [ ] T030 [US5] Implement provider switch detection and re-embedding in `internal/search/index_manager.go` — extend `DetectProviderChange()`: on startup, query the embeddings table for the most recent provider/model values, compare against current config. If changed: log "provider changed from X to Y, triggering re-embed", call `DeleteAllEmbeddings()` on repository, call `HNSWIndex.Rebuild(nil)` to clear the index (edge case: dimension mismatch requires new index), enqueue all message IDs to embedding_queue. Ensure full-text search remains available during backfill
|
||||
- [ ] T031 [US5] Add provider configuration validation in `internal/search/embedding/factory.go` — extend `NewProvider()` with validation: OpenAI requires non-empty APIKey, Gemini requires non-empty APIKey, Ollama requires valid endpoint URL (try HEAD request with 2s timeout). Return actionable error messages: "openai provider requires SYNAPBUS_EMBEDDING_API_KEY to be set". Add `ValidateConfig(cfg ProviderConfig) error` exported function for use during startup config validation
|
||||
- [ ] T032 [US5] Add embedding provider status to server info in `internal/search/service.go` — add method `Status() SearchStatus` returning current provider name/model, index size (vector count), queue depth (pending embeddings), provider health, and whether a backfill is in progress. This can be exposed via REST API for the Web UI dashboard and via a future MCP resource
|
||||
|
||||
**Checkpoint**: All three embedding providers are implemented, validated, and switchable. Provider changes trigger automatic re-embedding. Operators have visibility into embedding status.
|
||||
|
||||
---
|
||||
|
||||
## Phase 8: Polish & Cross-Cutting Concerns
|
||||
|
||||
**Purpose**: Improvements that affect multiple user stories
|
||||
|
||||
- [ ] T033 [P] Add structured logging throughout search package — ensure all operations in `internal/search/` use `slog` with consistent attributes: `component=search`, `provider`, `message_id`, `queue_depth`, `duration_ms`. Log at appropriate levels: Info for lifecycle events (start/stop/backfill), Warn for truncation and fallback, Error for provider failures
|
||||
- [ ] T034 [P] Add vector serialization helpers in `internal/search/vector.go` — `SerializeVector(v []float32) []byte` and `DeserializeVector(b []byte) []float32` using `encoding/binary.LittleEndian`. Add `CosineSimilarity(a, b []float32) float64` utility. Add `ValidateVector(v []float32, expectedDim int) error`. These support the repository and HNSW wrapper
|
||||
- [ ] T035 Implement HNSW index persistence and corruption recovery in `internal/search/hnsw.go` — extend `Save()` to write to `{dataDir}/hnsw.idx` with temp-file-then-rename for atomic writes. Extend `Load()` to detect corrupted files (invalid header, dimension mismatch) and trigger rebuild from SQLite. Add periodic auto-save (every 5 minutes or every N insertions) via background goroutine
|
||||
- [ ] T036 [P] Add search-related trace logging in `internal/search/service.go` — after each `Search()` call, write a trace record (agent_name, action="search_messages", details JSON with query/filters/mode/result_count/duration_ms) to the `traces` table per Constitution Principle VIII
|
||||
- [ ] T037 Code cleanup and edge case hardening — review all search package files for: empty body handling (skip embedding), duplicate message embedding (upsert semantics), concurrent HNSW access under load, queue stale item cleanup (items stuck in "processing" for > 5 minutes reset to "pending"), and message deletion cascading to embedding cleanup
|
||||
- [ ] T038 Run `make lint` and `make test` — ensure all new code passes linting and existing tests are not broken. Fix any issues
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & Execution Order
|
||||
|
||||
### Phase Dependencies
|
||||
|
||||
- **Setup (Phase 1)**: No dependencies — can start immediately
|
||||
- **Foundational (Phase 2)**: Depends on Setup completion — BLOCKS all user stories
|
||||
- **US4 - Background Embedding (Phase 3)**: Depends on Foundational — must complete before US1 (provides data to search)
|
||||
- **US1 - Semantic Search (Phase 4)**: Depends on Foundational + US4 (needs vectors in index)
|
||||
- **US2 - Filtered Search (Phase 5)**: Depends on US1 (extends search engines)
|
||||
- **US3 - Fallback (Phase 6)**: Depends on US1 (modifies search service); can proceed in parallel with US2
|
||||
- **US5 - Provider Config (Phase 7)**: Depends on US4 (extends provider factory and index manager); can proceed in parallel with US2/US3
|
||||
- **Polish (Phase 8)**: Depends on all user stories being complete
|
||||
|
||||
### User Story Dependencies
|
||||
|
||||
- **US4 (P1)**: Can start after Foundational (Phase 2) — provides the embedding pipeline that US1 depends on
|
||||
- **US1 (P1)**: Depends on US4 — search requires vectors to exist in the index
|
||||
- **US2 (P2)**: Depends on US1 — extends search engines with filter support
|
||||
- **US3 (P2)**: Depends on US1 — modifies search service fallback logic. Can be done in parallel with US2
|
||||
- **US5 (P3)**: Depends on US4 — extends provider factory and index manager. Can be done in parallel with US2/US3
|
||||
|
||||
### Within Each User Story
|
||||
|
||||
- Models/types before repositories
|
||||
- Repositories before services
|
||||
- Services before MCP tool handlers
|
||||
- MCP tool before wiring in main.go
|
||||
- Core implementation before integration
|
||||
|
||||
### Parallel Opportunities
|
||||
|
||||
- All Foundational tasks T005-T010 marked [P] can run in parallel (different files)
|
||||
- T011, T012, T013 (embedding providers) can run in parallel (different files)
|
||||
- US2 and US3 can proceed in parallel after US1 completes
|
||||
- US5 can proceed in parallel with US2/US3 after US4 completes
|
||||
- All Phase 8 tasks marked [P] can run in parallel
|
||||
|
||||
---
|
||||
|
||||
## Implementation Strategy
|
||||
|
||||
### MVP First (US4 + US1 Only)
|
||||
|
||||
1. Complete Phase 1: Setup (add dependencies)
|
||||
2. Complete Phase 2: Foundational (schema, interfaces, HNSW wrapper, repos, config)
|
||||
3. Complete Phase 3: US4 — Background Embedding (providers, pipeline, ingest hook)
|
||||
4. Complete Phase 4: US1 — Semantic Search (search engines, MCP tool, wiring)
|
||||
5. **STOP and VALIDATE**: Send messages, call `search_messages`, verify semantic results with access control
|
||||
|
||||
### Incremental Delivery
|
||||
|
||||
1. Setup + Foundational -> Foundation ready
|
||||
2. Add US4 (embedding pipeline) -> Messages are being embedded
|
||||
3. Add US1 (search) -> Test independently -> Deploy/Demo (MVP!)
|
||||
4. Add US2 (filters) -> Test independently -> Deploy/Demo
|
||||
5. Add US3 (fallback) -> Test independently -> Deploy/Demo
|
||||
6. Add US5 (multi-provider) -> Test independently -> Deploy/Demo
|
||||
7. Each story adds value without breaking previous stories
|
||||
|
||||
---
|
||||
|
||||
## Notes
|
||||
|
||||
- [P] tasks = different files, no dependencies
|
||||
- [Story] label maps task to specific user story for traceability
|
||||
- US4 is placed before US1 because search requires indexed vectors — the pipeline must exist first
|
||||
- The HNSW index uses `sync.RWMutex` for concurrent read-during-write safety (FR-012)
|
||||
- Vector serialization uses `encoding/binary.LittleEndian` for cross-platform consistency
|
||||
- All embedding providers are pure Go HTTP clients — zero CGO (Constitution Principle III)
|
||||
- FTS5 index already exists in `schema/001_initial.sql` — US3 fallback builds on existing infrastructure
|
||||
- Commit after each task or logical group
|
||||
@@ -0,0 +1,135 @@
|
||||
# Feature Specification: Attachments
|
||||
|
||||
**Feature Branch**: `009-attachments`
|
||||
**Created**: 2026-03-13
|
||||
**Status**: Draft
|
||||
**Input**: User description: "Upload files up to 50MB per message. Content-addressable storage with SHA-256 dedup. MCP tools for upload/download. Web UI inline preview. Garbage collection for orphaned files."
|
||||
|
||||
## User Scenarios & Testing *(mandatory)*
|
||||
|
||||
### User Story 1 - Agent Uploads and Attaches a File to a Message (Priority: P1)
|
||||
|
||||
An AI agent produces an artifact (e.g., a generated report, a CSV export, a log file) and needs to share it with another agent or a human owner. The agent calls `upload_attachment` via MCP with the file content and original filename. SynapBus computes the SHA-256 hash, stores the file in content-addressable storage, records metadata in SQLite, and returns the hash. The agent then includes the hash in a message sent via `send_message`. The recipient agent or human can later retrieve the file by hash.
|
||||
|
||||
**Why this priority**: Without upload capability, no other attachment feature is possible. This is the foundational building block that enables all downstream stories.
|
||||
|
||||
**Independent Test**: Can be fully tested by calling `upload_attachment` with a file payload via MCP and verifying the returned hash matches the SHA-256 of the content, the file exists on disk at the correct path, and the metadata row is created in SQLite.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** an authenticated agent, **When** it calls `upload_attachment` with a 1MB PNG file named "chart.png", **Then** the tool returns a JSON response containing the SHA-256 hash, the file is stored at `{data_dir}/attachments/{hash[0:2]}/{hash[2:4]}/{hash}`, and a metadata row is inserted with `original_filename="chart.png"`, `mime_type="image/png"`, `size=1048576`.
|
||||
2. **Given** an authenticated agent, **When** it calls `upload_attachment` with a file identical to one already stored (same SHA-256 hash), **Then** the tool returns the same hash, no duplicate file is written to disk, and a new metadata row is created linking this upload to the new `message_id`.
|
||||
3. **Given** an authenticated agent, **When** it calls `upload_attachment` with a 60MB file, **Then** the tool returns an error indicating the file exceeds the 50MB limit, and no file is written to disk.
|
||||
|
||||
---
|
||||
|
||||
### User Story 2 - Agent Downloads an Attachment by Hash (Priority: P1)
|
||||
|
||||
An agent receives a message containing an attachment hash. It calls `download_attachment` with the hash to retrieve the file content. SynapBus looks up the hash in storage, verifies the file exists, and streams the content back to the agent along with the original filename and MIME type.
|
||||
|
||||
**Why this priority**: Download is the counterpart to upload; together they form the minimum viable attachment feature. An upload without download is useless.
|
||||
|
||||
**Independent Test**: Can be fully tested by first uploading a file, then calling `download_attachment` with the returned hash and verifying the content matches byte-for-byte, the original filename is returned, and the MIME type is correct.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a file previously uploaded with hash `abc123...`, **When** an authenticated agent calls `download_attachment` with that hash, **Then** the tool returns the file content with `original_filename` and `mime_type` metadata.
|
||||
2. **Given** a hash that does not exist in storage, **When** an agent calls `download_attachment` with that hash, **Then** the tool returns a clear error: "attachment not found".
|
||||
3. **Given** an agent that does not own and has not received a message with the attachment, **When** it calls `download_attachment`, **Then** access is permitted (attachments are accessible by hash to any authenticated agent, like a content-addressable CDN).
|
||||
|
||||
---
|
||||
|
||||
### User Story 3 - Human Views Attachments in Web UI (Priority: P2)
|
||||
|
||||
A human owner browses a conversation in the Web UI that contains messages with attachments. For image attachments (JPEG, PNG, GIF, WebP, SVG), the UI renders an inline preview thumbnail. For all other file types, the UI displays the original filename, file size, and MIME type with a download link. Clicking the download link fetches the file via the REST API.
|
||||
|
||||
**Why this priority**: The Web UI is a first-class citizen (Principle X), but this story depends on upload/download (P1 stories) being implemented first. Inline preview is a significant usability improvement for human operators.
|
||||
|
||||
**Independent Test**: Can be tested by uploading image and non-image attachments, sending messages referencing them, then loading the conversation in the Web UI and verifying previews render for images and download links appear for other types.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a message with a PNG attachment in a conversation, **When** a human owner views the conversation in the Web UI, **Then** the image is rendered inline as a thumbnail (max 400px wide) with a "download original" link.
|
||||
2. **Given** a message with a PDF attachment (report.pdf, 2.3MB), **When** a human owner views the conversation, **Then** the UI shows a file card with icon, filename "report.pdf", size "2.3 MB", and a download button.
|
||||
3. **Given** a message with multiple attachments (2 images + 1 CSV), **When** viewing in the Web UI, **Then** all attachments are displayed: images as inline previews, CSV as a download card, in the order they were attached.
|
||||
|
||||
---
|
||||
|
||||
### User Story 4 - Garbage Collection of Orphaned Attachments (Priority: P3)
|
||||
|
||||
Over time, messages may be deleted or expire, leaving attachment files on disk that no message references. An administrator (or automated background process) runs garbage collection. SynapBus identifies attachment hashes present on disk but not referenced by any message's metadata, and removes them from both the filesystem and the SQLite metadata table.
|
||||
|
||||
**Why this priority**: This is an operational concern that only matters after the system has been running for a while with active attachment usage. It is not needed for the feature to be functional.
|
||||
|
||||
**Independent Test**: Can be tested by uploading an attachment, associating it with a message, deleting the message, then triggering garbage collection and verifying the orphaned file and metadata are removed.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** an attachment file on disk whose hash is not referenced by any message in SQLite, **When** garbage collection runs, **Then** the file is deleted from disk and its metadata row is removed from SQLite.
|
||||
2. **Given** an attachment file referenced by two messages and one message is deleted, **When** garbage collection runs, **Then** the file is NOT deleted because it is still referenced by the remaining message.
|
||||
3. **Given** an empty attachments directory with no orphans, **When** garbage collection runs, **Then** it completes without errors and reports zero files removed.
|
||||
|
||||
---
|
||||
|
||||
### User Story 5 - Deduplication Across Messages (Priority: P2)
|
||||
|
||||
Multiple agents independently upload the same file (identical content, possibly different filenames). SynapBus detects that the SHA-256 hash already exists on disk and skips writing a duplicate. Each upload still creates its own metadata row (with its own `original_filename` and `message_id`), but all point to the same physical file. This saves disk space and speeds up uploads of previously-seen content.
|
||||
|
||||
**Why this priority**: Deduplication is a core architectural property of content-addressable storage and should be implemented alongside the basic upload flow, but it is not strictly required for a first working prototype.
|
||||
|
||||
**Independent Test**: Can be tested by uploading the same file content twice with different filenames, verifying only one file exists on disk, and both metadata rows reference the same hash.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** an agent uploads "results_v1.csv" (hash: `aabbcc...`), **When** a different agent uploads "final_results.csv" with identical content, **Then** only one file exists at `{data_dir}/attachments/aa/bb/aabbcc...`, and two metadata rows exist with different `original_filename` values but the same hash.
|
||||
2. **Given** 10 agents upload the same 5MB image, **When** checking disk usage, **Then** approximately 5MB is used (not 50MB), confirming deduplication.
|
||||
|
||||
---
|
||||
|
||||
### Edge Cases
|
||||
|
||||
- What happens when the disk is full during an upload? The system MUST return a clear error ("insufficient disk space") and not leave partial files on disk. Any partially written file MUST be cleaned up.
|
||||
- What happens when an attachment file exists in the metadata table but is missing from disk (e.g., manual deletion)? `download_attachment` MUST return an error ("attachment file missing") rather than panicking, and the metadata row SHOULD be flagged for cleanup.
|
||||
- What happens when the SHA-256 hash collides (astronomically unlikely)? The system MUST overwrite with the new content since SHA-256 collision implies identical content for practical purposes. No special handling is required.
|
||||
- What happens when an upload request has no filename? The system MUST accept the upload and assign a default filename based on the MIME type (e.g., "untitled.bin" for `application/octet-stream`).
|
||||
- What happens when the MIME type cannot be detected? The system MUST fall back to `application/octet-stream` and store the file normally.
|
||||
- What happens when `upload_attachment` is called with zero-byte content? The system MUST reject the upload with an error ("empty file not allowed").
|
||||
- What happens when the data directory's `attachments/` subdirectory does not exist at startup? The system MUST create the directory tree automatically on first use.
|
||||
- What happens when garbage collection runs concurrently with an upload? The system MUST use locking or transaction isolation to prevent deleting a file that is actively being uploaded and linked to a message.
|
||||
|
||||
## Requirements *(mandatory)*
|
||||
|
||||
### Functional Requirements
|
||||
|
||||
- **FR-001**: System MUST accept file uploads up to 50MB per attachment via the `upload_attachment` MCP tool. Files exceeding this limit MUST be rejected with a descriptive error.
|
||||
- **FR-002**: System MUST compute the SHA-256 hash of uploaded file content and use it as the storage key (content-addressable storage).
|
||||
- **FR-003**: System MUST store attachment files on the local filesystem at `{data_dir}/attachments/{hash[0:2]}/{hash[2:4]}/{hash}`, creating intermediate directories as needed.
|
||||
- **FR-004**: System MUST store attachment metadata in SQLite with at minimum: `hash` (TEXT, primary key component), `original_filename` (TEXT), `size` (INTEGER, bytes), `mime_type` (TEXT), `message_id` (TEXT, foreign key to messages), and `created_at` (DATETIME).
|
||||
- **FR-005**: System MUST deduplicate files on disk: if a file with the same SHA-256 hash already exists, the upload MUST skip writing the file and only create a new metadata row.
|
||||
- **FR-006**: System MUST expose a `download_attachment` MCP tool that accepts a SHA-256 hash and returns the file content along with `original_filename` and `mime_type` from the metadata.
|
||||
- **FR-007**: System MUST expose a REST API endpoint (e.g., `GET /api/attachments/{hash}`) for the Web UI to fetch attachment content, with the `Content-Type` header set from the stored `mime_type`.
|
||||
- **FR-008**: Web UI MUST render inline preview thumbnails for image MIME types (`image/jpeg`, `image/png`, `image/gif`, `image/webp`, `image/svg+xml`) within message views.
|
||||
- **FR-009**: Web UI MUST display a download card (filename, size, download link) for non-image attachments.
|
||||
- **FR-010**: System MUST implement garbage collection that identifies and removes attachment files not referenced by any message, cleaning up both the filesystem and SQLite metadata.
|
||||
- **FR-011**: System MUST detect MIME type from file content (magic bytes) when not explicitly provided by the caller, falling back to `application/octet-stream` if detection fails.
|
||||
- **FR-012**: System MUST reject zero-byte uploads with a descriptive error.
|
||||
- **FR-013**: System MUST clean up partially written files if an upload fails mid-write (e.g., disk full, connection dropped).
|
||||
- **FR-014**: System MUST log all attachment operations (upload, download, garbage collection) via `slog` structured logging per Principle VIII.
|
||||
|
||||
### Key Entities *(include if feature involves data)*
|
||||
|
||||
- **Attachment**: Represents a stored file. Key attributes: `hash` (SHA-256 hex string, identifies the physical file), `original_filename` (user-provided name), `size` (bytes), `mime_type` (detected or provided), `message_id` (the message this attachment belongs to), `created_at` (upload timestamp). A single physical file (hash) may be referenced by multiple Attachment metadata rows (deduplication). Relationship: many-to-one with Message.
|
||||
- **AttachmentFile**: The physical file on disk at `{data_dir}/attachments/{hash[0:2]}/{hash[2:4]}/{hash}`. This is an implicit entity -- it has no SQLite row of its own beyond the Attachment metadata rows that reference it. Its existence is determined by whether at least one Attachment row references its hash.
|
||||
|
||||
## Success Criteria *(mandatory)*
|
||||
|
||||
### Measurable Outcomes
|
||||
|
||||
- **SC-001**: An agent can upload a 50MB file via `upload_attachment` and receive a hash response within 10 seconds on standard hardware (SSD, 8GB RAM).
|
||||
- **SC-002**: An agent can download a previously uploaded file via `download_attachment` and receive byte-identical content with correct `original_filename` and `mime_type`.
|
||||
- **SC-003**: Uploading the same file content N times results in exactly 1 file on disk and N metadata rows in SQLite, confirming deduplication works correctly.
|
||||
- **SC-004**: Garbage collection correctly removes 100% of orphaned files (those not referenced by any message) and 0% of referenced files.
|
||||
- **SC-005**: The Web UI renders inline image previews for all supported image MIME types and download cards for non-image types without JavaScript errors.
|
||||
- **SC-006**: All attachment operations (upload, download, delete via GC) produce structured log entries viewable via `slog` output.
|
||||
- **SC-007**: The system handles concurrent uploads of the same file without data corruption, race conditions, or duplicate file writes.
|
||||
- **SC-008**: The attachment storage directory structure is created automatically on first upload -- no manual setup required, consistent with Principle I (single binary, zero setup).
|
||||
@@ -0,0 +1,219 @@
|
||||
# Tasks: Attachments
|
||||
|
||||
**Input**: Design documents from `/specs/009-attachments/`
|
||||
**Prerequisites**: spec.md (required), constitution.md (required)
|
||||
|
||||
**Organization**: Tasks are grouped by user story to enable independent implementation and testing of each story.
|
||||
|
||||
## Format: `[ID] [P?] [Story] Description`
|
||||
|
||||
- **[P]**: Can run in parallel (different files, no dependencies)
|
||||
- **[Story]**: Which user story this task belongs to (e.g., US1, US2, US3)
|
||||
- Include exact file paths in descriptions
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Setup (Shared Infrastructure)
|
||||
|
||||
**Purpose**: Package scaffolding and directory structure for the attachments feature
|
||||
|
||||
- [ ] T001 Create the `internal/attachments/` package directory and `doc.go` with package documentation in `internal/attachments/doc.go`
|
||||
- [ ] T002 [P] Create attachment domain types (Attachment struct, AttachmentFile, error sentinels) in `internal/attachments/model.go`
|
||||
- [ ] T003 [P] Create configuration constants (MaxFileSize=50MB, HashAlgorithm=SHA-256, sharding depth, default MIME type, supported image MIME types) in `internal/attachments/config.go`
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Foundational (Blocking Prerequisites)
|
||||
|
||||
**Purpose**: Core storage layer and database operations that ALL user stories depend on
|
||||
|
||||
**CRITICAL**: No user story work can begin until this phase is complete
|
||||
|
||||
- [ ] T004 Verify the `attachments` table exists in `schema/001_initial.sql` and confirm the schema matches FR-004 (hash, original_filename, size, mime_type, message_id, uploaded_by, created_at). No migration needed since the table is already defined in the initial schema.
|
||||
- [ ] T005 Implement the `Store` interface (repository pattern) in `internal/attachments/store.go` with methods: `InsertMetadata(ctx, Attachment) error`, `GetByHash(ctx, hash) ([]Attachment, error)`, `GetByMessageID(ctx, messageID) ([]Attachment, error)`, `DeleteOrphans(ctx) (int64, error)`, `CountReferences(ctx, hash) (int64, error)`. Define the interface only.
|
||||
- [ ] T006 Implement `SQLiteStore` (the `Store` interface backed by `modernc.org/sqlite`) in `internal/attachments/sqlite_store.go`. All queries use parameterized statements. Context propagation on every method. Structured logging via `slog`.
|
||||
- [ ] T007 Implement content-addressable storage (CAS) engine in `internal/attachments/cas.go`: `Write(ctx, reader io.Reader) (hash string, size int64, err error)` computes SHA-256 while streaming to a temp file, then atomically renames to `{dataDir}/attachments/{hash[0:2]}/{hash[2:4]}/{hash}`. Creates shard directories as needed. Cleans up temp files on failure (FR-013). Returns error on zero-byte content (FR-012).
|
||||
- [ ] T008 [P] Implement `Exists(ctx, hash) (bool, error)` and `Read(ctx, hash) (io.ReadCloser, error)` and `Delete(ctx, hash) error` on the CAS engine in `internal/attachments/cas.go`. `Read` returns `ErrNotFound` if the file is missing from disk.
|
||||
- [ ] T009 [P] Implement MIME type detection from file content (magic bytes) using `net/http.DetectContentType` with fallback to `application/octet-stream` in `internal/attachments/mime.go`. Also implement default filename generation from MIME type (FR-011, edge case: no filename).
|
||||
- [ ] T010 Implement the `Service` struct in `internal/attachments/service.go` that composes `Store` (SQLite) + CAS engine + MIME detector. Constructor: `NewService(store Store, cas *CAS, logger *slog.Logger) *Service`. This is the main entry point for all attachment operations.
|
||||
- [ ] T011 Write table-driven unit tests for CAS engine (write, read, exists, delete, zero-byte rejection, duplicate write dedup, partial cleanup) in `internal/attachments/cas_test.go`
|
||||
- [ ] T012 [P] Write table-driven unit tests for SQLiteStore (insert metadata, get by hash, get by message, delete orphans) in `internal/attachments/sqlite_store_test.go`
|
||||
- [ ] T013 [P] Write table-driven unit tests for MIME detection and default filename generation in `internal/attachments/mime_test.go`
|
||||
|
||||
**Checkpoint**: Foundation ready -- CAS engine writes/reads files, SQLiteStore persists metadata, MIME detection works. User story implementation can now begin in parallel.
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: User Story 1 -- Agent Uploads and Attaches a File (Priority: P1) MVP
|
||||
|
||||
**Goal**: An agent can call `upload_attachment` via MCP to store a file and receive its SHA-256 hash. The file is stored in content-addressable storage with metadata in SQLite.
|
||||
|
||||
**Independent Test**: Call `upload_attachment` with a file payload via MCP, verify the returned hash matches SHA-256 of the content, the file exists on disk at the correct sharded path, and the metadata row is created in SQLite.
|
||||
|
||||
### Implementation for User Story 1
|
||||
|
||||
- [ ] T014 [US1] Implement `Upload(ctx, UploadRequest) (UploadResponse, error)` on the `Service` in `internal/attachments/service.go`. The method: validates size <= 50MB (FR-001), rejects zero-byte (FR-012), detects MIME type (FR-011), assigns default filename if missing, calls CAS.Write (FR-002, FR-003), calls Store.InsertMetadata (FR-004), logs the operation via slog (FR-014). Returns hash, size, mime_type, original_filename.
|
||||
- [ ] T015 [US1] [P] Define `UploadRequest` and `UploadResponse` structs in `internal/attachments/model.go`. UploadRequest: `Content io.Reader`, `Filename string`, `MIMEType string`, `MessageID int64`, `UploadedBy string`. UploadResponse: `Hash string`, `Size int64`, `MIMEType string`, `Filename string`.
|
||||
- [ ] T016 [US1] Register `upload_attachment` MCP tool in `internal/mcp/tools_attachments.go`. The tool accepts base64-encoded file content, original filename (optional), MIME type (optional), and message_id. It decodes the content, calls `Service.Upload`, and returns the hash and metadata as JSON. Include JSON Schema descriptions per Principle II.
|
||||
- [ ] T017 [US1] Write integration test for upload flow (MCP call -> CAS file on disk + SQLite metadata row) in `internal/attachments/service_test.go`. Cover: successful upload, 50MB limit rejection, zero-byte rejection, dedup (same content yields same hash, no duplicate file on disk), missing filename defaults.
|
||||
|
||||
**Checkpoint**: User Story 1 is fully functional. An agent can upload files via MCP and receive hashes.
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: User Story 2 -- Agent Downloads an Attachment by Hash (Priority: P1) MVP
|
||||
|
||||
**Goal**: An agent can call `download_attachment` with a SHA-256 hash to retrieve the file content, original filename, and MIME type.
|
||||
|
||||
**Independent Test**: Upload a file, then call `download_attachment` with the returned hash and verify byte-identical content, correct filename, and correct MIME type.
|
||||
|
||||
### Implementation for User Story 2
|
||||
|
||||
- [ ] T018 [US2] Implement `Download(ctx, hash string) (DownloadResponse, error)` on the `Service` in `internal/attachments/service.go`. The method: calls Store.GetByHash to retrieve metadata (returns first row's filename/mime_type), calls CAS.Read to get file content, logs the operation via slog (FR-014). Returns content reader, original_filename, mime_type, size. Returns `ErrNotFound` if hash not in metadata or file missing from disk.
|
||||
- [ ] T019 [US2] [P] Define `DownloadResponse` struct in `internal/attachments/model.go`: `Content io.ReadCloser`, `Hash string`, `Filename string`, `MIMEType string`, `Size int64`.
|
||||
- [ ] T020 [US2] Register `download_attachment` MCP tool in `internal/mcp/tools_attachments.go`. The tool accepts a SHA-256 hash string, calls `Service.Download`, base64-encodes the content, and returns JSON with content, original_filename, mime_type, and size. Returns clear "attachment not found" error for missing hashes (FR-006).
|
||||
- [ ] T021 [US2] Write integration test for download flow in `internal/attachments/service_test.go`. Cover: successful download (byte-identical content), hash not found, file missing from disk (metadata exists but file deleted).
|
||||
|
||||
**Checkpoint**: User Stories 1 and 2 are both functional. Upload + Download form the minimum viable attachment feature.
|
||||
|
||||
---
|
||||
|
||||
## Phase 5: User Story 5 -- Deduplication Across Messages (Priority: P2)
|
||||
|
||||
**Goal**: Multiple agents uploading identical file content results in a single physical file on disk with multiple metadata rows, saving disk space.
|
||||
|
||||
**Independent Test**: Upload the same file content twice with different filenames, verify only one file exists on disk, and both metadata rows reference the same hash.
|
||||
|
||||
### Implementation for User Story 5
|
||||
|
||||
- [ ] T022 [US5] Verify and harden dedup logic in CAS.Write (`internal/attachments/cas.go`): if file at target path already exists, skip write (do not overwrite), return existing hash. Add file-level locking or atomic rename to handle concurrent uploads of the same content safely (SC-007). Log dedup events.
|
||||
- [ ] T023 [US5] Write dedicated dedup integration tests in `internal/attachments/service_test.go`: upload same content with different filenames, verify single file on disk, two metadata rows with same hash but different original_filename. Verify disk usage does not grow with duplicate uploads.
|
||||
- [ ] T024 [US5] [P] Write concurrent upload test in `internal/attachments/cas_test.go`: launch N goroutines uploading identical content simultaneously, verify no race conditions (use `-race` flag), single file on disk, no errors.
|
||||
|
||||
**Checkpoint**: Deduplication is verified and hardened. Content-addressable storage is fully production-ready.
|
||||
|
||||
---
|
||||
|
||||
## Phase 6: User Story 3 -- Human Views Attachments in Web UI (Priority: P2)
|
||||
|
||||
**Goal**: The Web UI renders inline image previews for supported image types and download cards for all other file types.
|
||||
|
||||
**Independent Test**: Upload image and non-image attachments, send messages referencing them, load the conversation in the Web UI, verify previews render for images and download links appear for other types.
|
||||
|
||||
### Implementation for User Story 3
|
||||
|
||||
- [ ] T025 [US3] Implement REST API endpoint `GET /api/attachments/{hash}` in `internal/api/attachments_handler.go`. The handler calls `Service.Download`, sets `Content-Type` from stored mime_type, sets `Content-Disposition: inline` for images and `attachment` for others, streams the file content to the response (FR-007). Return 404 for missing hashes.
|
||||
- [ ] T026 [US3] [P] Implement REST API endpoint `GET /api/attachments/{hash}/meta` in `internal/api/attachments_handler.go` returning JSON metadata (hash, original_filename, size, mime_type) for the Web UI to decide rendering strategy without downloading full content.
|
||||
- [ ] T027 [US3] Register attachment routes on the chi router in `internal/api/router.go` (or wherever routes are registered). Ensure routes are behind authentication middleware consistent with existing API patterns.
|
||||
- [ ] T028 [US3] [P] Create Svelte attachment preview component in `web/src/lib/components/AttachmentPreview.svelte`. For image MIME types (image/jpeg, image/png, image/gif, image/webp, image/svg+xml): render inline `<img>` thumbnail (max 400px wide) with "download original" link (FR-008). For all other types: render a file card with icon, filename, human-readable size, and download button (FR-009).
|
||||
- [ ] T029 [US3] Integrate `AttachmentPreview` component into the message view component (wherever messages are rendered in the Web UI). For each message, fetch attachment metadata and render previews/cards. Support multiple attachments per message displayed in order.
|
||||
- [ ] T030 [US3] Write handler tests for `GET /api/attachments/{hash}` in `internal/api/attachments_handler_test.go`. Cover: successful image download (Content-Type set), successful non-image download (Content-Disposition: attachment), 404 for missing hash.
|
||||
|
||||
**Checkpoint**: Human owners can view attachments in the Web UI with inline previews for images and download cards for other file types.
|
||||
|
||||
---
|
||||
|
||||
## Phase 7: User Story 4 -- Garbage Collection of Orphaned Attachments (Priority: P3)
|
||||
|
||||
**Goal**: Orphaned attachment files (not referenced by any message) are identified and removed from both disk and SQLite, reclaiming storage space.
|
||||
|
||||
**Independent Test**: Upload an attachment, associate it with a message, delete the message, trigger GC, verify the orphaned file and metadata are removed.
|
||||
|
||||
### Implementation for User Story 4
|
||||
|
||||
- [ ] T031 [US4] Implement `Store.FindOrphanHashes(ctx) ([]string, error)` in `internal/attachments/sqlite_store.go`. Query returns hashes from the attachments table where no row references a valid message_id (message_id IS NULL or the referenced message no longer exists).
|
||||
- [ ] T032 [US4] Implement `GarbageCollect(ctx) (GCResult, error)` on the `Service` in `internal/attachments/service.go`. The method: acquires a lock (mutex) to prevent concurrent GC/upload conflicts (edge case), calls Store.FindOrphanHashes, for each orphan hash calls CAS.Delete then Store.DeleteByHash, logs each deletion and the summary via slog (FR-010, FR-014). Returns `GCResult{FilesRemoved int, BytesReclaimed int64}`.
|
||||
- [ ] T033 [US4] [P] Define `GCResult` struct in `internal/attachments/model.go`.
|
||||
- [ ] T034 [US4] Register `gc_attachments` MCP tool (admin-only) in `internal/mcp/tools_attachments.go`. The tool calls `Service.GarbageCollect` and returns the GC summary as JSON. Consider also exposing as a CLI subcommand (`synapbus gc-attachments`) in `cmd/synapbus/`.
|
||||
- [ ] T035 [US4] Write integration tests for GC in `internal/attachments/service_test.go`. Cover: orphan removed (file + metadata deleted), file referenced by one remaining message is NOT removed, empty state GC completes without error, concurrent GC + upload safety (edge case).
|
||||
|
||||
**Checkpoint**: Garbage collection works correctly. Orphaned files are cleaned up without affecting referenced attachments.
|
||||
|
||||
---
|
||||
|
||||
## Phase 8: Polish & Cross-Cutting Concerns
|
||||
|
||||
**Purpose**: Improvements that affect multiple user stories
|
||||
|
||||
- [ ] T036 [P] Add structured slog logging review across all attachment operations (upload, download, dedup, GC) -- verify log entries include agent name, hash, filename, size, duration, and error details per Principle VIII in `internal/attachments/service.go`
|
||||
- [ ] T037 [P] Add trace entries for attachment operations (upload, download, GC) in the traces table via the existing trace infrastructure in `internal/trace/`. Ensure owners can see attachment activity for their agents.
|
||||
- [ ] T038 [P] Verify the attachments directory tree (`{data_dir}/attachments/`) is auto-created on first use (SC-008, Principle I). Add initialization logic to CAS constructor or Service constructor in `internal/attachments/cas.go` if not already present.
|
||||
- [ ] T039 [P] Handle edge case: attachment metadata exists but file missing from disk. `Download` should return a clear error ("attachment file missing") and optionally flag the metadata row for GC cleanup. Implement in `internal/attachments/service.go`.
|
||||
- [ ] T040 [P] Handle edge case: disk full during upload. CAS.Write must clean up the temp file and return a descriptive error ("insufficient disk space"). Verify in `internal/attachments/cas.go`.
|
||||
- [ ] T041 Add 50MB size limit enforcement at the MCP tool level (pre-check before reading full content into memory) in `internal/mcp/tools_attachments.go`. Consider streaming validation to avoid buffering the full 50MB.
|
||||
- [ ] T042 [P] Run `make lint` and `make test` -- fix any linting errors, ensure all tests pass with `-race` flag
|
||||
- [ ] T043 Run `make build` -- verify the binary compiles cleanly for `linux/amd64` and `darwin/arm64` with zero CGO per Principle III
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & Execution Order
|
||||
|
||||
### Phase Dependencies
|
||||
|
||||
- **Setup (Phase 1)**: No dependencies -- can start immediately
|
||||
- **Foundational (Phase 2)**: Depends on Phase 1 completion -- BLOCKS all user stories
|
||||
- **User Story 1 (Phase 3)**: Depends on Phase 2 completion -- MVP upload
|
||||
- **User Story 2 (Phase 4)**: Depends on Phase 2 completion -- MVP download (can run in parallel with Phase 3)
|
||||
- **User Story 5 (Phase 5)**: Depends on Phase 3 completion -- hardens dedup on top of upload
|
||||
- **User Story 3 (Phase 6)**: Depends on Phases 3 + 4 completion -- Web UI needs both upload and download working
|
||||
- **User Story 4 (Phase 7)**: Depends on Phase 2 completion -- can run in parallel with other stories after foundation
|
||||
- **Polish (Phase 8)**: Depends on all desired user stories being complete
|
||||
|
||||
### User Story Dependencies
|
||||
|
||||
- **US1 (Upload, P1)**: Can start after Phase 2 -- no dependencies on other stories
|
||||
- **US2 (Download, P1)**: Can start after Phase 2 -- no dependencies on other stories (can parallel with US1)
|
||||
- **US5 (Dedup, P2)**: Depends on US1 being implemented (hardens the upload dedup path)
|
||||
- **US3 (Web UI, P2)**: Depends on US1 + US2 (needs working upload and download for REST endpoints)
|
||||
- **US4 (GC, P3)**: Can start after Phase 2 -- independent of other stories conceptually, but best tested after US1
|
||||
|
||||
### Within Each User Story
|
||||
|
||||
- Models/types before service methods
|
||||
- Service methods before MCP tools / API handlers
|
||||
- Core implementation before integration tests
|
||||
- Story complete before moving to next priority
|
||||
|
||||
### Parallel Opportunities
|
||||
|
||||
- All Phase 1 tasks marked [P] can run in parallel
|
||||
- All Phase 2 tasks marked [P] can run in parallel (within Phase 2)
|
||||
- US1 (Phase 3) and US2 (Phase 4) can run in parallel after Phase 2
|
||||
- US4 (Phase 7) can start as soon as Phase 2 completes, in parallel with other stories
|
||||
- All Polish tasks marked [P] can run in parallel
|
||||
|
||||
---
|
||||
|
||||
## Implementation Strategy
|
||||
|
||||
### MVP First (User Stories 1 + 2)
|
||||
|
||||
1. Complete Phase 1: Setup
|
||||
2. Complete Phase 2: Foundational (CRITICAL -- blocks all stories)
|
||||
3. Complete Phase 3: User Story 1 (Upload)
|
||||
4. Complete Phase 4: User Story 2 (Download)
|
||||
5. **STOP and VALIDATE**: Test upload + download end-to-end via MCP
|
||||
6. Deploy/demo if ready -- agents can share files
|
||||
|
||||
### Incremental Delivery
|
||||
|
||||
1. Setup + Foundational -> Foundation ready
|
||||
2. US1 (Upload) + US2 (Download) -> MVP! Agents can share files
|
||||
3. US5 (Dedup) -> Hardened storage, production-ready CAS
|
||||
4. US3 (Web UI) -> Humans can view attachments inline
|
||||
5. US4 (GC) -> Operational cleanup for long-running instances
|
||||
6. Polish -> Logging, edge cases, cross-compilation verification
|
||||
|
||||
---
|
||||
|
||||
## Notes
|
||||
|
||||
- [P] tasks = different files, no dependencies
|
||||
- [Story] label maps task to specific user story for traceability
|
||||
- Each user story should be independently completable and testable
|
||||
- The `attachments` table already exists in `schema/001_initial.sql` -- no new migration is needed
|
||||
- CAS path format: `{data_dir}/attachments/{hash[0:2]}/{hash[2:4]}/{hash}` (two levels of sharding)
|
||||
- All code must be pure Go, zero CGO (Principle III)
|
||||
- All agent-facing operations via MCP tools only; REST API is for Web UI only (Principle II)
|
||||
- Structured logging via `slog` for all operations (Principle VIII)
|
||||
- Commit after each task or logical group
|
||||
@@ -0,0 +1,127 @@
|
||||
# Feature Specification: Swarm Patterns
|
||||
|
||||
**Feature Branch**: `010-swarm-patterns`
|
||||
**Created**: 2026-03-13
|
||||
**Status**: Draft
|
||||
**Input**: User description: "Stigmergy (blackboard), task auction, agent discovery, channel types (standard/blackboard/auction), MCP tools for swarm coordination"
|
||||
|
||||
## User Scenarios & Testing *(mandatory)*
|
||||
|
||||
### User Story 1 - Stigmergy on a Blackboard Channel (Priority: P1)
|
||||
|
||||
A team of AI agents coordinates a multi-step research workflow without a central orchestrator. A human owner creates a blackboard channel called `#research-pipeline`. A research agent posts a `#finding` tagged message ("Found 3 CVEs in dependency X"). An analysis agent, which is watching for `#finding` tags, picks up the message, analyzes the CVEs, and posts a `#decision` tagged message ("CVE-2026-1234 is critical, others are informational"). An action agent watching for `#decision` tags picks up the decision and creates a patch, posting a `#trace` tagged message with the result. The human owner observes the entire emergent workflow in the Web UI without having orchestrated any of it.
|
||||
|
||||
**Why this priority**: Stigmergy is the foundational swarm pattern. It enables emergent multi-agent coordination without point-to-point wiring, which is the core value proposition of SynapBus swarm features. Without blackboard channels, the other swarm patterns lack their substrate.
|
||||
|
||||
**Independent Test**: Can be fully tested by creating a blackboard channel, posting tagged messages from multiple agents, and verifying that agents can filter and read messages by tag. Delivers value as a standalone coordination mechanism even without task auctions or agent discovery.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a registered agent and no blackboard channel exists, **When** the agent calls `create_channel` with `type: "blackboard"` and `name: "research-pipeline"`, **Then** the channel is created with type `blackboard` and appears in `list_channels` results with `type: "blackboard"`.
|
||||
2. **Given** an agent has joined a blackboard channel, **When** it calls `send_message` with `tags: ["#finding"]` and a message body, **Then** the message is stored with the tags and is retrievable by other agents filtering by `tag: "#finding"`.
|
||||
3. **Given** a blackboard channel contains messages with mixed tags (`#finding`, `#decision`, `#trace`), **When** an agent calls `read_inbox` or `search_messages` filtered to `tag: "#decision"` on that channel, **Then** only messages tagged `#decision` are returned, ordered by timestamp.
|
||||
4. **Given** a blackboard channel, **When** an agent posts a message without any recognized tag (`#finding`, `#task`, `#decision`, `#trace`), **Then** the message is accepted but stored with an empty tag set (tags are not mandatory, but the channel supports them).
|
||||
|
||||
---
|
||||
|
||||
### User Story 2 - Task Auction Workflow (Priority: P2)
|
||||
|
||||
A coordinator agent has a task ("translate document X into French") that it cannot perform itself. It posts the task to an auction channel with requirements (`language: french`, `domain: legal`) and a deadline (30 minutes from now). Two translation agents see the task and submit bids: Agent A bids 10 minutes with confidence 0.9, Agent B bids 20 minutes with confidence 0.95. The coordinator reviews the bids, selects Agent A as the winner using `accept_bid`, and Agent A receives the assignment. Agent A completes the work and calls `complete_task` with the result. The coordinator and the human owner can see the full auction lifecycle in the trace log.
|
||||
|
||||
**Why this priority**: Task auction is the second most important swarm pattern. It enables dynamic work distribution among agents with different capabilities. However, it depends on channel infrastructure (P1) and is more complex than stigmergy.
|
||||
|
||||
**Independent Test**: Can be fully tested by creating an auction channel, posting a task, having agents bid, accepting a bid, and completing the task. Delivers value as a standalone task delegation mechanism.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** an auction channel exists and an agent has joined it, **When** the agent calls `post_task` with `title`, `description`, `requirements` (JSON object), and `deadline` (ISO 8601 timestamp), **Then** the task is created with status `open` and is visible to all channel members.
|
||||
2. **Given** an open task exists on an auction channel, **When** a different agent calls `bid_task` with `task_id`, `time_estimate_seconds`, `confidence` (0.0-1.0), and `capabilities` (JSON object), **Then** the bid is recorded and the task poster is notified of the new bid.
|
||||
3. **Given** a task has received multiple bids, **When** the task poster calls `accept_bid` with `task_id` and `bid_id`, **Then** the task status changes to `assigned`, the winning bidder is notified, all other bidders are notified of rejection, and no further bids are accepted.
|
||||
4. **Given** a task is assigned to an agent, **When** the assigned agent calls `complete_task` with `task_id` and `result` (JSON object), **Then** the task status changes to `completed`, the result is stored, and the task poster is notified.
|
||||
5. **Given** a task has a deadline, **When** the deadline passes and no bid has been accepted, **Then** the task status changes to `expired` and a system message is posted to the channel.
|
||||
|
||||
---
|
||||
|
||||
### User Story 3 - Agent Discovery by Capability (Priority: P3)
|
||||
|
||||
A new orchestration agent joins SynapBus and needs to find agents that can help with sentiment analysis. It calls `discover_agents` with the keyword `"sentiment analysis"`. SynapBus searches through registered agents' capability cards and returns matching agents ranked by relevance. The orchestrator inspects the returned capability cards (which include skills, supported input/output formats, and availability status) and selects the best match. If semantic search is configured, the query also matches agents whose capability descriptions are semantically similar (e.g., an agent whose card says "opinion mining and emotional tone detection").
|
||||
|
||||
**Why this priority**: Agent discovery enables agents to find collaborators dynamically rather than being hardwired. It is lower priority because it requires the agent registry and capability cards to already be populated, and basic agent coordination can work with manually configured agent names.
|
||||
|
||||
**Independent Test**: Can be fully tested by registering several agents with different capability cards, then calling `discover_agents` with various queries and verifying the results are relevant. Delivers value as a standalone agent directory.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** multiple agents are registered with capability cards containing `skills` arrays, **When** an agent calls `discover_agents` with `query: "sentiment analysis"`, **Then** agents whose capability cards contain matching keywords are returned, sorted by relevance score.
|
||||
2. **Given** an embedding provider is configured, **When** an agent calls `discover_agents` with a query that has no exact keyword match but is semantically similar to an agent's capabilities, **Then** the semantically similar agent is still returned (with a lower score than an exact match would produce).
|
||||
3. **Given** no embedding provider is configured, **When** an agent calls `discover_agents`, **Then** the system falls back to full-text search over capability card fields and still returns useful results.
|
||||
|
||||
---
|
||||
|
||||
### User Story 4 - Channel Type Enforcement (Priority: P2)
|
||||
|
||||
A human owner creates three channels: a `standard` channel for general chat, a `blackboard` channel for a research project, and an `auction` channel for task delegation. The channel types enforce appropriate interaction patterns. On the standard channel, agents send and read messages normally. On the blackboard channel, messages support tag filtering. On the auction channel, only `post_task` creates new top-level items, and agents interact with tasks through `bid_task`, `accept_bid`, and `complete_task`. An agent attempting to call `post_task` on a standard channel receives an error indicating that task operations require an auction channel.
|
||||
|
||||
**Why this priority**: Channel type enforcement is essential infrastructure that supports both stigmergy (P1) and task auctions (P2). It shares P2 priority because it must be delivered alongside or before the task auction story.
|
||||
|
||||
**Independent Test**: Can be fully tested by creating one channel of each type and verifying that type-specific operations succeed on the correct channel type and fail with clear errors on incorrect types.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a channel of type `standard`, **When** an agent calls `post_task` targeting that channel, **Then** the system returns an error: "post_task requires a channel of type 'auction'".
|
||||
2. **Given** a channel of type `auction`, **When** an agent calls `send_message` (a regular message, not a task), **Then** the message is accepted (auction channels allow discussion alongside tasks).
|
||||
3. **Given** any channel type, **When** an agent calls `create_channel` with `type: "invalid_type"`, **Then** the system returns a validation error listing the valid types: `standard`, `blackboard`, `auction`.
|
||||
|
||||
---
|
||||
|
||||
### Edge Cases
|
||||
|
||||
- What happens when an agent bids on a task that has already been assigned or completed? The system MUST reject the bid with a clear error indicating the task's current status.
|
||||
- What happens when the task poster tries to accept a bid on an expired task? The system MUST reject the acceptance and return the task's `expired` status.
|
||||
- What happens when the assigned agent calls `complete_task` but the task was already completed? The system MUST return an error indicating the task is already completed (idempotency: if the same agent calls with the same result, it should succeed silently or return the existing completion).
|
||||
- What happens when an agent that did not win the auction calls `complete_task`? The system MUST reject the call — only the assigned agent can complete a task.
|
||||
- What happens when an agent posts a task with a deadline in the past? The system MUST reject the task with a validation error.
|
||||
- What happens when a blackboard channel has thousands of messages and an agent filters by a tag that matches none? The system MUST return an empty result set, not an error.
|
||||
- What happens when `discover_agents` is called but no agents have capability cards? The system MUST return an empty result set with a clear indication that no agents matched.
|
||||
- What happens when the task poster bids on their own task? The system MUST reject the bid — a poster cannot bid on their own task.
|
||||
- What happens when a channel is deleted while it has open tasks? All open tasks MUST transition to `cancelled` status before the channel is removed.
|
||||
|
||||
## Requirements *(mandatory)*
|
||||
|
||||
### Functional Requirements
|
||||
|
||||
- **FR-001**: System MUST support three channel types: `standard`, `blackboard`, and `auction`. Channel type is set at creation and MUST NOT be changed after creation.
|
||||
- **FR-002**: Messages on blackboard channels MUST support a `tags` field accepting an array of strings. Recognized tags are `#finding`, `#task`, `#decision`, and `#trace`, but arbitrary tags MUST also be accepted.
|
||||
- **FR-003**: System MUST allow filtering messages by one or more tags on blackboard channels via `read_inbox` and `search_messages`.
|
||||
- **FR-004**: System MUST expose an MCP tool `post_task` that creates a task on an auction channel with fields: `channel_id`, `title`, `description`, `requirements` (JSON object), and `deadline` (ISO 8601 timestamp).
|
||||
- **FR-005**: System MUST expose an MCP tool `bid_task` that records a bid on an open task with fields: `task_id`, `time_estimate_seconds` (integer), `confidence` (float 0.0-1.0), and `capabilities` (JSON object).
|
||||
- **FR-006**: System MUST expose an MCP tool `accept_bid` that allows the task poster (and only the task poster) to select a winning bid, transitioning the task to `assigned` status.
|
||||
- **FR-007**: System MUST expose an MCP tool `complete_task` that allows the assigned agent (and only the assigned agent) to mark a task as completed with a `result` (JSON object).
|
||||
- **FR-008**: System MUST expose an MCP tool `discover_agents` that accepts a `query` string and optional `limit` (default 10) and returns matching agents with their capability cards, sorted by relevance.
|
||||
- **FR-009**: Agent capability cards MUST include at minimum: `skills` (array of strings), `description` (string), and `availability` (enum: `available`, `busy`, `offline`).
|
||||
- **FR-010**: Task status transitions MUST follow the lifecycle: `open` -> `assigned` -> `completed`, with `expired` and `cancelled` as terminal states reachable from `open`.
|
||||
- **FR-011**: The system MUST enforce channel type constraints: `post_task`, `bid_task`, `accept_bid`, and `complete_task` MUST only work on `auction` type channels.
|
||||
- **FR-012**: `discover_agents` MUST fall back to full-text search when no embedding provider is configured (per Constitution Principle VI).
|
||||
- **FR-013**: All swarm operations (task posts, bids, acceptances, completions, tag-filtered reads) MUST be recorded in the trace log (per Constitution Principle VIII).
|
||||
- **FR-014**: Expired task detection MUST be handled by a background goroutine that periodically checks for tasks past their deadline and transitions them to `expired` status.
|
||||
- **FR-015**: System MUST prevent the task poster from bidding on their own task.
|
||||
|
||||
### Key Entities
|
||||
|
||||
- **Channel** (extended): Existing channel entity gains a `type` field with values `standard`, `blackboard`, or `auction`. The type is immutable after creation. Standard channels behave as they do today. Blackboard channels enable tag-based message filtering. Auction channels enable task lifecycle operations.
|
||||
- **Task**: Represents a unit of work posted to an auction channel. Key attributes: `id`, `channel_id`, `poster_agent_id`, `title`, `description`, `requirements` (JSON), `deadline` (timestamp), `status` (open/assigned/completed/expired/cancelled), `assigned_agent_id` (nullable), `result` (JSON, nullable), `created_at`, `updated_at`.
|
||||
- **Bid**: Represents an agent's offer to complete a task. Key attributes: `id`, `task_id`, `bidder_agent_id`, `time_estimate_seconds`, `confidence` (float), `capabilities` (JSON), `status` (pending/accepted/rejected), `created_at`.
|
||||
- **Capability Card** (extended): Extends the existing agent entity with structured capability metadata. Key attributes: `skills` (array of strings), `description` (free text), `availability` (available/busy/offline). Stored as a JSON column on the agent record or a related table. Used by `discover_agents` for keyword and semantic matching.
|
||||
- **Message Tags**: An extension to the message entity for blackboard channels. Tags are stored as an array of strings associated with a message. Indexed for efficient filtering.
|
||||
|
||||
## Success Criteria *(mandatory)*
|
||||
|
||||
### Measurable Outcomes
|
||||
|
||||
- **SC-001**: An agent can create a blackboard channel, post 3 tagged messages with different tags, and retrieve only messages matching a specific tag filter in under 500ms total for the filtered read operation.
|
||||
- **SC-002**: A complete task auction lifecycle (post_task -> bid_task x2 -> accept_bid -> complete_task) can be executed across 3 agents in under 5 seconds with all state transitions correctly reflected.
|
||||
- **SC-003**: `discover_agents` returns relevant results within 200ms for a corpus of up to 1000 registered agents when using full-text search fallback.
|
||||
- **SC-004**: All 4 MCP tools (`post_task`, `bid_task`, `accept_bid`, `complete_task`) reject operations on incorrect channel types with descriptive error messages that include the required channel type.
|
||||
- **SC-005**: Task expiration is detected and status is updated within 60 seconds of the deadline passing.
|
||||
- **SC-006**: Every swarm operation (task post, bid, accept, complete, discover, tag-filtered read) produces a trace entry visible in the Web UI and queryable via the trace API.
|
||||
- **SC-007**: The system functions correctly with swarm features when no embedding provider is configured — `discover_agents` falls back to full-text search and all other swarm operations work without semantic capabilities.
|
||||
@@ -0,0 +1,158 @@
|
||||
# Tasks: Swarm Patterns
|
||||
|
||||
**Input**: Design documents from `/specs/010-swarm-patterns/`
|
||||
**Prerequisites**: spec.md (required), constitution.md (reviewed)
|
||||
|
||||
**Organization**: Tasks are grouped by user story to enable independent implementation and testing of each story. Channel type enforcement (US4) is co-located with Phase 2 since it is foundational infrastructure required by US1 and US2.
|
||||
|
||||
## Format: `[ID] [P?] [Story] Description`
|
||||
|
||||
- **[P]**: Can run in parallel (different files, no dependencies)
|
||||
- **[Story]**: Which user story this task belongs to (e.g., US1, US2, US3, US4)
|
||||
- Include exact file paths in descriptions
|
||||
|
||||
## Phase 1: Setup (Shared Infrastructure)
|
||||
|
||||
**Purpose**: Schema migration, domain types, and shared constants for swarm patterns
|
||||
|
||||
- [ ] T001 Create migration `schema/002_swarm_patterns.sql` — add `tags` column (JSON text, default '[]') to `messages` table, add `expired` to `tasks.status` CHECK constraint, add `result` column (JSON text, nullable) to `tasks` table, add `confidence` column (REAL, nullable) to `task_bids` table, add `time_estimate_seconds` column (INTEGER, nullable) to `task_bids` table, add capability card columns to `agents` table (`skills` JSON text default '[]', `capability_description` TEXT default '', `availability` TEXT default 'available' with CHECK), create FTS index `agents_fts` over `name, capability_description, skills` for fallback keyword search, add index `idx_messages_tags` for tag-based filtering
|
||||
- [ ] T002 [P] Define channel type constants and validation in `internal/channels/types.go` — `ChannelTypeStandard`, `ChannelTypeBlackboard`, `ChannelTypeAuction` string constants; `ValidChannelType(t string) bool` function; `ChannelTypeError` struct with descriptive messages per FR-011
|
||||
- [ ] T003 [P] Define task domain types in `internal/channels/task.go` — `Task` struct (id, channel_id, posted_by, title, description, requirements JSON, deadline, status, assigned_to, result JSON, created_at, updated_at), `Bid` struct (id, task_id, agent_name, time_estimate_seconds int, confidence float64, capabilities JSON, status, created_at), `TaskStatus` constants (`open`, `assigned`, `completed`, `expired`, `cancelled`), `BidStatus` constants (`pending`, `accepted`, `rejected`), lifecycle validation functions per FR-010
|
||||
- [ ] T004 [P] Define capability card types in `internal/agents/capability.go` — `CapabilityCard` struct (skills []string, description string, availability string), `Availability` constants (`available`, `busy`, `offline`), JSON marshal/unmarshal helpers per FR-009
|
||||
- [ ] T005 [P] Define message tag types in `internal/messaging/tags.go` — `WellKnownTags` list (`#finding`, `#task`, `#decision`, `#trace`), `ParseTags(input []string) []string` normalization function, `TagsToJSON`/`TagsFromJSON` helpers per FR-002
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Foundational (Blocking Prerequisites)
|
||||
|
||||
**Purpose**: Storage layer, channel type enforcement, and shared services that ALL user stories depend on
|
||||
|
||||
**CRITICAL**: No user story work can begin until this phase is complete
|
||||
|
||||
- [ ] T006 Extend channel store in `internal/channels/store.go` — update `CreateChannel` to accept and persist `type` field; update `GetChannel`/`ListChannels` to return type; add `ValidateChannelType` check that rejects unknown types with descriptive error per FR-001 (listing valid types: standard, blackboard, auction); ensure type is immutable after creation (reject any update attempt to change type)
|
||||
- [ ] T007 [P] Extend message store in `internal/messaging/store.go` — update `CreateMessage` to accept and persist `tags` JSON array; update `GetMessages`/`ReadInbox` to support optional `tag` filter parameter; implement tag-based filtering query using JSON functions or LIKE matching against the tags column; return empty result set (not error) when no messages match a tag filter per edge case requirement
|
||||
- [ ] T008 [P] Extend agent store in `internal/agents/store.go` — update `RegisterAgent`/`UpdateAgent` to accept and persist capability card fields (`skills`, `capability_description`, `availability`); add `SearchAgents(ctx, query string, limit int) ([]AgentWithScore, error)` using FTS5 `agents_fts` table; return empty result set when no agents match per edge case requirement
|
||||
- [ ] T009 [P] [US4] Implement channel type enforcement middleware in `internal/channels/enforcement.go` — `RequireChannelType(channelID int64, requiredType string) error` function that loads the channel and returns a typed error if the channel type does not match (e.g., "post_task requires a channel of type 'auction'"); used by all auction and blackboard operations per FR-011 and US4 acceptance scenarios
|
||||
- [ ] T010 Create task store in `internal/channels/task_store.go` — `CreateTask`, `GetTask`, `ListTasksByChannel`, `UpdateTaskStatus`, `AssignTask`, `CompleteTask` with SQL operations against `tasks` table; `CreateBid`, `GetBid`, `ListBidsByTask`, `UpdateBidStatus` against `task_bids` table; enforce lifecycle transitions per FR-010 (open->assigned->completed, open->expired, open->cancelled); reject bids on non-open tasks per edge case; reject completion by non-assigned agent per edge case
|
||||
- [ ] T011 [P] Add swarm trace helpers in `internal/trace/swarm.go` — helper functions `TraceTaskPosted`, `TraceBidSubmitted`, `TraceBidAccepted`, `TraceTaskCompleted`, `TraceTaskExpired`, `TraceDiscoverAgents`, `TraceTagFilteredRead` that wrap the existing trace store with structured JSON details per FR-013 and Principle VIII
|
||||
|
||||
**Checkpoint**: Foundation ready — channel types enforced, storage layer supports tags/tasks/capability cards, user story implementation can begin
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: User Story 1 — Stigmergy on a Blackboard Channel (Priority: P1)
|
||||
|
||||
**Goal**: Agents coordinate via tagged messages on blackboard channels without a central orchestrator
|
||||
|
||||
**Independent Test**: Create a blackboard channel, post tagged messages from multiple agents, verify agents can filter and read messages by tag
|
||||
|
||||
### Implementation for User Story 1
|
||||
|
||||
- [ ] T012 [US1] Implement blackboard MCP tool extensions in `internal/mcp/blackboard.go` — extend `send_message` tool to accept optional `tags` parameter (array of strings) when targeting a blackboard channel; messages without tags are accepted with empty tag set per acceptance scenario 4; validate that tags parameter is stored correctly as JSON array per FR-002
|
||||
- [ ] T013 [US1] Implement tag-filtered read in `internal/mcp/blackboard.go` — extend `read_inbox` and `search_messages` MCP tools to accept optional `tag` filter parameter; when tag filter is provided on a blackboard channel, return only matching messages ordered by timestamp per acceptance scenario 3; trace every tag-filtered read via `TraceTagFilteredRead` per FR-013
|
||||
- [ ] T014 [US1] Register blackboard MCP tools in `internal/mcp/server.go` — wire up the extended `send_message` (with tags support) and tag-filtered `read_inbox`/`search_messages` into the MCP tool registry with JSON Schema descriptions per Principle II
|
||||
|
||||
**Checkpoint**: User Story 1 fully functional — agents can create blackboard channels, post tagged messages, and filter by tag
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: User Story 4 — Channel Type Enforcement (Priority: P2)
|
||||
|
||||
**Goal**: Channel types enforce appropriate interaction patterns; type-incorrect operations fail with clear errors
|
||||
|
||||
**Independent Test**: Create one channel of each type, verify type-specific operations succeed on correct types and fail with descriptive errors on incorrect types
|
||||
|
||||
### Implementation for User Story 4
|
||||
|
||||
- [ ] T015 [US4] Wire enforcement into existing MCP tools in `internal/mcp/server.go` — ensure `create_channel` tool validates `type` parameter against the three valid types (standard, blackboard, auction) and returns validation error listing valid types for invalid input per acceptance scenario 3; ensure channel type is included in `list_channels` results
|
||||
- [ ] T016 [US4] Add enforcement guards to auction tools in `internal/mcp/auction.go` — `post_task`, `bid_task`, `accept_bid`, and `complete_task` must call `RequireChannelType(channelID, "auction")` before proceeding per acceptance scenario 1 and FR-011; `send_message` must remain allowed on auction channels per acceptance scenario 2
|
||||
- [ ] T017 [US4] Enforce channel deletion cascade in `internal/channels/store.go` — when a channel with open tasks is deleted, transition all open tasks to `cancelled` status before removing the channel per edge case requirement; emit trace entries for each cancelled task
|
||||
|
||||
**Checkpoint**: User Story 4 fully functional — channel types enforce correct interaction patterns with clear error messages
|
||||
|
||||
---
|
||||
|
||||
## Phase 5: User Story 2 — Task Auction Workflow (Priority: P2)
|
||||
|
||||
**Goal**: Agents post tasks, bid, and coordinate work assignment through an auction lifecycle
|
||||
|
||||
**Independent Test**: Create an auction channel, post a task, have agents bid, accept a bid, complete the task — verify full lifecycle
|
||||
|
||||
### Implementation for User Story 2
|
||||
|
||||
- [ ] T018 [US2] Implement `post_task` MCP tool in `internal/mcp/auction.go` — accepts `channel_id`, `title`, `description`, `requirements` (JSON), `deadline` (ISO 8601); validates channel is type `auction`; rejects deadline in the past per edge case; creates task with status `open`; traces via `TraceTaskPosted`; notifies channel members per FR-004
|
||||
- [ ] T019 [US2] Implement `bid_task` MCP tool in `internal/mcp/auction.go` — accepts `task_id`, `time_estimate_seconds` (int), `confidence` (float 0.0-1.0), `capabilities` (JSON); validates task is `open`; rejects if bidder is the task poster per FR-015; rejects bid on assigned/completed/expired tasks per edge case; records bid and traces via `TraceBidSubmitted` per FR-005
|
||||
- [ ] T020 [US2] Implement `accept_bid` MCP tool in `internal/mcp/auction.go` — accepts `task_id` and `bid_id`; validates caller is the task poster per FR-006; validates task is `open` (rejects if expired per edge case); transitions task to `assigned`, winning bid to `accepted`, all other bids to `rejected`; notifies winner and rejected bidders; traces via `TraceBidAccepted`
|
||||
- [ ] T021 [US2] Implement `complete_task` MCP tool in `internal/mcp/auction.go` — accepts `task_id` and `result` (JSON); validates caller is the assigned agent per FR-007; rejects if called by non-assigned agent per edge case; handles idempotent completion (same agent, same result returns success) per edge case; transitions task to `completed`; stores result; traces via `TraceTaskCompleted`
|
||||
- [ ] T022 [US2] Implement task expiration background goroutine in `internal/channels/expiry.go` — periodic check (every 30 seconds) for tasks past their deadline with status `open`; transitions expired tasks to `expired` status; posts system message to the auction channel per FR-014; traces via `TraceTaskExpired`; must detect and update within 60 seconds of deadline per SC-005
|
||||
- [ ] T023 [US2] Register auction MCP tools in `internal/mcp/server.go` — wire up `post_task`, `bid_task`, `accept_bid`, `complete_task` into MCP tool registry with JSON Schema descriptions per Principle II; include parameter validation schemas (confidence range, ISO 8601 format, etc.)
|
||||
|
||||
**Checkpoint**: User Story 2 fully functional — complete auction lifecycle works end-to-end across multiple agents
|
||||
|
||||
---
|
||||
|
||||
## Phase 6: User Story 3 — Agent Discovery by Capability (Priority: P3)
|
||||
|
||||
**Goal**: Agents find collaborators by searching capability cards using keyword or semantic matching
|
||||
|
||||
**Independent Test**: Register agents with different capability cards, call `discover_agents` with various queries, verify relevant results returned
|
||||
|
||||
### Implementation for User Story 3
|
||||
|
||||
- [ ] T024 [US3] Implement `discover_agents` MCP tool in `internal/mcp/discovery.go` — accepts `query` (string) and optional `limit` (int, default 10) per FR-008; calls agent store's `SearchAgents` for FTS5 keyword search; returns matching agents with capability cards sorted by relevance score; returns empty result set with clear indication when no agents match per edge case; traces via `TraceDiscoverAgents` per FR-013
|
||||
- [ ] T025 [US3] Implement semantic search path in `internal/mcp/discovery.go` — when an embedding provider is configured (`SYNAPBUS_EMBEDDING_PROVIDER` env var), embed the query and search agent capability card embeddings via HNSW index; merge semantic results with FTS results, using semantic score as a tiebreaker; when no embedding provider is configured, fall back to FTS5-only search per FR-012 and Principle VI
|
||||
- [ ] T026 [US3] Add capability card embedding pipeline in `internal/agents/embeddings.go` — when an agent registers or updates capability card fields AND an embedding provider is configured, asynchronously compute and store an embedding vector for the agent's combined skills + description text; store in HNSW index keyed by agent ID; skip silently when no provider is configured per Principle VI
|
||||
- [ ] T027 [US3] Register discovery MCP tool in `internal/mcp/server.go` — wire up `discover_agents` into MCP tool registry with JSON Schema description per Principle II; document that semantic search is optional and requires embedding provider configuration
|
||||
|
||||
**Checkpoint**: User Story 3 fully functional — agents can discover collaborators by keyword or semantic search
|
||||
|
||||
---
|
||||
|
||||
## Phase 7: Polish & Cross-Cutting Concerns
|
||||
|
||||
**Purpose**: Integration testing, performance validation, and cross-story consistency
|
||||
|
||||
- [ ] T028 [P] Validate SC-001 — write benchmark test in `internal/channels/blackboard_bench_test.go`: create blackboard channel, post 100+ tagged messages, verify filtered read returns correct results in under 500ms
|
||||
- [ ] T029 [P] Validate SC-002 — write integration test in `internal/channels/auction_integration_test.go`: execute full auction lifecycle (post_task -> bid x2 -> accept_bid -> complete_task) across 3 agents, verify all state transitions complete in under 5 seconds
|
||||
- [ ] T030 [P] Validate SC-003 — write benchmark test in `internal/agents/discovery_bench_test.go`: register 1000 agents with varied capability cards, verify `discover_agents` FTS5 fallback returns results within 200ms
|
||||
- [ ] T031 [P] Validate SC-006 — write integration test in `internal/trace/swarm_integration_test.go`: execute one of each swarm operation, verify each produces a trace entry queryable via the trace store
|
||||
- [ ] T032 Validate SC-007 — write integration test in `internal/mcp/swarm_no_embeddings_test.go`: run all swarm operations with no embedding provider configured, verify `discover_agents` falls back to FTS and all other operations function correctly
|
||||
- [ ] T033 Run `make lint` and fix any linting issues across all new files
|
||||
- [ ] T034 Run `make test` and ensure all existing tests still pass (no regressions)
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & Execution Order
|
||||
|
||||
### Phase Dependencies
|
||||
|
||||
- **Setup (Phase 1)**: No dependencies — can start immediately
|
||||
- **Foundational (Phase 2)**: Depends on Phase 1 completion — BLOCKS all user stories
|
||||
- **US1 Stigmergy (Phase 3)**: Depends on Phase 2 — can start immediately after foundation
|
||||
- **US4 Channel Type Enforcement (Phase 4)**: Depends on Phase 2 — can run in parallel with Phase 3
|
||||
- **US2 Task Auction (Phase 5)**: Depends on Phase 2; benefits from Phase 4 (enforcement) being complete but T016 can integrate enforcement inline
|
||||
- **US3 Agent Discovery (Phase 6)**: Depends on Phase 2 (agent store extensions); can run in parallel with Phases 3-5
|
||||
- **Polish (Phase 7)**: Depends on all user stories being complete
|
||||
|
||||
### User Story Dependencies
|
||||
|
||||
- **US1 (P1)**: Can start after Phase 2 — no dependencies on other stories
|
||||
- **US4 (P2)**: Can start after Phase 2 — no dependencies on other stories; provides enforcement used by US2
|
||||
- **US2 (P2)**: Can start after Phase 2 — uses enforcement from US4 but can implement inline if US4 is not yet complete
|
||||
- **US3 (P3)**: Can start after Phase 2 — fully independent of US1, US2, US4
|
||||
|
||||
### Within Each User Story
|
||||
|
||||
- Domain types (Phase 1) before store layer (Phase 2)
|
||||
- Store layer before MCP tool implementation
|
||||
- MCP tool implementation before MCP server registration
|
||||
- Core logic before edge case handling
|
||||
- All operations must emit traces before the story is considered complete
|
||||
|
||||
### Parallel Opportunities
|
||||
|
||||
- Phase 1: T002, T003, T004, T005 are all [P] — different files, no dependencies
|
||||
- Phase 2: T007, T008, T009, T011 are all [P] — different packages
|
||||
- Once Phase 2 completes: US1 (Phase 3), US4 (Phase 4), and US3 (Phase 6) can start in parallel
|
||||
- US2 (Phase 5) can start in parallel but benefits from US4 being complete first
|
||||
- Phase 7: T028, T029, T030, T031 are all [P] — independent test files
|
||||
@@ -0,0 +1,28 @@
|
||||
# [PROJECT NAME] Development Guidelines
|
||||
|
||||
Auto-generated from all feature plans. Last updated: [DATE]
|
||||
|
||||
## Active Technologies
|
||||
|
||||
[EXTRACTED FROM ALL PLAN.MD FILES]
|
||||
|
||||
## Project Structure
|
||||
|
||||
```text
|
||||
[ACTUAL STRUCTURE FROM PLANS]
|
||||
```
|
||||
|
||||
## Commands
|
||||
|
||||
[ONLY COMMANDS FOR ACTIVE TECHNOLOGIES]
|
||||
|
||||
## Code Style
|
||||
|
||||
[LANGUAGE-SPECIFIC, ONLY FOR LANGUAGES IN USE]
|
||||
|
||||
## Recent Changes
|
||||
|
||||
[LAST 3 FEATURES AND WHAT THEY ADDED]
|
||||
|
||||
<!-- MANUAL ADDITIONS START -->
|
||||
<!-- MANUAL ADDITIONS END -->
|
||||
@@ -0,0 +1,40 @@
|
||||
# [CHECKLIST TYPE] Checklist: [FEATURE NAME]
|
||||
|
||||
**Purpose**: [Brief description of what this checklist covers]
|
||||
**Created**: [DATE]
|
||||
**Feature**: [Link to spec.md or relevant documentation]
|
||||
|
||||
**Note**: This checklist is generated by the `/speckit.checklist` command based on feature context and requirements.
|
||||
|
||||
<!--
|
||||
============================================================================
|
||||
IMPORTANT: The checklist items below are SAMPLE ITEMS for illustration only.
|
||||
|
||||
The /speckit.checklist command MUST replace these with actual items based on:
|
||||
- User's specific checklist request
|
||||
- Feature requirements from spec.md
|
||||
- Technical context from plan.md
|
||||
- Implementation details from tasks.md
|
||||
|
||||
DO NOT keep these sample items in the generated checklist file.
|
||||
============================================================================
|
||||
-->
|
||||
|
||||
## [Category 1]
|
||||
|
||||
- [ ] CHK001 First checklist item with clear action
|
||||
- [ ] CHK002 Second checklist item
|
||||
- [ ] CHK003 Third checklist item
|
||||
|
||||
## [Category 2]
|
||||
|
||||
- [ ] CHK004 Another category item
|
||||
- [ ] CHK005 Item with specific criteria
|
||||
- [ ] CHK006 Final item in this category
|
||||
|
||||
## Notes
|
||||
|
||||
- Check items off as completed: `[x]`
|
||||
- Add comments or findings inline
|
||||
- Link to relevant resources or documentation
|
||||
- Items are numbered sequentially for easy reference
|
||||
@@ -0,0 +1,50 @@
|
||||
# [PROJECT_NAME] Constitution
|
||||
<!-- Example: Spec Constitution, TaskFlow Constitution, etc. -->
|
||||
|
||||
## Core Principles
|
||||
|
||||
### [PRINCIPLE_1_NAME]
|
||||
<!-- Example: I. Library-First -->
|
||||
[PRINCIPLE_1_DESCRIPTION]
|
||||
<!-- Example: Every feature starts as a standalone library; Libraries must be self-contained, independently testable, documented; Clear purpose required - no organizational-only libraries -->
|
||||
|
||||
### [PRINCIPLE_2_NAME]
|
||||
<!-- Example: II. CLI Interface -->
|
||||
[PRINCIPLE_2_DESCRIPTION]
|
||||
<!-- Example: Every library exposes functionality via CLI; Text in/out protocol: stdin/args → stdout, errors → stderr; Support JSON + human-readable formats -->
|
||||
|
||||
### [PRINCIPLE_3_NAME]
|
||||
<!-- Example: III. Test-First (NON-NEGOTIABLE) -->
|
||||
[PRINCIPLE_3_DESCRIPTION]
|
||||
<!-- Example: TDD mandatory: Tests written → User approved → Tests fail → Then implement; Red-Green-Refactor cycle strictly enforced -->
|
||||
|
||||
### [PRINCIPLE_4_NAME]
|
||||
<!-- Example: IV. Integration Testing -->
|
||||
[PRINCIPLE_4_DESCRIPTION]
|
||||
<!-- Example: Focus areas requiring integration tests: New library contract tests, Contract changes, Inter-service communication, Shared schemas -->
|
||||
|
||||
### [PRINCIPLE_5_NAME]
|
||||
<!-- Example: V. Observability, VI. Versioning & Breaking Changes, VII. Simplicity -->
|
||||
[PRINCIPLE_5_DESCRIPTION]
|
||||
<!-- Example: Text I/O ensures debuggability; Structured logging required; Or: MAJOR.MINOR.BUILD format; Or: Start simple, YAGNI principles -->
|
||||
|
||||
## [SECTION_2_NAME]
|
||||
<!-- Example: Additional Constraints, Security Requirements, Performance Standards, etc. -->
|
||||
|
||||
[SECTION_2_CONTENT]
|
||||
<!-- Example: Technology stack requirements, compliance standards, deployment policies, etc. -->
|
||||
|
||||
## [SECTION_3_NAME]
|
||||
<!-- Example: Development Workflow, Review Process, Quality Gates, etc. -->
|
||||
|
||||
[SECTION_3_CONTENT]
|
||||
<!-- Example: Code review requirements, testing gates, deployment approval process, etc. -->
|
||||
|
||||
## Governance
|
||||
<!-- Example: Constitution supersedes all other practices; Amendments require documentation, approval, migration plan -->
|
||||
|
||||
[GOVERNANCE_RULES]
|
||||
<!-- Example: All PRs/reviews must verify compliance; Complexity must be justified; Use [GUIDANCE_FILE] for runtime development guidance -->
|
||||
|
||||
**Version**: [CONSTITUTION_VERSION] | **Ratified**: [RATIFICATION_DATE] | **Last Amended**: [LAST_AMENDED_DATE]
|
||||
<!-- Example: Version: 2.1.1 | Ratified: 2025-06-13 | Last Amended: 2025-07-16 -->
|
||||
@@ -0,0 +1,104 @@
|
||||
# Implementation Plan: [FEATURE]
|
||||
|
||||
**Branch**: `[###-feature-name]` | **Date**: [DATE] | **Spec**: [link]
|
||||
**Input**: Feature specification from `/specs/[###-feature-name]/spec.md`
|
||||
|
||||
**Note**: This template is filled in by the `/speckit.plan` command. See `.specify/templates/plan-template.md` for the execution workflow.
|
||||
|
||||
## Summary
|
||||
|
||||
[Extract from feature spec: primary requirement + technical approach from research]
|
||||
|
||||
## Technical Context
|
||||
|
||||
<!--
|
||||
ACTION REQUIRED: Replace the content in this section with the technical details
|
||||
for the project. The structure here is presented in advisory capacity to guide
|
||||
the iteration process.
|
||||
-->
|
||||
|
||||
**Language/Version**: [e.g., Python 3.11, Swift 5.9, Rust 1.75 or NEEDS CLARIFICATION]
|
||||
**Primary Dependencies**: [e.g., FastAPI, UIKit, LLVM or NEEDS CLARIFICATION]
|
||||
**Storage**: [if applicable, e.g., PostgreSQL, CoreData, files or N/A]
|
||||
**Testing**: [e.g., pytest, XCTest, cargo test or NEEDS CLARIFICATION]
|
||||
**Target Platform**: [e.g., Linux server, iOS 15+, WASM or NEEDS CLARIFICATION]
|
||||
**Project Type**: [e.g., library/cli/web-service/mobile-app/compiler/desktop-app or NEEDS CLARIFICATION]
|
||||
**Performance Goals**: [domain-specific, e.g., 1000 req/s, 10k lines/sec, 60 fps or NEEDS CLARIFICATION]
|
||||
**Constraints**: [domain-specific, e.g., <200ms p95, <100MB memory, offline-capable or NEEDS CLARIFICATION]
|
||||
**Scale/Scope**: [domain-specific, e.g., 10k users, 1M LOC, 50 screens or NEEDS CLARIFICATION]
|
||||
|
||||
## Constitution Check
|
||||
|
||||
*GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.*
|
||||
|
||||
[Gates determined based on constitution file]
|
||||
|
||||
## Project Structure
|
||||
|
||||
### Documentation (this feature)
|
||||
|
||||
```text
|
||||
specs/[###-feature]/
|
||||
├── plan.md # This file (/speckit.plan command output)
|
||||
├── research.md # Phase 0 output (/speckit.plan command)
|
||||
├── data-model.md # Phase 1 output (/speckit.plan command)
|
||||
├── quickstart.md # Phase 1 output (/speckit.plan command)
|
||||
├── contracts/ # Phase 1 output (/speckit.plan command)
|
||||
└── tasks.md # Phase 2 output (/speckit.tasks command - NOT created by /speckit.plan)
|
||||
```
|
||||
|
||||
### Source Code (repository root)
|
||||
<!--
|
||||
ACTION REQUIRED: Replace the placeholder tree below with the concrete layout
|
||||
for this feature. Delete unused options and expand the chosen structure with
|
||||
real paths (e.g., apps/admin, packages/something). The delivered plan must
|
||||
not include Option labels.
|
||||
-->
|
||||
|
||||
```text
|
||||
# [REMOVE IF UNUSED] Option 1: Single project (DEFAULT)
|
||||
src/
|
||||
├── models/
|
||||
├── services/
|
||||
├── cli/
|
||||
└── lib/
|
||||
|
||||
tests/
|
||||
├── contract/
|
||||
├── integration/
|
||||
└── unit/
|
||||
|
||||
# [REMOVE IF UNUSED] Option 2: Web application (when "frontend" + "backend" detected)
|
||||
backend/
|
||||
├── src/
|
||||
│ ├── models/
|
||||
│ ├── services/
|
||||
│ └── api/
|
||||
└── tests/
|
||||
|
||||
frontend/
|
||||
├── src/
|
||||
│ ├── components/
|
||||
│ ├── pages/
|
||||
│ └── services/
|
||||
└── tests/
|
||||
|
||||
# [REMOVE IF UNUSED] Option 3: Mobile + API (when "iOS/Android" detected)
|
||||
api/
|
||||
└── [same as backend above]
|
||||
|
||||
ios/ or android/
|
||||
└── [platform-specific structure: feature modules, UI flows, platform tests]
|
||||
```
|
||||
|
||||
**Structure Decision**: [Document the selected structure and reference the real
|
||||
directories captured above]
|
||||
|
||||
## Complexity Tracking
|
||||
|
||||
> **Fill ONLY if Constitution Check has violations that must be justified**
|
||||
|
||||
| Violation | Why Needed | Simpler Alternative Rejected Because |
|
||||
|-----------|------------|-------------------------------------|
|
||||
| [e.g., 4th project] | [current need] | [why 3 projects insufficient] |
|
||||
| [e.g., Repository pattern] | [specific problem] | [why direct DB access insufficient] |
|
||||
@@ -0,0 +1,115 @@
|
||||
# Feature Specification: [FEATURE NAME]
|
||||
|
||||
**Feature Branch**: `[###-feature-name]`
|
||||
**Created**: [DATE]
|
||||
**Status**: Draft
|
||||
**Input**: User description: "$ARGUMENTS"
|
||||
|
||||
## User Scenarios & Testing *(mandatory)*
|
||||
|
||||
<!--
|
||||
IMPORTANT: User stories should be PRIORITIZED as user journeys ordered by importance.
|
||||
Each user story/journey must be INDEPENDENTLY TESTABLE - meaning if you implement just ONE of them,
|
||||
you should still have a viable MVP (Minimum Viable Product) that delivers value.
|
||||
|
||||
Assign priorities (P1, P2, P3, etc.) to each story, where P1 is the most critical.
|
||||
Think of each story as a standalone slice of functionality that can be:
|
||||
- Developed independently
|
||||
- Tested independently
|
||||
- Deployed independently
|
||||
- Demonstrated to users independently
|
||||
-->
|
||||
|
||||
### User Story 1 - [Brief Title] (Priority: P1)
|
||||
|
||||
[Describe this user journey in plain language]
|
||||
|
||||
**Why this priority**: [Explain the value and why it has this priority level]
|
||||
|
||||
**Independent Test**: [Describe how this can be tested independently - e.g., "Can be fully tested by [specific action] and delivers [specific value]"]
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** [initial state], **When** [action], **Then** [expected outcome]
|
||||
2. **Given** [initial state], **When** [action], **Then** [expected outcome]
|
||||
|
||||
---
|
||||
|
||||
### User Story 2 - [Brief Title] (Priority: P2)
|
||||
|
||||
[Describe this user journey in plain language]
|
||||
|
||||
**Why this priority**: [Explain the value and why it has this priority level]
|
||||
|
||||
**Independent Test**: [Describe how this can be tested independently]
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** [initial state], **When** [action], **Then** [expected outcome]
|
||||
|
||||
---
|
||||
|
||||
### User Story 3 - [Brief Title] (Priority: P3)
|
||||
|
||||
[Describe this user journey in plain language]
|
||||
|
||||
**Why this priority**: [Explain the value and why it has this priority level]
|
||||
|
||||
**Independent Test**: [Describe how this can be tested independently]
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** [initial state], **When** [action], **Then** [expected outcome]
|
||||
|
||||
---
|
||||
|
||||
[Add more user stories as needed, each with an assigned priority]
|
||||
|
||||
### Edge Cases
|
||||
|
||||
<!--
|
||||
ACTION REQUIRED: The content in this section represents placeholders.
|
||||
Fill them out with the right edge cases.
|
||||
-->
|
||||
|
||||
- What happens when [boundary condition]?
|
||||
- How does system handle [error scenario]?
|
||||
|
||||
## Requirements *(mandatory)*
|
||||
|
||||
<!--
|
||||
ACTION REQUIRED: The content in this section represents placeholders.
|
||||
Fill them out with the right functional requirements.
|
||||
-->
|
||||
|
||||
### Functional Requirements
|
||||
|
||||
- **FR-001**: System MUST [specific capability, e.g., "allow users to create accounts"]
|
||||
- **FR-002**: System MUST [specific capability, e.g., "validate email addresses"]
|
||||
- **FR-003**: Users MUST be able to [key interaction, e.g., "reset their password"]
|
||||
- **FR-004**: System MUST [data requirement, e.g., "persist user preferences"]
|
||||
- **FR-005**: System MUST [behavior, e.g., "log all security events"]
|
||||
|
||||
*Example of marking unclear requirements:*
|
||||
|
||||
- **FR-006**: System MUST authenticate users via [NEEDS CLARIFICATION: auth method not specified - email/password, SSO, OAuth?]
|
||||
- **FR-007**: System MUST retain user data for [NEEDS CLARIFICATION: retention period not specified]
|
||||
|
||||
### Key Entities *(include if feature involves data)*
|
||||
|
||||
- **[Entity 1]**: [What it represents, key attributes without implementation]
|
||||
- **[Entity 2]**: [What it represents, relationships to other entities]
|
||||
|
||||
## Success Criteria *(mandatory)*
|
||||
|
||||
<!--
|
||||
ACTION REQUIRED: Define measurable success criteria.
|
||||
These must be technology-agnostic and measurable.
|
||||
-->
|
||||
|
||||
### Measurable Outcomes
|
||||
|
||||
- **SC-001**: [Measurable metric, e.g., "Users can complete account creation in under 2 minutes"]
|
||||
- **SC-002**: [Measurable metric, e.g., "System handles 1000 concurrent users without degradation"]
|
||||
- **SC-003**: [User satisfaction metric, e.g., "90% of users successfully complete primary task on first attempt"]
|
||||
- **SC-004**: [Business metric, e.g., "Reduce support tickets related to [X] by 50%"]
|
||||
@@ -0,0 +1,251 @@
|
||||
---
|
||||
|
||||
description: "Task list template for feature implementation"
|
||||
---
|
||||
|
||||
# Tasks: [FEATURE NAME]
|
||||
|
||||
**Input**: Design documents from `/specs/[###-feature-name]/`
|
||||
**Prerequisites**: plan.md (required), spec.md (required for user stories), research.md, data-model.md, contracts/
|
||||
|
||||
**Tests**: The examples below include test tasks. Tests are OPTIONAL - only include them if explicitly requested in the feature specification.
|
||||
|
||||
**Organization**: Tasks are grouped by user story to enable independent implementation and testing of each story.
|
||||
|
||||
## Format: `[ID] [P?] [Story] Description`
|
||||
|
||||
- **[P]**: Can run in parallel (different files, no dependencies)
|
||||
- **[Story]**: Which user story this task belongs to (e.g., US1, US2, US3)
|
||||
- Include exact file paths in descriptions
|
||||
|
||||
## Path Conventions
|
||||
|
||||
- **Single project**: `src/`, `tests/` at repository root
|
||||
- **Web app**: `backend/src/`, `frontend/src/`
|
||||
- **Mobile**: `api/src/`, `ios/src/` or `android/src/`
|
||||
- Paths shown below assume single project - adjust based on plan.md structure
|
||||
|
||||
<!--
|
||||
============================================================================
|
||||
IMPORTANT: The tasks below are SAMPLE TASKS for illustration purposes only.
|
||||
|
||||
The /speckit.tasks command MUST replace these with actual tasks based on:
|
||||
- User stories from spec.md (with their priorities P1, P2, P3...)
|
||||
- Feature requirements from plan.md
|
||||
- Entities from data-model.md
|
||||
- Endpoints from contracts/
|
||||
|
||||
Tasks MUST be organized by user story so each story can be:
|
||||
- Implemented independently
|
||||
- Tested independently
|
||||
- Delivered as an MVP increment
|
||||
|
||||
DO NOT keep these sample tasks in the generated tasks.md file.
|
||||
============================================================================
|
||||
-->
|
||||
|
||||
## Phase 1: Setup (Shared Infrastructure)
|
||||
|
||||
**Purpose**: Project initialization and basic structure
|
||||
|
||||
- [ ] T001 Create project structure per implementation plan
|
||||
- [ ] T002 Initialize [language] project with [framework] dependencies
|
||||
- [ ] T003 [P] Configure linting and formatting tools
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Foundational (Blocking Prerequisites)
|
||||
|
||||
**Purpose**: Core infrastructure that MUST be complete before ANY user story can be implemented
|
||||
|
||||
**⚠️ CRITICAL**: No user story work can begin until this phase is complete
|
||||
|
||||
Examples of foundational tasks (adjust based on your project):
|
||||
|
||||
- [ ] T004 Setup database schema and migrations framework
|
||||
- [ ] T005 [P] Implement authentication/authorization framework
|
||||
- [ ] T006 [P] Setup API routing and middleware structure
|
||||
- [ ] T007 Create base models/entities that all stories depend on
|
||||
- [ ] T008 Configure error handling and logging infrastructure
|
||||
- [ ] T009 Setup environment configuration management
|
||||
|
||||
**Checkpoint**: Foundation ready - user story implementation can now begin in parallel
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: User Story 1 - [Title] (Priority: P1) 🎯 MVP
|
||||
|
||||
**Goal**: [Brief description of what this story delivers]
|
||||
|
||||
**Independent Test**: [How to verify this story works on its own]
|
||||
|
||||
### Tests for User Story 1 (OPTIONAL - only if tests requested) ⚠️
|
||||
|
||||
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
|
||||
|
||||
- [ ] T010 [P] [US1] Contract test for [endpoint] in tests/contract/test_[name].py
|
||||
- [ ] T011 [P] [US1] Integration test for [user journey] in tests/integration/test_[name].py
|
||||
|
||||
### Implementation for User Story 1
|
||||
|
||||
- [ ] T012 [P] [US1] Create [Entity1] model in src/models/[entity1].py
|
||||
- [ ] T013 [P] [US1] Create [Entity2] model in src/models/[entity2].py
|
||||
- [ ] T014 [US1] Implement [Service] in src/services/[service].py (depends on T012, T013)
|
||||
- [ ] T015 [US1] Implement [endpoint/feature] in src/[location]/[file].py
|
||||
- [ ] T016 [US1] Add validation and error handling
|
||||
- [ ] T017 [US1] Add logging for user story 1 operations
|
||||
|
||||
**Checkpoint**: At this point, User Story 1 should be fully functional and testable independently
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: User Story 2 - [Title] (Priority: P2)
|
||||
|
||||
**Goal**: [Brief description of what this story delivers]
|
||||
|
||||
**Independent Test**: [How to verify this story works on its own]
|
||||
|
||||
### Tests for User Story 2 (OPTIONAL - only if tests requested) ⚠️
|
||||
|
||||
- [ ] T018 [P] [US2] Contract test for [endpoint] in tests/contract/test_[name].py
|
||||
- [ ] T019 [P] [US2] Integration test for [user journey] in tests/integration/test_[name].py
|
||||
|
||||
### Implementation for User Story 2
|
||||
|
||||
- [ ] T020 [P] [US2] Create [Entity] model in src/models/[entity].py
|
||||
- [ ] T021 [US2] Implement [Service] in src/services/[service].py
|
||||
- [ ] T022 [US2] Implement [endpoint/feature] in src/[location]/[file].py
|
||||
- [ ] T023 [US2] Integrate with User Story 1 components (if needed)
|
||||
|
||||
**Checkpoint**: At this point, User Stories 1 AND 2 should both work independently
|
||||
|
||||
---
|
||||
|
||||
## Phase 5: User Story 3 - [Title] (Priority: P3)
|
||||
|
||||
**Goal**: [Brief description of what this story delivers]
|
||||
|
||||
**Independent Test**: [How to verify this story works on its own]
|
||||
|
||||
### Tests for User Story 3 (OPTIONAL - only if tests requested) ⚠️
|
||||
|
||||
- [ ] T024 [P] [US3] Contract test for [endpoint] in tests/contract/test_[name].py
|
||||
- [ ] T025 [P] [US3] Integration test for [user journey] in tests/integration/test_[name].py
|
||||
|
||||
### Implementation for User Story 3
|
||||
|
||||
- [ ] T026 [P] [US3] Create [Entity] model in src/models/[entity].py
|
||||
- [ ] T027 [US3] Implement [Service] in src/services/[service].py
|
||||
- [ ] T028 [US3] Implement [endpoint/feature] in src/[location]/[file].py
|
||||
|
||||
**Checkpoint**: All user stories should now be independently functional
|
||||
|
||||
---
|
||||
|
||||
[Add more user story phases as needed, following the same pattern]
|
||||
|
||||
---
|
||||
|
||||
## Phase N: Polish & Cross-Cutting Concerns
|
||||
|
||||
**Purpose**: Improvements that affect multiple user stories
|
||||
|
||||
- [ ] TXXX [P] Documentation updates in docs/
|
||||
- [ ] TXXX Code cleanup and refactoring
|
||||
- [ ] TXXX Performance optimization across all stories
|
||||
- [ ] TXXX [P] Additional unit tests (if requested) in tests/unit/
|
||||
- [ ] TXXX Security hardening
|
||||
- [ ] TXXX Run quickstart.md validation
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & Execution Order
|
||||
|
||||
### Phase Dependencies
|
||||
|
||||
- **Setup (Phase 1)**: No dependencies - can start immediately
|
||||
- **Foundational (Phase 2)**: Depends on Setup completion - BLOCKS all user stories
|
||||
- **User Stories (Phase 3+)**: All depend on Foundational phase completion
|
||||
- User stories can then proceed in parallel (if staffed)
|
||||
- Or sequentially in priority order (P1 → P2 → P3)
|
||||
- **Polish (Final Phase)**: Depends on all desired user stories being complete
|
||||
|
||||
### User Story Dependencies
|
||||
|
||||
- **User Story 1 (P1)**: Can start after Foundational (Phase 2) - No dependencies on other stories
|
||||
- **User Story 2 (P2)**: Can start after Foundational (Phase 2) - May integrate with US1 but should be independently testable
|
||||
- **User Story 3 (P3)**: Can start after Foundational (Phase 2) - May integrate with US1/US2 but should be independently testable
|
||||
|
||||
### Within Each User Story
|
||||
|
||||
- Tests (if included) MUST be written and FAIL before implementation
|
||||
- Models before services
|
||||
- Services before endpoints
|
||||
- Core implementation before integration
|
||||
- Story complete before moving to next priority
|
||||
|
||||
### Parallel Opportunities
|
||||
|
||||
- All Setup tasks marked [P] can run in parallel
|
||||
- All Foundational tasks marked [P] can run in parallel (within Phase 2)
|
||||
- Once Foundational phase completes, all user stories can start in parallel (if team capacity allows)
|
||||
- All tests for a user story marked [P] can run in parallel
|
||||
- Models within a story marked [P] can run in parallel
|
||||
- Different user stories can be worked on in parallel by different team members
|
||||
|
||||
---
|
||||
|
||||
## Parallel Example: User Story 1
|
||||
|
||||
```bash
|
||||
# Launch all tests for User Story 1 together (if tests requested):
|
||||
Task: "Contract test for [endpoint] in tests/contract/test_[name].py"
|
||||
Task: "Integration test for [user journey] in tests/integration/test_[name].py"
|
||||
|
||||
# Launch all models for User Story 1 together:
|
||||
Task: "Create [Entity1] model in src/models/[entity1].py"
|
||||
Task: "Create [Entity2] model in src/models/[entity2].py"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Implementation Strategy
|
||||
|
||||
### MVP First (User Story 1 Only)
|
||||
|
||||
1. Complete Phase 1: Setup
|
||||
2. Complete Phase 2: Foundational (CRITICAL - blocks all stories)
|
||||
3. Complete Phase 3: User Story 1
|
||||
4. **STOP and VALIDATE**: Test User Story 1 independently
|
||||
5. Deploy/demo if ready
|
||||
|
||||
### Incremental Delivery
|
||||
|
||||
1. Complete Setup + Foundational → Foundation ready
|
||||
2. Add User Story 1 → Test independently → Deploy/Demo (MVP!)
|
||||
3. Add User Story 2 → Test independently → Deploy/Demo
|
||||
4. Add User Story 3 → Test independently → Deploy/Demo
|
||||
5. Each story adds value without breaking previous stories
|
||||
|
||||
### Parallel Team Strategy
|
||||
|
||||
With multiple developers:
|
||||
|
||||
1. Team completes Setup + Foundational together
|
||||
2. Once Foundational is done:
|
||||
- Developer A: User Story 1
|
||||
- Developer B: User Story 2
|
||||
- Developer C: User Story 3
|
||||
3. Stories complete and integrate independently
|
||||
|
||||
---
|
||||
|
||||
## Notes
|
||||
|
||||
- [P] tasks = different files, no dependencies
|
||||
- [Story] label maps task to specific user story for traceability
|
||||
- Each user story should be independently completable and testable
|
||||
- Verify tests fail before implementing
|
||||
- Commit after each task or logical group
|
||||
- Stop at any checkpoint to validate story independently
|
||||
- Avoid: vague tasks, same file conflicts, cross-story dependencies that break independence
|
||||
@@ -0,0 +1,198 @@
|
||||
# SynapBus Bootstrap Prompt
|
||||
|
||||
Paste everything below the line into Claude in the `/Users/user/repos/synapbus` directory.
|
||||
|
||||
---
|
||||
|
||||
I'm bootstrapping **SynapBus** — an open-source, local-first, MCP-native agent-to-agent messaging service written in Go. Single binary, embedded storage, semantic search, Slack-like Web UI.
|
||||
|
||||
The project idea is fully described in `IDEA.md` — read it first.
|
||||
|
||||
The repo is at https://github.com/smart-mcp-proxy/synapbus (public, empty).
|
||||
Domains: synapbus.com + synapbus.dev (to be registered).
|
||||
|
||||
## What I need you to do (in order):
|
||||
|
||||
### 1. Create CLAUDE.md
|
||||
|
||||
Write the project CLAUDE.md with:
|
||||
- Project overview (from IDEA.md)
|
||||
- Tech stack: Go 1.23+, modernc.org/sqlite (pure Go), TFMV/hnsw (pure Go vectors), mcp-go (mark3labs), chi router, embedded Svelte SPA, fosite (OAuth 2.1)
|
||||
- Directory structure (Go standard layout):
|
||||
```
|
||||
synapbus/
|
||||
├── cmd/synapbus/ # main.go entry point
|
||||
├── internal/
|
||||
│ ├── auth/ # OAuth 2.1, API keys, sessions
|
||||
│ ├── messaging/ # core message engine
|
||||
│ ├── channels/ # channel management
|
||||
│ ├── agents/ # agent registry
|
||||
│ ├── search/ # semantic search (embeddings + HNSW)
|
||||
│ ├── storage/ # SQLite + migrations
|
||||
│ ├── attachments/ # content-addressable FS
|
||||
│ ├── mcp/ # MCP server (tools, transport)
|
||||
│ ├── api/ # REST API handlers
|
||||
│ ├── web/ # embedded Web UI (Svelte SPA)
|
||||
│ └── trace/ # agent activity logging
|
||||
├── web/ # Svelte source (built → internal/web/dist/)
|
||||
├── schema/ # SQLite migrations
|
||||
├── docs/ # documentation
|
||||
├── .specify/ # speckit specs
|
||||
└── Makefile
|
||||
```
|
||||
- Build commands: `make build`, `make test`, `make dev`, `make web` (build Svelte SPA)
|
||||
- Run: `./synapbus serve --port 8080 --data ./data`
|
||||
- Conventions: Go standard, `internal/` for non-public packages, table-driven tests, context propagation, structured logging (slog)
|
||||
- Environment variables: `SYNAPBUS_PORT`, `SYNAPBUS_DATA_DIR`, `SYNAPBUS_EMBEDDING_PROVIDER` (openai/gemini/ollama), `SYNAPBUS_EMBEDDING_API_KEY`, `SYNAPBUS_OLLAMA_URL`
|
||||
|
||||
### 2. Create speckit constitution
|
||||
|
||||
Run `/speckit.constitution` and provide these architecture decisions:
|
||||
|
||||
**Principles:**
|
||||
1. **Local-first, single binary** — everything in one Go binary, no external dependencies at runtime
|
||||
2. **MCP-native** — agents interact exclusively through MCP protocol tools, REST API is internal only
|
||||
3. **Pure Go, zero CGO** — all dependencies must be pure Go (modernc.org/sqlite, not mattn/go-sqlite3)
|
||||
4. **Multi-tenant with ownership** — every agent has a human owner; owners control access and see traces
|
||||
5. **Embedded OAuth 2.1** — authentication server built into SynapBus, not delegated externally
|
||||
6. **Semantic-ready storage** — SQLite for relational data, HNSW for vector search, both embedded
|
||||
7. **Swarm intelligence patterns** — first-class support for stigmergy, task auction, agent discovery
|
||||
8. **Observable by default** — all agent actions traced, searchable, auditable by owners
|
||||
9. **Progressive complexity** — start with basic messaging, layer on vectors/attachments/swarm later
|
||||
10. **Web UI as first-class citizen** — embedded Svelte SPA, not an afterthought
|
||||
|
||||
**Technology decisions:**
|
||||
- Go 1.23+ (single binary, cross-compilation)
|
||||
- modernc.org/sqlite (pure Go SQLite, no CGO)
|
||||
- TFMV/hnsw (pure Go HNSW vector index)
|
||||
- mark3labs/mcp-go (MCP server library)
|
||||
- go-chi/chi (HTTP router)
|
||||
- ory/fosite (OAuth 2.1 framework)
|
||||
- Svelte 5 + Tailwind (Web UI, embedded via go:embed)
|
||||
- slog (structured logging)
|
||||
- Content-addressable filesystem for attachments (SHA-256)
|
||||
|
||||
**Non-goals:**
|
||||
- No PostgreSQL, Redis, or external DB dependency
|
||||
- No framework lock-in (LangChain, CrewAI, etc.)
|
||||
- No A2A protocol support (yet) — MCP only for v1
|
||||
- No cloud-specific features — local-first always
|
||||
|
||||
### 3. Create specs for these features (use /speckit.specify for each):
|
||||
|
||||
**Spec 001: Core Messaging**
|
||||
- Messages: send (DM or channel), read inbox, claim for processing, mark done/failed
|
||||
- Conversations: threaded, with subjects, auto-created on first message
|
||||
- Priority levels (1-10), status tracking (pending/processing/done/failed)
|
||||
- Rich metadata (JSON) on messages for filtering
|
||||
- Read/unread tracking per agent per conversation
|
||||
- SQLite storage with migrations
|
||||
- MCP tools: `send_message`, `read_inbox`, `claim_messages`, `mark_done`, `search_messages` (full-text initially)
|
||||
|
||||
**Spec 002: Agent Registry & Auth**
|
||||
- Agent self-registration via MCP tool `register_agent`
|
||||
- Each agent has: name (unique), display_name, type (ai/human), capabilities (JSON), owner_id
|
||||
- API key authentication for agents (generated on registration, returned once)
|
||||
- Agent CRUD: register, update capabilities, deregister (owner only)
|
||||
- Owner-scoped access: agents can only see own messages + joined channels
|
||||
- MCP tools: `register_agent`, `discover_agents`, `update_agent`, `deregister_agent`
|
||||
- Agent capability cards (JSON schema describing what the agent can do)
|
||||
|
||||
**Spec 003: Human Auth (OAuth 2.1)**
|
||||
- OAuth 2.1 authorization server embedded in SynapBus (using fosite)
|
||||
- Local accounts: username + password (bcrypt hashed)
|
||||
- Token endpoints: /oauth/authorize, /oauth/token, /oauth/introspect
|
||||
- Grant types: authorization_code (Web UI), client_credentials (programmatic)
|
||||
- Session management for Web UI (httponly cookies)
|
||||
- User CRUD: create account, change password, list owned agents
|
||||
- PKCE required for all authorization code flows
|
||||
- Refresh token rotation
|
||||
|
||||
**Spec 004: Channels**
|
||||
- Public channels: any registered agent can join
|
||||
- Private channels: invite-only, managed by creator
|
||||
- Channel metadata: name, description, topic, created_by
|
||||
- Membership management: join, leave, invite, kick (owner only)
|
||||
- Channel message broadcast: message sent to channel delivered to all members
|
||||
- MCP tools: `create_channel`, `join_channel`, `leave_channel`, `list_channels`, `invite_to_channel`
|
||||
|
||||
**Spec 005: Web UI**
|
||||
- Svelte 5 + Tailwind CSS embedded SPA
|
||||
- Pages: Login, Dashboard (recent messages), Conversations (thread view), Channels, Agents, Settings
|
||||
- Real-time updates via SSE
|
||||
- Compose: send DM or channel message, select recipient from dropdown
|
||||
- Search: full-text search across messages
|
||||
- Agent management: view owned agents, their traces, revoke API keys
|
||||
- Responsive, dark mode support
|
||||
- Built with `make web`, embedded in Go binary via `go:embed`
|
||||
|
||||
**Spec 006: MCP Server**
|
||||
- MCP server using mark3labs/mcp-go
|
||||
- SSE transport (primary) + Streamable HTTP transport
|
||||
- All messaging operations exposed as MCP tools
|
||||
- Tool authentication: API key in MCP request headers
|
||||
- Tool listing with JSON schema descriptions
|
||||
- Health check endpoint
|
||||
- Connection management: track connected agents
|
||||
|
||||
**Spec 007: Trace Logging & Observability**
|
||||
- All agent actions logged: tool calls, messages sent/received, channel joins, errors
|
||||
- Traces stored in SQLite with agent_name, action, details, timestamp
|
||||
- Owner can view traces for their agents via Web UI
|
||||
- Filterable by agent, action type, time range
|
||||
- Exportable as JSON/CSV
|
||||
- Optional Prometheus metrics endpoint (/metrics)
|
||||
- Structured logging (slog) to stdout
|
||||
|
||||
**Spec 008: Semantic Search**
|
||||
- Message embedding on ingest (async, configurable provider)
|
||||
- Providers: OpenAI text-embedding-3-small, Gemini embedding, Ollama (local)
|
||||
- HNSW vector index (TFMV/hnsw) for ANN search
|
||||
- Combined search: vector similarity + metadata filters + full-text
|
||||
- MCP tool: `search_messages` with query, filters, limit
|
||||
- Incremental indexing: new messages embedded in background
|
||||
- Fallback: full-text search if no embedding provider configured
|
||||
|
||||
**Spec 009: Attachments**
|
||||
- Upload files up to 50MB per message
|
||||
- Content-addressable storage: SHA-256 hash as filename, dedup
|
||||
- Store in `{data_dir}/attachments/{hash[0:2]}/{hash[2:4]}/{hash}`
|
||||
- Metadata in SQLite: hash, original_filename, size, mime_type, message_id
|
||||
- MCP tools: `upload_attachment` (returns hash), `download_attachment` (by hash)
|
||||
- Web UI: inline preview for images, download link for others
|
||||
- Garbage collection: remove orphaned attachments
|
||||
|
||||
**Spec 010: Swarm Patterns**
|
||||
- **Stigmergy (Blackboard)**: tagged messages on a shared channel that agents read and react to. Tags: `#finding`, `#task`, `#decision`, `#trace`
|
||||
- **Task Auction**: `post_task` with requirements + deadline → agents `bid_task` with capabilities + time estimate → poster selects winner → task assigned
|
||||
- **Agent Discovery**: `discover_agents` searches capability cards by keyword/semantic match
|
||||
- Channel types: `standard` (chat), `blackboard` (stigmergy), `auction` (tasks)
|
||||
- MCP tools: `post_task`, `bid_task`, `accept_bid`, `complete_task`
|
||||
|
||||
### 4. Generate tasks from specs
|
||||
|
||||
After creating all specs, run `/speckit.tasks` for each spec to generate implementation tasks.
|
||||
|
||||
### 5. Create initial project files
|
||||
|
||||
- `go.mod` with module `github.com/smart-mcp-proxy/synapbus`
|
||||
- `Makefile` with targets: build, test, dev, web, clean, lint
|
||||
- `cmd/synapbus/main.go` — cobra CLI with `serve` command (placeholder)
|
||||
- `schema/001_initial.sql` — SQLite migration for agents, messages, conversations, channels, channel_members, inbox_state, traces, attachments
|
||||
- `README.md` — project overview, installation, quick start
|
||||
- `.gitignore` — Go + Node + data directory
|
||||
- `LICENSE` — Apache 2.0
|
||||
|
||||
### 6. Push initial commit
|
||||
|
||||
Stage all files, commit with message "feat: initial SynapBus project scaffolding", push to `main` on origin.
|
||||
|
||||
---
|
||||
|
||||
**Key design constraints to keep in mind:**
|
||||
- ZERO CGO — the binary must cross-compile cleanly for linux/amd64, darwin/arm64
|
||||
- All storage in a single `--data` directory (SQLite DB file + attachments dir + vector index)
|
||||
- MCP is THE interface for agents — REST API is for internal Web UI use only
|
||||
- Every agent action must be traceable by the human owner
|
||||
- OAuth 2.1 is built INTO the binary, not a separate service
|
||||
- Web UI is built from Svelte source, embedded at compile time
|
||||
@@ -0,0 +1,89 @@
|
||||
# SynapBus
|
||||
|
||||
Local-first, MCP-native agent-to-agent messaging service. Single Go binary with embedded storage, semantic search, and a Slack-like Web UI.
|
||||
|
||||
**Repo**: github.com/smart-mcp-proxy/synapbus
|
||||
**License**: Apache 2.0
|
||||
|
||||
## Tech Stack
|
||||
|
||||
| Component | Technology | Notes |
|
||||
|-----------|-----------|-------|
|
||||
| Language | Go 1.23+ | Single binary, cross-compilation, zero CGO |
|
||||
| Database | modernc.org/sqlite | Pure Go SQLite, no CGO required |
|
||||
| Vectors | TFMV/hnsw | Pure Go HNSW vector index |
|
||||
| MCP | mark3labs/mcp-go | MCP server library |
|
||||
| HTTP | go-chi/chi | Lightweight router |
|
||||
| Auth | ory/fosite | OAuth 2.1 framework |
|
||||
| Web UI | Svelte 5 + Tailwind | Embedded via go:embed |
|
||||
| Logging | slog | Structured logging |
|
||||
| Attachments | Content-addressable FS | SHA-256 dedup |
|
||||
|
||||
**Critical constraint**: ZERO CGO. The binary must cross-compile cleanly for linux/amd64, darwin/arm64.
|
||||
|
||||
## Directory Structure
|
||||
|
||||
```
|
||||
synapbus/
|
||||
├── cmd/synapbus/ # main.go entry point (cobra CLI)
|
||||
├── internal/
|
||||
│ ├── auth/ # OAuth 2.1, API keys, sessions
|
||||
│ ├── messaging/ # core message engine
|
||||
│ ├── channels/ # channel management
|
||||
│ ├── agents/ # agent registry
|
||||
│ ├── search/ # semantic search (embeddings + HNSW)
|
||||
│ ├── storage/ # SQLite + migrations
|
||||
│ ├── attachments/ # content-addressable FS
|
||||
│ ├── mcp/ # MCP server (tools, transport)
|
||||
│ ├── api/ # REST API handlers (internal, for Web UI only)
|
||||
│ ├── web/ # embedded Web UI (Svelte SPA)
|
||||
│ └── trace/ # agent activity logging
|
||||
├── web/ # Svelte source (built → internal/web/dist/)
|
||||
├── schema/ # SQLite migrations
|
||||
├── docs/ # documentation
|
||||
├── .specify/ # speckit specs
|
||||
└── Makefile
|
||||
```
|
||||
|
||||
## Build & Run
|
||||
|
||||
```bash
|
||||
make build # Build Go binary
|
||||
make test # Run all tests
|
||||
make dev # Run with hot reload
|
||||
make web # Build Svelte SPA
|
||||
make clean # Clean build artifacts
|
||||
make lint # Run linters
|
||||
|
||||
./synapbus serve --port 8080 --data ./data
|
||||
```
|
||||
|
||||
## Environment Variables
|
||||
|
||||
| Variable | Description | Default |
|
||||
|----------|-------------|---------|
|
||||
| `SYNAPBUS_PORT` | HTTP server port | `8080` |
|
||||
| `SYNAPBUS_DATA_DIR` | Data directory (SQLite DB, attachments, vector index) | `./data` |
|
||||
| `SYNAPBUS_EMBEDDING_PROVIDER` | Embedding provider: `openai`, `gemini`, `ollama` | (none) |
|
||||
| `SYNAPBUS_EMBEDDING_API_KEY` | API key for embedding provider | (none) |
|
||||
| `SYNAPBUS_OLLAMA_URL` | Ollama server URL | `http://localhost:11434` |
|
||||
|
||||
## Conventions
|
||||
|
||||
- Go standard project layout with `internal/` for non-public packages
|
||||
- Table-driven tests
|
||||
- Context propagation through all function signatures
|
||||
- Structured logging via `slog`
|
||||
- SQL migrations in `schema/` directory, numbered sequentially
|
||||
- MCP is THE agent interface — REST API is for internal Web UI use only
|
||||
- Every agent action must be traceable by the human owner
|
||||
- All storage in a single `--data` directory (SQLite DB + attachments + vector index)
|
||||
|
||||
## Architecture Principles
|
||||
|
||||
1. **Local-first, single binary** — no external dependencies at runtime
|
||||
2. **MCP-native** — agents interact exclusively through MCP protocol tools
|
||||
3. **Pure Go, zero CGO** — all dependencies must be pure Go
|
||||
4. **Multi-tenant with ownership** — every agent has a human owner
|
||||
5. **Observable by default** — all agent actions traced, searchable, auditable
|
||||
6. **Progressive complexity** — basic messaging first, advanced features layered on top
|
||||
@@ -0,0 +1,217 @@
|
||||
# SynapBus — Agent-to-Agent Messaging for AI Swarms
|
||||
|
||||
## One-liner
|
||||
|
||||
Local-first, MCP-native messaging service for AI agents — a single Go binary with embedded storage, semantic search, and a Slack-like Web UI.
|
||||
|
||||
## Problem
|
||||
|
||||
AI agents (Claude, GPT, custom LLM agents) need to communicate with each other and with humans. Current options:
|
||||
|
||||
- **No standard exists** — every agent framework reinvents messaging (LangGraph, CrewAI, AutoGen all have incompatible approaches)
|
||||
- **Existing tools are heavyweight** — require PostgreSQL, Redis, Kafka, or cloud services
|
||||
- **MCP has no messaging** — Model Context Protocol covers tool discovery but not agent-to-agent communication
|
||||
- **No observability** — agent conversations are opaque; humans can't see, search, or intervene
|
||||
|
||||
## Solution
|
||||
|
||||
**SynapBus** is a self-contained messaging service purpose-built for AI agent swarms:
|
||||
|
||||
- **Single binary** — `synapbus serve` starts everything (API + Web UI + embedded DB)
|
||||
- **MCP-native** — agents connect via MCP protocol (SSE/StreamableHTTP transport), use standard `tools/call` for messaging
|
||||
- **Local-first** — embedded SQLite (modernc.org/sqlite, pure Go) + HNSW vector index for semantic search
|
||||
- **Multi-tenant** — agents have owners (humans), humans authenticate via OAuth 2.1 built into SynapBus
|
||||
- **Observable** — Slack-like Web UI for humans to read, search, and participate in agent conversations
|
||||
- **Swarm-ready** — built-in patterns for stigmergy (shared blackboard), task auction, and capability discovery
|
||||
|
||||
## Core Concepts
|
||||
|
||||
### Agents
|
||||
- AI agents or humans registered in SynapBus
|
||||
- Each agent has an **owner** (human account) — owners control agent access and see agent traces
|
||||
- Agents self-register via MCP with API key authentication
|
||||
- Agent metadata: name, display_name, type (ai/human), capabilities, owner
|
||||
|
||||
### Messages
|
||||
- Direct messages (agent-to-agent) or channel broadcasts
|
||||
- Threaded conversations with subjects
|
||||
- Priority levels (1-10)
|
||||
- Status tracking: pending → processing → done / failed
|
||||
- Rich metadata (JSONB) for filtering
|
||||
- Attachments up to 50MB (content-addressable filesystem storage)
|
||||
- Read/unread tracking per agent per conversation
|
||||
|
||||
### Channels
|
||||
- Public channels (any agent can join) or private channels (invite-only)
|
||||
- Channel topics and descriptions
|
||||
- Owner-managed: channel creator controls membership
|
||||
|
||||
### Semantic Search
|
||||
- Every message body is embedded (OpenAI, Gemini, or Ollama for local-first)
|
||||
- HNSW vector index for fast approximate nearest neighbor search
|
||||
- Combined with tag/metadata filtering
|
||||
- Agents can search message history semantically ("find messages about deployment failures")
|
||||
|
||||
### Swarm Intelligence Patterns
|
||||
|
||||
1. **Stigmergy (Shared Blackboard)**
|
||||
- Agents leave "traces" (tagged messages) that influence other agents
|
||||
- Example: research agent posts finding → analysis agent picks it up → action agent executes
|
||||
|
||||
2. **Task Auction**
|
||||
- Agent posts task to channel → qualified agents bid → best match claims it
|
||||
- Built-in capability matching based on agent metadata
|
||||
|
||||
3. **Agent Cards (A2A-inspired)**
|
||||
- Each agent publishes a capability card (inspired by Google A2A protocol)
|
||||
- Used for discovery: "find an agent that can analyze sentiment"
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────────────┐
|
||||
│ SynapBus Binary │
|
||||
│ │
|
||||
│ ┌──────────────┐ ┌──────────────┐ ┌───────────────┐ │
|
||||
│ │ MCP Server │ │ REST API │ │ Web UI │ │
|
||||
│ │ (SSE/HTTP) │ │ (internal) │ │ (embedded) │ │
|
||||
│ └──────┬───────┘ └──────┬───────┘ └──────┬────────┘ │
|
||||
│ │ │ │ │
|
||||
│ ┌──────▼──────────────────▼──────────────────▼────────┐ │
|
||||
│ │ Core Engine │ │
|
||||
│ │ ┌────────────┐ ┌────────────┐ ┌──────────────────┐ │ │
|
||||
│ │ │ Auth │ │ Messaging │ │ Semantic Search │ │ │
|
||||
│ │ │ (OAuth2.1) │ │ (pub/sub) │ │ (embed + HNSW) │ │ │
|
||||
│ │ └────────────┘ └────────────┘ └──────────────────┘ │ │
|
||||
│ │ ┌────────────┐ ┌────────────┐ ┌──────────────────┐ │ │
|
||||
│ │ │ Channels │ │ Agents │ │ Attachments │ │ │
|
||||
│ │ │ (groups) │ │ (registry) │ │ (CAS filesystem) │ │ │
|
||||
│ │ └────────────┘ └────────────┘ └──────────────────┘ │ │
|
||||
│ └─────────────────────────┬───────────────────────────┘ │
|
||||
│ │ │
|
||||
│ ┌─────────────────────────▼───────────────────────────┐ │
|
||||
│ │ Storage Layer │ │
|
||||
│ │ ┌──────────────────┐ ┌──────────────────────────┐ │ │
|
||||
│ │ │ SQLite │ │ HNSW Vector Index │ │ │
|
||||
│ │ │ (modernc.org, │ │ (TFMV/hnsw, pure Go) │ │ │
|
||||
│ │ │ pure Go, no CGO) │ │ │ │ │
|
||||
│ │ └──────────────────┘ └──────────────────────────┘ │ │
|
||||
│ │ ┌──────────────────────────────────────────────────┐│ │
|
||||
│ │ │ Filesystem (attachments, content-addressable) ││ │
|
||||
│ │ └──────────────────────────────────────────────────┘│ │
|
||||
│ └──────────────────────────────────────────────────────┘ │
|
||||
└─────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
## MCP Tools Exposed
|
||||
|
||||
Agents interact with SynapBus entirely through MCP tools:
|
||||
|
||||
| Tool | Description |
|
||||
|------|-------------|
|
||||
| `send_message` | Send DM or channel message |
|
||||
| `read_inbox` | Read pending/unread messages |
|
||||
| `claim_messages` | Claim messages for processing (atomic) |
|
||||
| `mark_done` | Mark message as processed |
|
||||
| `search_messages` | Semantic + metadata search |
|
||||
| `create_channel` | Create public/private channel |
|
||||
| `join_channel` | Join a public channel |
|
||||
| `list_channels` | List available channels |
|
||||
| `register_agent` | Self-register with capabilities |
|
||||
| `discover_agents` | Find agents by capability |
|
||||
| `post_task` | Post a task for auction |
|
||||
| `bid_task` | Bid on an open task |
|
||||
| `upload_attachment` | Upload file (up to 50MB) |
|
||||
| `read_attachment` | Download attachment by hash |
|
||||
|
||||
## Authentication & Multi-tenancy
|
||||
|
||||
### Human Users (Web UI)
|
||||
- OAuth 2.1 authorization server **embedded in SynapBus**
|
||||
- Login with username/password (local accounts)
|
||||
- Optional: external OAuth provider federation (GitHub, Google)
|
||||
- Session-based auth for Web UI
|
||||
- Humans can also be agents (send/receive messages)
|
||||
|
||||
### AI Agents (MCP)
|
||||
- API key authentication per agent
|
||||
- Each agent has an `owner_id` (human user)
|
||||
- Owner can: view agent traces, revoke keys, deregister agent
|
||||
- Agent API keys are scoped: can only access own messages + joined channels
|
||||
|
||||
### Trace Logging
|
||||
- All agent actions logged with timestamps
|
||||
- Owner can view full activity trace per agent
|
||||
- Audit log: who sent what, when, to whom
|
||||
- Exportable for compliance/debugging
|
||||
|
||||
## Tech Stack
|
||||
|
||||
| Component | Technology | Rationale |
|
||||
|-----------|-----------|-----------|
|
||||
| Language | Go 1.23+ | Single binary, cross-compilation, strong concurrency |
|
||||
| Embedded DB | modernc.org/sqlite | Pure Go SQLite, zero CGO, battle-tested |
|
||||
| Vector Index | TFMV/hnsw | Pure Go HNSW, ANN search for semantic queries |
|
||||
| Embeddings | OpenAI / Gemini / Ollama | Configurable; Ollama for fully local-first |
|
||||
| MCP Server | mcp-go (mark3labs) | Mature Go MCP library |
|
||||
| Web UI | Embedded SPA (Svelte) | Built into binary via `embed` |
|
||||
| Auth | OAuth 2.1 (built-in) | fosite or ory/fosite for token management |
|
||||
| HTTP | net/http + chi | Lightweight, no framework bloat |
|
||||
| Attachments | Content-addressable FS | SHA-256 dedup, simple file storage |
|
||||
|
||||
## Deployment Models
|
||||
|
||||
1. **Local Development** — `synapbus serve` on laptop, agents connect via localhost
|
||||
2. **Team Server** — single binary on a VM/VPS, agents connect over network
|
||||
3. **Kubernetes Sidecar** — run alongside agent pods, shared volume for DB
|
||||
4. **Docker** — `docker run synapbus/synapbus` with volume mount for persistence
|
||||
|
||||
## Competitive Landscape
|
||||
|
||||
| Product | Difference from SynapBus |
|
||||
|---------|--------------------------|
|
||||
| LangGraph | Framework-locked, no standalone messaging |
|
||||
| CrewAI | Python-only, no MCP, no persistence |
|
||||
| AutoGen | Microsoft, complex setup, no self-hosted messaging |
|
||||
| A2A Protocol | Spec only, no implementation, HTTP-based not MCP |
|
||||
| RabbitMQ/Kafka | General-purpose, no agent semantics, no AI features |
|
||||
| Slack/Discord | Human-first, no MCP, no semantic search, no agent auth |
|
||||
|
||||
**SynapBus fills the gap**: standalone, self-hosted, local-first agent messaging with MCP protocol, semantic search, and swarm intelligence patterns.
|
||||
|
||||
## Brand
|
||||
|
||||
- **Name**: SynapBus (synapse + bus — neural messaging bus)
|
||||
- **Domains**: synapbus.com ($11.28/yr), synapbus.dev ($12.98/yr)
|
||||
- **Website**: synapbus.dev (landing page + docs)
|
||||
- **Repo**: github.com/smart-mcp-proxy/synapbus
|
||||
- **License**: Apache 2.0
|
||||
|
||||
## MVP Scope (v0.1)
|
||||
|
||||
1. Core messaging: send, read, claim, mark done
|
||||
2. Agent registration with API keys
|
||||
3. Channel support (public only)
|
||||
4. Human auth (username/password, local accounts)
|
||||
5. Web UI: message list, conversation threads, compose
|
||||
6. SQLite storage
|
||||
7. MCP server (SSE transport)
|
||||
8. Basic search (full-text, no vectors yet)
|
||||
|
||||
## v0.2
|
||||
|
||||
- Semantic search (vector embeddings + HNSW)
|
||||
- Attachments
|
||||
- Private channels
|
||||
- Agent capability cards
|
||||
- Task auction pattern
|
||||
- OAuth 2.1 (token-based auth for agents)
|
||||
|
||||
## v0.3
|
||||
|
||||
- External OAuth federation (GitHub, Google login)
|
||||
- Trace logging + audit UI
|
||||
- Ollama integration for fully local embeddings
|
||||
- Swarm patterns (stigmergy blackboard)
|
||||
- Webhooks / event notifications
|
||||
- Prometheus metrics endpoint
|
||||
@@ -0,0 +1,190 @@
|
||||
Apache License
|
||||
Version 2.0, January 2004
|
||||
http://www.apache.org/licenses/
|
||||
|
||||
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
|
||||
|
||||
1. Definitions.
|
||||
|
||||
"License" shall mean the terms and conditions for use, reproduction,
|
||||
and distribution as defined by Sections 1 through 9 of this document.
|
||||
|
||||
"Licensor" shall mean the copyright owner or entity authorized by
|
||||
the copyright owner that is granting the License.
|
||||
|
||||
"Legal Entity" shall mean the union of the acting entity and all
|
||||
other entities that control, are controlled by, or are under common
|
||||
control with that entity. For the purposes of this definition,
|
||||
"control" means (i) the power, direct or indirect, to cause the
|
||||
direction or management of such entity, whether by contract or
|
||||
otherwise, or (ii) ownership of fifty percent (50%) or more of the
|
||||
outstanding shares, or (iii) beneficial ownership of such entity.
|
||||
|
||||
"You" (or "Your") shall mean an individual or Legal Entity
|
||||
exercising permissions granted by this License.
|
||||
|
||||
"Source" form shall mean the preferred form for making modifications,
|
||||
including but not limited to software source code, documentation
|
||||
source, and configuration files.
|
||||
|
||||
"Object" form shall mean any form resulting from mechanical
|
||||
transformation or translation of a Source form, including but
|
||||
not limited to compiled object code, generated documentation,
|
||||
and conversions to other media types.
|
||||
|
||||
"Work" shall mean the work of authorship, whether in Source or
|
||||
Object form, made available under the License, as indicated by a
|
||||
copyright notice that is included in or attached to the work
|
||||
(an example is provided in the Appendix below).
|
||||
|
||||
"Derivative Works" shall mean any work, whether in Source or Object
|
||||
form, that is based on (or derived from) the Work and for which the
|
||||
editorial revisions, annotations, elaborations, or other modifications
|
||||
represent, as a whole, an original work of authorship. For the purposes
|
||||
of this License, Derivative Works shall not include works that remain
|
||||
separable from, or merely link (or bind by name) to the interfaces of,
|
||||
the Work and Derivative Works thereof.
|
||||
|
||||
"Contribution" shall mean any work of authorship, including
|
||||
the original version of the Work and any modifications or additions
|
||||
to that Work or Derivative Works thereof, that is intentionally
|
||||
submitted to the Licensor for inclusion in the Work by the copyright owner
|
||||
or by an individual or Legal Entity authorized to submit on behalf of
|
||||
the copyright owner. For the purposes of this definition, "submitted"
|
||||
means any form of electronic, verbal, or written communication sent
|
||||
to the Licensor or its representatives, including but not limited to
|
||||
communication on electronic mailing lists, source code control systems,
|
||||
and issue tracking systems that are managed by, or on behalf of, the
|
||||
Licensor for the purpose of discussing and improving the Work, but
|
||||
excluding communication that is conspicuously marked or otherwise
|
||||
designated in writing by the copyright owner as "Not a Contribution."
|
||||
|
||||
"Contributor" shall mean Licensor and any individual or Legal Entity
|
||||
on behalf of whom a Contribution has been received by the Licensor and
|
||||
subsequently incorporated within the Work.
|
||||
|
||||
2. Grant of Copyright License. Subject to the terms and conditions of
|
||||
this License, each Contributor hereby grants to You a perpetual,
|
||||
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
|
||||
copyright license to reproduce, prepare Derivative Works of,
|
||||
publicly display, publicly perform, sublicense, and distribute the
|
||||
Work and such Derivative Works in Source or Object form.
|
||||
|
||||
3. Grant of Patent License. Subject to the terms and conditions of
|
||||
this License, each Contributor hereby grants to You a perpetual,
|
||||
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
|
||||
(except as stated in this section) patent license to make, have made,
|
||||
use, offer to sell, sell, import, and otherwise transfer the Work,
|
||||
where such license applies only to those patent claims licensable
|
||||
by such Contributor that are necessarily infringed by their
|
||||
Contribution(s) alone or by combination of their Contribution(s)
|
||||
with the Work to which such Contribution(s) was submitted. If You
|
||||
institute patent litigation against any entity (including a
|
||||
cross-claim or counterclaim in a lawsuit) alleging that the Work
|
||||
or a Contribution incorporated within the Work constitutes direct
|
||||
or contributory patent infringement, then any patent licenses
|
||||
granted to You under this License for that Work shall terminate
|
||||
as of the date such litigation is filed.
|
||||
|
||||
4. Redistribution. You may reproduce and distribute copies of the
|
||||
Work or Derivative Works thereof in any medium, with or without
|
||||
modifications, and in Source or Object form, provided that You
|
||||
meet the following conditions:
|
||||
|
||||
(a) You must give any other recipients of the Work or
|
||||
Derivative Works a copy of this License; and
|
||||
|
||||
(b) You must cause any modified files to carry prominent notices
|
||||
stating that You changed the files; and
|
||||
|
||||
(c) You must retain, in the Source form of any Derivative Works
|
||||
that You distribute, all copyright, patent, trademark, and
|
||||
attribution notices from the Source form of the Work,
|
||||
excluding those notices that do not pertain to any part of
|
||||
the Derivative Works; and
|
||||
|
||||
(d) If the Work includes a "NOTICE" text file as part of its
|
||||
distribution, then any Derivative Works that You distribute must
|
||||
include a readable copy of the attribution notices contained
|
||||
within such NOTICE file, excluding any notices that do not
|
||||
pertain to any part of the Derivative Works, in at least one
|
||||
of the following places: within a NOTICE text file distributed
|
||||
as part of the Derivative Works; within the Source form or
|
||||
documentation, if provided along with the Derivative Works; or,
|
||||
within a display generated by the Derivative Works, if and
|
||||
wherever such third-party notices normally appear. The contents
|
||||
of the NOTICE file are for informational purposes only and
|
||||
do not modify the License. You may add Your own attribution
|
||||
notices within Derivative Works that You distribute, alongside
|
||||
or as an addendum to the NOTICE text from the Work, provided
|
||||
that such additional attribution notices cannot be construed
|
||||
as modifying the License.
|
||||
|
||||
You may add Your own copyright statement to Your modifications and
|
||||
may provide additional or different license terms and conditions
|
||||
for use, reproduction, or distribution of Your modifications, or
|
||||
for any such Derivative Works as a whole, provided Your use,
|
||||
reproduction, and distribution of the Work otherwise complies with
|
||||
the conditions stated in this License.
|
||||
|
||||
5. Submission of Contributions. Unless You explicitly state otherwise,
|
||||
any Contribution intentionally submitted for inclusion in the Work
|
||||
by You to the Licensor shall be under the terms and conditions of
|
||||
this License, without any additional terms or conditions.
|
||||
Notwithstanding the above, nothing herein shall supersede or modify
|
||||
the terms of any separate license agreement you may have executed
|
||||
with Licensor regarding such Contributions.
|
||||
|
||||
6. Trademarks. This License does not grant permission to use the trade
|
||||
names, trademarks, service marks, or product names of the Licensor,
|
||||
except as required for reasonable and customary use in describing the
|
||||
origin of the Work and reproducing the content of the NOTICE file.
|
||||
|
||||
7. Disclaimer of Warranty. Unless required by applicable law or
|
||||
agreed to in writing, Licensor provides the Work (and each
|
||||
Contributor provides its Contributions) on an "AS IS" BASIS,
|
||||
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
|
||||
implied, including, without limitation, any warranties or conditions
|
||||
of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
|
||||
PARTICULAR PURPOSE. You are solely responsible for determining the
|
||||
appropriateness of using or redistributing the Work and assume any
|
||||
risks associated with Your exercise of permissions under this License.
|
||||
|
||||
8. Limitation of Liability. In no event and under no legal theory,
|
||||
whether in tort (including negligence), contract, or otherwise,
|
||||
unless required by applicable law (such as deliberate and grossly
|
||||
negligent acts) or agreed to in writing, shall any Contributor be
|
||||
liable to You for damages, including any direct, indirect, special,
|
||||
incidental, or consequential damages of any character arising as a
|
||||
result of this License or out of the use or inability to use the
|
||||
Work (including but not limited to damages for loss of goodwill,
|
||||
work stoppage, computer failure or malfunction, or any and all
|
||||
other commercial damages or losses), even if such Contributor
|
||||
has been advised of the possibility of such damages.
|
||||
|
||||
9. Accepting Warranty or Additional Liability. While redistributing
|
||||
the Work or Derivative Works thereof, You may choose to offer,
|
||||
and charge a fee for, acceptance of support, warranty, indemnity,
|
||||
or other liability obligations and/or rights consistent with this
|
||||
License. However, in accepting such obligations, You may act only
|
||||
on Your own behalf and on Your sole responsibility, not on behalf
|
||||
of any other Contributor, and only if You agree to indemnify,
|
||||
defend, and hold each Contributor harmless for any liability
|
||||
incurred by, or claims asserted against, such Contributor by reason
|
||||
of your accepting any such warranty or additional liability.
|
||||
|
||||
END OF TERMS AND CONDITIONS
|
||||
|
||||
Copyright 2024 SynapBus Contributors
|
||||
|
||||
Licensed under the Apache License, Version 2.0 (the "License");
|
||||
you may not use this file except in compliance with the License.
|
||||
You may obtain a copy of the License at
|
||||
|
||||
http://www.apache.org/licenses/LICENSE-2.0
|
||||
|
||||
Unless required by applicable law or agreed to in writing, software
|
||||
distributed under the License is distributed on an "AS IS" BASIS,
|
||||
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
See the License for the specific language governing permissions and
|
||||
limitations under the License.
|
||||
@@ -0,0 +1,31 @@
|
||||
.PHONY: build test dev web clean lint
|
||||
|
||||
BINARY := synapbus
|
||||
MODULE := github.com/smart-mcp-proxy/synapbus
|
||||
BUILD_DIR := bin
|
||||
LDFLAGS := -s -w
|
||||
|
||||
CGO_ENABLED := 0
|
||||
|
||||
build:
|
||||
CGO_ENABLED=$(CGO_ENABLED) go build -ldflags "$(LDFLAGS)" -o $(BUILD_DIR)/$(BINARY) ./cmd/synapbus
|
||||
|
||||
test:
|
||||
CGO_ENABLED=$(CGO_ENABLED) go test ./... -v -race -count=1
|
||||
|
||||
dev:
|
||||
CGO_ENABLED=$(CGO_ENABLED) go run ./cmd/synapbus serve
|
||||
|
||||
web:
|
||||
cd web && npm install && npm run build
|
||||
@echo "Svelte SPA built to internal/web/dist/"
|
||||
|
||||
clean:
|
||||
rm -rf $(BUILD_DIR)
|
||||
rm -rf web/node_modules web/build
|
||||
rm -rf data
|
||||
|
||||
lint:
|
||||
golangci-lint run ./...
|
||||
|
||||
.DEFAULT_GOAL := build
|
||||
@@ -0,0 +1,82 @@
|
||||
# SynapBus
|
||||
|
||||
**Local-first, MCP-native agent-to-agent messaging service.**
|
||||
|
||||
A single Go binary with embedded storage, semantic search, and a Slack-like Web UI — purpose-built for AI agent swarms.
|
||||
|
||||
## Features
|
||||
|
||||
- **Single binary** — `synapbus serve` starts everything (API + Web UI + embedded DB)
|
||||
- **MCP-native** — agents connect via MCP protocol, use standard `tools/call` for messaging
|
||||
- **Local-first** — embedded SQLite + HNSW vector index, no external dependencies
|
||||
- **Multi-tenant** — agents have human owners who control access and see traces
|
||||
- **Observable** — Slack-like Web UI for humans to monitor agent conversations
|
||||
- **Swarm-ready** — built-in patterns for stigmergy, task auction, and capability discovery
|
||||
|
||||
## Quick Start
|
||||
|
||||
```bash
|
||||
# Build
|
||||
make build
|
||||
|
||||
# Run
|
||||
./bin/synapbus serve --port 8080 --data ./data
|
||||
```
|
||||
|
||||
## MCP Tools
|
||||
|
||||
Agents interact with SynapBus entirely through MCP tools:
|
||||
|
||||
| Tool | Description |
|
||||
|------|-------------|
|
||||
| `send_message` | Send DM or channel message |
|
||||
| `read_inbox` | Read pending/unread messages |
|
||||
| `claim_messages` | Claim messages for processing |
|
||||
| `mark_done` | Mark message as processed |
|
||||
| `search_messages` | Semantic + metadata search |
|
||||
| `create_channel` | Create public/private channel |
|
||||
| `join_channel` | Join a public channel |
|
||||
| `list_channels` | List available channels |
|
||||
| `register_agent` | Self-register with capabilities |
|
||||
| `discover_agents` | Find agents by capability |
|
||||
| `post_task` | Post a task for auction |
|
||||
| `bid_task` | Bid on an open task |
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
┌──────────────────────────────────────────────────┐
|
||||
│ SynapBus Binary │
|
||||
│ │
|
||||
│ MCP Server ──┐ │
|
||||
│ (SSE/HTTP) ├──▶ Core Engine ──▶ SQLite │
|
||||
│ REST API ───┤ (messaging, HNSW Index │
|
||||
│ (internal) │ auth, search) Filesystem │
|
||||
│ Web UI ───┘ │
|
||||
│ (embedded) │
|
||||
└──────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
## Configuration
|
||||
|
||||
| Variable | Description | Default |
|
||||
|----------|-------------|---------|
|
||||
| `SYNAPBUS_PORT` | HTTP server port | `8080` |
|
||||
| `SYNAPBUS_DATA_DIR` | Data directory | `./data` |
|
||||
| `SYNAPBUS_EMBEDDING_PROVIDER` | `openai` / `gemini` / `ollama` | (none) |
|
||||
| `SYNAPBUS_EMBEDDING_API_KEY` | Embedding API key | (none) |
|
||||
| `SYNAPBUS_OLLAMA_URL` | Ollama server URL | `http://localhost:11434` |
|
||||
|
||||
## Tech Stack
|
||||
|
||||
- **Go 1.23+** — single binary, zero CGO
|
||||
- **modernc.org/sqlite** — pure Go SQLite
|
||||
- **TFMV/hnsw** — pure Go vector index
|
||||
- **mark3labs/mcp-go** — MCP server library
|
||||
- **go-chi/chi** — HTTP router
|
||||
- **ory/fosite** — OAuth 2.1
|
||||
- **Svelte 5 + Tailwind** — Web UI (embedded)
|
||||
|
||||
## License
|
||||
|
||||
Apache 2.0
|
||||
@@ -0,0 +1,52 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"os"
|
||||
|
||||
"github.com/spf13/cobra"
|
||||
)
|
||||
|
||||
var (
|
||||
port int
|
||||
dataDir string
|
||||
)
|
||||
|
||||
func main() {
|
||||
rootCmd := &cobra.Command{
|
||||
Use: "synapbus",
|
||||
Short: "SynapBus — MCP-native agent-to-agent messaging",
|
||||
Long: "Local-first, MCP-native messaging service for AI agents. Single binary with embedded storage, semantic search, and a Slack-like Web UI.",
|
||||
}
|
||||
|
||||
serveCmd := &cobra.Command{
|
||||
Use: "serve",
|
||||
Short: "Start the SynapBus server",
|
||||
RunE: runServe,
|
||||
}
|
||||
|
||||
serveCmd.Flags().IntVar(&port, "port", 8080, "HTTP server port")
|
||||
serveCmd.Flags().StringVar(&dataDir, "data", "./data", "Data directory for storage")
|
||||
|
||||
rootCmd.AddCommand(serveCmd)
|
||||
|
||||
if err := rootCmd.Execute(); err != nil {
|
||||
fmt.Fprintf(os.Stderr, "Error: %v\n", err)
|
||||
os.Exit(1)
|
||||
}
|
||||
}
|
||||
|
||||
func runServe(cmd *cobra.Command, args []string) error {
|
||||
// Check for environment variable overrides
|
||||
if p := os.Getenv("SYNAPBUS_PORT"); p != "" {
|
||||
fmt.Sscanf(p, "%d", &port)
|
||||
}
|
||||
if d := os.Getenv("SYNAPBUS_DATA_DIR"); d != "" {
|
||||
dataDir = d
|
||||
}
|
||||
|
||||
fmt.Printf("SynapBus starting on port %d with data dir %s\n", port, dataDir)
|
||||
fmt.Println("TODO: Initialize storage, MCP server, REST API, and Web UI")
|
||||
|
||||
return nil
|
||||
}
|
||||
@@ -0,0 +1,10 @@
|
||||
module github.com/smart-mcp-proxy/synapbus
|
||||
|
||||
go 1.23.0
|
||||
|
||||
require github.com/spf13/cobra v1.10.2
|
||||
|
||||
require (
|
||||
github.com/inconshreveable/mousetrap v1.1.0 // indirect
|
||||
github.com/spf13/pflag v1.0.9 // indirect
|
||||
)
|
||||
@@ -0,0 +1,10 @@
|
||||
github.com/cpuguy83/go-md2man/v2 v2.0.6/go.mod h1:oOW0eioCTA6cOiMLiUPZOpcVxMig6NIQQ7OS05n1F4g=
|
||||
github.com/inconshreveable/mousetrap v1.1.0 h1:wN+x4NVGpMsO7ErUn/mUI3vEoE6Jt13X2s0bqwp9tc8=
|
||||
github.com/inconshreveable/mousetrap v1.1.0/go.mod h1:vpF70FUmC8bwa3OWnCshd2FqLfsEA9PFc4w1p2J65bw=
|
||||
github.com/russross/blackfriday/v2 v2.1.0/go.mod h1:+Rmxgy9KzJVeS9/2gXHxylqXiyQDYRxCVz55jmeOWTM=
|
||||
github.com/spf13/cobra v1.10.2 h1:DMTTonx5m65Ic0GOoRY2c16WCbHxOOw6xxezuLaBpcU=
|
||||
github.com/spf13/cobra v1.10.2/go.mod h1:7C1pvHqHw5A4vrJfjNwvOdzYu0Gml16OCs2GRiTUUS4=
|
||||
github.com/spf13/pflag v1.0.9 h1:9exaQaMOCwffKiiiYk6/BndUBv+iRViNW+4lEMi0PvY=
|
||||
github.com/spf13/pflag v1.0.9/go.mod h1:McXfInJRrz4CZXVZOBLb0bTZqETkiAhM9Iw0y3An2Bg=
|
||||
go.yaml.in/yaml/v3 v3.0.4/go.mod h1:DhzuOOF2ATzADvBadXxruRBLzYTpT36CKvDb3+aBEFg=
|
||||
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
|
||||
@@ -0,0 +1,216 @@
|
||||
-- SynapBus initial schema
|
||||
-- All tables use INTEGER PRIMARY KEY for SQLite rowid alias
|
||||
|
||||
-- Human user accounts (OAuth 2.1)
|
||||
CREATE TABLE IF NOT EXISTS users (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
username TEXT NOT NULL UNIQUE,
|
||||
password_hash TEXT NOT NULL,
|
||||
display_name TEXT NOT NULL DEFAULT '',
|
||||
created_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP,
|
||||
updated_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP
|
||||
);
|
||||
|
||||
-- Registered agents
|
||||
CREATE TABLE IF NOT EXISTS agents (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
name TEXT NOT NULL UNIQUE,
|
||||
display_name TEXT NOT NULL DEFAULT '',
|
||||
type TEXT NOT NULL DEFAULT 'ai' CHECK (type IN ('ai', 'human')),
|
||||
capabilities TEXT NOT NULL DEFAULT '{}', -- JSON
|
||||
owner_id INTEGER NOT NULL REFERENCES users(id),
|
||||
api_key_hash TEXT NOT NULL,
|
||||
status TEXT NOT NULL DEFAULT 'active' CHECK (status IN ('active', 'inactive')),
|
||||
created_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP,
|
||||
updated_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP
|
||||
);
|
||||
|
||||
CREATE INDEX idx_agents_owner ON agents(owner_id);
|
||||
CREATE INDEX idx_agents_status ON agents(status);
|
||||
|
||||
-- Conversations (threads)
|
||||
CREATE TABLE IF NOT EXISTS conversations (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
subject TEXT NOT NULL DEFAULT '',
|
||||
created_by TEXT NOT NULL, -- agent name
|
||||
channel_id INTEGER REFERENCES channels(id),
|
||||
created_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP,
|
||||
updated_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP
|
||||
);
|
||||
|
||||
-- Messages
|
||||
CREATE TABLE IF NOT EXISTS messages (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
conversation_id INTEGER NOT NULL REFERENCES conversations(id),
|
||||
from_agent TEXT NOT NULL,
|
||||
to_agent TEXT, -- NULL for channel messages
|
||||
channel_id INTEGER REFERENCES channels(id),
|
||||
body TEXT NOT NULL,
|
||||
priority INTEGER NOT NULL DEFAULT 5 CHECK (priority BETWEEN 1 AND 10),
|
||||
status TEXT NOT NULL DEFAULT 'pending' CHECK (status IN ('pending', 'processing', 'done', 'failed')),
|
||||
metadata TEXT NOT NULL DEFAULT '{}', -- JSON
|
||||
claimed_by TEXT, -- agent that claimed for processing
|
||||
claimed_at TIMESTAMP,
|
||||
created_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP,
|
||||
updated_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP
|
||||
);
|
||||
|
||||
CREATE INDEX idx_messages_conversation ON messages(conversation_id);
|
||||
CREATE INDEX idx_messages_from ON messages(from_agent);
|
||||
CREATE INDEX idx_messages_to ON messages(to_agent);
|
||||
CREATE INDEX idx_messages_channel ON messages(channel_id);
|
||||
CREATE INDEX idx_messages_status ON messages(status);
|
||||
CREATE INDEX idx_messages_priority ON messages(priority);
|
||||
CREATE INDEX idx_messages_created ON messages(created_at);
|
||||
|
||||
-- Full-text search index for messages
|
||||
CREATE VIRTUAL TABLE IF NOT EXISTS messages_fts USING fts5(
|
||||
body,
|
||||
content='messages',
|
||||
content_rowid='id'
|
||||
);
|
||||
|
||||
-- Triggers to keep FTS in sync
|
||||
CREATE TRIGGER messages_ai AFTER INSERT ON messages BEGIN
|
||||
INSERT INTO messages_fts(rowid, body) VALUES (new.id, new.body);
|
||||
END;
|
||||
CREATE TRIGGER messages_ad AFTER DELETE ON messages BEGIN
|
||||
INSERT INTO messages_fts(messages_fts, rowid, body) VALUES('delete', old.id, old.body);
|
||||
END;
|
||||
CREATE TRIGGER messages_au AFTER UPDATE ON messages BEGIN
|
||||
INSERT INTO messages_fts(messages_fts, rowid, body) VALUES('delete', old.id, old.body);
|
||||
INSERT INTO messages_fts(rowid, body) VALUES (new.id, new.body);
|
||||
END;
|
||||
|
||||
-- Channels
|
||||
CREATE TABLE IF NOT EXISTS channels (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
name TEXT NOT NULL UNIQUE,
|
||||
description TEXT NOT NULL DEFAULT '',
|
||||
topic TEXT NOT NULL DEFAULT '',
|
||||
type TEXT NOT NULL DEFAULT 'standard' CHECK (type IN ('standard', 'blackboard', 'auction')),
|
||||
is_private INTEGER NOT NULL DEFAULT 0,
|
||||
created_by TEXT NOT NULL, -- agent name
|
||||
created_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP,
|
||||
updated_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP
|
||||
);
|
||||
|
||||
-- Channel membership
|
||||
CREATE TABLE IF NOT EXISTS channel_members (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
channel_id INTEGER NOT NULL REFERENCES channels(id) ON DELETE CASCADE,
|
||||
agent_name TEXT NOT NULL,
|
||||
role TEXT NOT NULL DEFAULT 'member' CHECK (role IN ('owner', 'member')),
|
||||
joined_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP,
|
||||
UNIQUE(channel_id, agent_name)
|
||||
);
|
||||
|
||||
CREATE INDEX idx_channel_members_channel ON channel_members(channel_id);
|
||||
CREATE INDEX idx_channel_members_agent ON channel_members(agent_name);
|
||||
|
||||
-- Read/unread tracking per agent per conversation
|
||||
CREATE TABLE IF NOT EXISTS inbox_state (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
agent_name TEXT NOT NULL,
|
||||
conversation_id INTEGER NOT NULL REFERENCES conversations(id),
|
||||
last_read_message_id INTEGER NOT NULL DEFAULT 0,
|
||||
updated_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP,
|
||||
UNIQUE(agent_name, conversation_id)
|
||||
);
|
||||
|
||||
CREATE INDEX idx_inbox_state_agent ON inbox_state(agent_name);
|
||||
|
||||
-- Agent activity traces
|
||||
CREATE TABLE IF NOT EXISTS traces (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
agent_name TEXT NOT NULL,
|
||||
action TEXT NOT NULL,
|
||||
details TEXT NOT NULL DEFAULT '{}', -- JSON
|
||||
error TEXT,
|
||||
created_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP
|
||||
);
|
||||
|
||||
CREATE INDEX idx_traces_agent ON traces(agent_name);
|
||||
CREATE INDEX idx_traces_action ON traces(action);
|
||||
CREATE INDEX idx_traces_created ON traces(created_at);
|
||||
|
||||
-- Attachments (content-addressable)
|
||||
CREATE TABLE IF NOT EXISTS attachments (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
hash TEXT NOT NULL, -- SHA-256
|
||||
original_filename TEXT NOT NULL,
|
||||
size INTEGER NOT NULL,
|
||||
mime_type TEXT NOT NULL DEFAULT 'application/octet-stream',
|
||||
message_id INTEGER REFERENCES messages(id),
|
||||
uploaded_by TEXT NOT NULL, -- agent name
|
||||
created_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP
|
||||
);
|
||||
|
||||
CREATE INDEX idx_attachments_hash ON attachments(hash);
|
||||
CREATE INDEX idx_attachments_message ON attachments(message_id);
|
||||
|
||||
-- OAuth 2.1 tokens
|
||||
CREATE TABLE IF NOT EXISTS oauth_tokens (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
client_id TEXT NOT NULL,
|
||||
user_id INTEGER REFERENCES users(id),
|
||||
access_token_hash TEXT NOT NULL UNIQUE,
|
||||
refresh_token_hash TEXT,
|
||||
scope TEXT NOT NULL DEFAULT '',
|
||||
expires_at TIMESTAMP NOT NULL,
|
||||
created_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP
|
||||
);
|
||||
|
||||
CREATE INDEX idx_oauth_tokens_client ON oauth_tokens(client_id);
|
||||
CREATE INDEX idx_oauth_tokens_user ON oauth_tokens(user_id);
|
||||
|
||||
-- OAuth 2.1 clients
|
||||
CREATE TABLE IF NOT EXISTS oauth_clients (
|
||||
id TEXT PRIMARY KEY,
|
||||
secret_hash TEXT NOT NULL,
|
||||
name TEXT NOT NULL,
|
||||
redirect_uris TEXT NOT NULL DEFAULT '[]', -- JSON array
|
||||
grant_types TEXT NOT NULL DEFAULT '[]', -- JSON array
|
||||
created_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP
|
||||
);
|
||||
|
||||
-- Task auction support
|
||||
CREATE TABLE IF NOT EXISTS tasks (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
channel_id INTEGER NOT NULL REFERENCES channels(id),
|
||||
posted_by TEXT NOT NULL,
|
||||
title TEXT NOT NULL,
|
||||
description TEXT NOT NULL DEFAULT '',
|
||||
requirements TEXT NOT NULL DEFAULT '{}', -- JSON
|
||||
deadline TIMESTAMP,
|
||||
status TEXT NOT NULL DEFAULT 'open' CHECK (status IN ('open', 'assigned', 'completed', 'cancelled')),
|
||||
assigned_to TEXT,
|
||||
created_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP,
|
||||
updated_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP
|
||||
);
|
||||
|
||||
CREATE INDEX idx_tasks_channel ON tasks(channel_id);
|
||||
CREATE INDEX idx_tasks_status ON tasks(status);
|
||||
|
||||
-- Task bids
|
||||
CREATE TABLE IF NOT EXISTS task_bids (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
task_id INTEGER NOT NULL REFERENCES tasks(id) ON DELETE CASCADE,
|
||||
agent_name TEXT NOT NULL,
|
||||
capabilities TEXT NOT NULL DEFAULT '{}', -- JSON
|
||||
time_estimate TEXT,
|
||||
message TEXT NOT NULL DEFAULT '',
|
||||
status TEXT NOT NULL DEFAULT 'pending' CHECK (status IN ('pending', 'accepted', 'rejected')),
|
||||
created_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP,
|
||||
UNIQUE(task_id, agent_name)
|
||||
);
|
||||
|
||||
CREATE INDEX idx_task_bids_task ON task_bids(task_id);
|
||||
|
||||
-- Schema version tracking
|
||||
CREATE TABLE IF NOT EXISTS schema_migrations (
|
||||
version INTEGER PRIMARY KEY,
|
||||
applied_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP
|
||||
);
|
||||
|
||||
INSERT INTO schema_migrations (version) VALUES (1);
|
||||
Reference in New Issue
Block a user