From 1e7c57ff58631270d9f8321457ffa969363cffd1 Mon Sep 17 00:00:00 2001 From: Claude Date: Fri, 15 May 2026 20:56:10 +0000 Subject: [PATCH] Research SPDD + Fowler Fragments + cc-connect; promote 2 rules Summary: - Logged Thoughtworks SPDD, Fowler's 29 Apr Fragments, and chenhg5/cc-connect as sources. - Promoted spec/code two-way sync and sensor-based verification into AGENTS.md + Copilot mirror. - Added two knowledge notes (SPDD / REASONS Canvas; agent-integration patterns). Changed files: - docs/source-repos.md: 3 rows (SPDD + Fragments promoted; cc-connect extracted). - docs/research-log.md: 3 dated observation blocks, extends/contradicts, promoted-rule rows, pending + anti-pattern entries. - AGENTS.md: spec/code sync bullet (House style) + sensor-verification check (Validation). - .github/copilot-instructions.md: mirrored both rules. - docs/knowledge/structured-prompt-driven-development.md: new descriptive note. - docs/knowledge/agent-integration-patterns.md: new descriptive note. - docs/knowledge/README.md: listed both new notes. - docs/ai-agent-coding-strategy.md: pointer to the SPDD note. Validation: - ./scripts/check.sh: Tier 1 (length, Boundaries) and Tier 2 (markdownlint) clean. - Tier 3/4 binaries absent locally; will verify in CI. - All new files end with a Boundaries section; all under 200 lines. Follow-ups: - Confirm Tier 3 (lychee) and Tier 4 (typos) pass in CI; fix if flagged. - cc-connect left as extracted (knowledge note only, no instruction-file rule). https://claude.ai/code/session_01AmmEaVLpyeJz7TZ3iyXtWK --- .github/copilot-instructions.md | 2 + AGENTS.md | 2 + docs/ai-agent-coding-strategy.md | 6 ++ docs/knowledge/README.md | 6 ++ docs/knowledge/agent-integration-patterns.md | 63 +++++++++++++++ .../structured-prompt-driven-development.md | 77 +++++++++++++++++++ docs/research-log.md | 39 ++++++++++ docs/source-repos.md | 3 + 8 files changed, 198 insertions(+) create mode 100644 docs/knowledge/agent-integration-patterns.md create mode 100644 docs/knowledge/structured-prompt-driven-development.md diff --git a/.github/copilot-instructions.md b/.github/copilot-instructions.md index 3daf592..fb65c35 100644 --- a/.github/copilot-instructions.md +++ b/.github/copilot-instructions.md @@ -16,6 +16,7 @@ The universal baseline is in `AGENTS.md`. This file restates the parts that matt - Use checklists, tables, and templates. - Keep each file under ~200 lines. - Markdown only. No HTML except block comments for human-only notes. +- When a tracked spec, prompt, or design doc governs code, update both in the same change: behavior changes update the spec first then the code; pure refactors change the code first then sync the spec back. ## Editing rules @@ -24,6 +25,7 @@ The universal baseline is in `AGENTS.md`. This file restates the parts that matt - Prefer append-and-refine over replacing whole documents. - Do not delete `docs/research-log.md` notes unless the user asks. - Do not invent file paths, function names, commands, URLs, or identifiers. +- Treat "verified" as checked by a deterministic sensor (tests, types, lint, CI gate), not merely read. Prefer adding a sensor over re-reading at scale. ## Safety diff --git a/AGENTS.md b/AGENTS.md index 14555ee..f32399c 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -28,6 +28,7 @@ Claude Code does not read this file natively; `CLAUDE.md` cross-references it vi - One topic per file. Per-path rules go in `.claude/rules/` (Claude) or `.github/instructions/` (Copilot). - End any prescriptive instruction file (rules, skills, commands) with a `Boundaries` section that lists what the file does *not* cover and which sibling file does. Forces explicit scope hand-off instead of silent ambiguity. - When a file uses YAML frontmatter with a `description:` field, the description must include *trigger phrases* (when to load this) and any competing-tool overrides ("Use this instead of X for Y"), not just a one-line summary. +- When a tracked spec, prompt, or design doc governs code, keep the two in sync within the same change: for behavior changes, update the spec first then the code; for pure refactors, change the code first then sync the spec back. Never let intent and implementation drift silently. ## Safety @@ -56,6 +57,7 @@ Before declaring a change done: - [ ] No secrets in the diff. `git diff --cached | grep -iE 'api[_-]?key|secret|token|password'` returns nothing. - [ ] Each new or edited rule is imperative and verifiable. - [ ] Each new source is recorded in `docs/source-repos.md` and `docs/research-log.md`. +- [ ] "Verified" means checked by a deterministic sensor (tests, type checker, linter, CI gate), not just read. At agent throughput, prefer adding a sensor over re-reading a diff; reserve human review for where judgement genuinely matters. ## Dev environment tips diff --git a/docs/ai-agent-coding-strategy.md b/docs/ai-agent-coding-strategy.md index a43cb84..5effa02 100644 --- a/docs/ai-agent-coding-strategy.md +++ b/docs/ai-agent-coding-strategy.md @@ -72,6 +72,12 @@ See [`docs/knowledge/agentic-pr-review-loop.md`](./knowledge/agentic-pr-review-l for the prompt templates, polling schedule, race-condition controls, and anti-patterns. +## Practical example: spec/prompt as a tracked artifact + +Treating the governing prompt or spec as a version-controlled, reviewed artifact — kept in two-way sync with the code — turns AI assistance from personal speed into governable team capability. Behavior changes update the spec first; pure refactors update the code first then sync the spec back. + +See [`docs/knowledge/structured-prompt-driven-development.md`](./knowledge/structured-prompt-driven-development.md) for the seven-part REASONS Canvas, the sync flows, and a fitness assessment. + ## Anti-patterns - Burying critical constraints inside long prose. Agents will skim past them. diff --git a/docs/knowledge/README.md b/docs/knowledge/README.md index a6a05ca..dc8e5f4 100644 --- a/docs/knowledge/README.md +++ b/docs/knowledge/README.md @@ -32,6 +32,12 @@ End each note with a `Boundaries` section listing what the note does *not* cover agent separation for PR quality gates, including compact prompts and scheduler guardrails. - [`azure-genai.md`](./azure-genai.md) — Azure / Azure OpenAI / Foundry / RAG / agents / LLMOps reference notes. +- [`structured-prompt-driven-development.md`](./structured-prompt-driven-development.md) + — REASONS Canvas, the behavior-vs-refactor spec/code sync flows, and a + fitness assessment. +- [`agent-integration-patterns.md`](./agent-integration-patterns.md) — + bridge-daemon, capability-interface, plugin self-registration, and preset-manifest + patterns for connecting agents and hosts. ## Planned topics (add when you have material) diff --git a/docs/knowledge/agent-integration-patterns.md b/docs/knowledge/agent-integration-patterns.md new file mode 100644 index 0000000..9caeea9 --- /dev/null +++ b/docs/knowledge/agent-integration-patterns.md @@ -0,0 +1,63 @@ +# Agent integration patterns + +## Purpose + +Capture reusable architecture patterns for connecting AI coding agents to each other and to external surfaces, extracted from `chenhg5/cc-connect` (a Go bridge daemon linking local coding agents to chat platforms). + +This is a descriptive note. The patterns are language-agnostic even though the source is Go. + +## When it applies + +Reach for these patterns when building any layer that bridges, multiplexes, or wraps multiple agents or hosts: a connector daemon, an MCP gateway, a multi-agent provider config, or a cross-tool installer. + +## Pattern body + +### Bridge daemon + +A single long-running process connects locally-running agents (Claude Code, Codex, Cursor, Gemini CLI, …) to external surfaces (chat platforms, webhooks) so the agent is drivable remotely with no public IP. The daemon is transport and orchestration only — it carries no instruction logic. + +### Capability-interface over identity-switch + +Behavior differences between integrations go through capability interfaces, never `if target.Name() == "feishu"`. Adding an integration means implementing an interface, not editing a switch. This keeps the core closed to modification. + +### Plugin self-registration with a strict dependency direction + +Adapters register themselves at init time (`RegisterAgent()` / `RegisterPlatform()`). The `core/` package must never import `agent/` or `platform/` — dependencies point inward only. The core is the stable nucleus; adapters are the volatile edge. + +### Uniform per-host adapter layout + +One folder per integration target (`agent/{claudecode,codex,cursor,gemini,…}/`) with the same internal file shape. Predictable layout makes a new adapter a copy-and-fill exercise. + +### Versioned, dual-language preset manifest + +Catalog-style extension manifests carry `version`, `updated_at`, and per-entry `name` / `display_name` / `description` / localized description / `tags` / `featured`. A provider abstraction nests per-agent-type config (`base_url`, `model`, `models`) so one provider serves heterogeneous agents. + +### Secrets and auth posture + +Secrets are injected via `${ENV_VAR}` interpolation, never inline. A redaction helper keeps tokens out of logs. Per-project authorization uses an admin allowlist, OS-user isolation for the run-as identity, and role-scoped command denylists and rate limits. + +### Session and context lifecycle as config + +Idle reset, auto-compression thresholds, and heartbeat intervals are explicit configuration keys, not hardcoded constants — context management is operator-tunable. + +## Trade-offs and decision criteria + +- Capability interfaces and self-registration add indirection; worth it once there are 3+ integrations, overkill for one. +- Selective compilation (build tags dropping unused integrations) trims binary size but is language-specific; treat as Go-flavored, not universal. +- A dual-language manifest is valuable only when the audience is genuinely multilingual. + +## Anti-patterns + +- Branching on integration identity instead of a capability interface. +- Letting the core import adapter packages (inverts the dependency arrow; every adapter change risks the core). +- Inlining credentials instead of env-var interpolation; logging un-redacted tokens. +- Shipping duplicated `AGENTS.md` and `CLAUDE.md` with no cross-reference — observed in cc-connect and the precise drift hazard the single-source-and-mirror discipline prevents. + +## References + +- `chenhg5/cc-connect` — +- `docs/research-log.md` — dated observations this note was promoted from. + +## Boundaries + +This note does not cover: the repository's instruction-layering rules (`AGENTS.md`, `CLAUDE.md`), the structured-prompt method (`docs/knowledge/structured-prompt-driven-development.md`), or Azure-specific guidance (`docs/knowledge/azure-genai.md`). diff --git a/docs/knowledge/structured-prompt-driven-development.md b/docs/knowledge/structured-prompt-driven-development.md new file mode 100644 index 0000000..97731eb --- /dev/null +++ b/docs/knowledge/structured-prompt-driven-development.md @@ -0,0 +1,77 @@ +# Structured Prompt-Driven Development (SPDD) + +## Purpose + +Describe SPDD, a method from Thoughtworks Global IT Services for making LLM-assisted changes governable, reviewable, and reusable by treating the prompt as a first-class delivery artifact kept in version control alongside the code. + +This is a descriptive note, not a rule. The promoted imperative ("keep spec and code in two-way sync") lives in `AGENTS.md` house style. + +## When it applies + +SPDD pays off when change must be governed and carried forward across iterations: + +- Scaled or standardized delivery with long-term maintainability. +- Hard-constraint or compliance systems (financial core, regulated domains). +- Multi-person delivery where every change must be traceable and reviewable. +- Cross-cutting refactors where logic must stay synchronized across services. + +It does not pay off for firefighting hotfixes, exploratory spikes, one-off scripts, ill-defined domains, or taste-driven creative work. + +## The REASONS Canvas + +A seven-part structured prompt that moves uncertainty left — from code review to design time. + +| Letter | Dimension | Captures | +|---|---|---| +| R | Requirements | The problem and the definition of done | +| E | Entities | Domain entities and their relationships | +| A | Approach | The strategy for meeting the requirements | +| S | Structure | Where the change fits; components and dependencies | +| O | Operations | Concrete, testable implementation steps | +| N | Norms | Cross-cutting norms (naming, observability, defensive coding) | +| S | Safeguards | Non-negotiable boundaries (invariants, performance, security) | + +R-E-A-S are abstract (intent and design); O is execution; N-S are governance. Because the full specification lives in one artifact, reviewers reason about a single document instead of scattered chat logs and partial diffs. + +## Pattern body: the closed loop + +The workflow is a closed loop: business input → abstraction → execution → validation → release, with prompt and code evolving together. + +Reference tooling (`openspdd`) implements the loop as discrete, versioned CLI steps — analysis, canvas generation, code generation, prompt update, and sync — so artifacts stay reviewable instead of trapped in chat. The structured prompt is never hand-edited; gaps are corrected by instructing the model to update only the affected Canvas sections. + +## Spec/code two-way sync + +The directional rule that distinguishes SPDD from one-shot spec-driven development: + +- **Behavior change** (new requirement, bug fix, logic correction): update the spec first, then regenerate or patch the code from it. "When reality diverges, fix the prompt first — then update the code." +- **Refactor** (no observable behavior change): change the code first, then sync the spec back so it stays an accurate record of the current code. + +Both directions keep intent and implementation from drifting apart. Birgitta Böckeler categorizes this as a *spec-anchored* approach: the spec is maintained, not generated once and discarded. + +## Trade-offs and decision criteria + +| Benefit | Nature | +|---|---| +| Determinism | A precise spec reduces hallucination and creative interpretation | +| Traceability | Every change traces back to the structured prompt; closes the audit loop | +| Faster reviews | Code arrives closer to standards; review focuses on logic, not cleanup | +| Safer evolution | Defined boundaries make targeted change lower-risk | + +The upfront cost is real: a design-first mindset shift, senior abstraction skill per feature, and automation tooling to keep prompts consistent. Without automation the method hits a throughput ceiling. + +## Anti-patterns + +- Treating the spec as write-once scaffolding and letting it rot after first generation. +- Hand-editing the structured prompt instead of regenerating affected sections. +- Applying SPDD's governance overhead to throwaway or exploratory work. +- Skipping the sync-back step after a refactor, so the spec lies about the code. + +## References + +- Structured-Prompt-Driven Development — +- Wei Zhang and Jessie Jie Xia, Thoughtworks, 28 April 2026. +- `docs/research-log.md` — dated observations this note was promoted from. + +## Boundaries + +This note does not cover: the repository's instruction-layering rules (`AGENTS.md`, `CLAUDE.md`), the research-handling pipeline (`docs/research-log.md`, `docs/source-repos.md`), or generic agent-integration plumbing (`docs/knowledge/agent-integration-patterns.md`). diff --git a/docs/research-log.md b/docs/research-log.md index 3f18e6b..c1bb94a 100644 --- a/docs/research-log.md +++ b/docs/research-log.md @@ -85,12 +85,44 @@ This log separates **observations** (what a source actually does) from **promote - Plugin reference uses `@` namespacing: `andrej-karpathy-skills@karpathy-skills` (package@skill). - Content is rule-shaped, not procedure-shaped: each principle is a problem→solution table, no shell snippets — counter-example to runnable-command-heavy skills like google/skills. +### 2026-05-15 — Thoughtworks SPDD (martinfowler.com/articles/structured-prompt-driven/) + +- Treats the prompt as a *first-class delivery artifact*: version-controlled, reviewed, reused, kept in sync with code — not an ad-hoc chat. +- **REASONS Canvas**: seven-part structured prompt — Requirements, Entities, Approach, Structure (abstract intent and design); Operations (execution steps); Norms, Safeguards (governance). Aligns intent and boundaries before code is generated. +- Spec and code stay synchronized via two named flows: requirements → prompt → code for behavior changes; code → prompt (`/spdd-sync`) for refactors that do not change observable behavior. Rule: "When reality diverges, fix the prompt first — then update the code." +- Structured prompts are never hand-edited; gaps are fixed by instructing the AI to update only the affected Canvas sections. +- `openspdd` CLI implements the workflow as discrete commands (`/spdd-analysis`, `/spdd-reasons-canvas`, `/spdd-generate`, `/spdd-prompt-update`, `/spdd-sync`) so artifacts stay versioned and reviewable instead of trapped in chat. +- Ships a fitness-assessment table rating where the method pays off (scaled or standardized delivery, hard-constraint or compliance systems, auditable team work) versus where it does not (hotfixes, spikes, one-off scripts, ill-defined domains, pure creative work). +- Categorized by Birgitta Böckeler as a *spec-anchored* approach: the spec is maintained, not generated once and discarded. + +### 2026-05-15 — Martin Fowler, Fragments (29 Apr 2026) + +- Roundup linking Chris Parsons' AI-coding guide, Birgitta Böckeler's "Harness Engineering," and Adam Tornhill on function length. +- Verification is the bottleneck, not generation: "The game is not 'how fast can we build' any more. It is 'how fast can we tell whether this is right.'" Implication: build better review surfaces, not better prompts. +- "Verified" has shifted from "read by you" to "checked by tests, by type checkers, by automated gates, or by you where your judgement matters" — at agent throughput, reading every diff in your head does not scale. +- Computational sensors (static analysis, tests) belong in the harness; agents address *every* warning where humans slack. Converting an objective rule into a formal, deterministic format gives more assurance than fuzzy prompt guidance. +- Naming matters more under LLMs, not less: models infer meaning from literal features (names, structure, local context); arbitrary identifiers measurably degrade model performance. Function boundaries are the first unit of structure. + +### 2026-05-15 — chenhg5/cc-connect + +- Go bridge daemon connecting local coding agents (Claude Code, Codex, Cursor, Gemini CLI, etc.) to chat platforms (Slack, Telegram, Discord, Feishu, …) so a local agent is drivable from a phone with no public IP. Integration and transport layer, not an instruction framework. +- Plugin self-registration via `init()` (`core.RegisterAgent()` / `core.RegisterPlatform()`); strict dependency direction — `core/` must never import `agent/` or `platform/`. +- Capability-interface over identity-switch: explicit prohibition of `if p.Name() == "feishu"`; behavior differences go through capability interfaces. +- Uniform per-host adapter layout: one folder per integration target (`agent/{claudecode,codex,cursor,gemini,…}/`). +- Versioned, dual-language preset manifests (`skill-presets.json`, `provider-presets.json`) with `version`, `updated_at`, and per-entry `name`/`display_name`/`description`/`tags`/`featured`; provider abstraction keyed by agent type. +- Secrets via `${ENV_VAR}` interpolation; `core.RedactToken()` keeps tokens out of logs. Per-project auth via `admin_from`, `run_as_user` OS isolation, role-scoped `disabled_commands` and `rate_limit`. +- Explicit session and context lifecycle as config: `reset_on_idle_mins`, `[projects.auto_compress]`, `[projects.heartbeat]`; selective compilation via build tags drops unused integrations (language-specific). +- Counter-example: ships both `AGENTS.md` and `CLAUDE.md` with largely duplicated content and no cross-reference — the drift hazard our single-source and mirror discipline exists to prevent. + ## Extends or contradicts existing observations - **Extends Anthropic Skills** — deer-flow shows `description` should embed *trigger phrases and override declarations*, not just "what + when". - **Extends huashu's anti-hallucination pattern** — deer-flow `deep-research` generalizes it as a hard-stop checklist gate inside any research-style skill. - **Extends gstack's "Iron Laws"** — SuperClaude's `Boundaries` section is the same pattern with a softer name and a fixed location at the end of the file. - **Sharper than our 200-line file rule** — deer-flow `skill-creator` uses a 300-line threshold for promoting prose to `references/`. Treat as a sibling rule for skill bodies (which can carry more domain detail than baseline rule files). +- **Extends our Validation step** — Fowler/Böckeler reframe "verified" as checked by deterministic sensors (tests, types, lint, CI gates), not read-by-human, at agent throughput. +- **New: spec/code two-way sync** — SPDD adds a directional rule we lacked: behavior change → spec first; refactor → code first then sync the spec back. +- **Reinforces single-source discipline (negative example)** — cc-connect's duplicated, un-cross-referenced `AGENTS.md`/`CLAUDE.md` is the exact failure mode our baseline-and-mirror rule prevents. ## Promoted rules @@ -115,6 +147,10 @@ These observations have been promoted into agent-readable rule files. Each row l | Hard-stop checklist for research-style work ("If any answer is NO, continue") | deer-flow `deep-research` | `AGENTS.md` §Anti-hallucination | | Single source, multi-host artifacts (one canonical content, many delivery files) | karpathy-skills, gstack multi-host installer | `docs/ai-agent-coding-strategy.md` §Single source, multi-host | | Append-mode install as a documented fallback for existing projects | karpathy-skills | `docs/ai-agent-coding-strategy.md` §Single source, multi-host | +| Spec/prompt and code kept in two-way sync (behavior → spec first; refactor → code first, then sync spec) | Thoughtworks SPDD | `AGENTS.md` §House style; `.github/copilot-instructions.md` §House style | +| "Verified" means checked by deterministic sensors (tests, types, lint, CI gates), not just read | Fowler Fragments (Parsons, Böckeler) | `AGENTS.md` §Validation; `.github/copilot-instructions.md` §Editing rules | +| REASONS Canvas as a structured-prompt template | Thoughtworks SPDD | `docs/knowledge/structured-prompt-driven-development.md` | +| Agent-integration plumbing patterns (bridge daemon, capability-interface, plugin self-registration, preset manifest) | chenhg5/cc-connect | `docs/knowledge/agent-integration-patterns.md` | ## Pending observations (not yet promoted) @@ -126,9 +162,12 @@ These observations have been promoted into agent-readable rule files. Each row l - Command frontmatter that declares wiring (`mcp-servers:`, `personas:`) — adopt when we add commands. - `evals/evals.json` co-located per skill with baseline-vs-skill parallel runs — defer until we have a skill runtime. - Bridge-skill pattern (e.g., `claude-to-deerflow`) — defer until we ship cross-platform integrations. +- `openspdd`-style per-step CLI for spec → code → sync — adopt only if we add a code-generation runtime. +- Provider abstraction keyed by agent type (cc-connect) — relevant only if we add multi-agent provider config. ## Anti-patterns observed (and rejected) - Hard-coding brand colors, IDs, or commands from memory (huashu). - Letting agents fix code without first investigating root cause (gstack `/investigate` Iron Law). - Mixing project-specific examples into general agent rules (CLAUDE.md guideline; observed in mixed instruction files). +- Shipping duplicated `AGENTS.md` and `CLAUDE.md` without a cross-reference (cc-connect) — guarantees drift; our baseline-and-mirror discipline forbids it. diff --git a/docs/source-repos.md b/docs/source-repos.md index 02c3b04..c3c64f4 100644 --- a/docs/source-repos.md +++ b/docs/source-repos.md @@ -28,6 +28,9 @@ Tracking sources studied to extract reusable AI-coding-agent practices. | SuperClaude-Org/SuperClaude_Framework | Repo | promoted | Three-axis taxonomy (`agents/` / `commands/` / `modes/`), command frontmatter that declares wiring (`mcp-servers:`, `personas:`), `Boundaries` as a required final section, slash-command namespacing | https://github.com/SuperClaude-Org/SuperClaude_Framework | | bytedance/deer-flow | Repo | promoted | `description:` as a trigger advertisement with competing-tool overrides, hard-stop checklists inside skill bodies, evals co-located per skill, 300-line threshold for promoting prose to `references/` | https://github.com/bytedance/deer-flow | | forrestchang/andrej-karpathy-skills | Repo | promoted | Single source, multi-host artifacts (`CLAUDE.md` + `CURSOR.md` + `.cursor/rules/*.mdc` + `SKILL.md` + plugin manifest); append-mode install (`curl … >> CLAUDE.md`) | https://github.com/forrestchang/andrej-karpathy-skills | +| Thoughtworks SPDD | Article | promoted | REASONS Canvas seven-part structured-prompt template; prompt as first-class artifact; behavior→spec-first / refactor→code-first two-way sync; fitness-assessment table | https://martinfowler.com/articles/structured-prompt-driven/ | +| Martin Fowler — Fragments, 29 Apr 2026 | Article | promoted | Verification is the bottleneck, not generation; computational sensors (static analysis, types, tests, gates) as the harness; identifiers matter more under LLMs | https://martinfowler.com/fragments/2026-04-29.html | +| chenhg5/cc-connect | Repo | extracted | Bridge-daemon pattern (local agents ↔ chat platforms); capability-interface over identity-switch; plugin self-registration with strict core→adapter dependency direction; versioned dual-language preset manifest; counter-example of duplicated `AGENTS.md`/`CLAUDE.md` without cross-reference | https://github.com/chenhg5/cc-connect | ## Adding a new source