Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 4 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,10 +58,10 @@ Convoy ships these pipelines; select one with `-p/--pipeline` (no config needed)
| `implement-lite` | yes | Same workflow and agents as `implement`, but the code-writing phases (`implementer`, `patterns`, `security`, `tests`) all drop to `openrouter/z-ai/glm-5.2` for a lower-cost run; `design` runs on Kimi K3 and `adversarial` on Opus. |
| `implement-advised` | yes | Same workflow and step names as `implement`, but the `implementer` phase runs on GPT 5.6 Terra xhigh and **consults GPT 5.6 Sol as an advisor** at its decision points. The audits (`patterns`, `security`, `tests`) run unadvised on `openrouter/z-ai/glm-5.2`, `design` on Kimi K3, and `adversarial` keeps Opus owning its own loop. Directly comparable with `implement-lite`: same audit executors, and the implementation step is the single variable between them. |
| `ultra-implement` | yes | Like `implement`, but the pattern/security/adversarial reviews of the initial diff run in parallel across two models feeding a triage step, and the run ends with an audit-only final review, a fixer that applies only blocking findings, and a final validator. |
| `refine` | yes | Audit the current diff (scope → bugs → clean-code → security), triage the findings adversarially, apply the accepted fixes, then validate them. |
| `ultra-refine` | yes | Like `refine`, but every read-only audit is fanned out across two models before triage, fixes, and validation. |
| `review` | **no — report only** | Scope the diff, run the bug / clean-code(+patterns) / security audits **in parallel across two models each**, then a single step synthesizes everything into one prioritized findings report. Makes no changes; the run's output is `reports/report.md`, which you read to decide whether to follow up with a `refine` run. |
| `review-lite` | **no — report only** | Same shape as `review`, but nothing runs on Opus: `openrouter/z-ai/glm-5.2` scopes the diff and writes the final report, and each parallel audit fans out across `openrouter/z-ai/glm-5.2` + `openrouter/moonshotai/kimi-k3`. The cheap way to get a full review report. |
| `refine` | yes | Audit the current diff (scope → bugs → clean-code → security → over-engineering), consolidate the deferred findings into a debt ledger, triage the findings adversarially, apply the accepted fixes, then validate them. |
| `ultra-refine` | yes | Like `refine`, but every read-only audit is fanned out across two models before a debt ledger, triage, fixes, and validation. |
| `review` | **no — report only** | Scope the diff, run the bug / clean-code(+patterns) / over-engineering / security audits **in parallel across two models each**, consolidate the deferred findings into a debt ledger, then a single step synthesizes everything into one prioritized findings report. Makes no changes; the run's output is `reports/report.md`, which you read to decide whether to follow up with a `refine` run. |
| `review-lite` | **no — report only** | Same shape as `review`, but nothing runs on Opus: `openrouter/z-ai/glm-5.2` scopes the diff, writes the debt ledger and the final report, and each parallel audit fans out across `openrouter/z-ai/glm-5.2` + `openrouter/moonshotai/kimi-k3`. The cheap way to get a full review report. |
| `fixer` | yes | The follow-up to a report-only run. Give it a set of findings (as the prompt or an attachment) and it proves each one with a focused regression test **before** touching production code, applies minimal fixes only for the findings that actually went red, independently validates them, then reports a final per-finding verdict (`fixed`, `already-resolved`, `not-reproducible`, `not-automatable`, `blocked`, `not-fixed`). Never promotes an unproven finding to fixed. |
| `review-cc` | **no — report only** | Same shape as `review`, but each audit is paired with a second run on the locally installed [`claude` CLI](https://code.claude.com) (`runner: claude-code`) instead of a second API model — cross-vendor diversity billed to a Claude subscription rather than per token. Requires `claude` on `PATH`. |
| `hunter` | **no — report only** | Repo-wide audit across six specialty tracks (correctness, memory, performance, security, reliability, supply chain), each run on GPT 5.6 Terra xhigh plus one specialty model, then reconciled into a single deduplicated, prioritized consensus report. |
Expand Down
31 changes: 31 additions & 0 deletions prompts/debt-auditor.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,31 @@
# Debt Auditor

You are the **debt-auditor** agent of Convoy's `review` and `refine` pipelines. This is an audit-only phase: do not modify the repository.

## Review scope

Default scope is the attached diff: this branch or pull request against the base ref, plus any uncommitted changes. Read the rest of the repository freely as *context* — to confirm a deferred observation, trace a caller, or judge whether a re-evaluation trigger is namingable — but every ledger entry you record must trace back to a finding the change ships with.

You do not re-litigate the change. The four audits already decided what is must-fix, should-fix, and deferred. Your scope is the **deferred / non-blocking / explicitly-dropped** tail of those audits — the debt the change accepts when it merges — plus `reports/scope.md`, the diff, and `prd.md` so you can name honest re-evaluation triggers.

## Objective

Consolidate every deferred / non-blocking / explicitly-dropped observation from the four audit reports into one tracked debt ledger, and name a concrete **trigger** for each entry: the condition, event, or threshold that should make a future change re-open it. Surface the deferrals that would otherwise rot silently — the ones with no namingable trigger — so the human sees at a glance which deferrals are honest and which are procrastination.

## Workflow

1. Read `prd.md`, `reports/scope.md`, the attached diff, and the four audit reports: `reports/clean-code.md`, `reports/over-engineering.md`, `reports/security.md`, `reports/bugs.md` (in `refine`/`ultra-refine` the single-model variants are named `reports/<step>.md`; in `review`/`review-lite` the two-model variants are named `reports/<step>__<model-slug>.md` — read every variant present).
2. Extract every deferred / non-blocking / explicitly-dropped observation from those reports. Look for sections named `Deferred`, `Non-blocking`, `No-finding notes`, `Deferred/non-blocking`, or equivalent — and for individual findings an auditor downgraded to a non-blocking note.
3. For each, decide whether it is **real debt worth tracking**. Skip taste-only noise and observations that are not debt (an auditor's "looks fine" note is not debt). The ledger is for debt the change ships with, not opinions.
4. Name a concrete **trigger** per entry: the condition, event, or threshold that should re-open it — e.g. "next time this function is touched", "when the second caller appears", "on the security audit of the auth boundary", "before promoting to v0.3.0", "when X exceeds N". If no honest trigger exists, the entry is deferred because it should not be — mark it **`no-trigger`**.
5. Do **not** re-litigate must-fix or should-fix findings — those belong to the audits and the report/triage, not the ledger. The ledger is strictly for what the change ships with.
6. Write the ledger to `reports/debt.md`.

## Report

Return Markdown with:

- **Ledger**: a table with columns `DEBT-N` (stable id, `DEBT-1`, `DEBT-2`, ...), `file:line`, `source` (which audit reported it — `clean-code`, `over-engineering`, `security`, or `bugs`), `debt` (what was deferred, one line), `trigger` (the re-evaluation condition, or `no-trigger`), `severity` (`high|medium|low`). Place `no-trigger` rows last within each severity, and tag the row itself with `no-trigger` in the trigger column.
- **Summary**: exactly three lines — total entries, how many have triggers, how many are `no-trigger` — followed by one line of recommendation (e.g. "schedule a debt pass before v0.3.0", "no deferred debt worth tracking", or "the N `no-trigger` entries should be re-opened or explicitly accepted").

Be decisive and concise. An empty ledger is a valid result — say so plainly when the audits deferred nothing worth tracking.
2 changes: 1 addition & 1 deletion prompts/review-adversary.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ Act as a skeptical second reviewer over the audit reports. Validate which findin

## Workflow

1. Read `prd.md`, `reports/scope.md`, `reports/bugs.md`, `reports/clean-code.md`, `reports/security.md`, `reports/over-engineering.md`, and the attached diff.
1. Read `prd.md`, `reports/scope.md`, `reports/bugs.md`, `reports/clean-code.md`, `reports/security.md`, `reports/over-engineering.md`, `reports/debt.md`, and the attached diff.
2. Challenge every finding:
- Is the evidence present in the diff or adjacent code?
- Is the severity justified?
Expand Down
5 changes: 3 additions & 2 deletions prompts/review-report.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,11 +10,11 @@ Widen scope only when `prd.md` explicitly asks for a repository-wide review.

## Objective

Synthesize every audit that ran before you — clean-code/pattern, security, bug, and over-engineering audits, each produced by two different models — into a single, concise, prioritized findings report. Decide which findings are real and worth acting on, and rank them so a maintainer can act (or defer) without re-reading the raw audits.
Synthesize every audit that ran before you — clean-code/pattern, security, bug, and over-engineering audits, each produced by two different models, plus the `debt-auditor` ledger of deferred findings — into a single, concise, prioritized findings report. Decide which findings are real and worth acting on, and rank them so a maintainer can act (or defer) without re-reading the raw audits.

## Workflow

1. Read `prd.md`, `reports/scope.md`, every attached audit report (both model variants of clean-code, security, bugs, and over-engineering; reads `reports/over-engineering.md` when present), and the attached diff.
1. Read `prd.md`, `reports/scope.md`, every attached audit report (both model variants of clean-code, security, bugs, and over-engineering; reads `reports/over-engineering.md` when present), the debt ledger `reports/debt.md` (when present), and the attached diff.
2. Cross-check the two models behind each audit:
- Where both models raise the same finding, treat it as **high-confidence**.
- Where they disagree, use your own judgment against the diff to keep or drop it.
Expand All @@ -32,6 +32,7 @@ Return Markdown with:
- **Verdict**: one line — overall risk of the change and whether a fix run is warranted (`fix recommended` / `optional cleanup` / `looks good`).
- **Must-fix**: findings that should be fixed before merge. For each: a short id, severity, source audit(s) and which model(s) raised it, file reference, evidence, impact, and the concrete fix.
- **Should-fix**: worthwhile but non-blocking findings, same shape, briefer.
- **Deferred debt**: when `reports/debt.md` is present, surface its `no-trigger` entries here — deferrals with no namingable re-evaluation condition, which would rot silently — and briefly note the entries that do have honest triggers so a maintainer knows what is already tracked.
- **Skip / rejected**: findings you dropped and why (disagreement, no evidence, out of scope, stylistic). Keep this tight — it exists so the human trusts nothing real was silently discarded.
- **Suggested fix scope**: if a fix run is worth it, the minimal ordered set of changes it should make; otherwise state that no run is needed.

Expand Down
2 changes: 2 additions & 0 deletions src/built-in-prompts.ts
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,7 @@ import advisorSystem from "../prompts/advisor-system.md" with { type: "text" }
import advisorTiming from "../prompts/advisor-timing.md" with { type: "text" }
import bugAuditor from "../prompts/bug-auditor.md" with { type: "text" }
import cleanCodeAuditor from "../prompts/clean-code-auditor.md" with { type: "text" }
import debtAuditor from "../prompts/debt-auditor.md" with { type: "text" }
import designPolisher from "../prompts/design-polisher.md" with { type: "text" }
import fixerImplementer from "../prompts/fixer-implementer.md" with { type: "text" }
import fixerReporter from "../prompts/fixer-reporter.md" with { type: "text" }
Expand Down Expand Up @@ -45,6 +46,7 @@ export const builtInPrompts: Record<string, string> = {
"advisor-timing": advisorTiming,
"bug-auditor": bugAuditor,
"clean-code-auditor": cleanCodeAuditor,
"debt-auditor": debtAuditor,
"design-polisher": designPolisher,
"fixer-implementer": fixerImplementer,
"fixer-reporter": fixerReporter,
Expand Down
24 changes: 18 additions & 6 deletions src/pipeline.ts
Original file line number Diff line number Diff line change
Expand Up @@ -117,6 +117,14 @@ export const builtInAgents: readonly AgentSpec[] = [
readOnly: true,
builtIn: true,
},
{
name: "debt-auditor",
description: "Audit-only consolidator that turns deferred audit findings into a tracked debt ledger with re-evaluation triggers",
defaultModel: fallbackModel,
temperature: 0.1,
readOnly: true,
builtIn: true,
},
{
name: "review-adversary",
description: "Adversarial reviewer that validates and filters audit findings before fixes",
Expand Down Expand Up @@ -388,7 +396,7 @@ export const builtInPipelines: Record<string, PipelineSpec> = {
},
review: {
description:
"Report-only PR review: scope, then parallel bug/clean-code/security/over-engineering audits across two models, then one prioritized findings report. Makes no changes.",
"Report-only PR review: scope, then parallel bug/clean-code/security/over-engineering audits across two models, a debt ledger of the deferred findings, then one prioritized findings report. Makes no changes.",
steps: [
{ agent: "review-scope", name: "scope", model: defaultOpusModel, reports: "none", diff: true },
{
Expand All @@ -399,12 +407,13 @@ export const builtInPipelines: Record<string, PipelineSpec> = {
{ agent: "bug-auditor", name: "bugs", models: [fallbackModel, defaultOpusModel], reports: ["scope"] },
],
},
{ agent: "debt-auditor", name: "debt", model: defaultOpusModel, reports: ["scope", "clean-code", "over-engineering", "security", "bugs"] },
{ agent: "review-report", name: "report", model: defaultOpusModel, reports: "all" },
],
},
"review-lite": {
description:
"Like review, but every phase runs on a low-cost model: GLM 5.2 scopes and writes the report, and the audit fan-out pairs GLM 5.2 with Kimi K3 instead of Opus.",
"Like review, but every phase runs on a low-cost model: GLM 5.2 scopes, audits, writes the debt ledger and the report, and the audit fan-out pairs GLM 5.2 with Kimi K3 instead of Opus.",
steps: [
{ agent: "review-scope", name: "scope", model: glmModel, reports: "none", diff: true },
{
Expand All @@ -415,24 +424,26 @@ export const builtInPipelines: Record<string, PipelineSpec> = {
{ agent: "bug-auditor", name: "bugs", models: [glmModel, kimiModel], reports: ["scope"] },
],
},
{ agent: "debt-auditor", name: "debt", model: glmModel, reports: ["scope", "clean-code", "over-engineering", "security", "bugs"] },
{ agent: "review-report", name: "report", model: glmModel, reports: "all" },
],
},
refine: {
description: "Audit-only PR review, adversarial finding triage, targeted fixes, and final validation — applies changes.",
description: "Audit-only PR review, a debt ledger of the deferred findings, adversarial finding triage, targeted fixes, and final validation — applies changes.",
steps: [
{ agent: "review-scope", name: "scope", model: glmModel, reports: "none", diff: true },
{ agent: "bug-auditor", name: "bugs", model: fallbackModel, reports: ["scope"] },
{ agent: "clean-code-auditor", name: "clean-code", model: fallbackModel, reports: ["scope"] },
{ agent: "security-reviewer", name: "security", model: fallbackModel, reports: ["scope"] },
{ agent: "over-engineering-auditor", name: "over-engineering", model: fallbackModel, reports: ["scope"] },
{ agent: "review-adversary", name: "triage", model: defaultOpusModel, reports: ["scope", "bugs", "clean-code", "security", "over-engineering"] },
{ agent: "debt-auditor", name: "debt", model: fallbackModel, reports: ["scope", "bugs", "clean-code", "security", "over-engineering"] },
{ agent: "review-adversary", name: "triage", model: defaultOpusModel, reports: ["scope", "bugs", "clean-code", "security", "over-engineering", "debt"] },
{ agent: "review-fixer", name: "fixes", model: fallbackModel, reports: ["triage"] },
{ agent: "review-validator", name: "validator", model: fallbackModel, reports: "all" },
],
},
"ultra-refine": {
description: "Like refine, but every read-only audit runs in parallel across two models before triage, targeted fixes, and validation.",
description: "Like refine, but every read-only audit runs in parallel across two models, a debt ledger consolidates the deferred findings before triage, and validation runs on Opus.",
steps: [
{ agent: "review-scope", name: "scope", models: [sonnetModel, fallbackModel], reports: "none", diff: true },
{
Expand All @@ -443,7 +454,8 @@ export const builtInPipelines: Record<string, PipelineSpec> = {
{ agent: "over-engineering-auditor", name: "over-engineering", models: [sonnetModel, fallbackModel], reports: ["scope"] },
],
},
{ agent: "review-adversary", name: "triage", model: defaultOpusModel, reports: ["scope", "bugs", "clean-code", "security", "over-engineering"] },
{ agent: "debt-auditor", name: "debt", model: fallbackModel, reports: ["scope", "bugs", "clean-code", "security", "over-engineering"] },
{ agent: "review-adversary", name: "triage", model: defaultOpusModel, reports: ["scope", "bugs", "clean-code", "security", "over-engineering", "debt"] },
{ agent: "review-fixer", name: "fixes", model: sonnetModel, reports: ["triage"] },
{ agent: "review-validator", name: "validator", model: defaultOpusModel, reports: "all" },
],
Expand Down
16 changes: 16 additions & 0 deletions test/agents.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -44,6 +44,22 @@ describe("opencode config", () => {
expect(prompt).toContain("# Convoy Runtime Safety")
})

test("loads debt-auditor prompt with the ledger contract and safety guard rails", () => {
const prompt = loadAgentPrompt("debt-auditor", "/tmp/non-existent-convoy-target")

expect(prompt).toContain("# Debt Auditor")
expect(prompt).toContain("audit-only phase: do not modify the repository")
// The ledger is sourced only from deferred / non-blocking audit findings.
expect(prompt).toContain("deferred")
// Each entry must name a re-evaluation trigger, or be marked no-trigger.
expect(prompt).toContain("trigger")
expect(prompt).toContain("no-trigger")
// Ledger columns and the report path.
expect(prompt).toContain("DEBT-N")
expect(prompt).toContain("reports/debt.md")
expect(prompt).toContain("# Convoy Runtime Safety")
})

test("project agent prompts replace built-ins but keep runtime safety", async () => {
const dir = await mkdtemp(join(tmpdir(), "convoy-agents-"))
try {
Expand Down
Loading