From 00dc234be636fce3a507211c72b46009e6f70eaf Mon Sep 17 00:00:00 2001 From: Josemi Date: Fri, 31 Jul 2026 18:10:31 +0200 Subject: [PATCH] =?UTF-8?q?feat:=20debt-auditor=20=E2=80=94=20consolidate?= =?UTF-8?q?=20deferred=20findings=20into=20a=20tracked=20debt?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - tests — feat(review): debt-auditor - Pattern auditor report — feat(review): debt-auditor - Implementer report — feat(review): debt-auditor - debt-auditor work in progress (interrupted by quota) --- README.md | 8 ++++---- prompts/debt-auditor.md | 31 +++++++++++++++++++++++++++++++ prompts/review-adversary.md | 2 +- prompts/review-report.md | 5 +++-- src/built-in-prompts.ts | 2 ++ src/pipeline.ts | 24 ++++++++++++++++++------ test/agents.test.ts | 16 ++++++++++++++++ test/config.test.ts | 21 ++++++++++++++++++++- test/pipeline.test.ts | 29 +++++++++++++++++++++++++++-- 9 files changed, 122 insertions(+), 16 deletions(-) create mode 100644 prompts/debt-auditor.md diff --git a/README.md b/README.md index d854899..bd813da 100644 --- a/README.md +++ b/README.md @@ -58,10 +58,10 @@ Convoy ships these pipelines; select one with `-p/--pipeline` (no config needed) | `implement-lite` | yes | Same workflow and agents as `implement`, but the code-writing phases (`implementer`, `patterns`, `security`, `tests`) all drop to `openrouter/z-ai/glm-5.2` for a lower-cost run; `design` runs on Kimi K3 and `adversarial` on Opus. | | `implement-advised` | yes | Same workflow and step names as `implement`, but the `implementer` phase runs on GPT 5.6 Terra xhigh and **consults GPT 5.6 Sol as an advisor** at its decision points. The audits (`patterns`, `security`, `tests`) run unadvised on `openrouter/z-ai/glm-5.2`, `design` on Kimi K3, and `adversarial` keeps Opus owning its own loop. Directly comparable with `implement-lite`: same audit executors, and the implementation step is the single variable between them. | | `ultra-implement` | yes | Like `implement`, but the pattern/security/adversarial reviews of the initial diff run in parallel across two models feeding a triage step, and the run ends with an audit-only final review, a fixer that applies only blocking findings, and a final validator. | -| `refine` | yes | Audit the current diff (scope → bugs → clean-code → security), triage the findings adversarially, apply the accepted fixes, then validate them. | -| `ultra-refine` | yes | Like `refine`, but every read-only audit is fanned out across two models before triage, fixes, and validation. | -| `review` | **no — report only** | Scope the diff, run the bug / clean-code(+patterns) / security audits **in parallel across two models each**, then a single step synthesizes everything into one prioritized findings report. Makes no changes; the run's output is `reports/report.md`, which you read to decide whether to follow up with a `refine` run. | -| `review-lite` | **no — report only** | Same shape as `review`, but nothing runs on Opus: `openrouter/z-ai/glm-5.2` scopes the diff and writes the final report, and each parallel audit fans out across `openrouter/z-ai/glm-5.2` + `openrouter/moonshotai/kimi-k3`. The cheap way to get a full review report. | +| `refine` | yes | Audit the current diff (scope → bugs → clean-code → security → over-engineering), consolidate the deferred findings into a debt ledger, triage the findings adversarially, apply the accepted fixes, then validate them. | +| `ultra-refine` | yes | Like `refine`, but every read-only audit is fanned out across two models before a debt ledger, triage, fixes, and validation. | +| `review` | **no — report only** | Scope the diff, run the bug / clean-code(+patterns) / over-engineering / security audits **in parallel across two models each**, consolidate the deferred findings into a debt ledger, then a single step synthesizes everything into one prioritized findings report. Makes no changes; the run's output is `reports/report.md`, which you read to decide whether to follow up with a `refine` run. | +| `review-lite` | **no — report only** | Same shape as `review`, but nothing runs on Opus: `openrouter/z-ai/glm-5.2` scopes the diff, writes the debt ledger and the final report, and each parallel audit fans out across `openrouter/z-ai/glm-5.2` + `openrouter/moonshotai/kimi-k3`. The cheap way to get a full review report. | | `fixer` | yes | The follow-up to a report-only run. Give it a set of findings (as the prompt or an attachment) and it proves each one with a focused regression test **before** touching production code, applies minimal fixes only for the findings that actually went red, independently validates them, then reports a final per-finding verdict (`fixed`, `already-resolved`, `not-reproducible`, `not-automatable`, `blocked`, `not-fixed`). Never promotes an unproven finding to fixed. | | `review-cc` | **no — report only** | Same shape as `review`, but each audit is paired with a second run on the locally installed [`claude` CLI](https://code.claude.com) (`runner: claude-code`) instead of a second API model — cross-vendor diversity billed to a Claude subscription rather than per token. Requires `claude` on `PATH`. | | `hunter` | **no — report only** | Repo-wide audit across six specialty tracks (correctness, memory, performance, security, reliability, supply chain), each run on GPT 5.6 Terra xhigh plus one specialty model, then reconciled into a single deduplicated, prioritized consensus report. | diff --git a/prompts/debt-auditor.md b/prompts/debt-auditor.md new file mode 100644 index 0000000..258d082 --- /dev/null +++ b/prompts/debt-auditor.md @@ -0,0 +1,31 @@ +# Debt Auditor + +You are the **debt-auditor** agent of Convoy's `review` and `refine` pipelines. This is an audit-only phase: do not modify the repository. + +## Review scope + +Default scope is the attached diff: this branch or pull request against the base ref, plus any uncommitted changes. Read the rest of the repository freely as *context* — to confirm a deferred observation, trace a caller, or judge whether a re-evaluation trigger is namingable — but every ledger entry you record must trace back to a finding the change ships with. + +You do not re-litigate the change. The four audits already decided what is must-fix, should-fix, and deferred. Your scope is the **deferred / non-blocking / explicitly-dropped** tail of those audits — the debt the change accepts when it merges — plus `reports/scope.md`, the diff, and `prd.md` so you can name honest re-evaluation triggers. + +## Objective + +Consolidate every deferred / non-blocking / explicitly-dropped observation from the four audit reports into one tracked debt ledger, and name a concrete **trigger** for each entry: the condition, event, or threshold that should make a future change re-open it. Surface the deferrals that would otherwise rot silently — the ones with no namingable trigger — so the human sees at a glance which deferrals are honest and which are procrastination. + +## Workflow + +1. Read `prd.md`, `reports/scope.md`, the attached diff, and the four audit reports: `reports/clean-code.md`, `reports/over-engineering.md`, `reports/security.md`, `reports/bugs.md` (in `refine`/`ultra-refine` the single-model variants are named `reports/.md`; in `review`/`review-lite` the two-model variants are named `reports/__.md` — read every variant present). +2. Extract every deferred / non-blocking / explicitly-dropped observation from those reports. Look for sections named `Deferred`, `Non-blocking`, `No-finding notes`, `Deferred/non-blocking`, or equivalent — and for individual findings an auditor downgraded to a non-blocking note. +3. For each, decide whether it is **real debt worth tracking**. Skip taste-only noise and observations that are not debt (an auditor's "looks fine" note is not debt). The ledger is for debt the change ships with, not opinions. +4. Name a concrete **trigger** per entry: the condition, event, or threshold that should re-open it — e.g. "next time this function is touched", "when the second caller appears", "on the security audit of the auth boundary", "before promoting to v0.3.0", "when X exceeds N". If no honest trigger exists, the entry is deferred because it should not be — mark it **`no-trigger`**. +5. Do **not** re-litigate must-fix or should-fix findings — those belong to the audits and the report/triage, not the ledger. The ledger is strictly for what the change ships with. +6. Write the ledger to `reports/debt.md`. + +## Report + +Return Markdown with: + +- **Ledger**: a table with columns `DEBT-N` (stable id, `DEBT-1`, `DEBT-2`, ...), `file:line`, `source` (which audit reported it — `clean-code`, `over-engineering`, `security`, or `bugs`), `debt` (what was deferred, one line), `trigger` (the re-evaluation condition, or `no-trigger`), `severity` (`high|medium|low`). Place `no-trigger` rows last within each severity, and tag the row itself with `no-trigger` in the trigger column. +- **Summary**: exactly three lines — total entries, how many have triggers, how many are `no-trigger` — followed by one line of recommendation (e.g. "schedule a debt pass before v0.3.0", "no deferred debt worth tracking", or "the N `no-trigger` entries should be re-opened or explicitly accepted"). + +Be decisive and concise. An empty ledger is a valid result — say so plainly when the audits deferred nothing worth tracking. diff --git a/prompts/review-adversary.md b/prompts/review-adversary.md index 284d117..a06f4fe 100644 --- a/prompts/review-adversary.md +++ b/prompts/review-adversary.md @@ -14,7 +14,7 @@ Act as a skeptical second reviewer over the audit reports. Validate which findin ## Workflow -1. Read `prd.md`, `reports/scope.md`, `reports/bugs.md`, `reports/clean-code.md`, `reports/security.md`, `reports/over-engineering.md`, and the attached diff. +1. Read `prd.md`, `reports/scope.md`, `reports/bugs.md`, `reports/clean-code.md`, `reports/security.md`, `reports/over-engineering.md`, `reports/debt.md`, and the attached diff. 2. Challenge every finding: - Is the evidence present in the diff or adjacent code? - Is the severity justified? diff --git a/prompts/review-report.md b/prompts/review-report.md index f9a3ba8..68cd853 100644 --- a/prompts/review-report.md +++ b/prompts/review-report.md @@ -10,11 +10,11 @@ Widen scope only when `prd.md` explicitly asks for a repository-wide review. ## Objective -Synthesize every audit that ran before you — clean-code/pattern, security, bug, and over-engineering audits, each produced by two different models — into a single, concise, prioritized findings report. Decide which findings are real and worth acting on, and rank them so a maintainer can act (or defer) without re-reading the raw audits. +Synthesize every audit that ran before you — clean-code/pattern, security, bug, and over-engineering audits, each produced by two different models, plus the `debt-auditor` ledger of deferred findings — into a single, concise, prioritized findings report. Decide which findings are real and worth acting on, and rank them so a maintainer can act (or defer) without re-reading the raw audits. ## Workflow -1. Read `prd.md`, `reports/scope.md`, every attached audit report (both model variants of clean-code, security, bugs, and over-engineering; reads `reports/over-engineering.md` when present), and the attached diff. +1. Read `prd.md`, `reports/scope.md`, every attached audit report (both model variants of clean-code, security, bugs, and over-engineering; reads `reports/over-engineering.md` when present), the debt ledger `reports/debt.md` (when present), and the attached diff. 2. Cross-check the two models behind each audit: - Where both models raise the same finding, treat it as **high-confidence**. - Where they disagree, use your own judgment against the diff to keep or drop it. @@ -32,6 +32,7 @@ Return Markdown with: - **Verdict**: one line — overall risk of the change and whether a fix run is warranted (`fix recommended` / `optional cleanup` / `looks good`). - **Must-fix**: findings that should be fixed before merge. For each: a short id, severity, source audit(s) and which model(s) raised it, file reference, evidence, impact, and the concrete fix. - **Should-fix**: worthwhile but non-blocking findings, same shape, briefer. +- **Deferred debt**: when `reports/debt.md` is present, surface its `no-trigger` entries here — deferrals with no namingable re-evaluation condition, which would rot silently — and briefly note the entries that do have honest triggers so a maintainer knows what is already tracked. - **Skip / rejected**: findings you dropped and why (disagreement, no evidence, out of scope, stylistic). Keep this tight — it exists so the human trusts nothing real was silently discarded. - **Suggested fix scope**: if a fix run is worth it, the minimal ordered set of changes it should make; otherwise state that no run is needed. diff --git a/src/built-in-prompts.ts b/src/built-in-prompts.ts index a4096ec..d102118 100644 --- a/src/built-in-prompts.ts +++ b/src/built-in-prompts.ts @@ -3,6 +3,7 @@ import advisorSystem from "../prompts/advisor-system.md" with { type: "text" } import advisorTiming from "../prompts/advisor-timing.md" with { type: "text" } import bugAuditor from "../prompts/bug-auditor.md" with { type: "text" } import cleanCodeAuditor from "../prompts/clean-code-auditor.md" with { type: "text" } +import debtAuditor from "../prompts/debt-auditor.md" with { type: "text" } import designPolisher from "../prompts/design-polisher.md" with { type: "text" } import fixerImplementer from "../prompts/fixer-implementer.md" with { type: "text" } import fixerReporter from "../prompts/fixer-reporter.md" with { type: "text" } @@ -45,6 +46,7 @@ export const builtInPrompts: Record = { "advisor-timing": advisorTiming, "bug-auditor": bugAuditor, "clean-code-auditor": cleanCodeAuditor, + "debt-auditor": debtAuditor, "design-polisher": designPolisher, "fixer-implementer": fixerImplementer, "fixer-reporter": fixerReporter, diff --git a/src/pipeline.ts b/src/pipeline.ts index 3b5e15a..4dccf93 100644 --- a/src/pipeline.ts +++ b/src/pipeline.ts @@ -117,6 +117,14 @@ export const builtInAgents: readonly AgentSpec[] = [ readOnly: true, builtIn: true, }, + { + name: "debt-auditor", + description: "Audit-only consolidator that turns deferred audit findings into a tracked debt ledger with re-evaluation triggers", + defaultModel: fallbackModel, + temperature: 0.1, + readOnly: true, + builtIn: true, + }, { name: "review-adversary", description: "Adversarial reviewer that validates and filters audit findings before fixes", @@ -388,7 +396,7 @@ export const builtInPipelines: Record = { }, review: { description: - "Report-only PR review: scope, then parallel bug/clean-code/security/over-engineering audits across two models, then one prioritized findings report. Makes no changes.", + "Report-only PR review: scope, then parallel bug/clean-code/security/over-engineering audits across two models, a debt ledger of the deferred findings, then one prioritized findings report. Makes no changes.", steps: [ { agent: "review-scope", name: "scope", model: defaultOpusModel, reports: "none", diff: true }, { @@ -399,12 +407,13 @@ export const builtInPipelines: Record = { { agent: "bug-auditor", name: "bugs", models: [fallbackModel, defaultOpusModel], reports: ["scope"] }, ], }, + { agent: "debt-auditor", name: "debt", model: defaultOpusModel, reports: ["scope", "clean-code", "over-engineering", "security", "bugs"] }, { agent: "review-report", name: "report", model: defaultOpusModel, reports: "all" }, ], }, "review-lite": { description: - "Like review, but every phase runs on a low-cost model: GLM 5.2 scopes and writes the report, and the audit fan-out pairs GLM 5.2 with Kimi K3 instead of Opus.", + "Like review, but every phase runs on a low-cost model: GLM 5.2 scopes, audits, writes the debt ledger and the report, and the audit fan-out pairs GLM 5.2 with Kimi K3 instead of Opus.", steps: [ { agent: "review-scope", name: "scope", model: glmModel, reports: "none", diff: true }, { @@ -415,24 +424,26 @@ export const builtInPipelines: Record = { { agent: "bug-auditor", name: "bugs", models: [glmModel, kimiModel], reports: ["scope"] }, ], }, + { agent: "debt-auditor", name: "debt", model: glmModel, reports: ["scope", "clean-code", "over-engineering", "security", "bugs"] }, { agent: "review-report", name: "report", model: glmModel, reports: "all" }, ], }, refine: { - description: "Audit-only PR review, adversarial finding triage, targeted fixes, and final validation — applies changes.", + description: "Audit-only PR review, a debt ledger of the deferred findings, adversarial finding triage, targeted fixes, and final validation — applies changes.", steps: [ { agent: "review-scope", name: "scope", model: glmModel, reports: "none", diff: true }, { agent: "bug-auditor", name: "bugs", model: fallbackModel, reports: ["scope"] }, { agent: "clean-code-auditor", name: "clean-code", model: fallbackModel, reports: ["scope"] }, { agent: "security-reviewer", name: "security", model: fallbackModel, reports: ["scope"] }, { agent: "over-engineering-auditor", name: "over-engineering", model: fallbackModel, reports: ["scope"] }, - { agent: "review-adversary", name: "triage", model: defaultOpusModel, reports: ["scope", "bugs", "clean-code", "security", "over-engineering"] }, + { agent: "debt-auditor", name: "debt", model: fallbackModel, reports: ["scope", "bugs", "clean-code", "security", "over-engineering"] }, + { agent: "review-adversary", name: "triage", model: defaultOpusModel, reports: ["scope", "bugs", "clean-code", "security", "over-engineering", "debt"] }, { agent: "review-fixer", name: "fixes", model: fallbackModel, reports: ["triage"] }, { agent: "review-validator", name: "validator", model: fallbackModel, reports: "all" }, ], }, "ultra-refine": { - description: "Like refine, but every read-only audit runs in parallel across two models before triage, targeted fixes, and validation.", + description: "Like refine, but every read-only audit runs in parallel across two models, a debt ledger consolidates the deferred findings before triage, and validation runs on Opus.", steps: [ { agent: "review-scope", name: "scope", models: [sonnetModel, fallbackModel], reports: "none", diff: true }, { @@ -443,7 +454,8 @@ export const builtInPipelines: Record = { { agent: "over-engineering-auditor", name: "over-engineering", models: [sonnetModel, fallbackModel], reports: ["scope"] }, ], }, - { agent: "review-adversary", name: "triage", model: defaultOpusModel, reports: ["scope", "bugs", "clean-code", "security", "over-engineering"] }, + { agent: "debt-auditor", name: "debt", model: fallbackModel, reports: ["scope", "bugs", "clean-code", "security", "over-engineering"] }, + { agent: "review-adversary", name: "triage", model: defaultOpusModel, reports: ["scope", "bugs", "clean-code", "security", "over-engineering", "debt"] }, { agent: "review-fixer", name: "fixes", model: sonnetModel, reports: ["triage"] }, { agent: "review-validator", name: "validator", model: defaultOpusModel, reports: "all" }, ], diff --git a/test/agents.test.ts b/test/agents.test.ts index ea51281..e0bc650 100644 --- a/test/agents.test.ts +++ b/test/agents.test.ts @@ -44,6 +44,22 @@ describe("opencode config", () => { expect(prompt).toContain("# Convoy Runtime Safety") }) + test("loads debt-auditor prompt with the ledger contract and safety guard rails", () => { + const prompt = loadAgentPrompt("debt-auditor", "/tmp/non-existent-convoy-target") + + expect(prompt).toContain("# Debt Auditor") + expect(prompt).toContain("audit-only phase: do not modify the repository") + // The ledger is sourced only from deferred / non-blocking audit findings. + expect(prompt).toContain("deferred") + // Each entry must name a re-evaluation trigger, or be marked no-trigger. + expect(prompt).toContain("trigger") + expect(prompt).toContain("no-trigger") + // Ledger columns and the report path. + expect(prompt).toContain("DEBT-N") + expect(prompt).toContain("reports/debt.md") + expect(prompt).toContain("# Convoy Runtime Safety") + }) + test("project agent prompts replace built-ins but keep runtime safety", async () => { const dir = await mkdtemp(join(tmpdir(), "convoy-agents-")) try { diff --git a/test/config.test.ts b/test/config.test.ts index 8ac3324..67235b1 100644 --- a/test/config.test.ts +++ b/test/config.test.ts @@ -367,6 +367,7 @@ describe("agent registry", () => { "clean-code-auditor", "over-engineering-auditor", "security-reviewer", + "debt-auditor", "review-adversary", "review-fixer", "review-validator", @@ -389,6 +390,24 @@ describe("agent registry", () => { "hunter-max-report", ]) }) + + test("debt-auditor carries the audit shape and sits between security-reviewer and review-adversary", () => { + const names = builtInAgents.map((agent) => agent.name) + const debtIndex = names.indexOf("debt-auditor") + expect(debtIndex).toBeGreaterThan(-1) + expect(names[debtIndex - 1]).toBe("security-reviewer") + expect(names[debtIndex + 1]).toBe("review-adversary") + + const debt = builtInAgents[debtIndex]! + // The audit family shape: read-only, low temperature, built-in, on the fallback model. + expect(debt).toMatchObject({ + temperature: 0.1, + readOnly: true, + builtIn: true, + defaultModel: "openai/gpt-5.6-terra#xhigh", + }) + expect(debt.description).toContain("debt ledger") + }) }) describe("pipeline selection", () => { @@ -926,7 +945,7 @@ describe("materializing built-in pipelines", () => { if (originalGroup === undefined || !isParallelSpec(originalGroup)) throw new Error("expected a parallel block") const originalMember = originalGroup.parallel[0] if (typeof originalMember === "string") throw new Error("expected a member object") - expect(original.steps).toHaveLength(3) + expect(original.steps).toHaveLength(4) expect(originalGroup.parallel).toHaveLength(4) expect(originalMember.models).toHaveLength(2) expect(originalMember.name).toBe("clean-code") diff --git a/test/pipeline.test.ts b/test/pipeline.test.ts index 3cafbb8..5e29d19 100644 --- a/test/pipeline.test.ts +++ b/test/pipeline.test.ts @@ -206,7 +206,7 @@ describe("built-in review pipeline", () => { expect(pipeline.steps.some((step) => step.type === "human")).toBe(false) }) - test("fans each audit across GPT 5.6 Terra xhigh + opus and feeds a single report step with every audit", () => { + test("fans each audit across GPT 5.6 Terra xhigh + opus, consolidates deferred findings into a debt ledger, and feeds every audit plus the ledger to a single report step", () => { const pipeline = review() expect(stepNames(pipeline)).toEqual([ "scope", @@ -218,9 +218,27 @@ describe("built-in review pipeline", () => { "security__anthropic-claude-opus-5", "bugs__openai-gpt-5-6-terra-xhigh", "bugs__anthropic-claude-opus-5", + "debt", "report", ]) + const byName = Object.fromEntries( + pipeline.steps.filter((step): step is AgentStep => step.type === "agent").map((step) => [step.name, step]), + ) + expect(byName.debt).toMatchObject({ model: defaultOpusModel, readOnly: true }) + expect(byName.debt?.inputFiles).toEqual([ + "prd.md", + "reports/scope.md", + "reports/clean-code__openai-gpt-5-6-terra-xhigh.md", + "reports/clean-code__anthropic-claude-opus-5.md", + "reports/over-engineering__openai-gpt-5-6-terra-xhigh.md", + "reports/over-engineering__anthropic-claude-opus-5.md", + "reports/security__openai-gpt-5-6-terra-xhigh.md", + "reports/security__anthropic-claude-opus-5.md", + "reports/bugs__openai-gpt-5-6-terra-xhigh.md", + "reports/bugs__anthropic-claude-opus-5.md", + ]) + const report = pipeline.steps.find((step): step is AgentStep => step.type === "agent" && step.stepName === "report") expect(report?.inputFiles).toEqual([ "prd.md", @@ -233,6 +251,7 @@ describe("built-in review pipeline", () => { "reports/security__anthropic-claude-opus-5.md", "reports/bugs__openai-gpt-5-6-terra-xhigh.md", "reports/bugs__anthropic-claude-opus-5.md", + "reports/debt.md", ]) }) }) @@ -248,7 +267,7 @@ describe("built-in review-lite pipeline", () => { expect(pipeline.steps.some((step) => step.type === "human")).toBe(false) }) - test("runs entirely on low-cost models: GLM 5.2 scopes and reports, and the fan-out pairs GLM 5.2 with Kimi K3", () => { + test("runs entirely on low-cost models: GLM 5.2 scopes, audits, writes the debt ledger and the report, and the fan-out pairs GLM 5.2 with Kimi K3", () => { const pipeline = reviewLite() expect(stepNames(pipeline)).toEqual([ "scope", @@ -260,6 +279,7 @@ describe("built-in review-lite pipeline", () => { "security__openrouter-moonshotai-kimi-k3", "bugs__openrouter-z-ai-glm-5-2", "bugs__openrouter-moonshotai-kimi-k3", + "debt", "report", ]) @@ -267,6 +287,7 @@ describe("built-in review-lite pipeline", () => { pipeline.steps.filter((step): step is AgentStep => step.type === "agent").map((step) => [step.name, step]), ) expect(byName.scope?.model).toBe("openrouter/z-ai/glm-5.2") + expect(byName.debt?.model).toBe("openrouter/z-ai/glm-5.2") expect(byName.report?.model).toBe("openrouter/z-ai/glm-5.2") expect(byName.report?.inputFiles).toEqual([ "prd.md", @@ -279,6 +300,7 @@ describe("built-in review-lite pipeline", () => { "reports/security__openrouter-moonshotai-kimi-k3.md", "reports/bugs__openrouter-z-ai-glm-5-2.md", "reports/bugs__openrouter-moonshotai-kimi-k3.md", + "reports/debt.md", ]) }) @@ -300,6 +322,7 @@ describe("built-in refine pipeline", () => { expect(byName["clean-code"]).toMatchObject({ model: "openai/gpt-5.6-terra", variant: "xhigh" }) expect(byName.security).toMatchObject({ model: "openai/gpt-5.6-terra", variant: "xhigh" }) expect(byName["over-engineering"]).toMatchObject({ model: "openai/gpt-5.6-terra", variant: "xhigh" }) + expect(byName.debt).toMatchObject({ model: "openai/gpt-5.6-terra", variant: "xhigh" }) expect(byName.triage).toMatchObject({ model: "anthropic/claude-opus-5" }) expect(byName.triage.inputFiles).toEqual([ "prd.md", @@ -308,6 +331,7 @@ describe("built-in refine pipeline", () => { "reports/clean-code.md", "reports/security.md", "reports/over-engineering.md", + "reports/debt.md", ]) expect(byName.fixes).toMatchObject({ model: "openai/gpt-5.6-terra", variant: "xhigh" }) expect(byName.validator).toMatchObject({ model: "openai/gpt-5.6-terra", variant: "xhigh" }) @@ -332,6 +356,7 @@ describe("built-in ultra-refine pipeline", () => { "reports/security__openai-gpt-5-6-terra-xhigh.md", "reports/over-engineering__openrouter-anthropic-claude-sonnet-5.md", "reports/over-engineering__openai-gpt-5-6-terra-xhigh.md", + "reports/debt.md", ]) }) })