Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
64 changes: 63 additions & 1 deletion .github/workflows/python-quality.yml
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,7 @@ jobs:
skill_authoring: ${{ steps.scopes.outputs.skill_authoring }}
observability_langfuse: ${{ steps.scopes.outputs.observability_langfuse }}
discovery: ${{ steps.scopes.outputs.discovery }}
verification: ${{ steps.scopes.outputs.verification }}
steps:
- uses: actions/checkout@v4
with:
Expand Down Expand Up @@ -120,6 +121,30 @@ jobs:
enable-cache: true
- run: scripts/check-python --package plugins/capability/darrow-discovery/skills/plan-implementation/backend

verification-windows:
name: Verification Python ${{ matrix.python-version }} on windows-latest
needs: changes
if: needs.changes.outputs.verification == 'true'
strategy:
fail-fast: false
matrix:
python-version: ["3.10", "3.11", "3.12", "3.13"]
runs-on: windows-latest
defaults:
run:
shell: bash
env:
UV_PYTHON: ${{ matrix.python-version }}
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: ${{ matrix.python-version }}
- uses: astral-sh/setup-uv@v6
with:
enable-cache: true
- run: scripts/check-python --package plugins/capability/darrow-verification/skills/verify-change/backend

inventory:
name: Python inventory guard
runs-on: ubuntu-latest
Expand Down Expand Up @@ -201,6 +226,36 @@ jobs:
shell: pwsh
run: "& '${{ matrix.test }}'"

verification-fresh-install:
name: Fresh install verification on ${{ matrix.os }}
needs: changes
if: needs.changes.outputs.verification == 'true'
strategy:
fail-fast: false
matrix:
include:
- os: ubuntu-latest
shell: bash
test: plugins/capability/darrow-verification/tests/fresh-install.test.sh
- os: macos-latest
shell: bash
test: plugins/capability/darrow-verification/tests/fresh-install.test.sh
- os: windows-latest
shell: pwsh
test: plugins/capability/darrow-verification/tests/fresh-install.test.ps1
runs-on: ${{ matrix.os }}
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.13"
- uses: astral-sh/setup-uv@v6
- if: matrix.shell == 'bash'
run: bash "${{ matrix.test }}"
- if: matrix.shell == 'pwsh'
shell: pwsh
run: "& '${{ matrix.test }}'"

performance:
name: Python performance invariants
needs: changes
Expand All @@ -222,10 +277,12 @@ jobs:
- package
- skill-authoring-windows
- discovery-windows
- verification-windows
- inventory
- skill-authoring-fresh-install
- observability-fresh-install
- discovery-fresh-install
- verification-fresh-install
- performance
runs-on: ubuntu-latest
steps:
Expand All @@ -238,12 +295,15 @@ jobs:
SKILL_AUTHORING_CHANGED: ${{ needs.changes.outputs.skill_authoring }}
OBSERVABILITY_CHANGED: ${{ needs.changes.outputs.observability_langfuse }}
DISCOVERY_CHANGED: ${{ needs.changes.outputs.discovery }}
VERIFICATION_CHANGED: ${{ needs.changes.outputs.verification }}
WINDOWS_RESULT: ${{ needs.skill-authoring-windows.result }}
DISCOVERY_WINDOWS_RESULT: ${{ needs.discovery-windows.result }}
VERIFICATION_WINDOWS_RESULT: ${{ needs.verification-windows.result }}
INVENTORY_RESULT: ${{ needs.inventory.result }}
SKILL_INSTALL_RESULT: ${{ needs.skill-authoring-fresh-install.result }}
OBSERVABILITY_INSTALL_RESULT: ${{ needs.observability-fresh-install.result }}
DISCOVERY_INSTALL_RESULT: ${{ needs.discovery-fresh-install.result }}
VERIFICATION_INSTALL_RESULT: ${{ needs.verification-fresh-install.result }}
PERFORMANCE_RESULT: ${{ needs.performance.result }}
run: |
scripts/verify-python-quality-results \
Expand All @@ -253,4 +313,6 @@ jobs:
"$OBSERVABILITY_CHANGED" "$OBSERVABILITY_INSTALL_RESULT" \
"$PERFORMANCE_RESULT" \
"$DISCOVERY_CHANGED" "$DISCOVERY_WINDOWS_RESULT" \
"$DISCOVERY_INSTALL_RESULT"
"$DISCOVERY_INSTALL_RESULT" \
"$VERIFICATION_CHANGED" "$VERIFICATION_WINDOWS_RESULT" \
"$VERIFICATION_INSTALL_RESULT"
9 changes: 9 additions & 0 deletions docs/specs/verification.md
Original file line number Diff line number Diff line change
Expand Up @@ -85,6 +85,15 @@ delivery under #155. It does not execute QA or create reviewer-facing evidence p
any next-action field in a no-progress result carries the stop reason or
needed investigation, never an implementation instruction or another repair
followed by reassessment.
- **VF-C8 — Portable assessment renderer.** The retained-report renderer MUST
be a skill-contained, UV-locked Python package with no runtime dependencies.
Its public command MUST preserve the accepted argument order, assessment
bytes, report-link escaping and placement, validation diagnostics,
stdout/stderr separation, and exit statuses on Linux, macOS, and native
Windows. Claude Code and Codex MUST resolve the backend from the installed
skill directory and invoke the same frozen package entrypoint directly.
Fresh copied-artifact checks MUST exercise that installed layout without
repository-relative or sibling-plugin references.

## Consumer handoff

Expand Down
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "darrow-verification",
"version": "0.1.1",
"version": "0.2.0",
"description": "Bounded acceptance verification through replaceable independent code review",
"license": "BUSL-1.1",
"author": { "name": "Björn Rochel", "email": "bjoern@bjro.de" }
Expand Down
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "darrow-verification",
"version": "0.1.1",
"version": "0.2.0",
"description": "Bounded acceptance verification through replaceable independent code review",
"author": { "name": "Björn Rochel", "email": "bjoern@bjro.de" },
"skills": "./skills/",
Expand Down
11 changes: 7 additions & 4 deletions plugins/capability/darrow-verification/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,7 +36,8 @@ The plugin is self-contained and supports Claude Code and Codex. A compatible
review provider needs native fresh-agent support and its own check tools.
The host must allow a bounded provider assessment context and that provider's
independent readers; unavailable depth or authority blocks the operation.
The small report renderer uses Bash 3.2 or Bash 5 and baseline Unix utilities.
The report renderer requires UV and a UV-managed Python 3.10–3.13 runtime. It is
installed from this plugin's locked, dependency-free Python package.
Discover the
review provider by host-advertised intent, then check prerequisites, authorized
effects, result evidence and stop conditions. The provider must independently
Expand Down Expand Up @@ -83,9 +84,11 @@ A compatible review capability must be separately available when invoked.
## Troubleshooting

Missing review, incompatible effects, stale target identity, unreadable original
findings or incomplete required evidence blocks assessment. Supply the exact
missing input or compatible provider, preserving prior evidence for follow-up.
Do not replace the blocked result with self-review or a new repair loop.
findings, incomplete required evidence, a missing backend or lock, UV, or a
supported Python runtime blocks assessment. Supply the exact missing input or
compatible provider, preserving prior evidence for follow-up. Do not replace
the blocked result with self-review, another checkout's renderer, or a new
repair loop.

## License

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -243,25 +243,49 @@ or a lifecycle ledger.

## 5. Render a retained-report handoff

When the selected provider returned a retained local report, use the bundled
[assessment renderer](scripts/render-assessment) for the final handoff. It owns
When the selected provider returned a retained local report, use the packaged
[assessment renderer](backend/src/darrow_verification/assessment.py) for the final handoff. It owns
absolute-reference validation and rendering; do not reproduce its output by
hand. It does not interpret the provider's format or decide findings.

First finish the semantic assessment above. Obtain a unique temporary draft
file with `mktemp` outside the product scope and write that assessment using
the host's file-writing tool. This temporary draft is assessment output, not
an implementation edit or a required evidence package. Then run:
file outside the product scope with the host's native temporary-file facility
(`mktemp` in a POSIX shell or `[System.IO.Path]::GetTempFileName()` in
PowerShell), then write that assessment using the host's file-writing tool.
This temporary draft is assessment output, not an implementation edit or a
required evidence package. Resolve the renderer backend from the loaded skill:

```sh
bash <skill-dir>/scripts/render-assessment --assessment <absolute-draft-file> --provider-report <absolute-provider-report>
- Claude Code: resolve `backend` from the absolute skill directory supplied in
`CLAUDE_SKILL_DIR`; use the host shell's environment-variable and path syntax.
- Codex: take the absolute `SKILL.md` path supplied in the selected skill's
catalog entry and resolve `backend` relative to that file's directory.

Then run:

```text
uv run --quiet --isolated --frozen --no-dev --project "<absolute-backend-path>" darrow-render-assessment --assessment "<absolute-draft-file>" --provider-report "<absolute-provider-report>"
```

Substitute the actual paths; resolve the script from this installed skill,
never another plugin. On success, copy the command's complete stdout unchanged
as the final response. Do not shorten, rewrite or append to it: the renderer's
absolute provider reference is part of the result. Complete semantic checks
before this final command so no later tool or commentary displaces its output.
Substitute the actual paths and invoke this frozen entrypoint directly; do not
look for or create a host-specific launcher. Resolve the backend from this
installed skill, never the user's repository or another plugin. On success,
copy the command's complete stdout unchanged as the final response. Do not
shorten, rewrite or append to it: the renderer's absolute provider reference is
part of the result. Complete semantic checks before this final command so no
later tool or commentary displaces its output.

The final nonempty line must remain the renderer's angle-delimited Markdown
link: the literal prefix `Complete provider result: [report]` immediately
followed by `(<absolute-path>)`. A rewritten destination without the surrounding
angle delimiters is not the renderer's exact stdout and does not complete this
handoff.

That exact link is necessary but not sufficient. Before invocation, read the
temporary draft back and confirm it contains the complete step 4 assessment,
including the criterion-by-criterion evidence and conclusion. After invocation,
return the assessment prefix and final link together as the renderer emitted
them. A response containing only the link, only a conclusion, or a summary of
the draft is an incomplete handoff.

A renderer refusal is a blocked handoff. Return the precise missing/unreadable
evidence gap rather than a partial rendering or a passing summary. Do not
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,57 @@
[project]
name = "darrow-verification"
version = "0.2.0"
description = "Deterministic helpers for Darrow verification"
requires-python = ">=3.10,<3.14"
dependencies = []

[project.scripts]
darrow-render-assessment = "darrow_verification.assessment:entrypoint"

[dependency-groups]
dev = [
"coverage[toml]>=7.10,<8",
"hypothesis>=6.138,<7",
"mypy>=1.18,<2",
"pytest>=8.4,<9",
"ruff>=0.12,<1",
]

[build-system]
requires = ["hatchling"]
build-backend = "hatchling.build"

[tool.uv]
default-groups = []

[tool.hatch.build.targets.wheel]
packages = ["src/darrow_verification"]

[tool.ruff]
target-version = "py310"
line-length = 88
src = ["src", "tests"]

[tool.ruff.lint]
select = ["B", "C4", "C90", "E4", "E7", "E9", "F", "I", "N", "RUF", "SIM", "UP"]
mccabe.max-complexity = 5

[tool.ruff.format]
docstring-code-format = true

[tool.mypy]
python_version = "3.10"
strict = true
files = ["src", "tests"]

[tool.pytest.ini_options]
addopts = ["--strict-config", "--strict-markers"]
testpaths = ["tests"]

[tool.coverage.run]
branch = true
source = ["darrow_verification"]

[tool.coverage.report]
show_missing = true
skip_covered = false
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
"""Deterministic helpers for Darrow verification."""
Loading
Loading