Skip to content

refactor: move plugin validation mechanics to Python - #204

Merged
BjRo merged 4 commits into
mainfrom
refactor/python-validation-mechanics
Sep 19, 2026
Merged

BjRo merged 4 commits into
mainfrom
refactor/python-validation-mechanics

Conversation

@BjRo

@BjRo BjRo commented Sep 19, 2026 •

Copy link
Copy Markdown
Owner

Why

Substantial plugin validation still relied on duplicated shell fixtures, text parsing, and separate Unix/Windows installation probes despite the Python mechanics policy. Test drivers also remained in shell after their fixture mechanics had moved to Python.

What changed

  • Move GitHub ticket and PR-evidence fixtures, review route checks, information-architecture checks, and observability package/hook tests into their owning Python packages.
  • Replace the PR-evidence publication shell suite with 29 pytest cases using the console API, real Git, and an independent fake GitHub process. Replace the exporter-failure, refusal, and strict-mode shell suites with 30 cases that invoke both bash and /bin/bash.
  • Consolidate discovery, verification, and skill-authoring fresh-install checks into Python; convert Langfuse's fresh-install check, remove obsolete Git installation wrappers, and update CI callers.
  • Keep extracted grading helpers hidden by excluding nested eval directories from participant mounts and denying agent access to .git/eval-checks/.
  • Update the mechanics and evaluation invariants, documentation, and both manifests for all eight affected plugins. Clarify that invoking a shell launcher does not require a shell test driver.

Verification

  • CI follow-up: fixed the fresh-install inspector assertion to accept native Windows CRLF output; five regression cases cover LF/CRLF success and invalid-status rejection. The updated skill-authoring package gate passed all 80 tests locally.
  • Hosted CI for e4fba44: all 163 Python quality jobs passed, including native Windows fresh-install validation; documentation CI passed.
  • bun run check:python: passed across all registered packages, 1,448 tests, including 189 Git and 141 Langfuse tests; formatting, lint, strict typing, and statement/branch coverage gates passed.
  • bun test evals/runner/sandbox.test.ts evals/runner/fixture.test.ts evals/runner/review-route-eval-checks.test.ts: 38 passed for the eval-isolation change.
  • bun run lint, bun run lint:ts, bun run lint:shell, bun run typecheck, bun run check:decisions, and bun run check:docs: passed.
  • Fresh copied-install probes passed on macOS. The latest Git and Langfuse probes ran directly with uv run --quiet --frozen --no-dev --project <plugin>/backend python <plugin>/backend/tests/fresh_install.py; Langfuse exercised its registered hook command. Hook execution tests passed with both bash and /bin/bash.
  • Ticket fixture differential checks matched the prior shell implementation across 23 operations; all 18 skills across the eight changed plugins passed structural inspection during the initial conversion.
  • Live single-trial evals during the initial conversion: ticket filter-readonly and IA move-procedure-to-skill passed on Codex and Claude; review reviewer-route-override passed on Codex, including the final protected layout. All changed eval fixtures prepared successfully; dry preparation is not behavioral pass evidence.
  • Two Claude review trials failed because the candidate omitted required route/report artifacts; activation passed. These failures remain visible and checks were not relaxed. The subsequent test-driver conversions do not change skill instructions or eval cases.

Review notes

Focus review on fixture behavior parity and the separation between candidate access and grading access. Production review instructions and routing implementation are unchanged.

Native Windows support for the Langfuse backend and hook entrypoint is explicitly separate in #205; the Unix launcher and locking implementation remain unchanged. The README now states that limitation. Other plugins retain their Linux/macOS/Windows CI matrices; native Linux/Windows and Bash 5 were not verified locally. Live evals are small smoke samples, not evidence of a reliability improvement. Raw eval evidence remains under ignored local results paths.

Checklist

  • I have read and followed CONTRIBUTING.md, including the contribution
    licensing terms.
  • I added or updated the applicable invariant before implementation, or
    this change does not affect a capability invariant.
  • I added or updated colocated evals, or this change does not affect skill
    behavior.
  • I confirmed that each changed plugin remains self-contained, or this
    change does not affect plugin content.
  • I ran bun run check:python, or this change does not affect registered
    Python packages or their repository quality infrastructure.

@BjRo
BjRo merged commit a03b1da into main Sep 19, 2026
165 checks passed
@BjRo
BjRo deleted the refactor/python-validation-mechanics branch September 19, 2026 11:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant