Assurance harness targeting golf-web-app — a portfolio demonstration of modern, automation-led quality engineering.
Strategy and risk register are the source of truth — docs/test-strategy.md tracks layer-by-layer status and findings, and docs/risk-register.md tracks the risks and mitigations.
Currently in place across the two repos:
- Per-PR CI gates (
assurance-harness/.github/workflows/assurance.yml): lint (ruff), harness pytest, contract (Schemathesis), functional (Playwright), TypeScript E2E (Playwright/TS — the polyglote2e_ts/layer, B20), accessibility (axe-core), performance (k6), data quality (pandera), and a shift-left security gate — Bandit + CodeQL (SAST), pip-audit (SCA), gitleaks (secrets), OWASP ZAP baseline (DAST) undernonfunctional/security/(B1) with ratchet gating. Every gate runs against an ephemeral SUT brought up fromgolf-web-app's source. - Cross-repo shift-left (B19, F-034):
assurance.ymlis also a reusableworkflow_callthatgolf-web-app's own PR pipeline invokes against its head SHA —build-and-publishis gated behind it and a branch-protection ruleset makes the result a required status check, so no SUT change merges or ships unassured. The harness is pinned to a tag (@v1) so its evolution can't surprise-break the SUT pipeline. - Local on-demand agent layers (Ollama-backed): AI evaluation harness — two methods over the booking assistant: golden-set accuracy and metamorphic invariance (
ai_evaluation/, phase 8 + B3/F-035), risk-prioritisation agent with a deterministic register pre-filter (risk_agent/, phase 9 + phase 13), triage agent (triage_agent/, phase 10), exploratory agent — API + UI surfaces plus spec-aware auth-bypass probing (explore_agent/, phase 12), and security agent — judges the B1 scanner findings (FP-vs-real + disposition + register R-ID) and reconciles the SCA allowlist (security_agent/, B1c). All five use a local Ollama runtime so the per-PR path stays fast and reproducible; each carries a deterministic golden-set eval tier, and their evidence artefacts are committed under each module'sreports/dir. - Production-style observability: Prometheus + Grafana stack (
observability/, phase 11) scraping the SUT's/metrics; SLO thresholds on the dashboard match the k6 perf gate's pre-merge budget. Closes R-013. - Resilience / chaos testing (B4, F-037 + F-038): a local-on-demand
chaos/pillar that fault-injects the compose stack — one representative fault per failure axis — and asserts graceful degradation + automatic recovery. v1: a real DB-outage fail-fast gap (a paused DB makes/coursehang — no client-side timeout, tracked as R-021) and process-death auto-recovery via the restart policy (RestartCount-proven;docker killis an operator-stop the policy ignores, so the crash is injected in-container). v2 (grey failure): toxiproxy injects ~600ms on the DB path —/coursestays 200 but breaches the 500ms p95 SLO, confirmed live in Prometheus (p95 2425ms), proving the observability stack catches a slow dependency. Scenario logic is gate-tested with fakes; the live run is local-only. - Tests of the harness's own agents: a gated regression suite (
tests/agents/, phase 12 v2 v2 + F-024) runningrisk_agent,triage_agent, and theexplore_agentjudge N times against fixed inputs, asserting schema/vocab invariants and stability under LLM jitter — including a non-blinding positive control on the explore judge. - Documented findings: F-001 through F-038 captured in the strategy with diagnosis, fix, and generalisation — including the full security lifecycle (F-029 detect → F-030 judge → F-031 reconcile → F-032 remediate + re-arm → F-033 SARIF-native + secrets), the B19 cross-repo shift-left wiring (F-034), B3 metamorphic testing of the AI feature — invariance (F-035) + directional (F-036), and B4 resilience/chaos testing — v1 outage + process-death (F-037) and v2 grey-failure latency with a live Prometheus SLO-breach (F-038).
- Python 3.12, managed with uv
- pytest as the test runner, with JUnit + HTML reporting
- Schemathesis for property-based API contract tests
- Playwright for UI / E2E functional tests
- TypeScript + Playwright for a second, polyglot E2E layer (
e2e_ts/) — the same journeys in TS, on a Node 22 / npm toolchain (B20) - axe-core (axe-playwright-python) for WCAG 2.1 A/AA accessibility checks
- k6 for performance budgets (thresholds-as-code)
- pandera for data-quality checks on the live database
- Bandit + CodeQL (SAST), pip-audit (SCA), gitleaks (secrets), OWASP ZAP (DAST baseline) for the security gate
- Ollama for the five agent layers — AI evaluation, risk-prioritisation, triage, exploration, and security triage (local on-demand only — not in CI)
- ruff for lint + format
- GitHub Actions for CI
# Install everything (uv creates and manages the venv)
uv sync --dev
# Run the harness's own pytest suite (no SUT needed)
uv run pytest
# Lint and format check
uv run ruff check .
uv run ruff format --check .assurance-harness/
├── pyproject.toml
├── .python-version
├── schemathesis.toml # contract-test check config
├── docs/
│ ├── test-strategy.md # how we assure, layer status, findings to date
│ └── risk-register.md # risks tracked and their mitigations
├── contract/ # phase 4: Schemathesis API contract tests
│ ├── conftest.py
│ └── test_api_contract.py
├── functional/ # phase 3 + 7: Playwright UI / E2E journeys
│ ├── conftest.py # Playwright config — see F-009 for expect timeout
│ ├── test_public_pages.py
│ ├── test_member_journey.py
│ ├── test_access_control.py
│ └── test_booking_assistant.py
├── e2e_ts/ # B20: TypeScript / Playwright E2E (polyglot twin of functional/)
│ ├── package.json # pinned: Node 22 LTS, exact @playwright/test
│ ├── playwright.config.ts # baseURL from env, CI retries, gating
│ ├── tsconfig.json
│ ├── fixtures.ts # creds, login, memberPage, R-018 scroll shim
│ ├── components/ # component objects (NavBar — shared nav)
│ ├── pages/ # page objects (Home, Login, MemberDashboard, Booking)
│ └── tests/ # member-journey, public-pages, access-control
├── nonfunctional/
│ ├── accessibility/ # phase 5a: axe-core WCAG 2.1 A/AA sweep
│ │ ├── conftest.py
│ │ └── test_accessibility.py
│ ├── performance/ # phase 5b: k6 thresholds-as-code
│ │ └── api_load.js
│ ├── security/ # B1: shift-left security gate
│ │ ├── scan.py # SAST (Bandit) + SCA (pip-audit) + secrets (gitleaks), ratchet gating
│ │ └── sca_allowlist.txt # accepted CVEs (re-armed to empty after F-032)
│ └── reports/ # CI-only evidence (a11y + perf), gitignored,
│ # uploaded as GitHub Actions artefacts
├── data_quality/ # phase 6: pandera schemas + invariants
│ ├── conftest.py
│ └── test_data_quality.py
├── ai_evaluation/ # phase 8: black-box golden-set scoring
│ ├── evaluator.py # field equality + safety scoring
│ ├── judge.py # LLM-judge tier (holistic + fuzzy)
│ ├── run.py # CLI: single / compare / --with-judge
│ ├── golden_set.yaml # ground truth (40 labelled cases)
│ └── reports/ # committed evidence
├── risk_agent/ # phase 9: PR diff → ranked test plan
│ ├── register.py # parses docs/risk-register.md
│ ├── diff.py # gh pr diff / --diff file
│ ├── agent.py # Ollama structured-output call
│ ├── prefilter.py # phase 13: deterministic register pre-filter (path + content)
│ ├── render.py + run.py # CLI + markdown
│ ├── golden_set.yaml # expected ranks per historic PR
│ ├── eval.py # deterministic scorer (precision/recall/F1)
│ └── reports/ # committed evidence (per-PR + eval-report)
├── triage_agent/ # phase 10: cluster failed CI runs by root cause
│ ├── fetcher.py # gh CLI wrappers (cached log dumps)
│ ├── parser.py # pytest + step-failure extraction
│ ├── cluster.py # heuristic group + LLM category + R-ID xref
│ ├── render.py + run.py # CLI + markdown
│ ├── golden_set.yaml # v1 v2: expected (category, R-ID) per cluster
│ ├── eval.py # v1 v2: deterministic scorer
│ └── reports/ # committed evidence (report.md + eval-report.md)
├── observability/ # phase 11: Prometheus + Grafana stack
│ ├── docker-compose.yml # stack (Prometheus + Grafana)
│ ├── prometheus/prometheus.yml # scrape config (host.docker.internal:5000)
│ ├── grafana/ # provisioned datasource + dashboard
│ └── evidence/ # committed screenshots
├── explore_agent/ # phase 12: LLM-driven exploration (API + UI)
│ ├── spec.py + probe.py # v1 v1: API surface — OpenAPI-driven + auth-bypass probing
│ ├── judge.py + render.py + run.py
│ ├── tours.py + ui_probe.py # v1 v3: UI surface — adaptive Playwright tours (policy)
│ ├── ui_judge.py + ui_run.py
│ ├── eval.py # v2 v1: deterministic golden-set scorer
│ └── reports/ # committed evidence (report.md + ui/report.md + screenshots)
├── security_agent/ # B1c: judges the B1 security findings
│ ├── findings.py # normalise Bandit + pip-audit + gitleaks/any SARIF (SARIF-native)
│ ├── judge.py # LLM: verdict + disposition + R-ID xref
│ ├── writeback.py # F-031: reconcile + propose SCA allowlist diff (--apply)
│ ├── render.py + run.py # CLI + markdown
│ ├── golden_set.yaml # expected (verdict, disposition, R-ID) per finding
│ ├── eval.py # deterministic scorer
│ └── reports/ # committed evidence (report + eval + writeback)
├── chaos/ # B4: resilience / chaos testing (local-only run)
│ ├── faults.py # probe() classifier + ComposeController + wait_for_recovery
│ ├── latency.py # v2: Toxiproxy + Prometheus clients (grey-failure axis)
│ ├── compose.latency.yml # v2: toxiproxy sidecar override (web -> toxiproxy -> db)
│ ├── scenarios.py # DB-outage + process-death + DB-latency hypotheses
│ ├── run.py # CLI + markdown/JSON report; owns the v2 latency lifecycle
│ └── reports/ # committed evidence (report.md + report.json)
├── tests/ # tests OF the harness itself
│ ├── test_smoke.py
│ ├── test_prefilter.py # phase 13: risk_agent register pre-filter unit tests
│ ├── test_chaos_scenarios.py # B4: chaos scenario logic (fakes, no Docker — gated)
│ ├── test_auth_finding.py # F-020: explore_agent spec-aware auth finding unit tests
│ └── agents/ # phase 12 v2 v2 + F-024: agent regression (gated)
│ ├── fixtures/ # cached PR diffs + synthetic clusters
│ ├── _runner.py # run-N-times harness + jitter metrics
│ ├── test_risk_agent_invariants.py
│ ├── test_triage_agent_invariants.py
│ ├── test_explore_judge_nonblinding.py # F-024: judge non-blinding positive control
│ ├── render_report.py # combined markdown from JSON dumps
│ └── reports/ # committed evidence
└── .github/workflows/
└── assurance.yml # the per-PR gates above
The contract, functional, TypeScript E2E, accessibility, performance, and data-quality layers need the SUT running. Bring it up first:
cd ../golf-web-app && docker compose up -d && docker compose exec web python seed.pyThen, from this repo:
# API contract tests
uv run pytest contract/
# UI / E2E functional tests (one-time browser download first)
uv run playwright install chromium
uv run pytest functional/
# Accessibility sweep (axe-core, WCAG 2.1 A/AA)
uv run pytest nonfunctional/accessibility/
# Data-quality checks (pandera against the live database)
uv run pytest data_quality/The TypeScript E2E layer (e2e_ts/) uses its own Node toolchain (not uv/pytest):
cd e2e_ts
npm ci # install pinned deps
npx playwright install chromium # one-time browser download
npm test # gating; SUT_BASE_URL overrides the targetPerformance is run by k6 (not pytest). With k6 installed:
k6 run nonfunctional/performance/api_load.jsOr via Docker, with no local k6 install:
SUT_BASE_URL=http://host.docker.internal:5000 \
docker run --rm -i -e SUT_BASE_URL -v "$PWD:/work" -w /work \
grafana/k6 run nonfunctional/performance/api_load.jsThe security gate's static scans (SAST + SCA + secrets) need only the SUT source checked out as a sibling, not a running SUT:
# Shift-left security scan (Bandit + pip-audit + gitleaks), ratchet gating
uv run python nonfunctional/security/scan.py --sut ../golf-web-appThe five agent layers (phases 8, 9, 10, 12 + B1c) use a local Ollama runtime — they are deliberately not in CI so the per-PR gate stays fast and reproducible, and the model isn't a moving budget on the critical path.
# AI evaluation harness — score the booking assistant across a model list
# (assumes the SUT is up and pointed at Ollama; see ai_evaluation/README.md)
uv run python -m ai_evaluation.run --models "qwen3:8b-fp16,qwen3.6:27b-q4_K_M"
# Risk-prioritisation agent — rank risks raised by a PR diff
uv run python -m risk_agent.run --pr 12 --repo ayyadam/golf-web-app
# Risk-prioritisation agent — eval against the golden set
uv run python -m risk_agent.eval # score against cached reports
uv run python -m risk_agent.eval --refresh # re-run agent on each case first
# Triage agent — cluster failed CI runs over a time window
uv run python -m triage_agent.run # default: this repo, last 30 days
uv run python -m triage_agent.run --since-days 7 # narrower window
uv run python -m triage_agent.run --no-llm # heuristic clusters only
# Triage agent — eval against the golden set
uv run python -m triage_agent.eval # score against cached report
uv run python -m triage_agent.eval --refresh # re-run the triage agent first
# Exploratory agent — probe every API endpoint with LLM-generated payload variants
# (SUT must be up; see explore_agent/README.md)
uv run python -m explore_agent.run # default model, full LLM run
uv run python -m explore_agent.run --no-llm # deterministic empty-body probes only
# Exploratory agent — UI tours via Playwright (adaptive: LLM decides one step at a time, LLM judges per step)
uv run python -m explore_agent.ui_run # all tours, headless
uv run python -m explore_agent.ui_run --tour booking-assistant # single tour
uv run python -m explore_agent.ui_run --headed # show the browser
# Exploratory agent — eval against the golden set (API surface)
uv run python -m explore_agent.eval # score against cached report
uv run python -m explore_agent.eval --refresh # re-run the agent first
# Security agent — judge the B1 scanner findings (verdict + disposition + R-ID)
uv run python -m security_agent.run --refresh # re-scan + judge
uv run python -m security_agent.eval # score against the golden set
uv run python -m security_agent.writeback # reconcile + propose allowlist diff
uv run python -m security_agent.writeback --apply # write additions + stale removals
# Agent regression suite (phase 12 v2 v2 + F-024) — runs risk_agent,
# triage_agent, and the explore_agent judge N times against fixed inputs;
# asserts schema/vocab invariants, top-result stability under LLM jitter,
# and a non-blinding positive control on the explore judge. ~3 min local.
RUN_AGENT_REGRESSION=1 uv run pytest tests/agents/ -v
uv run python tests/agents/render_report.py # refresh the markdown report (risk + triage)Local Prometheus + Grafana scraping the SUT's /metrics — see observability/README.md. Bring up the SUT first, then:
cd observability && docker compose up -d
# Grafana: http://localhost:3000 (anonymous viewer enabled)
# Dashboard: http://localhost:3000/d/sut-overview
# Prometheus: http://localhost:9090See each module's README — ai_evaluation/, risk_agent/, triage_agent/, explore_agent/, security_agent/, nonfunctional/security/, observability/ — for the full design notes and committed evidence.
- System under test: golf-web-app — Flask golf-club app, GHCR-published, CI on its own pipeline.