diff --git a/.pre-commit-config.yaml b/.pre-commit-config.yaml new file mode 100644 index 0000000..fcd91fe --- /dev/null +++ b/.pre-commit-config.yaml @@ -0,0 +1,9 @@ +repos: + - repo: local + hooks: + - id: deferred-validation + name: deferred validation catalog + entry: python shared_scripts/deferred_validation.py --repo-root . + language: python + pass_filenames: false + always_run: true diff --git a/README.md b/README.md index a8d5577..a8de197 100644 --- a/README.md +++ b/README.md @@ -13,6 +13,11 @@ provenance and existing-source research work. Issues #8 and #18 remain open for those executable slices. Plans provide no measurements, permission or consent; private data continues to live only in the governed private authority. +After installing `.[dev]`, run +`python -m pre_commit run deferred-validation --all-files` to validate the owner +catalog. The existing Python suite exercises this same command and verifies +shared-file hashes. Passing software checks supply no measurement or permission. + ## Authorized setup ```powershell diff --git a/docs/development/DEVELOPMENT_LOG.md b/docs/development/DEVELOPMENT_LOG.md index 6cd4114..c93d447 100644 --- a/docs/development/DEVELOPMENT_LOG.md +++ b/docs/development/DEVELOPMENT_LOG.md @@ -18,9 +18,22 @@ reachable from any live state and `abandoned` from `parked`. ## Active +### DL-#63 - Deferred Catalog Enforcement + +- **State:** in_review +- **Owner:** codex +- **Issue:** #63; full rollout Repository_Management#1687 +- **Branch:** `chore/63-deferred-catalog-guard` +- **PR:** #64 +- **Paths:** `shared_scripts/`, `.pre-commit-config.yaml`, `pyproject.toml`, `tests/test_deferred_catalog_hook.py`, `docs/development/`, `AGENTS.md`, `CLAUDE.md`, `README.md` +- **Started:** 2026-09-23 +- **Last verified:** SELF (five RED absent-hook/receipt failures; 54 GREEN tests including configured command and exact bundle hashes; actual hook, Ruff lint/format and new-test strict mypy pass; unchanged corpus.py:193 full-mypy limitation retained) +- **Summary:** Enforces the two published catalogs without changing original scope, source snapshots, private data pins or rights/consent boundaries. +- **Next step:** Complete repository checks, publish and verify exact default-branch adoption. + ### DL-#59 · Deferred Validation Project Projection -- **State:** in_review +- **State:** shipped - **Owner:** codex - **Issue:** #59; parents Repository_Management#1687 and Runner_Dashboard#1248 - **Branch:** `docs/deferred-project-projection` @@ -29,7 +42,7 @@ reachable from any live state and `abandoned` from `parked`. - **Started:** 2026-09-23 - **Last verified:** 2026-09-23 (265bcc7 plus docs; strict catalog and both parsers pass, two parked plans/six decisions; 49 tests and Ruff pass; unchanged local corpus.py:193 typing limitation retained) - **Summary:** Exposes both published external dependencies and pending decisions without inventing evidence or authority. Existing client, rights and data-custody contracts are preserved. -- **Next step:** Protected publication, default-branch verification, then actual deployed Projects/staff adoption under the fleet parents. +- **Next step:** #60 merged as 16001c7; both parked owner plans are verified in the running dashboard. Enforcement continues in DL-#63. ### DL-#8 - Deferred Paired Data Dependency @@ -42,7 +55,7 @@ reachable from any live state and `abandoned` from `parked`. - **Started:** 2026-09-22 - **Last verified:** 2026-09-22 (base `93d6435` plus planning changes; 49 tests, catalog and Ruff pass; local mypy has one unchanged corpus.py:193 unused-ignore error) - **Summary:** External obligations are deferred for Board consideration; the mixed source issue stays active for software and source research. -- **Next step:** Publish the planning PR and link the scope split on #8. +- **Next step:** Owner plan is published; executable #8 work stays active and unavailable physical collection remains deferred. ### DL-#18 - Deferred External Rights Decisions @@ -55,7 +68,7 @@ reachable from any live state and `abandoned` from `parked`. - **Started:** 2026-09-22 - **Last verified:** 2026-09-22 (base `93d6435` plus planning changes; 49 tests, catalog and Ruff pass; local mypy has one unchanged corpus.py:193 unused-ignore error) - **Summary:** External obligations are deferred for Board consideration; the mixed source issue stays active for software and source research. -- **Next step:** Publish the planning PR and link the scope split on #18. +- **Next step:** Owner plan is published; executable #18 source research stays active and rights-holder decisions remain deferred. ### DL-0001 · Chore Bump Private Lock 1672D2D diff --git a/docs/development/HANDOFF.md b/docs/development/HANDOFF.md index 591c276..a3827c8 100644 --- a/docs/development/HANDOFF.md +++ b/docs/development/HANDOFF.md @@ -1,3 +1,33 @@ +# Deferred Catalog Enforcement - #63 + +- Worktree: `C:/Users/diete/Repositories/Worktrees/Launch-Monitor-Data-deferred-guard`. + Branch `chore/63-deferred-catalog-guard`; base `308a1ed`; commit `SELF`; PR [#64](https://github.com/D-sorganization/Launch-Monitor-Data/pull/64), protected auto-merge armed. +- Governing issue #63; full fleet rollout Repository_Management#1687 remains open. +- Exact three-file validator bundle from central `0a104101`, pinned SHA-256 receipt, + always-run local hook and five configured-command/digest tests in the existing + Python CI suite. Invalid activation, missing checker and competing catalog fail. + Both original plans, source snapshots and the private-data lock stay unchanged. +- Synchronizes only the approved deferred-validation managed block from central + `a59cb194`; v1 fields remain distinct from new-adopter fields. Developer tooling + includes pre-commit, PyYAML and its typing stubs. No external data is downloaded. +- TDD: five missing-hook/receipt RED failures, then all 54 Python tests pass. + Root Ruff lint/format (20 files), the actual pre-commit hook and strict typing + of the new consumer test pass. Full local mypy retains the previously documented + unused-ignore failure at unchanged corpus.py:193 (17 files checked). No ignore + or typing gate was relaxed; hosted Python 3.12 qualification remains required. + All original planning bytes, private-data lock and three bundle hashes match. +- Commands: `python -m pytest -q`, `python -m ruff check .`, + `python -m ruff format --check .`, `python -m mypy`, + `python -m pre_commit run deferred-validation --all-files`. +- Previous view #60 merged at `16001c7`; both parked owner plans are visible in + the deployed Projects API/UI. Source issues #8/#18 retain executable software + and existing-source research; resource/rights decisions stay deferred. +- Next: publish through protected CI, and verify default-branch + bundle/hook/rule bytes. No rights-holder contact, consent, calibration or + physical/perceptual validation is authorized or supplied by this deployment. + +--- + # Deferred Validation Project Projection — #59 - Worktree: `C:/Users/diete/Repositories/Worktrees/Launch-Monitor-Data-deferred-project`. diff --git a/docs/development/deferred-catalog-bundle.json b/docs/development/deferred-catalog-bundle.json new file mode 100644 index 0000000..894051a --- /dev/null +++ b/docs/development/deferred-catalog-bundle.json @@ -0,0 +1,9 @@ +{ + "source_repository": "D-sorganization/Repository_Management", + "source_commit": "0a1041018e737173e49ff97ed4b82283e3cb672f", + "files": { + "shared_scripts/deferred_validation.py": "152ab473b23daf5b3a32618563ef0cddcdb957d71580cf09aeeb03c17f9d5ee5", + "shared_scripts/deferred_planning.py": "521f1a98fecf36990f1f31696144163c85d023be34a465dfbde9aacf81f33f94", + "shared_scripts/handoff_validator.py": "eac9039b72c2a2d7f5708b0cb098f84d2a1e78568ed1ec6d17f6756ca6274cd5" + } +} diff --git a/pyproject.toml b/pyproject.toml index 3bbedec..39eca2e 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -18,7 +18,9 @@ dev = [ "mypy>=1.11", "pandas-stubs>=2.0", "pytest>=8.0", - "ruff>=0.6", + "pre-commit>=3.0", + "PyYAML>=6.0", + "types-PyYAML>=6.0", "ruff>=0.6", ] [project.scripts] diff --git a/shared_scripts/deferred_planning.py b/shared_scripts/deferred_planning.py new file mode 100644 index 0000000..c9785c3 --- /dev/null +++ b/shared_scripts/deferred_planning.py @@ -0,0 +1,280 @@ +"""Validate repository-owned future plans; never infer completion from deferral. + +This module is network-free. A migration caller must fetch published bytes from +the owning repository's default branch, then call ``closure_payload`` immediately +before applying the reviewed GitHub disposition. The module never closes issues. +""" + +from __future__ import annotations + +import argparse +import json +import re +from pathlib import Path +from typing import Any + +ENTRY_FIELDS = { + "id", + "title", + "source_issue", + "source_snapshot", + "record", + "disposition", + "kind", + "status", + "rationale", + "prerequisites", + "acceptance", + "board", + "activation_issue", +} +REPOSITORY = re.compile(r"D-sorganization/[A-Za-z0-9_.-]+\Z") + + +class PlanningError(ValueError): + """The plan is incomplete, unsafe to migrate, or violates its contract.""" + + +def _require(condition: bool, message: str) -> None: + if not condition: + raise PlanningError(message) + + +def _text(value: object) -> bool: + return isinstance(value, str) and bool(value.strip()) + + +def validate_migration_pr_body(body: str) -> None: + """Reject empty text and issue-closing phrases, including negated phrases. + + GitHub can interpret a closing keyword inside prose such as "do not + auto-close". A migration must use references and separately record a + not-planned disposition after verifying default-branch publication. + """ + _require(_text(body), "migration PR body must be nonempty text") + closing = re.search( + r"\b(?:close[sd]?|fix(?:e[sd])?|resolve[sd]?)\s+" + r"(?:(?:[\w.-]+/[\w.-]+)?#\d+" + r"|https://github\.com/[\w.-]+/[\w.-]+/issues/\d+)", + body, + re.IGNORECASE, + ) + _require(closing is None, "migration PR must reference issues without closing them") + + +def _object(value: object, keys: set[str], context: str) -> dict[str, Any]: + _require(isinstance(value, dict), f"{context}: expected object") + assert isinstance(value, dict) # validated above; type narrowing only + _require(set(value) == keys, f"{context}: missing or unknown fields") + return value + + +def _file(root: Path, name: object) -> Path: + _require(_text(name), "artifact path must be text") + assert isinstance(name, str) + _require(not any(c in name for c in ("\\", ":")), "nonportable artifact path") + relative = Path(name) + _require(not relative.is_absolute() and ".." not in relative.parts, "unsafe path") + path = (root / relative).resolve() + _require(path.is_relative_to(root.resolve()), "artifact escapes planning root") + _require(path.is_file(), f"missing artifact: {name}") + return path + + +def _read_json(path: Path) -> Any: + try: + return json.loads(path.read_text(encoding="utf-8-sig")) + except (OSError, UnicodeError, json.JSONDecodeError) as exc: + raise PlanningError(f"cannot read JSON: {path.name}") from exc + + +def activation_ready(entry: dict[str, Any]) -> bool: + """Resource-ready for Board-approved execution, not scientifically validated. + + Evidence references still require human/Board review. Their presence does not + authenticate an artifact or authorize a coding agent to perform a human study. + """ + board = entry.get("board", {}) + resources = entry.get("prerequisites", []) + return bool( + isinstance(board, dict) + and board.get("decision") == "approved" + and _text(board.get("evidence")) + and isinstance(resources, list) + and resources + and all(isinstance(r, dict) and _text(r.get("evidence")) for r in resources) + ) + + +def _validate_entry(root: Path, entry: object, repository: str) -> dict[str, Any]: + item = _object(entry, ENTRY_FIELDS, "entry") + for field in ("id", "title", "rationale"): + _require(_text(item[field]), f"{field}: nonempty text required") + _require(bool(re.fullmatch(r"DV-[1-9][0-9]*", item["id"])), "invalid plan id") + prefix = f"https://github.com/{repository}/issues/" + issue = item["source_issue"] + _require(isinstance(issue, str) and issue.startswith(prefix), "wrong source repo") + _require( + bool(re.fullmatch(r"[1-9][0-9]*", issue[len(prefix) :])), "invalid issue URL" + ) + _require(item["disposition"] in ("defer", "split"), "invalid disposition") + _require( + item["kind"] + in ( + "physical-measurement", + "human-validation", + "external-review", + "research-idea", + ), + "invalid kind", + ) + _require( + item["status"] in ("deferred", "ready", "activated", "declined"), + "invalid status", + ) + for field in ("acceptance", "prerequisites"): + _require( + isinstance(item[field], list) and bool(item[field]), f"{field}: required" + ) + _require(all(_text(x) for x in item["acceptance"]), "empty acceptance criterion") + for prerequisite in item["prerequisites"]: + resource = _object(prerequisite, {"name", "evidence"}, "prerequisite") + _require(_text(resource["name"]), "empty prerequisite") + _require( + resource["evidence"] is None or _text(resource["evidence"]), + "invalid evidence", + ) + board = _object(item["board"], {"decision", "evidence"}, "board") + _require( + board["decision"] in ("pending", "approved", "rejected"), + "invalid Board decision", + ) + _require( + board["evidence"] is None or _text(board["evidence"]), "invalid Board evidence" + ) + if board["decision"] != "pending": + _require(_text(board["evidence"]), "Board decision requires a durable receipt") + if item["status"] in ("ready", "activated"): + _require(activation_ready(item), "activation prerequisites not satisfied") + if item["status"] == "declined": + _require( + board["decision"] == "rejected", "declined plan requires Board decision" + ) + activation = item["activation_issue"] + if item["status"] == "activated": + _require( + isinstance(activation, str) + and bool( + re.fullmatch( + re.escape(prefix) + r"[1-9][0-9]*", + activation, + ) + ), + "activated plan requires owning-repository issue", + ) + else: + _require(activation is None, "inactive plan cannot declare activation issue") + record = _file(root, item["record"]) + _require( + record.suffix == ".md" and bool(record.read_text(encoding="utf-8").strip()), + "empty record", + ) + snapshot = _read_json(_file(root, item["source_snapshot"])) + _require(isinstance(snapshot, dict), "source snapshot must be an object") + for field in ("title", "body", "updated_at", "html_url"): + _require(_text(snapshot.get(field)), f"source snapshot missing {field}") + _require(snapshot["html_url"] == issue, "snapshot issue mismatch") + _require( + str(snapshot.get("number")) == issue[len(prefix) :], "snapshot number mismatch" + ) + return item + + +def validate_catalog(root: Path) -> dict[str, Any]: + """Read the strict v1 catalog, contained files, and preserved source snapshots.""" + catalog = _object( + _read_json(root / "catalog.json"), + { + "schema_version", + "repository", + "entries", + }, + "catalog", + ) + _require( + type(catalog["schema_version"]) is int and catalog["schema_version"] == 1, + "unsupported schema", + ) + repository = catalog["repository"] + _require( + isinstance(repository, str) and bool(REPOSITORY.fullmatch(repository)), + "invalid repository", + ) + _require(isinstance(catalog["entries"], list), "entries must be a list") + ids: set[str] = set() + sources: set[str] = set() + for value in catalog["entries"]: + item = _validate_entry(root, value, repository) + _require(item["id"] not in ids, "duplicate plan id") + _require(item["source_issue"] not in sources, "duplicate source issue") + ids.add(item["id"]) + sources.add(item["source_issue"]) + return catalog + + +def closure_payload( + root: Path, + plan_id: str, + live_issue: dict[str, Any], + published: dict[str, bytes], +) -> dict[str, str]: + """Fail closed unless an unchanged external-only issue has a durable plan. + + ``published`` must contain bytes fetched from the owning repository's default + branch at a recorded commit, keyed by paths relative to the planning root. + Callers must apply the exempt ``roadmap`` label and post the durable record + link before this payload is PATCHed; never use a closing keyword in the PR. + GitHub has no conditional issue PATCH here, so recheck immediately before it. + """ + catalog = validate_catalog(root) + entries = [e for e in catalog["entries"] if e["id"] == plan_id] + _require(len(entries) == 1, "unknown plan") + entry = entries[0] + _require(entry["disposition"] == "defer", "split issue retains software work") + _require(entry["status"] == "deferred", "only deferred plans can migrate") + for name in ("catalog.json", entry["record"], entry["source_snapshot"]): + _require( + published.get(name) == _file(root, name).read_bytes(), + f"unverified published artifact: {name}", + ) + source = _read_json(_file(root, entry["source_snapshot"])) + _require(live_issue.get("state") == "open", "issue is not open") + # Comments and labels change updated_at without changing scope. Preserve the + # snapshot timestamp for audit, but compare the actual source text/identity. + for field in ("number", "html_url", "title", "body"): + _require(live_issue.get(field) == source.get(field), f"source changed: {field}") + labels = live_issue.get("labels", []) + _require(isinstance(labels, list), "invalid issue labels") + names = {x.get("name", "") if isinstance(x, dict) else x for x in labels} + _require( + not names.intersection({"do-not-automate", "claim:user", "claim:local"}), + "protected issue", + ) + return {"state": "closed", "state_reason": "not_planned"} + + +def main() -> int: + """Validate only. This command does not mutate GitHub or local files.""" + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("root", type=Path, help="docs/development/planning directory") + args = parser.parse_args() + try: + catalog = validate_catalog(args.root) + except PlanningError as exc: + parser.exit(1, f"Invalid planning catalog: {exc}\n") + print(f"Valid: {catalog['repository']} ({len(catalog['entries'])} plans)") + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/shared_scripts/deferred_validation.py b/shared_scripts/deferred_validation.py new file mode 100644 index 0000000..add7784 --- /dev/null +++ b/shared_scripts/deferred_validation.py @@ -0,0 +1,678 @@ +#!/usr/bin/env python3 +"""Validate a repository's deferred external-validation catalog. + +Some accepted work cannot be finished by any coding agent: it ends in a +measurement, a laboratory booking, a field collection or a human trial. Left in +the issue queue it is noise that every sweep re-reads and no sweep can act on. +Deleted, its scientific obligation is lost. + +The catalog is the third option. `docs/planning/deferred-validation.json` holds +one draft planning record per deferred task — the original issue URL, the +acceptance criteria as written, why the work is external, what resources it +needs, and what must become true to reopen it — so the Board can act on it +later and nothing is silently dropped. + +The safeguards enforced here are the ones that make the deferral honest: + +* a record can never report validation as done (`validation_state` has no + "complete"), so deferring is never a way to appear finished; +* evidence must be a checkable reference, never prose, so no measurement is + conjured into the record; +* the original issue may only be marked closed **after** the record is + published at a durable URL and a named reviewer signed it off, so discovery + by keyword never becomes authority to close; +* a `split` disposition must name the implementable remainder that stays open. + +This module is portable: it is copied fleet-wide alongside +`handoff_validator.py` and `development_log.py` and must not import anything +outside the standard library. + +CLI:: + + python -m shared_scripts.deferred_validation --repo-root . + +Exit status is 1 when the catalog is present and invalid, 0 when it is valid or +absent — a repository with nothing deferred has nothing to validate. +""" + +from __future__ import annotations + +import argparse +import json +import re +import sys +from dataclasses import dataclass +from pathlib import Path +from typing import Any + + +def _load_sibling(name: str) -> Any: + """Load a shipped sibling without relying on the caller's working directory.""" + import importlib.util + + sibling = Path(__file__).with_name(f"{name}.py") + spec = importlib.util.spec_from_file_location(f"_fleet_{name}", sibling) + if spec is None or spec.loader is None: + raise ImportError(f"Cannot load required catalog dependency: {sibling}") + module = importlib.util.module_from_spec(spec) + sys.modules[spec.name] = module + spec.loader.exec_module(module) + return module + + +SECRET_PATTERNS = _load_sibling("handoff_validator").SECRET_PATTERNS +CANONICAL_RELATIVE_PATH = Path("docs") / "planning" / "deferred-validation.json" +PUBLISHED_RELATIVE_PATH = Path("docs/development/planning/catalog.json") +SCHEMA_RELATIVE_PATH = Path("docs") / "planning" / "deferred-validation.schema.json" + +SCHEMA_VERSION = 1 + +REQUIRED_TOP_LEVEL_FIELDS: tuple[str, ...] = ( + "schema_version", + "repository", + "updated", + "records", +) + +# Record ids are keyed by the governing issue, exactly like development-log +# entry ids (Repository_Management#1520): an issue number is unique by +# construction, so concurrent pull requests never pick the same id. +RECORD_ID = re.compile(r"^DV-#\d+$") +ISSUE_URL = re.compile( + r"^https://github\.com/[A-Za-z0-9_.-]+/[A-Za-z0-9_.-]+/issues/\d+$" +) +HTTPS_URL = re.compile(r"^https://\S+$") +ISO_DATE = re.compile(r"^\d{4}-\d{2}-\d{2}$") +REPO_PATH = re.compile(r"^[A-Za-z0-9_.][A-Za-z0-9_./-]*\.[A-Za-z0-9_]+(:\d+)?$") + +REQUIRED_RECORD_FIELDS: tuple[str, ...] = ( + "id", + "title", + "origin_issue_url", + "origin_state", + "disposition", + "retained_issue_url", + "blocked_on", + "validation_state", + "justification", + "evidence", + "resource_prerequisites", + "acceptance_criteria", + "reopen_criteria", + "discovery", + "reviewed_by", + "reviewed_on", + "record_url", +) +OPTIONAL_RECORD_FIELDS: tuple[str, ...] = ("notes",) + +DISPOSITIONS = frozenset({"migrate", "split", "retain"}) +ORIGIN_STATES = frozenset({"open", "closed_deferred"}) +# Deliberately closed and deliberately without a "complete" member. +VALIDATION_STATES = frozenset({"not_started", "blocked_external"}) +DISCOVERY_MODES = frozenset({"keyword_match", "manual_review", "board_referral"}) +EXTERNAL_BLOCKERS = frozenset( + { + "physical_measurement", + "laboratory_access", + "field_data_collection", + "specialized_hardware", + "human_trial", + "external_funding", + "third_party_service", + } +) + +# Obligation fields: a deferral that carries none of these has dropped the work +# rather than parked it. +NON_EMPTY_LIST_FIELDS: tuple[str, ...] = ( + "blocked_on", + "evidence", + "resource_prerequisites", + "acceptance_criteria", + "reopen_criteria", +) +MIN_JUSTIFICATION_CHARS = 40 + + +@dataclass(frozen=True) +class CatalogFinding: + """A single governance finding against a deferred-validation catalog.""" + + path: Path + pointer: str + kind: str + message: str + remediation: str + + +def _finding( + path: Path, pointer: str, kind: str, message: str, remediation: str +) -> CatalogFinding: + return CatalogFinding( + path=path, + pointer=pointer, + kind=kind, + message=message, + remediation=remediation, + ) + + +def _is_checkable_reference(value: str) -> bool: + """True when evidence points at something a reader can open.""" + text = value.strip() + return bool(HTTPS_URL.match(text) or REPO_PATH.match(text)) + + +def _validate_enum( + record: dict[str, Any], + field: str, + allowed: frozenset[str], + path: Path, + pointer: str, +) -> list[CatalogFinding]: + value = record.get(field) + if isinstance(value, str) and value in allowed: + return [] + return [ + _finding( + path, + f"{pointer}/{field}", + "bad_value", + f"{field} is {value!r}; allowed values are {', '.join(sorted(allowed))}.", + f"Set {field} to one of the allowed values.", + ) + ] + + +def _validate_record_shape( + record: dict[str, Any], path: Path, pointer: str +) -> list[CatalogFinding]: + """Required fields, unknown fields, and closed value sets.""" + findings: list[CatalogFinding] = [] + known = set(REQUIRED_RECORD_FIELDS) | set(OPTIONAL_RECORD_FIELDS) + + for field in sorted(set(record) - known): + findings.append( + _finding( + path, + f"{pointer}/{field}", + "unknown_field", + f"Unknown record field {field!r}.", + "Remove the field; the catalog schema is closed. Measurements " + "belong in the repository's results, never in a deferral record.", + ) + ) + for field in REQUIRED_RECORD_FIELDS: + if field not in record: + findings.append( + _finding( + path, + f"{pointer}/{field}", + "missing_field", + f"Required record field {field!r} is missing.", + f"Add {field}; write null where the schema allows it.", + ) + ) + if findings: + return findings + + record_id = record["id"] + if not (isinstance(record_id, str) and RECORD_ID.match(record_id)): + findings.append( + _finding( + path, + f"{pointer}/id", + "bad_id", + f"Record id {record_id!r} is not of the form DV-#.", + "Key the record by its governing issue number, never a serial.", + ) + ) + if not (isinstance(record["title"], str) and record["title"].strip()): + findings.append( + _finding( + path, + f"{pointer}/title", + "empty_field", + "title is empty.", + "Give the deferred task a one-line title.", + ) + ) + origin = record["origin_issue_url"] + if not (isinstance(origin, str) and ISSUE_URL.match(origin)): + findings.append( + _finding( + path, + f"{pointer}/origin_issue_url", + "bad_value", + f"origin_issue_url {origin!r} is not a GitHub issue URL.", + "Preserve the original issue URL in full; it is the audit trail.", + ) + ) + + findings.extend(_validate_enum(record, "disposition", DISPOSITIONS, path, pointer)) + findings.extend( + _validate_enum(record, "origin_state", ORIGIN_STATES, path, pointer) + ) + findings.extend( + _validate_enum(record, "validation_state", VALIDATION_STATES, path, pointer) + ) + findings.extend(_validate_enum(record, "discovery", DISCOVERY_MODES, path, pointer)) + + for field in NON_EMPTY_LIST_FIELDS: + value = record[field] + if not isinstance(value, list) or not value: + findings.append( + _finding( + path, + f"{pointer}/{field}", + "empty_field", + f"{field} is empty.", + "The obligation survives the deferral: record what the " + "original issue demanded and what would bring it back.", + ) + ) + continue + for index, item in enumerate(value): + if not isinstance(item, str) or not item.strip(): + findings.append( + _finding( + path, + f"{pointer}/{field}/{index}", + "empty_field", + f"{field}[{index}] is not a non-empty string.", + f"Write each {field} entry as text.", + ) + ) + + blockers = record["blocked_on"] + if isinstance(blockers, list): + for index, item in enumerate(blockers): + if isinstance(item, str) and item in EXTERNAL_BLOCKERS: + continue + findings.append( + _finding( + path, + f"{pointer}/blocked_on/{index}", + "bad_value", + f"blocked_on[{index}] is {item!r}; allowed values are " + f"{', '.join(sorted(EXTERNAL_BLOCKERS))}.", + "A deferral is for work blocked on the physical world, not " + "for work that is merely hard or unscheduled.", + ) + ) + + evidence = record["evidence"] + if isinstance(evidence, list): + for index, item in enumerate(evidence): + if isinstance(item, str) and _is_checkable_reference(item): + continue + findings.append( + _finding( + path, + f"{pointer}/evidence/{index}", + "unverifiable_evidence", + f"evidence[{index}] is not a URL or repo-relative path.", + "Cite something a reader can open. Prose is not evidence, " + "and a measurement that was never taken is not evidence.", + ) + ) + + justification = record["justification"] + if ( + not isinstance(justification, str) + or len(justification.strip()) < MIN_JUSTIFICATION_CHARS + ): + findings.append( + _finding( + path, + f"{pointer}/justification", + "thin_justification", + f"justification is shorter than {MIN_JUSTIFICATION_CHARS} characters.", + "State why no coding agent can finish this work.", + ) + ) + + for field in ("reviewed_on",): + value = record[field] + if value is not None and not (isinstance(value, str) and ISO_DATE.match(value)): + findings.append( + _finding( + path, + f"{pointer}/{field}", + "bad_value", + f"{field} {value!r} is not an ISO date (YYYY-MM-DD) or null.", + f"Write {field} as YYYY-MM-DD.", + ) + ) + for field in ("retained_issue_url", "record_url"): + value = record[field] + pattern = ISSUE_URL if field == "retained_issue_url" else HTTPS_URL + if value is not None and not (isinstance(value, str) and pattern.match(value)): + findings.append( + _finding( + path, + f"{pointer}/{field}", + "bad_value", + f"{field} {value!r} is not a valid URL or null.", + f"Write {field} as a full https URL, or null.", + ) + ) + return findings + + +def _validate_record_safeguards( + record: dict[str, Any], path: Path, pointer: str +) -> list[CatalogFinding]: + """Publish-then-close ordering, split remainders, retained work.""" + findings: list[CatalogFinding] = [] + origin_state = record.get("origin_state") + disposition = record.get("disposition") + + if origin_state == "closed_deferred": + if not record.get("record_url"): + findings.append( + _finding( + path, + f"{pointer}/record_url", + "closed_without_record", + "The original issue is marked closed but this record has no " + "durable published link.", + "Publish the record, verify the link resolves on the default " + "branch, and only then close the original as not planned.", + ) + ) + if not (record.get("reviewed_by") and record.get("reviewed_on")): + findings.append( + _finding( + path, + f"{pointer}/reviewed_by", + "closed_without_review", + "The original issue is marked closed but no reviewer and date " + "are recorded.", + "Keyword matching is candidate discovery, never authority to " + "close. Record the human or Board review that approved it.", + ) + ) + if disposition == "retain": + findings.append( + _finding( + path, + f"{pointer}/disposition", + "retained_but_closed", + "Disposition is 'retain' but the original issue is marked " + "closed as deferred.", + "Implementable software and CI work stays in the queue; " + "record it as retained and leave the issue open.", + ) + ) + + if disposition == "split" and not record.get("retained_issue_url"): + findings.append( + _finding( + path, + f"{pointer}/retained_issue_url", + "split_without_remainder", + "Disposition is 'split' but no retained issue is named.", + "Split means the implementable remainder stays open: file or " + "name it, then link it here.", + ) + ) + return findings + + +def _scan_secrets(value: Any, path: Path, pointer: str) -> list[CatalogFinding]: + """Reject credentials anywhere in the catalog.""" + findings: list[CatalogFinding] = [] + if isinstance(value, str): + for pattern, secret_type in SECRET_PATTERNS: + if pattern.search(value): + findings.append( + _finding( + path, + pointer, + "secret_detected", + f"Potential {secret_type} detected in the catalog.", + "Remove credentials, tokens, and secret keys.", + ) + ) + elif isinstance(value, dict): + for key, item in value.items(): + findings.extend(_scan_secrets(item, path, f"{pointer}/{key}")) + elif isinstance(value, list): + for index, item in enumerate(value): + findings.extend(_scan_secrets(item, path, f"{pointer}/{index}")) + return findings + + +def validate_catalog(data: Any, path: Path) -> list[CatalogFinding]: + """Validate a parsed catalog against the published schema and safeguards.""" + if not isinstance(data, dict): + return [ + _finding( + path, + "", + "bad_value", + "The catalog root is not a JSON object.", + "See docs/planning/deferred-validation.schema.json.", + ) + ] + + findings: list[CatalogFinding] = [] + for field in sorted(set(data) - set(REQUIRED_TOP_LEVEL_FIELDS)): + findings.append( + _finding( + path, + f"/{field}", + "unknown_field", + f"Unknown top-level field {field!r}.", + "Remove the field; the catalog schema is closed.", + ) + ) + for field in REQUIRED_TOP_LEVEL_FIELDS: + if field not in data: + findings.append( + _finding( + path, + f"/{field}", + "missing_field", + f"Required field {field!r} is missing.", + f"Add {field} to the catalog.", + ) + ) + if findings: + return findings + + if data["schema_version"] != SCHEMA_VERSION: + findings.append( + _finding( + path, + "/schema_version", + "bad_value", + f"schema_version must be {SCHEMA_VERSION}.", + "Migrate the catalog to the published schema version.", + ) + ) + if not (isinstance(data["repository"], str) and data["repository"].strip()): + findings.append( + _finding( + path, + "/repository", + "empty_field", + "repository is empty.", + "Name the repository the catalog belongs to.", + ) + ) + updated = data["updated"] + if not (isinstance(updated, str) and ISO_DATE.match(updated)): + findings.append( + _finding( + path, + "/updated", + "bad_value", + f"updated {updated!r} is not an ISO date (YYYY-MM-DD).", + "Refresh `updated` whenever a record changes.", + ) + ) + + records = data["records"] + if not isinstance(records, list): + findings.append( + _finding( + path, + "/records", + "bad_value", + "records is not a list.", + "Write records as a JSON array; an empty array is valid.", + ) + ) + return findings + + seen: set[str] = set() + for index, record in enumerate(records): + pointer = f"/records/{index}" + if not isinstance(record, dict): + findings.append( + _finding( + path, + pointer, + "bad_value", + "Record is not a JSON object.", + "See docs/planning/deferred-validation.schema.json.", + ) + ) + continue + shape = _validate_record_shape(record, path, pointer) + findings.extend(shape) + if not shape: + findings.extend(_validate_record_safeguards(record, path, pointer)) + record_id = record.get("id") + if isinstance(record_id, str): + if record_id in seen: + findings.append( + _finding( + path, + f"{pointer}/id", + "duplicate_id", + f"Record id {record_id} appears more than once.", + "One record per deferred task, updated in place.", + ) + ) + seen.add(record_id) + + findings.extend(_scan_secrets(data, path, "")) + return findings + + +def _validate_published_catalog(path: Path) -> list[CatalogFinding]: + """Reuse the deployed v1 contract, including its local evidence artifacts.""" + try: + checker = _load_sibling("deferred_planning") + except (ImportError, OSError) as exc: + return [ + _finding( + path, + "", + "missing_checker", + str(exc), + "Install the complete deferred-validation validator bundle.", + ) + ] + try: + data = checker.validate_catalog(path.parent) + except (ValueError, OSError, UnicodeError) as exc: + return [ + _finding( + path, + "", + "invalid_published_catalog", + str(exc), + "Restore the v1 catalog and its original plan/source artifacts.", + ) + ] + return _scan_secrets(data, path, "") + + +def validate_repository_catalog(repo_root: Path) -> list[CatalogFinding]: + """Validate the single authority, including v1 catalogs already published. + + Never choose between two catalogs silently or infer Board review during + schema recognition. This read-only operation does not migrate or close work. + """ + path = repo_root / CANONICAL_RELATIVE_PATH + published = repo_root / PUBLISHED_RELATIVE_PATH + if (path.exists() or path.is_symlink()) and ( + published.exists() or published.is_symlink() + ): + return [ + _finding( + path, + "", + "conflicting_catalogs", + "Two planning catalogs exist.", + "Reconcile into one reviewed authority without losing evidence.", + ) + ] + if published.exists() or published.is_symlink(): + return _validate_published_catalog(published) + if not path.exists() and not path.is_symlink(): + return [] + try: + raw = path.read_text(encoding="utf-8") + except (OSError, UnicodeError) as exc: + return [ + _finding( + path, + "", + "unreadable", + f"Cannot read the catalog: {exc}.", + "Restore the file or remove it.", + ) + ] + try: + data = json.loads(raw) + except json.JSONDecodeError as exc: + return [ + _finding( + path, + "", + "unparseable", + f"The catalog is not valid JSON: {exc}.", + "Fix the JSON syntax; the catalog is machine-read.", + ) + ] + return validate_catalog(data, path) + + +def format_findings(findings: list[CatalogFinding]) -> list[str]: + """Render findings as one line each.""" + return [ + f"{finding.path.as_posix()}{finding.pointer} [{finding.kind}]: " + f"{finding.message} (Remediation: {finding.remediation})" + for finding in findings + ] + + +def main(argv: list[str] | None = None) -> int: + """CLI entry point.""" + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument( + "--repo-root", + type=Path, + default=Path.cwd(), + help="Repository root containing docs/planning/deferred-validation.json.", + ) + args = parser.parse_args(argv) + + findings = validate_repository_catalog(args.repo_root.resolve()) + if not findings: + print("Deferred-validation catalog OK.") + return 0 + print("Deferred-validation catalog failed validation:") + for line in format_findings(findings): + print(f" - {line}") + return 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/shared_scripts/handoff_validator.py b/shared_scripts/handoff_validator.py new file mode 100644 index 0000000..b49d90e --- /dev/null +++ b/shared_scripts/handoff_validator.py @@ -0,0 +1,590 @@ +#!/usr/bin/env python3 +"""Canonical handoff schema validation and enforcement for the repository fleet. + +Ensures implementation state survives agent replacement and context exhaustion +by validating canonical handoff schema, detecting unedited placeholders, ensuring +implementation commits update continuation state, and protecting against secrets. +""" + +from __future__ import annotations + +import argparse +import os +import re +import subprocess +import sys +from collections.abc import Iterable, Sequence +from dataclasses import dataclass +from pathlib import Path + +REQUIRED_HEADINGS = ( + ("## Identity", re.compile(r"^##\s+Identity\s*$", re.MULTILINE)), + ( + "## Objective and status", + re.compile( + r"^##\s+Objective\s+and\s+status\s*$", + re.MULTILINE | re.IGNORECASE, + ), + ), + ( + "## Files and decisions", + re.compile( + r"^##\s+Files\s+and\s+decisions\s*$", + re.MULTILINE | re.IGNORECASE, + ), + ), + ("## Validation", re.compile(r"^##\s+Validation\s*$", re.MULTILINE)), + ( + "## Blockers and risks", + re.compile( + r"^##\s+Blockers\s+and\s+risks\s*$", + re.MULTILINE | re.IGNORECASE, + ), + ), + ( + "## Next steps", + re.compile( + r"^##\s+Next\s+steps\s*$", + re.MULTILINE | re.IGNORECASE, + ), + ), + ( + "## Change log", + re.compile( + r"^##\s+Change\s*log\s*$", + re.MULTILINE | re.IGNORECASE, + ), + ), +) + +REQUIRED_IDENTITY_FIELDS = ( + "Repository", + "Working directory", + "Branch", + "Baseline commit", + "Implementation commit", + "Pull request", + "Governing issue/epic", +) + +PLACEHOLDER_PATTERN = re.compile(r"<[^>\n]+>") +HEX_COMMIT_PATTERN = re.compile(r"^[0-9a-fA-F]{7,40}$") + +SECRET_PATTERNS = ( + (re.compile(r"(?:ghp|gho|ghu|ghs|ghr)_[a-zA-Z0-9]{36}"), "GitHub token"), + (re.compile(r"github_pat_[a-zA-Z0-9_]{82}"), "GitHub fine-grained PAT"), + (re.compile(r"\bsk-[a-zA-Z0-9]{20,}\b"), "API secret key"), + ( + re.compile(r"-----BEGIN (?:RSA |EC |DSA |OPENSSH )?PRIVATE KEY-----"), + "Private cryptographic key", + ), +) + +IMPLEMENTATION_SUFFIXES = { + ".c", + ".cc", + ".cpp", + ".cs", + ".go", + ".h", + ".hpp", + ".js", + ".jsx", + ".m", + ".ps1", + ".py", + ".rs", + ".sh", + ".ts", + ".tsx", + ".yaml", + ".yml", + ".toml", +} + +IMPLEMENTATION_PREFIXES = ( + "src/", + "app/", + "backend/", + "frontend/", + "scripts/", + "shared_scripts/", + "conductor/", + "forgejo/", + ".github/workflows/", + "tests/", +) + +SAFE_EXEMPT_SUFFIXES = { + ".md", + ".rst", + ".txt", + ".lock", + ".json", + ".log", + ".tmp", + ".bak", + ".svg", + ".png", + ".jpg", + ".jpeg", + ".gif", +} + +SAFE_EXEMPT_PATHS = { + ".gitignore", + ".gitattributes", + ".claudeignore", + ".prettierignore", + "LICENSE", + "SPEC.md", + "AGENTS.md", + "CLAUDE.md", + "AGENT_HANDOFF.md", + "requirements-lock.txt", +} + +OVERRIDE_PATTERN = re.compile( + r"Canonical handoff(?:\s+location)?\s+is\s+[`\"']?([a-zA-Z0-9_\-./\\]+\.md)[`\"']?", + re.IGNORECASE, +) + + +@dataclass(frozen=True) +class HandoffFinding: + """A single governance finding against a handoff document or repository.""" + + path: Path + line: int | None + kind: str + message: str + remediation: str + + +def resolve_canonical_handoff_path(repo_root: Path) -> Path: + """Resolve canonical handoff path, defaulting to docs/development/HANDOFF.md.""" + agents_path = repo_root / "AGENTS.md" + if agents_path.is_file(): + try: + agents_text = agents_path.read_text(encoding="utf-8", errors="ignore") + explicit_marker = re.search( + r"", + agents_text, + ) + if explicit_marker: + override_rel = explicit_marker.group(1).strip() + return repo_root / override_rel + + override_match = OVERRIDE_PATTERN.search(agents_text) + if override_match: + override_rel = override_match.group(1).strip() + return repo_root / override_rel + except OSError: + pass + + return repo_root / "docs" / "development" / "HANDOFF.md" + + +def is_implementation_file(path_str: str) -> bool: + """Return True if path_str is a source, workflow, or configuration file.""" + posix = path_str.replace("\\", "/").strip() + if not posix: + return False + + name = Path(posix).name + if name in SAFE_EXEMPT_PATHS: + return False + if name == "HANDOFF.md" or posix.endswith("/HANDOFF.md"): + return False + + if any( + posix.startswith(prefix) + for prefix in ( + ".codemap/", + ".jules/", + "docs/", + "reports/", + "archive/", + "node_modules/", + ".venv/", + "venv/", + ) + ): + return False + + suffix = Path(posix).suffix.lower() + if suffix in SAFE_EXEMPT_SUFFIXES: + return False + + if suffix in IMPLEMENTATION_SUFFIXES: + return True + + return any(posix.startswith(prefix) for prefix in IMPLEMENTATION_PREFIXES) + + +def requires_handoff_update(changed_paths: Iterable[str]) -> bool: + """Return True if any changed path is an implementation file.""" + return any(is_implementation_file(path) for path in changed_paths) + + +def validate_handoff_content( + content: str, + path: Path, + is_template: bool = False, +) -> list[HandoffFinding]: + """Validate handoff content against the canonical schema.""" + findings: list[HandoffFinding] = [] + lines = content.splitlines() + + # 1. Level 1 title check + if not content.startswith("# ") and not re.search( + r"^#\s+.*Handoff", content, re.MULTILINE + ): + findings.append( + HandoffFinding( + path=path, + line=1, + kind="missing_title", + message="Document must start with '# Implementation Handoff'.", + remediation="Add '# Implementation Handoff' as the first heading.", + ) + ) + + # 2. Required section headings + for heading_title, pattern in REQUIRED_HEADINGS: + if not pattern.search(content): + findings.append( + HandoffFinding( + path=path, + line=None, + kind="missing_section", + message=f"Missing required section '{heading_title}'.", + remediation=( + f"Add '{heading_title}' section per docs/templates/HANDOFF.md." + ), + ) + ) + + # 3. Required identity fields under ## Identity + identity_match = re.search( + r"^##\s+Identity\s*\n(.*?)(?=\n##|\Z)", content, re.DOTALL | re.MULTILINE + ) + if identity_match: + identity_text = identity_match.group(1) + for field in REQUIRED_IDENTITY_FIELDS: + field_re = re.compile( + rf"^-\s+{re.escape(field)}:", re.MULTILINE | re.IGNORECASE + ) + if not field_re.search(identity_text): + findings.append( + HandoffFinding( + path=path, + line=None, + kind="missing_field", + message=f"Missing required Identity field '- {field}:'.", + remediation=f"Add '- {field}: ' under '## Identity'.", + ) + ) + + # Validate Implementation commit value + commit_match = re.search( + r"^-\s+Implementation commit:\s*([^\n]+)", + identity_text, + re.MULTILINE | re.IGNORECASE, + ) + if commit_match and not is_template: + commit_val = commit_match.group(1).strip() + first_token = commit_val.split()[0].strip("`'\",") + if first_token != "SELF" and not HEX_COMMIT_PATTERN.match(first_token): + findings.append( + HandoffFinding( + path=path, + line=None, + kind="invalid_commit", + message=( + f"Implementation commit '{first_token}' is not " + "'SELF' or a valid SHA." + ), + remediation=( + "Set 'Implementation commit: `SELF`' or the exact SHA." + ), + ) + ) + + # 4. Check for unedited placeholders when not validating the template itself + if not is_template: + for idx, line in enumerate(lines, start=1): + if line.strip().startswith("