Skip to content

The one-command test: python -m cw.testing parity over self-contained fixtures #22

Description

@thorwhalen

This issue IS the definition of v1. v1 is done when this prints identical and exits 0.

Why it is re-cut from the spec's version

As specified — "for each of the seven hard-case repos … rebuilds that repo's command set
through cw" — it cannot run in cw's CI, for five independent reasons:

  1. It contradicts D4: cw/testing.py is specified to import no cw, yet parity must
    "rebuild the command set through cw".
  2. It contradicts its own headline ("argh is NOT a runtime or test dependency of cw"): to
    rebuild a repo's command set you must import that repo, and theremin/script_utils.py:12,
    coact/__main__.py:26 and epythet/cli.py:15 are module-scope import argh.
  3. It needs seven fleet packages installed in the CI of a package whose selling point is zero
    dependencies — including t/theremin, which depends on cw. Circular.
  4. "214 cases" was an invented number (see the corpus issue).
  5. It cannot pass on the Windows runner pyproject.toml:130 turns on.

The fix

Parity runs over the seven hard-case shapes in cw/tests/fixtures.py (from the corpus
issue), against the committed argh-recorded goldens (from the goldens issue). It then
needs stdlib + cw only, pulls no argh, installs no fleet package, and runs on Windows.

The seven real repos become the migration gate — cw.testing.replay per repo, run in each
repo's own CI — which is what D4 actually asks for.

Structure: parity() lives in cw/testing.py but imports cw lazily inside the function,
so the standalone-copy guarantee (D4) survives — assert this with the AST test.

What is asserted

For each case, with every seam on its default (convention=cw.ARGH,
decode=cw.argh_decode, egress=cw.argh_egress): exit code, stdout, stderr, and the
normalised usage: line, byte-identical to the golden. The full --help body is diffed
advisorily into a human review queue (diff_help), never asserted — --help wraps to
COLUMNS and priv --help is 489 lines at default width, 365 at COLUMNS=100.

Acceptance criteria

  • python -m cw.testing parity exits 0 and prints N shapes / M cases: identical with
    real, derived N and M.
  • It runs in cw's CI on Linux, macOS and Windows (per the Windows policy issue).
  • The CI environment contains no argh — asserted by the job, not by convention.
  • It installs no fleet package.
  • cw/testing.py's module-level imports are still stdlib-only (AST test still green).
  • A deliberately introduced grammar bug (e.g. disable short-flag collision suppression)
    makes it fail with a readable diff — demonstrated in the PR.
  • All 20 argh-contract rows are exercised, per the corpus coverage test.

Depends on


Research & rationale: the canonical spec and the full fleet audit live in the private research ledger (https://github.com/thorwhalen/priv/discussions/65). Effort: medium.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions