Migrate the CLI from argh to cw - #25
Merged
Merged
Conversation
31 argv-level cases recorded from the current argh-based console script: the whole help surface (top level + every subcommand), a happy path per subcommand, the short/long/`=` option spellings argh derives, seven usage errors that must keep exit code 2, and argh's own no-argument behaviour (usage on *stdout*, exit 0 -- which plain argparse does not do). Recorded first, on purpose: this is the baseline the migration is measured against, so it has to exist in history before any source changes. The corpus deliberately excludes `--format json` and every `install-skills` happy path. Both print the *resolved* project root, and a committed golden that embeds an absolute path cannot be replayed on another machine. Their grammar is still pinned by the tier-2 usage line of the corresponding `--help` case. Claude-Session: https://claude.ai/code/session_01K6LB3AwUmKDxaFNZ2NqPGr
`from cw import compat as argh` -- nothing else changes. Replayed against the
golden recorded in the previous commit:
31/31 identical (--strict-help, so the --help bodies are asserted too)
no --help changes
135 passed
Kept as its own commit because it is the evidence for the claim: the compat
shim is a drop-in, and any behaviour difference in the next commit is
attributable to the move to cw's own API rather than to leaving argh.
Claude-Session: https://claude.ai/code/session_01K6LB3AwUmKDxaFNZ2NqPGr
`__main__.py` now calls `cw.dispatch(_dispatch_funcs)` directly rather than
going through the compat shim, and `argh` is replaced by `cw>=0.1,<0.2` in
`dependencies`. `rg -n argh opsward/` returns nothing.
Replay against the pre-migration golden, with the --help bodies asserted:
31/31 identical
no --help changes
136 passed
`prog=` is deliberately NOT passed. Letting argparse derive the program name
from sys.argv[0] is what keeps `python -m opsward` reporting `__main__.py` and
the console script reporting `opsward` -- pinning `prog='opsward'` would have
changed the `-m` form's usage line, which is the invocation the module
docstring documents and the whole existing test suite uses.
New `tests/test_cli_parity.py` replays the golden as part of the ordinary test
run, so this cannot rot: a refactor here, or a future cw release, that changes
any exit code, any byte of stdout/stderr, or any `usage:` line fails the suite.
Verified it can fail -- forcing a different naming convention produced a red
test naming the exact flags that moved. The full --help *body* is asserted only
when the running CPython matches the one that recorded the golden, because
CPython rewrites its own help rendering between versions and the matrix here is
3.10 + 3.12.
Docs updated to describe the pinned surface and how to re-record it.
Claude-Session: https://claude.ai/code/session_01K6LB3AwUmKDxaFNZ2NqPGr
CI went red on Python 3.10, and it was not cw. argparse is stdlib and rewrites
its own text between versions:
* 3.12 stopped listing `nargs='*'` positionals among "the following arguments
are required", so bare `opsward find` says `required: query` on 3.12 and
`required: query, project-roots` on 3.10;
* 3.12 also changed how `invalid choice` quotes the choices.
Verified against plain argparse on 3.10/3.11/3.12/3.13 with no argh and no cw in
the picture. Those two cases are the ONLY difference between the two recordings;
argh produces them identically. A single golden asserted across a matrix fails
for something nobody caused, which is the fastest way to teach a team to ignore
a red parity test.
So: `misc/cli_golden_py310.json` and `misc/cli_golden_py312.json`, and the test
picks the one matching the running interpreter. A version with no recording
fails loudly with instructions -- a parity test that quietly skips is worse than
none.
The 3.10 golden is recorded from *argh* too, not from the migrated code: a
throwaway 3.10 venv over a worktree pinned at the pre-migration commit. It is a
migration proof on both matrix versions, not a baseline taken after the fact.
Replaying it against the cw code on 3.10 gives 31/31 identical, same as 3.12.
`strict_help=True` unconditionally now. The gate on the recording interpreter is
what makes that safe, and it upgrades the --help body from advisory to asserted.
Claude-Session: https://claude.ai/code/session_01K6LB3AwUmKDxaFNZ2NqPGr
argparse takes its `prog` from `basename(sys.argv[0])`, and on Windows the console
script is installed as `opsward.EXE`, so 22 of the 31 recorded cases differed on
the runner:
- usage: opsward [-h] {diagnose,generate,maintain,recommend,install-skills,find} ...
+ usage: opsward.EXE [-h] {diagnose,generate,maintain,recommend,install-skills,find} ...
That is a fact about packaging, not about the command line. cw 0.1.1 scrubs it
(i2mint/cw#34), a defect this migration found. The floor is raised rather than left
at 0.1 because with `cw>=0.1` a resolver could pick 0.1.0 and this repo's own parity
test would fail on Windows for a reason nobody could act on.
Claude-Session: https://claude.ai/code/session_01K6LB3AwUmKDxaFNZ2NqPGr
The Windows runner is down to six differing cases, and none is about the command line. opsward prints `misc\docs\...` where POSIX prints `misc/docs/...`, and it reports file sizes inflated by CRLF checkout (a 37-byte stub measures 40). Both are the pre-existing Windows bugs in #21 -- the same run fails test_generate.py::test_python_docs_path for the same reason -- and argh printed exactly the same thing. They are listed as `expect_diff` rather than skipping the test on Windows, so everything else there stays asserted: exit codes, stdout, stderr, usage lines and the whole --help body. That distinction is not academic. The `.exe` defect this migration found in cw (i2mint/cw#34) lived exactly in the part a platform-wide skip would have stopped checking, and it took a Windows run to see it. If one of the six starts matching, the test fails with `unexpected-match`. That is the correct outcome: it means #21 was fixed and the entry should be deleted. Claude-Session: https://claude.ai/code/session_01K6LB3AwUmKDxaFNZ2NqPGr
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #24.
Replaces
argh(LGPL-3.0-or-later) withcw0.1.0 (MIT, zero runtime dependencies). Nothing a user sees changes.The proof
misc/cli_golden.jsonwas recorded from the argh CLI before any source edit (its own commit, first in the branch) and replayed after each wave:--helpfrom cw import compat as arghcw.dispatch(_dispatch_funcs)--helpdiff: none. Not "close enough" —cw.testing replay --strict-helpasserts the full help body of the top level and all six subcommands, andpython -m cw.testing diff-helpprintsno --help changes. Noexpect_diffentries.136 passed(135 pre-existing + the new parity test).The corpus — 31 cases, in
misc/cli_cases.txt--helpplus every one of the six subcommands'--help(theirusage:lines are asserted at tier 2, so a lost flag or a changednargsis caught width-independently);-m 0), long (--min-score 0),=form (--min-score=0), boolean flags, repeated positionals;argparsedoes not do this, so it is the case that would catch a dispatcher that only looked compatible;sys.exit(2)runtime error path.Two case classes are deliberately absent, with the reason in the corpus file:
--format jsonand everyinstall-skillshappy path print the resolved project root. A committed golden holding an absolute path cannot replay on another machine. Their grammar is still pinned by the tier-2 usage line of the matching--helpcase, andtests/test_cli.pyalready covers the JSON behaviour.It cannot rot
tests/test_cli_parity.pyreplays the golden in the ordinary test run, so a futurecwrelease cannot silently change this CLI. Verified it can fail: forcingconvention=cw.MODERNturned it red with the exact flags that moved. The full--helpbody is asserted only when the running CPython matches the recording's, because CPython rewrites its own help rendering between versions and the matrix here is 3.10 + 3.12; exit codes, stdout, stderr andusage:lines are asserted on every version.One deliberate deviation from the issue
The issue suggested
cw.dispatch(_dispatch_funcs, prog='opsward').prog=is not passed. Lettingargparsederive the program name fromsys.argv[0]is what keepspython -m opswardreporting__main__.pyand the console script reportingopsward— pinningprog='opsward'would have changed the-mform's usage line, which is the invocation the module docstring documents and the entire existing test suite uses. That is a CLI surface change, and this migration promised none.Also
dependencies:argh>=0.31→cw>=0.1,<0.2(the fleet-wide spelling).rg -n 'argh' opsward/returns nothing.CLAUDE.md/AGENTS.md/misc/docs/architecture.mdupdated, including how to re-record the golden and the rule that a case must never print an absolute path.Wave 0 is kept as its own commit on purpose: it is the evidence that the compat shim is a genuine drop-in, so any difference in the following commit would have been attributable to cw's own API rather than to leaving argh. (There was none.)
Two cw defects this migration found (both fixed and released)
Fixed in i2mint/cw#34, shipped as cw 0.1.1, which is why the floor here is
cw>=0.1.1:.exewas not scrubbed from recorded text.argparsetakesprogfrombasename(sys.argv[0]), so the runner reportedusage: opsward.EXE ...— 22 of 31 cases red, none of them about this CLI.cw.testing's docstring explicitly promised "a golden recorded on a Mac asserts cleanly on a Windows runner", and even named the.execonsole-script shim as a reasonparityavoids subprocesses — whilecharacterize/replay, which do spawn subprocesses, had no defence against it.cw.testinginherited the recorder's stdin, so any CLI with an interactive path recorded a timeout at a terminal and an EOF under CI. Not hit by opsward (nothing here reads stdin) but hit hard by the siblinggrubmigration.One CI failure that was not cw
Python 3.10 went red on two cases, and the cause is CPython. 3.12 stopped listing
nargs='*'positionals among "the following arguments are required", and changed howinvalid choicequotes the choices — verified against plainargparseon 3.10/3.11/3.12/3.13 with neither argh nor cw in the picture. argh produces those messages identically.So: one golden per CPython version in the matrix, each recorded from argh (a throwaway 3.10 venv over a worktree pinned at the pre-migration commit) — a migration proof on both versions, not a baseline taken after the fact. Those two cases are the only difference between the two recordings. A version with no golden fails loudly with instructions rather than skipping; a parity test that quietly does nothing is worse than none.
strict_help=Trueis now unconditional. Gating on the recording interpreter is what makes that safe, and it promotes the--helpbody from advisory to asserted.Windows: asserted, minus six documented cases
After cw 0.1.1, the Windows runner was down to six differing cases, and none of them is about the command line:
generateprintsmisc\docs\...where POSIX printsmisc/docs/...;maintainanddiagnose --verbosereport file sizes inflated by CRLF checkout (a 37-byte stub measures 40).Both are the pre-existing Windows bugs in #21 — the same run fails
test_generate.py::test_python_docs_pathfor the identical reason — and argh printed exactly the same thing.They are listed as
expect_diffrather than skipping the test on Windows, so everything else there stays asserted: exit codes, stdout, stderr, usage lines, the whole--helpbody. That distinction is not academic — the.exedefect this migration found in cw lived precisely in the part a platform-wide skip would have stopped checking, and only a Windows run could see it. If one of the six starts matching, the test fails withunexpected-match, which is the correct outcome: it means #21 was fixed and the entry should be deleted.test_cli_paritynow passes on Windows. The Windows job still fails on the same three pre-existingtest_generate.pycases that fail onmaintoday — no new failures, and Windows is already non-blocking here per #21.Two follow-ups for #21, both out of scope for this PR but now precisely characterised: a
.gitattributeswitheol=lfwould fix the byte-count half, and printingPurePosixPath-style separators would fix the other.Checks
test_cli_paritypasses; the 3 pre-existing CI: Windows tests failing (pre-existing, non-blocking) #21 failures remain, unchanged frommain