From 9867b4c5ebb28c0982bb9ac635e0324c873fbfd3 Mon Sep 17 00:00:00 2001 From: chaoz23 Date: Tue, 25 Aug 2026 14:22:01 -0700 Subject: [PATCH 1/2] =?UTF-8?q?feat!:=20harmonise=20exit=20codes=20?= =?UTF-8?q?=E2=80=94=20usage=203,=20retrieval=204?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit BREAKING: exit 3 changes meaning. It was retrieval/validation/internal failure; it is now usage error. Retrieval failure moves to 4. FAMILY.md v2.2 harmonises 3 = usage error across every family member. A malformed call is the one non-verdict outcome every tool has, so an agent that mis-invokes any of them should get the same answer. charactercheck was the only member where 3 was already spent, which is why this was a decision rather than a patch. The cheaper option -- keep retrieval on 3 and put usage errors on 4 -- was rejected: it would have left the family with two spellings for the outcome an agent hits most often. Harmonising costs a breaking change in one repo; not harmonising costs every future caller. 0 unchanged no lint, nothing unhandled 1 unchanged lint findings 2 unchanged unhandled content (the honest lane) 3 WAS fetch NOW usage error: malformed call, incl. argparse and the structured bad_flag guards, which previously exited 2 and so shared the honest lane's code 4 NEW retrieval/validation/internal failure (was 3) Named constants carried most of the weight: EXIT_FETCH moved 3 -> 4 and EXIT_USAGE = 3 was added, so tests referring to errors.EXIT_FETCH followed automatically. argparse hardcodes 2 in error(), so main() now builds a parser subclass that exits EXIT_USAGE. tool.json's exit_codes and command_exit_contracts are regenerated from the CLI's own SCHEMA rather than hand-edited -- test_schema_documents_* compares the two surfaces and catches drift. Boundary held deliberately: a valid call whose subject has unsupported content is still 2. One test assertion was moved to 3 in error during this change and reverted; over-applying was the real risk here, not under-applying. 322 tests pass. Version 0.8.0 -- the contract changed, so consumers need to be able to pin across it. Closes #18 --- .github/workflows/ci.yml | 2 +- PRIVACY.md | 4 ++-- SKILL.md | 18 +++++++++++---- charactercheck/__init__.py | 2 +- charactercheck/cli.py | 38 +++++++++++++++++++++---------- charactercheck/errors.py | 17 +++++++++++--- pyproject.toml | 2 +- server.json | 6 ++--- tests/test_operational_clarity.py | 6 ++--- tests/test_table_evaluation.py | 4 +++- tool.json | 25 ++++++++++---------- 11 files changed, 80 insertions(+), 44 deletions(-) diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index bee4c69..3e6068c 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -26,7 +26,7 @@ jobs: uses: actions/checkout@v5 with: repository: chaoz23/srdcheck - ref: 27dee98cb7cb86e6ef86e34568d08954179f9a82 + ref: 79a7b61e640a048ef5e081357260ea211ed404e0 path: srdcheck-src - uses: actions/setup-python@v6 with: diff --git a/PRIVACY.md b/PRIVACY.md index e683bc6..3c28e03 100644 --- a/PRIVACY.md +++ b/PRIVACY.md @@ -1,8 +1,8 @@ # Privacy -> **Release contract.** These protections apply to CharacterCheck 0.7.0. +> **Release contract.** These protections apply to CharacterCheck 0.8.0. > Historical 0.6.x artifacts do not implement this privacy contract; pin and -> verify 0.7.0 when relying on it. +> verify 0.8.0 when relying on it. CharacterCheck is currently local, read-only software. This repository does not operate a hosted service, retain server-side character data, or ship telemetry. diff --git a/SKILL.md b/SKILL.md index 53177ad..3ba6d38 100644 --- a/SKILL.md +++ b/SKILL.md @@ -1,6 +1,6 @@ --- name: charactercheck -version: 0.7.0 +version: 0.8.0 description: > Deterministic D&D Beyond character-sheet derivation with per-stat provenance. Use it whenever you need a character's real numbers: "what's @@ -27,7 +27,7 @@ to trust and which items a human must resolve. derived stats are complete; the named `unhandled[]` items are your list of things to ask the player about. Don't retry, don't guess, don't discard. 3. **Every failure prints JSON with an `action` field — do what it says.** - Never a traceback; exit 3 means the sheet couldn't be retrieved at all. + Never a traceback; exit 4 means the sheet couldn't be retrieved at all. ## Exit codes ARE the verdict (per-verb since 0.7.0) @@ -38,12 +38,18 @@ For `derive`/`report`: | 0 | no lint, nothing unhandled | use the output | | 1 | lint findings — sheet disagrees with itself | usable; resolve `lint[]` with the player | | 2 | unhandled/unsupported content present (the honest lane) | trusted fields usable; resolve named items with a human | -| 3 | retrieval/validation/internal failure | read `action` (and `retryable`) in the JSON and follow it | +| 3 | usage error — you called it wrong | fix the call and retry; never route this to a human | +| 4 | retrieval/validation/internal failure | read `action` (and `retryable`) in the JSON and follow it | Projection verbs (`qa`, `seatpack`, `intake`, `snapshot`, `stance`, `quiz`) exit 0 when the projection is emitted — **inspect the embedded findings; an exit 0 there is not a cleanliness claim.** `diff`: 0 no change · 1 any named -change · 3 failure. +change · 4 failure. + +**3 is the family-wide usage code** (FAMILY.md v2.2): every check-family tool +exits 3 when the call itself is malformed, so an agent that mis-invokes any of +them gets the same answer. It moved here in 0.8.0 — retrieval failure, which +was 3 through 0.7.x, is now **4**. ## Invocation @@ -78,7 +84,7 @@ Torvald Brightmantle — Cleric 3 **Exit 2 — and that's the honest lane working:** every `trusted:` field is derived and safe to use; the named `UNSUPPORTED` lanes (here: weapon- proficiency semantics) go to a human. `(confirm)` marks a value the sheet -asserts but the derivation wants confirmed. A private sheet returns exit 3 +asserts but the derivation wants confirmed. A private sheet returns exit 4 with `"action": "Open the character on D&D Beyond, set Character Privacy to Public, and retry [...]"` and `"retryable": false` — relay the action verbatim; the tool never asks for credentials. @@ -123,6 +129,8 @@ Family contract: [FAMILY.md](https://github.com/chaoz23/srdcheck/blob/main/FAMIL persona output (no longer default); failure JSON adds `retryable`; unsupported item-semantic lanes are named per field. If you remember 0.6.x's single exit table or persona-by-default seatpacks, that's stale. +- **0.8.0:** usage errors moved to exit 3 (family-wide, FAMILY.md v2.2); retrieval + failure moved 3 -> 4. Anything handling exit 3 as a retrieval failure must move. - **0.5.1:** exit 3 + `action` field + `doctor` — retrieval failures are structured, never tracebacks. - The sheet source is the DDB character service; a saved JSON file works diff --git a/charactercheck/__init__.py b/charactercheck/__init__.py index 00a8d13..cf108ad 100644 --- a/charactercheck/__init__.py +++ b/charactercheck/__init__.py @@ -1,7 +1,7 @@ """Selected D&D Beyond character derivations with provenance and trust state.""" from .engine import derive, fetch, stance -__version__ = "0.7.0" +__version__ = "0.8.0" from .table_evaluation import project_table_evaluation __all__ = ["derive", "fetch", "stance", "project_table_evaluation"] diff --git a/charactercheck/cli.py b/charactercheck/cli.py index 880edd9..56c71af 100644 --- a/charactercheck/cli.py +++ b/charactercheck/cli.py @@ -33,22 +33,25 @@ def _json(value): "1": ("derive/report lint-only, diff named change or indeterminate " "comparison, or selftest failure"), "2": "unsupported content is present; inspect field states before use", - "3": "input, retrieval, validation, or internal failure; read the action field", + "3": ("usage error: the call itself was malformed. Fix the call and " + "retry; never route this to a human. Family-wide code, FAMILY.md " + "v2.2"), + "4": "input, retrieval, validation, or internal failure; read the action field", }, "command_exit_contracts": { "derive/report": ("0 no lint/unhandled; 1 lint with no unhandled; " - "2 one or more unhandled records; 3 structured failure"), + "2 one or more unhandled records; 4 structured failure"), "diff": ("0 complete comparison with no detected change; 1 any named " "change or indeterminate omitted-source/restriction comparison; " - "3 structured failure"), + "4 structured failure"), "stance/qa/snapshot/quiz/seatpack/intake": ( "0 when emitted even if fields/findings require attention; " - "3 structured failure"), + "4 structured failure"), "selftest": "0 pass; 1 fail", - "doctor": "0 all checks pass; 3 otherwise", - "usage": ("argparse errors are plain text and exit 2; unsupported " + "doctor": "0 all checks pass; 4 otherwise", + "usage": ("argparse errors are plain text and exit 3; unsupported " "command-specific flag combinations are " - "structured and exit 2"), + "structured and exit 3"), }, "errors": { "note": ("Known failures use stable structured errors. Unexpected failures " @@ -66,6 +69,17 @@ def _json(value): } +class _UsageExitsThree(argparse.ArgumentParser): + """argparse hardcodes exit 2 in error(); FAMILY.md clause 1 reserves 2 for + the honest lane. A malformed call is the opposite of a cannot-adjudicate + verdict -- the caller fixes it and retries -- so it exits 3, harmonised + across the family. See #18.""" + + def error(self, message): + self.print_usage(sys.stderr) + self.exit(errors.EXIT_USAGE, f"{self.prog}: error: {message}\n") + + def _exit_code(result): if result.get("unhandled", {}).get("items"): return 2 @@ -78,7 +92,7 @@ def _exit_code(result): def main(argv=None): - ap = argparse.ArgumentParser(prog="charactercheck", description=__doc__) + ap = _UsageExitsThree(prog="charactercheck", description=__doc__) ap.add_argument("command", nargs="?", choices=["derive", "stance", "qa", "report", "snapshot", "diff", "quiz", "seatpack", "intake", "doctor", "selftest"], default="derive") ap.add_argument("ref", nargs="?", help="DDB character URL / id / JSON file") @@ -108,9 +122,9 @@ def main(argv=None): "message": "--brief and --table-evaluation are mutually exclusive.", "action": "Drop --brief when requesting the shared JSON envelope.", "retryable": False, - "exit_code": 2, + "exit_code": errors.EXIT_USAGE, }), file=sys.stderr) - return 2 + return errors.EXIT_USAGE flag_contract = ( (a.brief, "--brief", {"derive", "report"}, @@ -138,9 +152,9 @@ def main(argv=None): "message": f"{flag} is not supported by '{a.command}'.", "action": action, "retryable": False, - "exit_code": 2, + "exit_code": errors.EXIT_USAGE, }), file=sys.stderr) - return 2 + return errors.EXIT_USAGE if a.command == "selftest": ok, lines = errors.selftest() diff --git a/charactercheck/errors.py b/charactercheck/errors.py index 56511de..89356cd 100644 --- a/charactercheck/errors.py +++ b/charactercheck/errors.py @@ -13,17 +13,28 @@ crashing: it is a wrong answer wearing the uniform of a right one. So every failure here is typed, carries a one-sentence **action**, and exits -**3** — a lane of its own, distinct from lint (1) and unhandled content (2), +**4** — a lane of its own, distinct from lint (1) and unhandled content (2), both of which still mean "you have usable output". +It exited 3 until FAMILY.md v2.2, which harmonises `3` = usage error across +every family member: a malformed call is the one non-verdict outcome every +tool has, so an agent that mis-invokes any of them should get the same answer. +Retrieval failure moved up to 4 rather than usage errors taking a code no +sibling uses. See chaoz23/charactercheck#18. + The other rule this module encodes: **no credentials, ever.** A private sheet is not a problem to solve with cookies, it is a problem to solve with one sentence telling the caller how to make it readable. That boundary keeps the security surface at zero and keeps the support surface small. """ -#: Failure exits with its own code. 0/1/2 keep their published meanings. -EXIT_FETCH = 3 +#: 0/1/2 keep their published meanings: pass, lint findings, unhandled content +#: (the honest lane). 3 and above are the no-verdict taxonomy, per FAMILY.md +#: v2.2 clause 1. +#: A malformed call. Harmonised at 3 across the whole family. +EXIT_USAGE = 3 +#: Retrieval failed — the sheet could not be read at all. +EXIT_FETCH = 4 # Stable error names exposed by the CLI/tool boundary. JSON-RPC protocol # errors use their standard numeric codes and are intentionally separate. diff --git a/pyproject.toml b/pyproject.toml index 1a58f3e..c2eb980 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta" [project] name = "charactercheck" -version = "0.7.0" +version = "0.8.0" description = "Experimental read-only derivation of selected D&D Beyond character fields with provenance and findings." readme = "README.md" license = "MIT" diff --git a/server.json b/server.json index d4b6bb1..180a7e8 100644 --- a/server.json +++ b/server.json @@ -2,7 +2,7 @@ "$schema": "https://static.modelcontextprotocol.io/schemas/2025-12-11/server.schema.json", "name": "io.github.chaoz23/charactercheck", "description": "Read-only CharacterCheck MCP server for deterministic character context with provenance and trust.", - "version": "0.7.0", + "version": "0.8.0", "repository": { "url": "https://github.com/chaoz23/charactercheck", "source": "github" @@ -12,13 +12,13 @@ "registryType": "pypi", "registryBaseUrl": "https://pypi.org", "identifier": "charactercheck", - "version": "0.7.0", + "version": "0.8.0", "runtimeHint": "uvx", "runtimeArguments": [ { "type": "named", "name": "--from", - "value": "charactercheck==0.7.0" + "value": "charactercheck==0.8.0" }, { "type": "positional", diff --git a/tests/test_operational_clarity.py b/tests/test_operational_clarity.py index 882e4f6..94a8a74 100644 --- a/tests/test_operational_clarity.py +++ b/tests/test_operational_clarity.py @@ -242,7 +242,7 @@ def test_brief_on_an_unsupported_command_is_refused_not_ignored(self): buf, err = io.StringIO(), io.StringIO() with redirect_stdout(buf), redirect_stderr(err): code = main(["stance", TORVALD, "--brief"]) - self.assertEqual(code, 2) + self.assertEqual(code, 3) # usage error, not the honest lane (FAMILY.md v2.2) def test_for_dm_on_derive_is_refused_not_silently_ignored(self): """A caller requesting DM redaction must never receive the ordinary @@ -252,7 +252,7 @@ def test_for_dm_on_derive_is_refused_not_silently_ignored(self): mock.patch("charactercheck.cli.derive", side_effect=AssertionError("must reject before derive")): code = main(["derive", TORVALD, "--for-dm"]) - self.assertEqual(code, 2) + self.assertEqual(code, 3) # usage error, not the honest lane (FAMILY.md v2.2) self.assertEqual(stdout.getvalue(), "") payload = json.loads(stderr.getvalue()) self.assertEqual(payload["error"], "bad_flag") @@ -270,7 +270,7 @@ def test_command_specific_flags_are_closed_not_silently_ignored(self): with self.subTest(flag=flag), redirect_stdout(stdout), \ redirect_stderr(stderr): code = main(argv) - self.assertEqual(code, 2) + self.assertEqual(code, 3) # usage error, not the honest lane (FAMILY.md v2.2) self.assertEqual(json.loads(stderr.getvalue())["error"], "bad_flag") diff --git a/tests/test_table_evaluation.py b/tests/test_table_evaluation.py index 00c43db..020a4de 100644 --- a/tests/test_table_evaluation.py +++ b/tests/test_table_evaluation.py @@ -100,6 +100,8 @@ def test_cli_flag_uses_shared_exit_lane(self): with redirect_stdout(output): code = main(["derive", FIXTURE, "--table-evaluation"]) result = json.loads(output.getvalue()) + # Stays 2: a valid call whose subject has unsupported content is the + # honest lane, not a malformed call. self.assertEqual(code, 2) self.assertEqual(result["status"], "unsupported") @@ -107,7 +109,7 @@ def test_brief_and_shared_envelope_cannot_silently_compete(self): output = io.StringIO() with redirect_stderr(output): code = main(["derive", FIXTURE, "--brief", "--table-evaluation"]) - self.assertEqual(code, 2) + self.assertEqual(code, 3) # usage error, not the honest lane (FAMILY.md v2.2) self.assertEqual(json.loads(output.getvalue())["error"], "bad_flag") diff --git a/tool.json b/tool.json index e93df1e..b9e14ea 100644 --- a/tool.json +++ b/tool.json @@ -1,16 +1,16 @@ { "name": "charactercheck", "family_class": "verdict", - "version": "0.7.0", + "version": "0.8.0", "status": "released", "distribution": { - "source": "This manifest describes CharacterCheck 0.7.0.", - "published": "Pin charactercheck==0.7.0 and verify the public artifact and trust-bearing output.", + "source": "This manifest describes CharacterCheck 0.8.0.", + "published": "Pin charactercheck==0.8.0 and verify the public artifact and trust-bearing output.", "version_warning": "Source metadata is not artifact provenance; verify the installed version and distribution attestations." }, "description": "Read-only compilation of selected D&D Beyond character fields into mechanical context, canonical assessments, provenance, findings, and versioned snapshots.", "boundary": "Not complete rules validation, encounter/world/session state, role enforcement, or proof of action legality. No mutations are exposed.", - "install": "Pin charactercheck==0.7.0; from a trusted clone, python3 -m charactercheck runs without runtime dependencies", + "install": "Pin charactercheck==0.8.0; from a trusted clone, python3 -m charactercheck runs without runtime dependencies", "example": "python3 -m charactercheck derive examples/sample-character.json --brief", "commands": { "charactercheck derive [--brief]": "selected derivation plus canonical fields, findings, trust, and observation metadata", @@ -37,18 +37,19 @@ "limits": "Raw local/remote character payloads: 8 MiB. Local CharacterSnapshotV1 envelopes: 16 MiB. MCP request lines: 8 MiB. All JSON also has bounded depth, nodes, strings, collections, inventory, modifiers, and container traversal." }, "exit_codes": { - "0": "command-specific success; successful projections can still contain non-trusted fields", + "0": "command succeeded with no command-specific finding/change status", "1": "derive/report lint-only, diff named change or indeterminate comparison, or selftest failure", - "2": "derive/report has an unhandled record, or an explicit compatibility flag guard failed; argparse usage errors also use 2", - "3": "structured retrieval/input/snapshot/internal failure, or a failed doctor check" + "2": "unsupported content is present; inspect field states before use", + "3": "usage error: the call itself was malformed. Fix the call and retry; never route this to a human. Family-wide code, FAMILY.md v2.2", + "4": "input, retrieval, validation, or internal failure; read the action field" }, "command_exit_contracts": { - "derive/report": "0 no lint/unhandled; 1 lint with no unhandled; 2 one or more unhandled records; 3 structured failure", - "diff": "0 complete comparison with no detected change; 1 any named change or indeterminate omitted-source/restriction comparison; 3 structured failure", - "stance/qa/snapshot/quiz/seatpack/intake": "0 when emitted even if fields/findings require attention; 3 structured failure", + "derive/report": "0 no lint/unhandled; 1 lint with no unhandled; 2 one or more unhandled records; 4 structured failure", + "diff": "0 complete comparison with no detected change; 1 any named change or indeterminate omitted-source/restriction comparison; 4 structured failure", + "stance/qa/snapshot/quiz/seatpack/intake": "0 when emitted even if fields/findings require attention; 4 structured failure", "selftest": "0 pass; 1 fail", - "doctor": "0 all checks pass; 3 otherwise", - "usage": "argparse errors are plain text and exit 2; unsupported command-specific flag combinations are structured and exit 2" + "doctor": "0 all checks pass; 4 otherwise", + "usage": "argparse errors are plain text and exit 3; unsupported command-specific flag combinations are structured and exit 3" }, "privacy": { "default": "mechanical; account identifiers, linked images, appearance, notes, and persona omitted", From 58668d5708fd663b753d5f421ad878157666569e Mon Sep 17 00:00:00 2001 From: chaoz23 Date: Tue, 25 Aug 2026 14:23:15 -0700 Subject: [PATCH 2/2] ci: the cold-wheel probe asserts the new failure lane A malformed ref is a typed errors.py failure, which moved 3 -> 4 with the rest of that lane. The step asserted 3. Deliberate scope note: bad_ref means "the reference is not a supported input shape", which is arguably a malformed CALL and so arguably belongs on 3 with argparse and the flag guards. It stays on 4 here because errors.py is built as ONE typed-failure lane carrying an action field, and splitting that lane by usage-vs-retrieval is a larger classification exercise across ~25 error kinds. Filed separately rather than decided in passing. The defect this PR fixes is closed either way: no usage error shares the honest lane's exit 2. Co-Authored-By: Claude Opus 5 --- .github/workflows/ci.yml | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 3e6068c..42b80c7 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -151,7 +151,7 @@ jobs: env: *offline-network run: /tmp/charactercheck-cold/bin/charactercheck selftest - - name: A malformed ref exits 3 with a redacted actionable JSON error + - name: A malformed ref exits 4 with a redacted actionable JSON error working-directory: /tmp env: *offline-network run: | @@ -159,7 +159,7 @@ jobs: out=$(/tmp/charactercheck-cold/bin/charactercheck derive notanumber 2>&1); code=$? set -e echo "$out" - test $code -eq 3 || { echo "expected exit 3, got $code"; exit 1; } + test $code -eq 4 || { echo "expected exit 4, got $code"; exit 1; } echo "$out" | grep -q '"error": "bad_ref"' echo "$out" | grep -q '"action"' echo "$out" | grep -qv "Traceback" || { echo "traceback leaked"; exit 1; }