diff --git a/AGENTS.md b/AGENTS.md index faaed3b..ed60543 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -8,7 +8,7 @@ A Python package that makes a TypeSafe Jev judgment usable as control flow: a ye `if`, a choice is an exhaustive `match`, a score is a comparison. One distribution, `guideme`, published to PyPI under `MIT OR Apache-2.0`. The public surface has two tiers: -- the 33 names in `__all__` in `src/guideme/__init__.py`, imported from `guideme` itself; +- the 35 names in `__all__` in `src/guideme/__init__.py`, imported from `guideme` itself; - `guideme.api` and `guideme.policy` as whole modules, imported by their own path and not re-exported at the top level: `guideme.api` is the wire mirror and `guideme.api.client` holds `Client` and `AsyncClient`, and `guideme.policy` holds `resolve`. @@ -40,7 +40,7 @@ two pages: `https://docs.typesafe.ai/api.md` covers `POST /v1/systemone` and | `src/guideme/scalars.py` | `Probability`, `Confidence`, `Key`, `Rank`, `ApiKey`, `Model` | validation happens once, here; `ApiKey` never prints | | `src/guideme/errors.py` | the `GuidemeError` tree and `kind` | `kind` is the cross-SDK name and the `error.type` value; imports nothing from `guideme` | | `src/guideme/policy.py` | `resolve`, `Policy`, `Thresholds`, `Verdict`, the answer and outcome dataclasses | pure: no I/O, no caller enums, keys and level indices only | -| `src/guideme/enums.py` | `Choice`, `Levels`, `fallback` | a member's name is its wire key and its value is its rubric; both validate at class definition, and a repeated rubric text is refused there | +| `src/guideme/enums.py` | `Choice`, `Levels`, `option`, `level`, `fallback`, and the internals `render` and the `require_*` checks | a member's name is its wire key and its value is its rubric; both validate at class definition, and a repeated rubric text is refused there. `render` is the one place a rubric's examples become wire text, and its output is a cross-SDK contract item. `render`, `require_unshared_examples`, `require_no_counterexamples` and `require_no_fallback` are internal despite their names: they are imported by `question.py` and are in no tier, like `question.validate` | | `src/guideme/question.py` | question kinds, constructors, `Ranked`, `Scored`, the unsure ladder | a question is inert until asked; the reader travels with it | | `src/guideme/ask.py` | shapes: `encode`, `decode`, `Plan` | ids are `q0..qN` in encounter order, insertion order for a dict | | `src/guideme/_ask_overloads.py` | the typed `ask` surfaces | GENERATED; edit `scripts/gen_ask_overloads.py` and run `mise run gen` | @@ -106,6 +106,32 @@ the rest are checked in review. is a `ConfigError` on the class statement; so is `fallback(…)` on a `Levels`, which has `.otherwise(level)` instead. Both fire where the enum is written, like the Rust derive's compile errors. +- **A rubric's examples are composed into its text at the wire, and nowhere else.** A rubric's + value stays the bare text; `enums.render` is what turns an `option(…)`, `level(…)` or + `fallback(…)` into what the model reads, and every site that puts a rubric on the wire — + `Choice.rubric()`, `Levels.levels()`, `choose_among`, `score_levels` and + `NoulQuestion.criteria` — goes through it, so a rubric cannot arrive having quietly lost what + it carries. All three kinds of question take examples: a yes and a no are as confusable as + two options, and `option(…)` is the one constructor for a described alternative, so there is + no fourth. A bare string and a rubric with no examples render to their own bytes, so an + existing caller's request does not move. The rendered string is a contract item shared with + every other SDK, stated in `docs/contract.md` and pinned by `spec/vectors/rubric.json`, and + so is the order: examples and counterexamples render in the order written, never sorted and + never de-duplicated into a set. +- **A rubric's examples must be consistent, and that is checked where it is written.** A blank + or whitespace-only entry, a repeat within one clause, a clause given as one string rather + than a sequence of them, an empty clause written out (`examples` and `counterexamples` + default to `None`, so any empty sequence was typed), examples attached to a blank rubric, + a `fallback(…)` on a runtime path, a `U+000A` or `U+000D` inside an entry, and a + counterexample on a level are each a `ConfigError`. The newline test is the literal + codepoint, never `str.splitlines()`, which splits on eight and would refuse declarations the + Rust SDK accepts; `docs/contract.md` records why. + A blank rubric that carries no examples is **not** an error: it means what it meant in 0.1.0, + and this release does not redefine it. So are the two contradictions: one + string as an example of two alternatives of the same question, and one string as both an + example and a counterexample of the same alternative. One string as an example of one + alternative and a counterexample of another is **legal and required** — it is the confusable + pattern the feature exists for, and `tests/test_rubrics.py` sends it on the wire. - **Nothing outside `api/` may see a `pydantic` exception.** A caller's state or instructions that pydantic refuses leaves `api/` as a `ConfigError`, and a `NaN` or an infinity is refused rather than serialised as `null`. @@ -159,13 +185,19 @@ is made in `guideme-rust` first, not here. ## Tests -Few tests, high grade. The ceiling is 43 test functions; a parametrised function counts once. +Few tests, high grade. The ceiling is 47 test functions; a parametrised function counts once. It was 40 before the logs signal, which is user-requested scope that the span assertions could not cover: correlation, severity and routing each need a record to look at. The forty-third is the pre-publish proof that a log sink which raises reaches neither the caller nor the ask span: it asserts the absence of a failure on a path where every other test asserts a presence, so no -existing test could carry it. Everything else that pass added went into a parameter of a test -that was already there. A new test must be one of: +existing test could carry it. Rubric examples added the last four, also user-requested scope: +the golden rendering table, the passthrough property that proves 0.1.0's bytes have not moved, +the wire proof that a rendered rubric reaches the request, and a live proof that examples move +the distribution — one function covering a choice and a noul, because both are the same +invariant and a second function would buy nothing. Each asserts a different thing about a +string no existing test looks at. +Everything else those two passes added went into a parameter of a test that was already there. +A new test must be one of: - a property test (`hypothesis`) over a law of `policy.resolve`, the shapes, or the wire types; - a wire or contract check through a real local HTTP server (`pytest-httpserver`), asserting on @@ -280,9 +312,12 @@ green before the tag, not after. 1. Bump `version` in `pyproject.toml`. Then `uv lock` at the root and `uv lock` inside `examples/otlp`: both lock files record the version, and both are resolved with `--locked`. 2. Move the `Unreleased` notes in `CHANGELOG.md` under the new version with today's date. -3. Refresh the **What arrives** capture in `docs/observability.md`, or elide the version in it. - Its `InstrumentationScope guideme X.Y.Z` lines carry the version the capture was taken at, - so they go stale on the first bump and a reader cannot tell a stale capture from a real one. +3. Leave the **What arrives** capture in `docs/observability.md` alone unless you re-take it. + Its `InstrumentationScope guideme X.Y.Z` lines carry the version the run actually emitted, + and the prose above it names that version, so a reader can tell the capture's age from a + claim about today. Do not edit the version forward to match a release no run produced, and + do not delete it either: that loses the provenance permanently. Re-take the capture and + update both together, or change nothing. 4. `mise run check`, then open a pull request and squash-merge it with CI green. 5. On `main`, at that commit: `git tag -a vX.Y.Z -m vX.Y.Z` and `git push origin vX.Y.Z`. The tag ruleset refuses a tag that is later moved or deleted, so tag the commit you mean. diff --git a/CHANGELOG.md b/CHANGELOG.md index aedd5d8..3511e0a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -4,6 +4,52 @@ Nothing yet. +## 0.1.1 — 2026-09-22 + +Additive. Nothing that worked in 0.1.0 sends different bytes. + +- `option(rubric, examples=…, counterexamples=…)` and `level(rubric, examples=…)` join + `fallback(…)`, which now takes the same keywords. All three are exported from `guideme`, + bringing `__all__` to 35 names. A rubric written as a bare string keeps working everywhere. +- The parts are composed into the rubric where it becomes wire text: clauses joined with a + newline, items within one joined with `"; "`, in the order written, the text verbatim. A + rubric with no examples renders to its own text, byte for byte, so an existing request is + unchanged. The rendered string and its order are cross-SDK contract items, stated in + `docs/contract.md` and pinned by `spec/vectors/rubric.json`. +- All three kinds of question take them. `noul("…").criteria(yes, no)` accepts an `option(…)` + for either side: a yes and a no are as confusable as two options, and there is no fourth + constructor for them. +- Every site that puts a rubric on the wire renders — `Choice.rubric()`, `Levels.levels()`, + `choose_among`, `score_levels` and `NoulQuestion.criteria` — so examples cannot be silently + dropped by reaching a runtime constructor. +- `level(…)` takes no counterexamples: "not this option" means nothing on an ordered scale. An + `option(…)` carrying counterexamples written where a level belongs is a `ConfigError`, as is + a blank or whitespace-only entry, a repeat within one clause, examples attached to a blank + rubric, and an empty clause written out: `examples` and `counterexamples` default to `None`, + so any empty sequence that arrives was typed on purpose and says nothing. A blank rubric + carrying no examples is untouched — it means what it meant in 0.1.0. +- Contradictory examples are a `ConfigError` too: one string as an example of two alternatives + of the same question, or as both an example and a counterexample of the same alternative. One + string as an example of one alternative and a counterexample of another stays legal — that is + the confusable-options pattern the feature exists for. +- A clause given as one string is refused rather than shredded. A `str` is a `Sequence[str]` of + its own characters, so `examples="refund"` would have become six one-letter examples and no + type checker would have said so; it is a `ConfigError` naming the mistake, the same way + `score_levels` already refuses a scale given as one string. +- `fallback(…)` is refused by `choose_among`, `score_levels` and `noul(…).criteria(…)`. Those + answer in a `Key`, a `Rank` and a `bool`, none of which has a member to fall back to, so the + marking had nothing to act on and was being dropped in silence. Use `.otherwise(…)` on the + question, which is what the `Levels` rule has always said. +- An example or counterexample containing `U+000A` or `U+000D` is refused. Items are joined + onto one line, so a newline inside one would read as a clause the rubric never declared. The + rubric text itself is unrestricted; only the entries are. `"; "` inside an entry stays legal, + because it changes how many examples a reader sees rather than which clause they are in. +- A rubric built by `option(…)`, `level(…)` or `fallback(…)` can be copied and pickled again. + Carrying the parts meant `__new__` took four arguments where `str` hands back one, so + `copy.copy`, `copy.deepcopy` and `pickle` raised a `TypeError` — including on the + `copy.deepcopy({"key": option(…)})` a caller writes before `choose_among`. A bare string did + this in 0.1.0 and does it again. + ## 0.1.0 — 2026-09-21 First release. Everything below is new, so this entry lists the surface rather than the diff --git a/README.md b/README.md index 71a79d3..8929bde 100644 --- a/README.md +++ b/README.md @@ -104,6 +104,87 @@ A noul can carry `.criteria("what yes means", "what no means")`. Instructions ac or any JSON-shaped value, so a question can reference structured data by field name the way the TypeSafe docs describe. +### Examples in a rubric + +Two alternatives that read alike are told apart by showing inputs rather than by describing +harder. `option(…)` takes the inputs that belong to an alternative and the ones that belong +somewhere else, `level(…)` takes the inputs that score at that level, and `fallback(…)` is an +`option(…)` that also marks the unsure member. All three kinds of question take them: a noul's +`.criteria(…)` accepts an `option(…)` for the yes and for the no. + +```python +from guideme import Choice, Levels, fallback, level, option + + +class Department(Choice): + billing = option( + "Payments, invoicing, refunds", + examples=["My card was charged twice", "Where is my refund?"], + counterexamples=["The dashboard is down"], + ) + technical = option("Bugs, outages, integrations", examples=["502 on every request"]) + sales = fallback("Pricing, upgrades, new accounts", examples=["Do you have a team plan?"]) + + +class Severity(Levels): + cosmetic = level("No impact to functionality", examples=["typo in a label"]) + degraded = level("Broken feature, workaround exists", examples=["export fails in one browser"]) + blocking = level("No workaround exists", examples=["cannot log in", "data loss"]) +``` + +The member's value is still the bare rubric; the examples are composed into it only in the +request, as + +```text +Payments, invoicing, refunds +Examples: My card was charged twice; Where is my refund? +Not this option: The dashboard is down +``` + +So a rubric with no examples sends exactly what it sent before, and the same strings work in +`choose_among("…", {"billing": option(…)})` and `score_levels("…", [level(…), …])`. Examples +and counterexamples render in the order they are written, always: that order is part of the +published contract. + +A yes and a no are two alternatives of one question, so they take examples too, and this is +where they pay best — a vague pair is the easiest thing to get wrong: + +```python +urgent = noul("Is this ticket urgent?").criteria( + option("Urgent", examples=["customers cannot log in", "money is moving to the wrong place"]), + option("Not urgent", examples=["a broken job with a manual workaround", "a cosmetic bug"]), +) +``` + +Asked about a nightly export job that has been failing since Tuesday while the numbers are +pulled by hand, a plain `Urgent` / `Not urgent` answers yes at 0.75. The criteria above answer +no at 0.17, because one of the not-urgent examples is what the ticket describes. + +A string may be an example of one option and a counterexample of another. That is the point +when two options are confusable, and it is the one overlap that stays legal. Offering the same +string as an example of two options, or as both an example and a counterexample of the same +option, says an input belongs where it cannot, so each is refused. + +`level(…)` has no counterexamples, because "not this option" means nothing on an ordered +scale — an input that does not belong at one level scores at another. + +Leave a clause out to say there is none. An empty one written out — `examples=[]` or +`examples=()` — says nothing, so it is refused as the mistake it is, along with a clause given +as one string rather than a list of them (`examples="refund"` would otherwise be six one-letter +examples), a blank entry, a repeat within one clause, a newline or carriage return inside an +entry, a counterexample on a level, a `fallback(…)` given to `choose_among`, `score_levels` or +`.criteria(…)`, and the two contradictions above. Each is a `ConfigError` where the rubric is +written. + +Entries go on one line each, so a newline inside one would read as a clause you never wrote. +`"; "` inside an entry is fine — `"card declined; retry failed"` is ordinary prose, and it +changes how many examples a reader sees rather than which clause they are in. The rubric text +itself may still contain newlines; only the entries are restricted. + +Attaching examples to a blank rubric is refused too, because they describe something that is not +there. A blank rubric on its own is not: it means what it has always meant, and adding examples +to the language does not make an old declaration an error. + The state is anything JSON-shaped: a text literal, a `dict`, a list of them. A dataclass goes through `dataclasses.asdict`, a pydantic model through `.model_dump()`. diff --git a/docs/contract.md b/docs/contract.md index 26ff89e..7d4a85a 100644 --- a/docs/contract.md +++ b/docs/contract.md @@ -16,6 +16,166 @@ every one of which is a parametrised case of `tests/test_policy_vectors.py`: an and another SDK is either a bug here or an ambiguity to resolve upstream; it is never fixed by changing the copy. +The rendered rubric string is a contract item too. An alternative written with `option(…)`, +`level(…)` or `fallback(…)` carries its examples beside its text, and the string those compose +into is what goes on the wire. Every guideme SDK composes the same string from the same parts, +and `spec/vectors/rubric.json` is where the renderers are held to each other. All three kinds +of question carry them, a noul's yes and no included. + +The algorithm, in full, because this is what a new SDK implements from: + +``` +render(what, examples, counterexamples) -> string: + if examples is empty and counterexamples is empty: + return what # unchanged, byte for byte + lines = [what] + if examples: + lines.append("Examples: " + join(examples, "; ")) + if counterexamples: + lines.append("Not this option: " + join(counterexamples, "; ")) + return join(lines, "\n") +``` + +The two labels are literal and exact: `"Examples: "` and `"Not this option: "`, each with its +trailing space. They are not formatting to be chosen locally — `Not this option` was measured +against `Not` and `Counterexamples` on the same confusable case and reached the correct option +with the highest mean probability of the three, so a different label is a different, and worse, +contract. The renderer appends the examples clause before the counterexamples clause — a +statement about the renderer, not about every string that reaches the wire, since a caller +who writes the labels into `what` by hand can put them in any order they like. + +`what` is used verbatim: never trimmed, never re-punctuated. Newline separation is what makes +that safe, since examples often end in `?` and a space-joined format would need a trailing `.` +that produces `Where is my refund?.`. + +**Order within a clause is part of it too.** Examples and counterexamples render in the order +they were written, never sorted and never collapsed into a set, because two SDKs ordering +differently would send different bytes for the same declaration. + +`spec/vectors/rubric.json` carries one object per case with five fields: `kind`, one of +`choice`, `levels` or `noul`, which says what surface the case came from; `what`, the rubric +text; `examples` and `counterexamples`, the two clauses in declaration order; and `rendered`, +the exact bytes the algorithm above must produce. The `noul` cases come from the runtime +renderer rather than a derive, so they are the vector's only cover for that path. + +An absent clause appears as `[]`, not as `null` and not by omitting the key. A consumer must +read that as *no clause was declared* and rebuild the case without the argument, because an +empty clause written out is itself a refused declaration and so can never be a vector case. +`kind` is what picks the constructor to rebuild with: a `levels` case has to go through the +level constructor and a `choice` or `noul` case through the option one, which is how the vector +exercises the surface a caller writes rather than the renderer alone. + +Items are inserted **verbatim**. Nothing is escaped, and the rendering is not required to be +reversible — no SDK parses a rendered rubric back into its parts, and none should be written +to. Rubric text is trusted: an SDK does not sanitise it, and it is the caller's to get right. + +What the rules below do guarantee is narrower and worth stating exactly: **a string an SDK +renders from declared parts carries exactly the clauses those parts declared.** That is +rendering integrity, not input trust. + +### Validation is over the declared items, not the rendered text + +One sentence that decides every case a third implementer will hit, and the reason none of them +needs a special rule: + +- `["a; b"]` renders identically to `["a", "b"]`, and is still **one** item. It passes the + duplicate and shared-example checks that two items would meet. +- `" a"` and `"a"` are two items, because they are two strings. +- `"A"` and `"a"` are two items, for the same reason. There is no case folding. + +An implementation that normalised, trimmed or split before comparing would refuse declarations +these SDKs accept. Compare the strings the caller declared, and nothing else. + +`"; "` inside an item is therefore legal, and a newline is not — which looks inconsistent until +you see what each one does. `"; "` changes how many examples a reader sees, and +`"card declined; retry failed"` is ordinary prose a caller is entitled to write. A newline +changes **which clause** a reader thinks an item is in, which forges a clause the rubric never +declared. Only the second breaks the guarantee above. An item reading +`"Not this option: x"` with no newline in it is the same class as `"; "`: legal, because +without a line break it cannot become a clause. + +### Which declarations are legal + +These rules hold in every SDK, and each is refused where the rubric is written rather than at +ask time. + +- An **empty** clause written out is refused. Leaving a clause off is how you say there is + none; an empty sequence is a clause the caller wrote that says nothing. +- An **empty** entry within a clause is refused. +- A **duplicate** entry within one clause is refused. +- One string as an example of **two alternatives of the same question** is refused: it says one + input belongs to both, which cannot be true. +- One string as **both an example and a counterexample of one alternative** is refused: it says + the input does and does not belong there. +- One string as an example of one alternative and a **counterexample of another must be + allowed**. This is the confusable-alternatives pattern the feature exists to serve, and an + implementation that refused it would break the main use case. +- Examples are **attached to a blank rubric** is refused; a blank rubric with no examples is + not. See below. +- A counterexample on a **level** is refused: an ordered scale has no "not this option". +- An item containing `U+000A` or `U+000D` is refused. `U+000A` is what the renderer joins + clauses with, so an item carrying one would read as a clause that was never declared; + `U+000D` is refused as hygiene, being pasted-text residue that breaks the `"; "`-joined line. + The rule applies to **items only**. A newline in `what` stays legal: a bare `what` has to + remain legal whatever it contains, so refusing it in the clause-bearing case would stop a + caller writing ordinary multi-line prose without closing any path. + + The test is the literal codepoint — `"\n" in item` in Python, `item.contains('\n')` in Rust. + Never `str.splitlines()`, never `str::lines()`, never an `is_control` predicate. Python's + `splitlines()` splits on eight codepoints, adding `U+000B`, `U+000C`, `U+001C`, `U+001D`, + `U+001E`, `U+0085`, `U+2028` and `U+2029`; Rust's `lines()` splits on `U+000A` alone. An + implementer reaching for the idiomatic call in either language writes a rule the other does + not have, and neither version looks wrong read on its own. `U+2028`, `U+2029` and `U+0085` + are deliberately **not** refused: they are `White_Space`, so an item made only of them is + already refused as empty, and embedded they cannot produce a clause boundary in the bytes an + SDK emits. What a model's tokenizer makes of them is unmeasured and stays out of the + contract, the same way the broader line-break set does. + +### What counts as blank, and what counts as a duplicate + +Two rules a third implementer would otherwise have to guess, and would guess differently. + +**A blank rubric is only an error when examples are attached to it.** A rubric that carries no +examples is never refused for its text, whatever that text is: it means what it meant before +this feature existed, and a patch release does not get to redefine it. Attaching examples to a +blank rubric is the error, because they describe something that is not there. Both SDKs draw +the line in the same place — Python inside `option()`, `level()` and `fallback()`, Rust in the +derive and in `Rubric::into_wire()` — so a declaration is legal in both or in neither. + +**"Blank" means Unicode `White_Space`.** Rust's `str::trim` is exactly that property. Python's +`str.strip()` is a superset: measured against the current runtime it strips 29 codepoints to +`White_Space`'s 25, and the four extra are `U+001C`, `U+001D`, `U+001E` and `U+001F` — file, +group, record and unit separator. Nothing goes the other way: every `White_Space` codepoint is +one Python strips. + +Those four are **C0** controls. The C0 block is `U+0000`–`U+001F`; C1 is `U+0080`–`U+009F` and +contains none of them. The distinction is worth stating because the point of writing this rule +down is that two SDKs must not describe the boundary differently, and naming the wrong block +would do exactly that. + +So a rubric made only of those four is blank to Python and not to Rust. The difference is +documented rather than hidden, and it is harmless: C0 controls are never valid rubric text, so +no real declaration reaches it. + +**A duplicate is an exact string match**, with no normalisation, no case folding and no +trimming. `"a"` and `" a"` are two different examples and may sit in the same clause. An +implementation that normalised before comparing would refuse declarations these SDKs accept, +which is the same divergence as a different label by another route. + +**The two checks disagree about `" a"`, and they are meant to.** The emptiness check trims +before deciding, so `" a"` is not blank and `" "` is; the duplicate check does not trim, so +`" a"` and `"a"` are two entries. That reads like a bug until you see what each one is asking. +Emptiness asks whether the caller wrote anything at all, and a leading space does not change +the answer. Duplication asks whether two entries would put the same bytes in front of the +model, and a leading space does change that, because the text is rendered verbatim. Trimming +for one and not the other is the only pairing that keeps both questions honest. + +One asymmetry is deliberate. `choose_among` and `score_levels` here take an `option(…)` or a +`level(…)` value; Rust's equivalents keep taking a plain string, because widening their +signatures risks inference breakage for existing callers on a path that can already pass a +string its own renderer composed. It is revisited at 0.2.0. Equivalent inputs put identical +bytes on the wire either way, which is what the contract actually promises. + Drift is caught rather than trusted. `mise run spec-check` clones guideme-rust, diffs its `spec/` against this one and fails on any difference except `spec/SOURCE`, which is provenance and has no counterpart upstream. It runs on every pull request as the `spec-drift` job of diff --git a/docs/design.md b/docs/design.md index 4385cf2..9fe13d9 100644 --- a/docs/design.md +++ b/docs/design.md @@ -97,6 +97,45 @@ Each module survives the test. an error span status rather than an error-level record. Anything without a convention is namespaced `guideme.`. The package depends on `opentelemetry-api` only and installs no provider; `docs/observability.md` shows the exporter side. +- **A rubric's examples are flattened into its text, not sent as structured criteria.** The + TypeSafe API takes structured `criteria`, and `docs.typesafe.ai/primitives/choice.md` + documents exactly the `what` / `not_for` / `examples` object this surface wants. It is not + used, for three measured reasons. A score answer echoes its criteria back in `legend`, which + is `dict[str, str]` here and `BTreeMap` in Rust; object criteria come back as + objects and fail to parse, so sending them means a breaking change to a public type — in a + field neither SDK reads beyond its length. Flattening is as good: on the docs' own worked + example, flattened scored 1.01 against structured's 1.03 at a higher confidence, and on an + ambiguous choice both reached the option that bare strings miss, inside run-to-run variance. + And flattening is cheaper: identical content billed 400 input tokens flattened against 450 + structured. The gain comes from the examples being present, not from the JSON shape. So + `option(…)`, `level(…)` and `fallback(…)` carry the parts on a `str` subclass and `render` + composes them where the rubric becomes wire text — the wire schema, `spec/`, and every + existing golden vector untouched. +- **The rendered rubric is a contract item, and the renderer has one entry per wire site.** + Clauses join with a newline and items with `"; "`, in the order written, and the text is used + verbatim: a newline rather than a space is what removes the need for a punctuation rule, + since an example ending in `?` would otherwise render as `Where is my refund?.`. The + `Not this option` label was measured against `Not` and `Counterexamples` and won. + `Choice.rubric()`, `Levels.levels()`, `choose_among`, `score_levels` and + `NoulQuestion.criteria` all render, so a value that reached a runtime constructor cannot + silently lose its examples. Rust renders the same string, which leaves one deliberate + asymmetry: `choose_among` here takes an `option(…)`, while Rust's equivalent keeps taking a + plain string rather than risk inference breakage for existing callers. Equivalent inputs put + identical bytes on the wire. +- **A noul's criteria take examples too, and through the same `option(…)`.** A yes and a no are + two described alternatives of one question, exactly as confusable as two options of a choice, + and measurement says so: asked whether a nightly export job with a manual workaround is + urgent, plain `Urgent` / `Not urgent` answers yes at 0.75 four times running, and the same + criteria carrying examples answer no at 0.17. The examples move it to the correct answer, + because one of the no examples is the situation the state describes. So there is no fourth + constructor — `option(…)` is "a described alternative" and serves both — and `.criteria(…)` + renders like every other wire site. +- **Contradictory examples are refused where they are written.** One string offered as an + example of two alternatives of the same question says an input belongs to both, which cannot + be true; one string offered as both an example and a counterexample of the same alternative + says it does and does not belong. Both are a `ConfigError`. The overlap that looks similar + and is the whole point stays legal: the same string as an example of one alternative and a + counterexample of another is how two confusable ones are told apart. - **An answer is a span event and a log record, and the caller picks.** Rust emits one `tracing` event and lets the subscriber fan it out, so its example filters events off the span exporter to store each one once. There is no subscriber here, so the library makes both @@ -123,6 +162,26 @@ Each module survives the test. statement, naming the members that repeat. - **`Key` and `Rank`** are only meaningful through `choose_among` and `score_levels`. They are `NewType`s over `str` and `int`, so nothing else hands you one. +- **A rubric's value is its bare text, examples or not.** `option("x", examples=[…])` still + equals `"x"`, so two options whose text matches are still one member however their examples + differ, and a member's `.value` still reads as it was written. The expansion happens only in + the request. +- **A blank rubric is an error only when examples are attached to it.** `option(" ")` on its + own is accepted and means exactly what a bare `""` member has always meant; + `option(" ", examples=[…])` is a `ConfigError`. The first version of this refused any blank + rubric written through the new constructors, which read as tidy and was wrong: it made a + declaration that was legal in 0.1.0 illegal in a patch release, on a degenerate input that + was already meaningless. "You attached examples to nothing" is the real mistake; "your + description is blank" is not one this release gets to invent. `guideme-rust` narrowed the + same rule from the same starting point, so a reader of both finds one rule: strict where + examples are, untouched where they are not. Revisit at a major bump, together. +- **A clause left out is `None`; an empty one was written on purpose.** `examples` and + `counterexamples` default to `None`, which is how "there are none" is said, so any empty + sequence that arrives was typed by the caller and says nothing — a `ConfigError`, whatever + its type, `[]` and `()` alike. The alternative, defaulting to `()` and telling the two apart + by identity, would have rested on CPython interning the empty tuple: correct today, and + silently wrong on a runtime that does not, with a real caller mistake quietly no longer + caught. `is None` needs no such assumption. - **The four scalars are brands, not validated types.** `Probability`, `Confidence`, `Key` and `Rank` are `NewType`s, so `Probability(2.0)` and `Rank(99)` are accepted by the checker and by the interpreter alike. What makes them trustworthy is that only the wire mints them, and it diff --git a/docs/observability.md b/docs/observability.md index 065015c..44ac878 100644 --- a/docs/observability.md +++ b/docs/observability.md @@ -328,7 +328,12 @@ a different backend. Captured from `otel/opentelemetry-collector-contrib` with the debug exporter and the `collector.yaml` above, running `examples/otlp` under `events("log")` against the live API. -Timestamps, `Flags` and the resource block are trimmed throughout. The first two blocks are +Timestamps, `Flags` and the resource block are trimmed throughout. The capture was taken at +**0.1.0** and the `InstrumentationScope` lines carry that version, which is the one the run +actually emitted; it is left as it was rather than edited forward, so nothing here is a number +no run produced. Read a version older than the current release as the age of the capture, not +as a claim about today: the scope names are the contract, and their version is not. The first +two blocks are otherwise complete; the later ones are excerpts, cut to the lines each is making a point about, so a missing `Parent ID`, `Kind`, scope line or attribute there means it was cut, not that it was absent. Nothing is reworded, and no value is invented. diff --git a/examples/otlp/uv.lock b/examples/otlp/uv.lock index 5820788..88d00b9 100644 --- a/examples/otlp/uv.lock +++ b/examples/otlp/uv.lock @@ -103,7 +103,7 @@ wheels = [ [[package]] name = "guideme" -version = "0.1.0" +version = "0.1.1" source = { editable = "../../" } dependencies = [ { name = "httpx" }, diff --git a/pyproject.toml b/pyproject.toml index 1cb67d9..0fc7092 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "hatchling.build" [project] name = "guideme" -version = "0.1.0" +version = "0.1.1" description = "Type-safe inline judgments from TypeSafe Jev: a yes/no is an if, a choice is an exhaustive match, a score is a comparison." readme = "README.md" requires-python = ">=3.12" diff --git a/spec/SOURCE b/spec/SOURCE index 6f6df11..f6b8d6e 100644 --- a/spec/SOURCE +++ b/spec/SOURCE @@ -1 +1 @@ -28e3103a874a79922fbf275ec83fbe12d9de07ab +d908ecfec39f267802781b10be74323286e661af diff --git a/spec/vectors/rubric.json b/spec/vectors/rubric.json new file mode 100644 index 0000000..67f4823 --- /dev/null +++ b/spec/vectors/rubric.json @@ -0,0 +1,96 @@ +[ + { + "kind": "choice", + "what": "Payments, invoicing, refunds", + "examples": [], + "counterexamples": [], + "rendered": "Payments, invoicing, refunds" + }, + { + "kind": "choice", + "what": "Bugs, outages, integrations", + "examples": [ + "502 on every request" + ], + "counterexamples": [], + "rendered": "Bugs, outages, integrations\nExamples: 502 on every request" + }, + { + "kind": "choice", + "what": "Payments, invoicing, refunds", + "examples": [ + "My card was charged twice", + "Where is my refund?" + ], + "counterexamples": [], + "rendered": "Payments, invoicing, refunds\nExamples: My card was charged twice; Where is my refund?" + }, + { + "kind": "choice", + "what": "Payments, invoicing, refunds", + "examples": [ + "My card was charged twice" + ], + "counterexamples": [ + "The dashboard is down" + ], + "rendered": "Payments, invoicing, refunds\nExamples: My card was charged twice\nNot this option: The dashboard is down" + }, + { + "kind": "choice", + "what": "Payments, invoicing, refunds", + "examples": [], + "counterexamples": [ + "The dashboard is down" + ], + "rendered": "Payments, invoicing, refunds\nNot this option: The dashboard is down" + }, + { + "kind": "choice", + "what": "Payments, invoicing, refunds", + "examples": [ + "Why was I billed twice?", + "Cancel and refund me" + ], + "counterexamples": [ + "The status page is red", + "Are you hiring?" + ], + "rendered": "Payments, invoicing, refunds\nExamples: Why was I billed twice?; Cancel and refund me\nNot this option: The status page is red; Are you hiring?" + }, + { + "kind": "levels", + "what": "No impact to functionality", + "examples": [ + "typo in a label", + "misaligned icon" + ], + "counterexamples": [], + "rendered": "No impact to functionality\nExamples: typo in a label; misaligned icon" + }, + { + "kind": "levels", + "what": "No workaround exists", + "examples": [], + "counterexamples": [], + "rendered": "No workaround exists" + }, + { + "kind": "noul", + "what": "Something is broken now and nobody can work around it", + "examples": [ + "the checkout page is down" + ], + "counterexamples": [ + "a nightly job failed and the numbers are pulled by hand for now" + ], + "rendered": "Something is broken now and nobody can work around it\nExamples: the checkout page is down\nNot this option: a nightly job failed and the numbers are pulled by hand for now" + }, + { + "kind": "noul", + "what": "It can wait for the next working day", + "examples": [], + "counterexamples": [], + "rendered": "It can wait for the next working day" + } +] diff --git a/src/guideme/__init__.py b/src/guideme/__init__.py index 17cae35..7c34762 100644 --- a/src/guideme/__init__.py +++ b/src/guideme/__init__.py @@ -1,6 +1,6 @@ """Type-safe inline judgments from TypeSafe Jev.""" -from guideme.enums import Choice, Levels, fallback +from guideme.enums import Choice, Levels, fallback, level, option from guideme.errors import ( AuthError, ConfigError, @@ -57,7 +57,9 @@ "choose", "choose_among", "fallback", + "level", "noul", + "option", "score", "score_levels", ] diff --git a/src/guideme/enums.py b/src/guideme/enums.py index 9d86269..900e50e 100644 --- a/src/guideme/enums.py +++ b/src/guideme/enums.py @@ -1,11 +1,15 @@ """The two enum bases a caller declares their options and levels with. A member's name is its wire key and its value is its rubric, the sentence the -model reads. Docstrings on members are not the rubric; the value is. Both bases -validate at class-definition time, so a rubric that could not be asked is an -error where it is written rather than on the first request. +model reads. Docstrings on members are not the rubric; the value is. A rubric is +a bare string, or an `option(...)`, `level(...)` or `fallback(...)` that carries +examples beside the text and has them composed into it at the one point the +rubric becomes wire text. Both bases validate at class-definition time, so a +rubric that could not be asked is an error where it is written rather than on +the first request. """ +from collections.abc import Iterable, Sequence from enum import Enum from typing import Self, final @@ -15,26 +19,262 @@ MIN_OPTIONS = 1 """Fewest options a choice may carry.""" +EXAMPLES = "Examples: " +"""Opens the clause naming inputs that belong to this option or level.""" + +NOT_THIS = "Not this option: " +"""Opens the clause naming inputs that belong to some other option.""" + @final -class _Fallback(str): - """A rubric marked as the member to fall back to when the policy says unsure. +class _Rubric(str): # noqa: SLOT000 -- a str subclass may not carry non-empty __slots__ + """A rubric held as its parts: the text, its examples, its counterexamples. + + A `str` subclass whose `str` value is the bare text, so a member's value reads + exactly like a bare string's and nothing downstream meets a new type. The parts + ride along as attributes until `render` composes them. `str` is a variable-length + built-in, so its subclasses cannot have slots and the parts live in the instance + dictionary; they are set in `__new__` because a `str` is immutable and carries its + text from there. + """ - A `str` subclass, so the member's value is its rubric exactly like every - other member's and the marking lives in the type rather than in a second - attribute the caller could see. + examples: tuple[str, ...] = () + counterexamples: tuple[str, ...] = () + is_fallback: bool = False + + def __new__( + cls, + rubric: str, + examples: tuple[str, ...], + counterexamples: tuple[str, ...], + *, + is_fallback: bool, + ) -> Self: + """Carry the parts on the string without changing what the string says.""" + self = super().__new__(cls, rubric) + self.examples = examples + self.counterexamples = counterexamples + self.is_fallback = is_fallback + return self + + def __getnewargs_ex__( + self, + ) -> tuple[tuple[str, tuple[str, ...], tuple[str, ...]], dict[str, bool]]: + """Hand `copy` and `pickle` every argument `__new__` needs, marking included. + + `str`'s own `__getnewargs__` supplies the text alone, which is one argument + where this takes four, so without this a copied or pickled rubric raised a + `TypeError`. A bare `str` rubric copied fine in 0.1.0 and has to keep doing so: + `copy.deepcopy({"key": option(...)})` is exactly the shape a caller builds + before handing it to `choose_among`. + """ + return (str(self), self.examples, self.counterexamples), {"is_fallback": self.is_fallback} + + +def _items(what: str, items: Sequence[str] | None, named: str) -> tuple[str, ...]: + """Check one clause's items: written means non-empty, each says something, no repeats. + + `None` is the clause left out. Anything else that is empty was written on purpose + and says nothing, which is the caller's mistake rather than a default. + + A `str` is a `Sequence[str]` of its own characters, so `examples="refund"` would + quietly become six one-letter examples. It is a `ConfigError` instead, the same + way `score_levels` refuses a scale given as one string. + """ + if items is None: + return () + if isinstance(items, str | bytes): + detail = ( + f"{what!r}: {named} must be a sequence of strings, got a single " + f"{type(items).__name__}; wrap it in a list" + ) + raise ConfigError(detail) + if not items: + detail = f"{what!r}: empty {named}=; a clause written on purpose must say something" + raise ConfigError(detail) + values = tuple(items) + blank = [item for item in values if not item.strip()] + if blank: + detail = f"{what!r}: empty {named[:-1]} {blank[0]!r}; every entry must say something" + raise ConfigError(detail) + # Blank wins over this, which strip-first already gives: an item of one newline is + # empty, not broken. The test is the literal codepoint, never `splitlines()`, which + # here would also split on CR, VT, FF, FS, NEL, U+2028 and U+2029 and refuse seven + # declarations the Rust SDK accepts. + broken = [item for item in values if "\n" in item or "\r" in item] + if broken: + detail = ( + f"{what!r}: {named[:-1]} {broken[0]!r} may not contain a line break " + f"(U+000A or U+000D); items are joined with '; ' onto one line, and a newline " + f"in one would read as a clause the rubric never declared" + ) + raise ConfigError(detail) + seen: set[str] = set() + for item in values: + if item in seen: + detail = f"{what!r}: duplicate {named[:-1]} {item!r}; each entry must be distinct" + raise ConfigError(detail) + seen.add(item) + return values + + +def _rubric( + rubric: str, + examples: Sequence[str] | None, + counterexamples: Sequence[str] | None, + *, + is_fallback: bool, +) -> str: + """Check the parts and hold them on the rubric, unrendered. + + A blank rubric is refused only when examples are attached to it. Attaching them + to nothing is the mistake; a blank rubric on its own was legal in 0.1.0, means + exactly what a bare `""` member means, and a patch release does not get to make + it an error. + """ + shown = _items(rubric, examples, "examples") + excluded = _items(rubric, counterexamples, "counterexamples") + if (shown or excluded) and not rubric.strip(): + detail = ( + f"examples were attached to a blank rubric ({rubric!r}); they say what the " + f"rubric covers, so there has to be something for them to say it about" + ) + raise ConfigError(detail) + ruled_out = set(excluded) + both = [item for item in shown if item in ruled_out] + if both: + detail = ( + f"{rubric!r}: {both[0]!r} is both an example and a counterexample of it, which " + f"says the input does and does not belong here; it may be an example of one " + f"option and a counterexample of another, but not of the same one" + ) + raise ConfigError(detail) + return _Rubric(rubric, shown, excluded, is_fallback=is_fallback) + + +def option( + rubric: str, + *, + examples: Sequence[str] | None = None, + counterexamples: Sequence[str] | None = None, +) -> str: + """A described alternative: what it covers, inputs that belong to it, inputs that do not. + + Use it for a `Choice` member, for a runtime option of `choose_among`, and for either + side of a noul's `.criteria(...)`; a yes and a no are as confusable as two options. + + The value is still the rubric text; the examples are composed into it where the + rubric goes on the wire, so an option written as a bare string and one written as + `option("…")` send the same bytes, and both clauses keep the order they were + written in. A string may be an example of one alternative and a counterexample of + another: that is how two confusable ones are told apart. + + Leave a clause out to say there is none. An empty one written out says nothing and + is refused: an `examples=[]`, a clause given as one string rather than a sequence of + them, a blank entry, a repeat within one clause, a string given as both an example + and a counterexample of this one alternative, and examples attached to a blank + rubric are each a `ConfigError` where the option is written. A blank rubric with no + examples is not: that is what it has always meant. """ + return _rubric(rubric, examples, counterexamples, is_fallback=False) + - __slots__ = () +def level(rubric: str, *, examples: Sequence[str] | None = None) -> str: + """A score level: what it means, and inputs that score here. + + An example listed under a level is the statement that such an input scores that + level; its place in the ordered scale is what says which score. There are no + counterexamples: "not this option" means nothing on an ordered scale, so the + level below or above is what an input that does not belong here scores. + """ + return _rubric(rubric, examples, None, is_fallback=False) -def fallback(rubric: str) -> str: +def fallback( + rubric: str, + *, + examples: Sequence[str] | None = None, + counterexamples: Sequence[str] | None = None, +) -> str: """Mark the member to use when the policy says unsure. At most one per `Choice`. - A `.otherwise(...)` on the question beats it; with neither, an unsure answer - raises `UnsureError`. + An `option(...)` in every other respect, examples included. A `.otherwise(...)` on + the question beats it; with neither, an unsure answer raises `UnsureError`. """ - return _Fallback(rubric) + return _rubric(rubric, examples, counterexamples, is_fallback=True) + + +def render(rubric: str) -> str: + """Compose a rubric's parts into the text the model reads. + + Clauses are joined with a newline and the items of one with `"; "`, and the text + is used verbatim: never trimmed, never re-punctuated, so an example ending in `?` + does not become `Where is my refund?.`. A bare string, and a rubric carrying no + parts, come back byte for byte, which is what leaves an existing caller's request + unchanged. The rendered string is a cross-SDK contract item; `guideme-rust` + composes the same one in its derive macro. + """ + if not isinstance(rubric, _Rubric): + return rubric + if not rubric.examples and not rubric.counterexamples: + return str(rubric) + lines = [str(rubric)] + if rubric.examples: + lines.append(EXAMPLES + "; ".join(rubric.examples)) + if rubric.counterexamples: + lines.append(NOT_THIS + "; ".join(rubric.counterexamples)) + return "\n".join(lines) + + +def require_unshared_examples(where: str, rubrics: Iterable[tuple[str, str | None]]) -> None: + """Refuse one string offered as an example of two alternatives of the same question. + + It would say the input belongs to both, which cannot be true. The reverse is legal + and is the point of the feature: the same string as an example of one alternative + and a counterexample of another is how two confusable ones are told apart. + """ + seen: dict[str, str] = {} + for name, rubric in rubrics: + if not isinstance(rubric, _Rubric): + continue + for example in rubric.examples: + first = seen.setdefault(example, name) + if first != name: + detail = ( + f"{where}: {example!r} is an example of both {first} and {name}, so it " + f"says one input belongs to two alternatives; make it a counterexample " + f"of one of them instead" + ) + raise ConfigError(detail) + + +def require_no_fallback(where: str, rubric: str | None) -> None: + """Refuse a `fallback(...)` where there is no `Choice` member for it to mark. + + The runtime constructors answer in a `Key`, a `Rank` or a `bool`, none of which has + a member to fall back to, so the marking has nothing to act on. Dropping it would + leave a caller believing an unsure answer is handled when it raises instead. + """ + if isinstance(rubric, _Rubric) and rubric.is_fallback: + detail = ( + f"{where}: fallback(...) marks a member of a Choice and there is none here; " + f"use .otherwise(value) on the question" + ) + raise ConfigError(detail) + + +def require_no_counterexamples(where: str, rubric: str) -> None: + """Refuse counterexamples on a level, rather than rendering or dropping them. + + `level(...)` does not offer them, so this is an `option(...)` written where a level + belongs. Silently losing what it carries is the failure this raises instead. + """ + if isinstance(rubric, _Rubric) and rubric.counterexamples: + detail = ( + f"{where}: counterexamples are not allowed on a level, which is a position on " + f"an ordered scale; use level(...), which takes examples only" + ) + raise ConfigError(detail) def _require_distinct(cls: type[Enum]) -> None: @@ -81,21 +321,30 @@ def __init_subclass__(cls) -> None: ) raise ConfigError(detail) _require_text(cls, "choice option") - marked = [member.name for member in cls if isinstance(member.value, _Fallback)] + marked = [ + member.name + for member in cls + if isinstance(member.value, _Rubric) and member.value.is_fallback + ] if len(marked) > 1: detail = f"{cls.__name__}: only one member may be marked fallback, got {marked}" raise ConfigError(detail) + require_unshared_examples(cls.__name__, ((m.name, m.value) for m in cls)) @classmethod def rubric(cls) -> tuple[tuple[str, str], ...]: - """`(key, rubric)` pairs in declaration order.""" - return tuple((member.name, member.value) for member in cls) + """`(key, rendered rubric)` pairs in declaration order. + + This is where an `option(...)`'s examples are composed into its text, so a + member cannot reach the wire having quietly lost them. + """ + return tuple((member.name, render(member.value)) for member in cls) @classmethod def fallback_member(cls) -> Self | None: """The `fallback(...)` member, if the rubric marks one.""" for member in cls: - if isinstance(member.value, _Fallback): + if isinstance(member.value, _Rubric) and member.value.is_fallback: return member return None @@ -131,12 +380,14 @@ def __init_subclass__(cls) -> None: raise ConfigError(detail) _require_text(cls, "level") for member in cls: - if isinstance(member.value, _Fallback): + if isinstance(member.value, _Rubric) and member.value.is_fallback: detail = ( f"{cls.__name__}.{member.name}: fallback(...) is not allowed on Levels; " f"use .otherwise(level) on the question" ) raise ConfigError(detail) + require_no_counterexamples(f"{cls.__name__}.{member.name}", member.value) + require_unshared_examples(cls.__name__, ((m.name, m.value) for m in cls)) @property def index(self) -> int: @@ -145,8 +396,11 @@ def index(self) -> int: @classmethod def levels(cls) -> tuple[str, ...]: - """Level descriptions, low to high.""" - return tuple(member.value for member in cls) + """Rendered level descriptions, low to high. + + This is where a `level(...)`'s examples are composed into its text. + """ + return tuple(render(member.value) for member in cls) @classmethod def from_index(cls, index: int) -> Self | None: diff --git a/src/guideme/question.py b/src/guideme/question.py index 1c9e8f5..dc42be0 100644 --- a/src/guideme/question.py +++ b/src/guideme/question.py @@ -10,7 +10,14 @@ from typing import Self, final from guideme._json import Json -from guideme.enums import Choice, Levels +from guideme.enums import ( + Choice, + Levels, + render, + require_no_counterexamples, + require_no_fallback, + require_unshared_examples, +) from guideme.errors import ConfigError, ProtocolError, UnsureError from guideme.policy import ( MAX_LEVELS, @@ -325,8 +332,16 @@ class NoulQuestion(_Binary[bool], _Fallible[bool]): """A yes/no question read as a `bool`.""" def criteria(self, yes: str, no: str) -> Self: - """Describe what a yes and a no mean.""" - return replace(self, spec=NoulSpec(NoulCriteria(yes, no))) + """Describe what a yes and a no mean. + + Either may be an `option(...)` carrying examples, which are composed into it + here: a yes and a no are as confusable as two options of a choice, and showing + an input that belongs to each is what tells them apart. + """ + require_no_fallback("criteria yes", yes) + require_no_fallback("criteria no", no) + require_unshared_examples("criteria", (("yes", yes), ("no", no))) + return replace(self, spec=NoulSpec(NoulCriteria(render(yes), render(no)))) def detail(self) -> DetailedNoul: """Read the full `Verdict` instead. Any `.otherwise(...)` is dropped.""" @@ -432,10 +447,18 @@ def choose_among(instructions: Json, options: Mapping[str, str | None]) -> Choic outside it is a `ConfigError` raised here, where the options are written, never later at `ask`. Keys cannot collide: a `Mapping` has already made them unique. + + A rubric is a string, `None`, or an `option(...)` carrying examples, which + are composed into it here: the same text a `Choice` member would send. A + `fallback(...)` is a `ConfigError`: this answers in a `Key`, so there is no + member for the marking to name; use `.otherwise(...)` on the question. """ + for key, text in options.items(): + require_no_fallback(f"choose_among option {key!r}", text) + require_unshared_examples("choose_among", options.items()) return _choice( instructions, - tuple(options.items()), + tuple((key, None if text is None else render(text)) for key, text in options.items()), Key, lambda: None, ) @@ -468,6 +491,9 @@ def score_levels(instructions: Json, levels: Sequence[str]) -> ScoreQuestion[Ran Takes 2 to 10 levels, the same range a `Levels` enum takes; outside it is a `ConfigError`. + A level is a string or a `level(...)` carrying examples, which are composed + into it here: the same text a `Levels` member would send. + A `str` is a `Sequence[str]` of its own characters, so `score_levels("…", "abc")` would quietly ask about a three-letter scale. It is a `ConfigError` instead. """ @@ -476,4 +502,10 @@ def score_levels(instructions: Json, levels: Sequence[str]) -> ScoreQuestion[Ran f"levels must be a sequence of level descriptions, got a single {type(levels).__name__}" ) raise ConfigError(detail) - return _score(instructions, tuple(levels), Rank) + for index, text in enumerate(levels): + require_no_counterexamples(f"level {index}", text) + require_no_fallback(f"level {index}", text) + require_unshared_examples( + "score_levels", ((f"level {index}", text) for index, text in enumerate(levels)) + ) + return _score(instructions, tuple(render(text) for text in levels), Rank) diff --git a/tests/test_enums.py b/tests/test_enums.py index 4565dfa..c93fa17 100644 --- a/tests/test_enums.py +++ b/tests/test_enums.py @@ -1,3 +1,4 @@ +import copy from collections.abc import Callable from typing import cast @@ -6,7 +7,7 @@ from hypothesis import strategies as st from guideme import ConfigError -from guideme.enums import Choice, Levels, fallback +from guideme.enums import Choice, Levels, fallback, level, option def _compare(low: Levels, high: Levels) -> bool: @@ -71,6 +72,105 @@ class OneLevel(Levels): return OneLevel +def _examples_attached_to_a_blank_rubric() -> type[Choice]: + # The blank rubric alone is legal and means what it always meant. Attaching + # examples to it is the mistake: they describe something that is not there. + class Blank(Choice): + a = option(" ", examples=["My card was charged twice"]) + + return Blank + + +def _an_examples_clause_written_empty() -> type[Choice]: + class Empty(Choice): + a = option("Payments", examples=[]) + + return Empty + + +def _examples_given_as_one_string() -> type[Choice]: + # A `str` is a `Sequence[str]`, so the checker allows it and the clause would be + # the six letters of "refund". + class Shredded(Choice): + a = option("Payments", examples="refund") + + return Shredded + + +def _an_empty_tuple_written_out() -> type[Choice]: + # `None` is how a clause is left out, so an empty sequence is always a clause + # written on purpose that says nothing -- whatever type the caller reached for. + class Empty(Choice): + a = option("Payments", counterexamples=()) + + return Empty + + +def _a_blank_example() -> type[Choice]: + class Blank(Choice): + a = option("Payments", examples=["My card was charged twice", " "]) + + return Blank + + +def _a_newline_in_an_example() -> type[Choice]: + # Items are joined onto one line, so a newline in one reads as a clause the rubric + # never declared -- here, a counterexample clause that was never written. + class Forged(Choice): + a = option("Payments", examples=["x\nNot this option: anything at all"]) + + return Forged + + +def _a_carriage_return_in_a_counterexample() -> type[Choice]: + class Stray(Choice): + a = option("Payments", counterexamples=["the dashboard is down\r"]) + + return Stray + + +def _a_repeated_example() -> type[Choice]: + class Repeated(Choice): + a = option("Payments", examples=["Where is my refund?", "Where is my refund?"]) + + return Repeated + + +def _one_example_of_two_options() -> type[Choice]: + class Shared(Choice): + billing = option("Payments", examples=["Where is my refund?"]) + technical = option("Bugs", examples=["Where is my refund?"]) + + return Shared + + +def _one_example_of_two_levels() -> type[Levels]: + class Shared(Levels): + cosmetic = level("No impact", examples=["a typo in a label"]) + blocking = level("No workaround exists", examples=["a typo in a label"]) + + return Shared + + +def _an_example_that_is_also_a_counterexample() -> type[Choice]: + class Both(Choice): + billing = option( + "Payments", + examples=["Where is my refund?"], + counterexamples=["Where is my refund?"], + ) + + return Both + + +def _a_counterexample_on_a_level() -> type[Levels]: + class Marked(Levels): + cosmetic = option("No impact", counterexamples=["cannot log in"]) + blocking = level("No workaround exists") + + return Marked + + def _eleven_levels() -> type[Levels]: class ElevenLevels(Levels): l0 = "a" @@ -104,6 +204,19 @@ class Department(Choice): assert Department.from_key("technical") is Department.technical assert Department.from_key("marketing") is None + # A rubric survives being copied, parts and fallback marking alike. `copy.deepcopy` + # of a dict of options is what a caller builds before `choose_among`, and a bare + # string copied fine before this feature existed. + class Copied(Choice): + billing = copy.deepcopy(option("Payments", examples=["My card was charged twice"])) + sales = copy.deepcopy(fallback("Pricing")) + + assert Copied.rubric() == ( + ("billing", "Payments\nExamples: My card was charged twice"), + ("sales", "Pricing"), + ) + assert Copied.fallback_member() is Copied.sales + def test_levels_are_totally_ordered_by_declaration() -> None: class Frustration(Levels): @@ -136,6 +249,18 @@ class Urgency(Levels): _duplicate_choice_rubric, _duplicate_level_rubric, _fallback_on_a_level, + _examples_attached_to_a_blank_rubric, + _an_examples_clause_written_empty, + _examples_given_as_one_string, + _an_empty_tuple_written_out, + _a_blank_example, + _a_newline_in_an_example, + _a_carriage_return_in_a_counterexample, + _a_repeated_example, + _a_counterexample_on_a_level, + _one_example_of_two_options, + _one_example_of_two_levels, + _an_example_that_is_also_a_counterexample, ] REFUSED_IDS = [ @@ -147,6 +272,18 @@ class Urgency(Levels): "two_options_with_the_same_rubric", "two_levels_with_the_same_rubric", "a_fallback_marker_on_a_levels", + "examples_attached_to_a_blank_rubric", + "an_examples_clause_written_empty", + "examples_given_as_one_string", + "an_empty_tuple_written_out", + "a_blank_example", + "a_newline_in_an_example", + "a_carriage_return_in_a_counterexample", + "a_repeated_example", + "a_counterexample_on_a_level", + "one_example_of_two_options", + "one_example_of_two_levels", + "an_example_that_is_also_a_counterexample", ] diff --git a/tests/test_live.py b/tests/test_live.py index 01290f5..a6427a4 100644 --- a/tests/test_live.py +++ b/tests/test_live.py @@ -7,11 +7,14 @@ AsyncGuide, Choice, Guide, + Key, Levels, Scored, choose, + choose_among, fallback, noul, + option, score, ) from guideme.question import Question @@ -110,3 +113,77 @@ async def run() -> tuple[bool, Department, Scored[Frustration], dict[str, bool]] urgent, dept, mood, flags = asyncio.run(run()) check(urgent, dept, mood, flags) assert billed(spans) > 0 + + +AMBIGUOUS = "About those shoes - what is the situation with the money side of things?" +"""A ticket two options both half fit, where the examples are what tells them apart.""" + +POLICY = "The shop's rules" +STATUS = "One customer's open case" + +POLICY_EXAMPLES = ["How many days do I have to send it back?", "Can I return a sale item?"] +STATUS_EXAMPLES = ["Where is my refund?", "I posted the shoes back last week and heard nothing"] + +BARE = {"return_policy": POLICY, "return_status": STATUS} +"""The rubrics as bare strings. Deliberately vague: the ticket is near a coin flip.""" + +DESCRIBED = { + "return_policy": option(POLICY, examples=POLICY_EXAMPLES, counterexamples=[STATUS_EXAMPLES[1]]), + "return_status": option(STATUS, examples=STATUS_EXAMPLES, counterexamples=[POLICY_EXAMPLES[1]]), +} +"""The same two rubrics, with the examples that tell the two apart.""" + + +WORKAROUND = ( + "Our nightly export job has been failing since Tuesday. We pull the numbers by hand for now." +) +"""A ticket a vague `Urgent` / `Not urgent` calls urgent, and the examples call otherwise.""" + +URGENT = "Urgent" +NOT_URGENT = "Not urgent" + +URGENT_EXAMPLES = ["customers cannot log in", "money is moving to the wrong place"] +NOT_URGENT_EXAMPLES = ["a broken job with a manual workaround", "a cosmetic bug"] + + +def _return_status(guide: Guide, options: dict[str, str]) -> float: + ranked = guide.ask( + choose_among("What is the customer asking about?", options).detail(), AMBIGUOUS + ) + return dict(ranked.probabilities)[Key("return_status")] + + +def _urgent(guide: Guide, yes: str, no: str) -> float: + verdict = guide.ask(noul("Is this ticket urgent?").criteria(yes, no).detail(), WORKAROUND) + return verdict.p + + +def test_examples_move_the_distribution_towards_the_alternative_they_describe() -> None: + """The invariant the feature exists for, not a number the model is not stable to. + + Both halves hold the rubric text constant across their two asks, so the examples + are the only thing that changed. Measured on 2026-09-21 against `jev-1.13.0`. + The choice half, three runs: 0.50 / 0.49 / 0.53 bare, 0.87 / 0.89 / 0.89 described. + The noul half, four runs: 0.75 / 0.75 / 0.74 / 0.76 plain, 0.17 every time with + examples. The noul half is the one where the plain rubric is outright wrong: one + of the not-urgent examples is what this ticket describes. + + The cross-SDK design measured the same noul case at 0.25 rather than 0.17. The + difference is that this ask gives each side counterexamples as well as examples, + which the design's probe did not; it is a stronger rubric, not a disagreement + between the two SDKs. + """ + guide = Guide.from_env() + try: + bare = _return_status(guide, BARE) + described = _return_status(guide, DESCRIBED) + plain = _urgent(guide, URGENT, NOT_URGENT) + told = _urgent( + guide, + option(URGENT, examples=URGENT_EXAMPLES, counterexamples=NOT_URGENT_EXAMPLES), + option(NOT_URGENT, examples=NOT_URGENT_EXAMPLES, counterexamples=URGENT_EXAMPLES), + ) + finally: + guide.close() + assert described > bare + assert told < plain diff --git a/tests/test_rubrics.py b/tests/test_rubrics.py new file mode 100644 index 0000000..8653f8a --- /dev/null +++ b/tests/test_rubrics.py @@ -0,0 +1,203 @@ +import copy +import pickle +from pathlib import Path + +import pytest +from hypothesis import given +from hypothesis import strategies as st +from pytest_httpserver import HTTPServer + +from guideme import Choice, Levels, choose, fallback, level, noul, option, score +from guideme.enums import render + +from .conftest import ( + JSON, + REPO_ROOT, + TICKET, + Json, + Runner, + as_list, + as_object, + as_str, + expect_post, + load_json, + narrow, + reply, + validator, +) + +check_request = validator("request") + +BILLING = "Payments, invoicing, refunds" +TECHNICAL = "Bugs, outages, integrations" +COSMETIC = "No impact to functionality" + +VECTOR_PATH = REPO_ROOT / "spec" / "vectors" / "rubric.json" + + +def _declared(raw: Json) -> tuple[str, str, str]: + """One vector case as the rubric it declares, its bare text, and its expected bytes. + + `kind` is what picks the constructor, which is what that field is for: a `levels` + case has to go through `level(...)`, and a `choice` or `noul` case through + `option(...)`, so the vector exercises the constructors a caller writes rather than + `render` alone. + + An absent clause arrives as `[]` because that is how the generator serialises an + empty `Vec`, and it is mapped back to "no argument". Passing `[]` through would be + a different declaration: these constructors refuse an empty clause written out. + """ + case = as_object(raw) + what = as_str(case["what"]) + examples = [as_str(item) for item in as_list(case["examples"])] or None + counterexamples = [as_str(item) for item in as_list(case["counterexamples"])] or None + kind = as_str(case["kind"]) + if kind == "levels": + declared = level(what, examples=examples) + elif kind in {"choice", "noul"}: + declared = option(what, examples=examples, counterexamples=counterexamples) + else: + message = f"{VECTOR_PATH}: unknown kind {kind!r}" + raise AssertionError(message) + return declared, what, as_str(case["rendered"]) + + +def _cases(path: Path) -> list[Json]: + """The vector's cases, refusing an empty file. + + An empty parametrisation is a test that passes without ever running, so a vector + that arrived empty would silence the drift this file exists to catch. + """ + cases = as_list(load_json(path)) + if not cases: + message = f"{path} carries no cases" + raise AssertionError(message) + return cases + + +CASES = _cases(VECTOR_PATH) +VECTOR = [_declared(raw) for raw in CASES] +VECTOR_IDS = [f"{index}_{as_str(as_object(raw)['kind'])}" for index, raw in enumerate(CASES)] + + +@pytest.mark.parametrize(("rubric", "bare", "expected"), VECTOR, ids=VECTOR_IDS) +def test_a_rubric_renders_the_bytes_the_contract_names( + rubric: str, bare: str, expected: str +) -> None: + assert render(rubric) == expected + # The value itself stays the bare text: a rubric only expands where it becomes + # wire text, so a member's value reads exactly as it was written. + assert str(rubric) == bare + # And it survives the three ways a value gets duplicated. `__new__` takes four + # arguments where `str` hands back one, so each of these raised a `TypeError` + # until the rubric said how to rebuild itself. + for clone in (copy.copy(rubric), copy.deepcopy(rubric), pickle.loads(pickle.dumps(rubric))): # noqa: S301 -- the payload is this test's own value + assert render(clone) == expected + + +@given(st.text()) +def test_a_rubric_with_no_parts_renders_byte_for_byte(what: str) -> None: + # The load-bearing invariant: 0.1.0's bytes do not move. A bare string and an + # `option(...)` with nothing attached both render to the text itself. + # + # The strategy is unrestricted on purpose, blank and whitespace-only text + # included. A rubric carrying no examples is refused for nothing at all: it + # meant whatever it meant in 0.1.0 and a patch release does not redefine it. + # Attaching examples to a blank rubric is the error, and that case is in + # `tests/test_enums.py`'s refusal list. + assert render(what) == what + assert render(option(what)) == what + assert render(level(what)) == what + assert render(fallback(what)) == what + + +DASHBOARD = "The dashboard is down" +"""An example of one option and a counterexample of another: the confusable pattern.""" + + +class Department(Choice): + """A choice whose every member carries examples, one of them the fallback.""" + + billing = option( + BILLING, + examples=["My card was charged twice", "Where is my refund?"], + counterexamples=[DASHBOARD], + ) + technical = option(TECHNICAL, examples=[DASHBOARD, "502 on every request"]) + sales = fallback("Pricing, upgrades, new accounts", examples=["Do you have a team plan?"]) + + +class Severity(Levels): + """Ordered levels, each with the inputs that score there.""" + + cosmetic = level(COSMETIC, examples=["typo in a label"]) + degraded = level("Broken feature, workaround exists", examples=["export fails in one browser"]) + blocking = level("No workaround exists", examples=["cannot log in", "data loss"]) + + +ANSWERS: dict[str, Json] = { + "q0": { + "type": "choice", + "choice": "technical", + "probabilities": {"billing": 0.1, "technical": 0.8, "sales": 0.1}, + "confidence": 0.9, + }, + "q1": { + "type": "score", + "score": 1.0, + "legend": {"0": COSMETIC, "1": "Broken feature", "2": "No workaround"}, + "probabilities": {"0": 0.1, "1": 0.8, "2": 0.1}, + "confidence": 0.9, + }, + "q2": {"type": "noul", "noul": 0.2}, +} +"""One answer per question of the batch below, over its own keys and levels.""" + +CHOICE_CRITERIA: Json = { + "billing": ( + f"{BILLING}\nExamples: My card was charged twice; Where is my refund?" + f"\nNot this option: {DASHBOARD}" + ), + "technical": f"{TECHNICAL}\nExamples: {DASHBOARD}; 502 on every request", + "sales": "Pricing, upgrades, new accounts\nExamples: Do you have a team plan?", +} +"""What `Department` must put on the wire, key by key.""" + +SCORE_CRITERIA: Json = [ + f"{COSMETIC}\nExamples: typo in a label", + "Broken feature, workaround exists\nExamples: export fails in one browser", + "No workaround exists\nExamples: cannot log in; data loss", +] +"""What `Severity` must put on the wire, low to high.""" + +NOUL_CRITERIA: Json = { + "true": "Needs a person now\nExamples: the whole site is down", + "false": "Can wait\nExamples: a broken job someone has a manual workaround for", +} +"""What a noul's described criteria must put on the wire, under the wire's own names.""" + + +def test_examples_reach_the_wire_as_the_rendered_criteria( + httpserver: HTTPServer, runner: Runner +) -> None: + expect_post(httpserver).respond_with_data(reply(ANSWERS), content_type=JSON) + batch = ( + choose(Department, "Which team should handle this?"), + score(Severity, "How bad is it?"), + noul("Is this urgent?").criteria( + option("Needs a person now", examples=["the whole site is down"]), + option("Can wait", examples=["a broken job someone has a manual workaround for"]), + ), + ) + assert runner.ask(batch, TICKET) == (Department.technical, Severity.degraded, False) + + request, _ = httpserver.log[-1] + body = as_object(narrow(request.get_json())) + check_request(body) + questions = as_object(body["questions"]) + assert as_object(questions["q0"])["criteria"] == CHOICE_CRITERIA + assert as_object(questions["q1"])["criteria"] == SCORE_CRITERIA + assert as_object(questions["q2"])["criteria"] == NOUL_CRITERIA + # A fallback marked with examples is still the fallback, and `The dashboard is down` + # went out as an example of one option and a counterexample of another. + assert Department.fallback_member() is Department.sales diff --git a/tests/test_surface.py b/tests/test_surface.py index dd5624e..966e471 100644 --- a/tests/test_surface.py +++ b/tests/test_surface.py @@ -38,7 +38,9 @@ "choose", "choose_among", "fallback", + "level", "noul", + "option", "score", "score_levels", } diff --git a/tests/test_wire.py b/tests/test_wire.py index 3898d92..f6201c5 100644 --- a/tests/test_wire.py +++ b/tests/test_wire.py @@ -37,6 +37,7 @@ UnsureError, choose, fallback, + option, score, ) from guideme.api import NoulAnswer, Request, Response, question_to_wire, request_to_wire @@ -559,6 +560,48 @@ def _levels_given_as_one_string(_monkeypatch: pytest.MonkeyPatch) -> None: _ = score_levels("How cross?", "abc") +def _a_runtime_level_with_counterexamples(_monkeypatch: pytest.MonkeyPatch) -> None: + # `level(...)` offers no counterexamples, so this is an `option(...)` written where a + # level belongs. Rendering "Not this option" onto an ordered scale is meaningless and + # dropping what it carries is silent, so the constructor refuses it. + _ = score_levels("How cross?", [option("Calm", counterexamples=["shouting"]), "Cross"]) + + +def _a_fallback_as_a_runtime_option(_monkeypatch: pytest.MonkeyPatch) -> None: + # `choose_among` answers in a `Key`, so there is no member for the marking to name. + # Accepting it and dropping it would leave the caller believing an unsure answer is + # handled when it raises instead. Same for the two below. + _ = choose_among("Which team?", {"sales": fallback("Pricing"), "billing": "Money"}) + + +def _a_fallback_as_a_runtime_level(_monkeypatch: pytest.MonkeyPatch) -> None: + _ = score_levels("How cross?", [fallback("Calm"), "Cross"]) + + +def _a_fallback_as_a_noul_criterion(_monkeypatch: pytest.MonkeyPatch) -> None: + _ = noul("Urgent?").criteria(fallback("Needs a person now"), "Can wait") + + +def _one_example_of_two_runtime_options(_monkeypatch: pytest.MonkeyPatch) -> None: + # One input cannot belong to two options. The reverse -- an example of one and a + # counterexample of another -- is legal and is what tells confusable options apart. + _ = choose_among( + "Which team?", + { + "billing": option("Money", examples=["Where is my refund?"]), + "technical": option("Bugs", examples=["Where is my refund?"]), + }, + ) + + +def _one_example_of_both_noul_criteria(_monkeypatch: pytest.MonkeyPatch) -> None: + # A yes and a no are two alternatives of one question, so the same rule holds there. + _ = noul("Urgent?").criteria( + option("Needs a person now", examples=["the export is broken"]), + option("Can wait", examples=["the export is broken"]), + ) + + def _events_log_without_the_logs_api(monkeypatch: pytest.MonkeyPatch) -> None: # Asking for log records where `opentelemetry-api` has no logs API is refused where # it is asked for. The alternative is a guide that emits none and never says so. @@ -597,6 +640,12 @@ def _events_given_an_unknown_mode(_monkeypatch: pytest.MonkeyPatch) -> None: _an_api_key_of_spaces, _an_empty_model, _levels_given_as_one_string, + _a_runtime_level_with_counterexamples, + _a_fallback_as_a_runtime_option, + _a_fallback_as_a_runtime_level, + _a_fallback_as_a_noul_criterion, + _one_example_of_two_runtime_options, + _one_example_of_both_noul_criteria, _events_log_without_the_logs_api, _events_both_without_the_logs_api, _events_log_with_a_drifted_log_record, @@ -610,6 +659,12 @@ def _events_given_an_unknown_mode(_monkeypatch: pytest.MonkeyPatch) -> None: "api_key_of_spaces", "empty_model", "levels_given_as_one_string", + "a_runtime_level_with_counterexamples", + "a_fallback_as_a_runtime_option", + "a_fallback_as_a_runtime_level", + "a_fallback_as_a_noul_criterion", + "one_example_of_two_runtime_options", + "one_example_of_both_noul_criteria", "events_log_without_the_logs_api", "events_both_without_the_logs_api", "events_log_with_a_drifted_log_record", diff --git a/uv.lock b/uv.lock index 399cd96..8562ffc 100644 --- a/uv.lock +++ b/uv.lock @@ -421,7 +421,7 @@ wheels = [ [[package]] name = "guideme" -version = "0.1.0" +version = "0.1.1" source = { editable = "." } dependencies = [ { name = "httpx" },