Carry examples and counterexamples on a rubric - #12
Merged
Merged
Conversation
A choice option and a score level could only say what they cover. Callers want to show inputs that belong there and, for a choice, inputs that belong to some other option: measured against live Jev, a deliberately ambiguous ticket goes from a 0.51/0.49 coin flip on the wrong option to 0.91 on the right one once the options carry examples. The parts ride on the rubric string and are composed into it at the four places a rubric becomes wire text, so an option cannot reach the wire having quietly lost them. The string value stays the bare text, and a rubric with no parts renders to that text byte for byte, which is what leaves every 0.1.0 caller's request unchanged. They are flattened into the text rather than sent as the API's structured criteria because a score answer echoes its criteria back in `legend`, which is `dict[str, str]` here: objects would not deserialise, so a documented API feature would cost a breaking change to a field nothing reads, for no measured gain and about 11% more billed input tokens. `level(...)` takes no counterexamples: "not this option" means nothing on an ordered scale, where the level above or below is what an input that does not belong here scores. An `option(...)` written where a level belongs is a `ConfigError` rather than a clause silently dropped.
The next reader will ask why a documented API feature was ignored, so docs/design.md records the measurement: object criteria come back in a score answer's legend and would not parse against a public type nothing reads, for no gain and about 11% more billed input tokens. docs/contract.md states the rendered string as a contract item and names the one deliberate asymmetry with Rust, whose renderer lives in the derive macro. The README shows option, level and fallback in code, which is also what tests/test_surface.py asks of every exported name, and AGENTS.md gains the invariant and the new test ceiling.
Purely additive, so a patch. Both lock files record the version and both are resolved with --locked, so the root and examples/otlp are relocked here. The observability capture carried the version on its InstrumentationScope lines and would now be stale. It was taken once and says that nothing in it is invented, so the version is elided rather than edited to a number no run produced.
A yes and a no are two described alternatives of one question, as confusable as two options of a choice. Measured against live Jev: asked whether a nightly export job with a manual workaround is urgent, plain `Urgent` / `Not urgent` answers yes at 0.75 four runs running, and the same criteria carrying examples answer no at 0.17 four runs running. The plain criteria are wrong, and one of the no examples is exactly what the state describes. So `.criteria(...)` renders like every other wire site and `option(...)` serves it, rather than a fourth constructor for the same idea. Two contradictions are now refused where the rubric is written: one string as an example of two alternatives of the same question, which says an input belongs to both, and one string as both an example and a counterexample of the same alternative. The overlap that reads alike and is the whole point stays legal -- an example of one alternative and a counterexample of another -- and the wire test now sends exactly that, so a future tightening cannot take it away unnoticed. Rendering order was already declaration order; it is now stated as contract and pinned by a golden case whose clauses are written out of alphabetical order, so a renderer that sorted or took a set would fail it.
Telling `examples=[]` from an argument left out used to be identity against the `()` default, which works only because CPython hands out one empty tuple. That is an implementation detail rather than a language guarantee, and the failure mode is silent: on a runtime that does not intern it, the caller mistake this refuses stops being caught and nothing says so. `None` is now the clause left out, so the distinction is a plain `is None` and rests on nothing. Any empty sequence that arrives was typed by the caller, which makes `examples=()` refused as well -- consistent, because an empty clause written out says nothing whatever type it is, and now pinned by its own case in the refusal list rather than left as an accident of the rule.
The noul criteria path was proven to reach the wire by the wire assertion, but its effect on the answer was only ever measured by hand. It is the same invariant as the choice half -- examples move the distribution towards the alternative they describe -- so it folds into the existing live function rather than becoming a second one, and the entry count does not move. It is also the stronger half: a vague `Urgent` / `Not urgent` calls a broken nightly job with a manual workaround urgent at 0.75, and the same criteria carrying examples call it 0.17, because one of the not-urgent examples is what the ticket describes. Asserted as `told < plain`, never a float.
Two ways a rubric could be quietly wrong, both found in review. A `str` is a `Sequence[str]` of its own characters, so `examples="refund"` type-checked and rendered as six one-letter examples, and a multi-word string failed with the actively misleading "every entry in examples must say something, got ' '". This package already guards the identical trap two functions away, in `score_levels`, so the guard and its message now mirror that one. One check at the head of the clause validator covers all three constructors. `fallback(...)` was accepted and dropped by `choose_among`, `score_levels` and `noul(...).criteria(...)`, while the declarative `Levels` twin raises. Those three answer in a `Key`, a `Rank` and a `bool`, none of which has a member to fall back to, so a caller could believe an unsure answer was handled when it would raise. All three now refuse and point at `.otherwise(...)`, which is the rule `Levels` has always stated. Refusing rather than honouring it in `choose_among` keeps one story: `fallback(...)` marks a Choice member, and everywhere else the question carries the value.
…sion
docs/contract.md gave the separators and the verbatim rule but never the two
literal labels or the order the clauses come in, and that file is what a third
SDK implements from: as written, its author would pick their own labels and
produce different bytes for the same declaration. The algorithm is now there in
full, with `Examples: ` and `Not this option: ` quoted exactly and the reason
the latter is not a formatting choice -- it was measured against `Not` and
`Counterexamples` and reached the correct option most often.
The observability capture gets its `0.1.0` back. Eliding it avoided inventing a
version no run produced, which was the right instinct and the wrong fix: it also
threw away the provenance. The version is the one the run emitted, the prose now
says so, and the release procedure says to re-take the capture or change
nothing rather than edit the number forward.
Also recorded: that a blank bare rubric stays legal while `option(" ")` is
refused, which is a version boundary rather than an oversight, and guideme-rust
draws the same line from the other side. And that `render` and the `require_*`
checks are internal despite living beside the three exported constructors.
Design 9.7 supersedes 4 here, and this build was on the wrong side of it:
`option("")` and `option(" ")` raised whether or not anything was attached.
That makes a declaration which was legal in 0.1.0 illegal in a patch release,
on a degenerate input that was already meaningless, and it diverges from
guideme-rust, which narrowed the same rule from the same starting point. The
contract says one declaration is legal in both SDKs or in neither.
So the check now fires only when a clause is present, and says what the mistake
actually is: examples were attached to nothing, rather than the description
being blank -- which is not an error this release gets to invent.
The passthrough property drops its non-blank filter as a result. It now runs
over unrestricted text, which is the honest statement of the invariant: a rubric
carrying no examples renders to its own bytes, whatever those bytes are.
Three things a third implementer would otherwise have to guess, and would guess differently, closed identically in both SDKs per design 9.8. The vector's `kind` field is named alongside `what`, `examples`, `counterexamples` and `rendered`, including that the `noul` cases come from the runtime renderer rather than a derive and are the only cover for that path. Blankness is Unicode `White_Space`, which is exactly Rust's `str::trim`. Python strips a superset, and the difference is stated rather than left to be discovered: measured against this runtime, 29 codepoints to 25, the four extra being `U+001C`-`U+001F`. They are C0 controls that are never valid rubric text, so no real declaration reaches the boundary -- but an implementer comparing the two SDKs would otherwise find it and have no way to tell deliberate from accidental. Duplicate detection is exact-string equality. `"a"` and `" a"` may coexist. An implementation that normalised or trimmed before comparing would refuse declarations these SDKs accept, which is the same divergence as a different clause label reached by another route.
The whitespace paragraph already said C0, but only in passing at the end. It now gives both block ranges, because the whole reason this rule is written down is that two SDKs must not describe the boundary differently, and naming the wrong block would be precisely that failure. It also records that nothing goes the other way: every Unicode White_Space codepoint is one Python strips. The emptiness check trims and the duplicate check does not, so `" a"` is not blank yet is a different entry from `"a"`. That reads like a bug until the two questions are separated: emptiness asks whether the caller wrote anything, where a leading space changes nothing, and duplication asks whether two entries put the same bytes in front of the model, where it changes everything, because the text is rendered verbatim.
Carrying the parts on the string meant `__new__` took four arguments where
`str.__getnewargs__` hands back one, so `copy.copy`, `copy.deepcopy` and
`pickle` raised `TypeError: _Rubric.__new__() missing 2 required positional
arguments`. A bare string did all three in 0.1.0, so this was a regression in a
release whose whole premise is that nothing a 0.1.0 caller does changes.
Enum members hid it: `Enum.__deepcopy__` returns self, so a declared `Choice`
was fine and only bare values broke -- including
`copy.deepcopy({"key": option(...)})`, which is exactly the dict a caller builds
before handing it to `choose_among`.
`__getnewargs_ex__` supplies all four, `is_fallback` included, so a copied
fallback is still the fallback. Covered where the shapes already are: every
golden row now round-trips through all three operations, and the choice test
builds an enum from copied rubrics and checks both the rendering and
`fallback_member()`.
Error vocabulary aligned with the Rust SDK while here: "empty" and "duplicate"
as the nouns, and the duplicate message now names the repeated string rather
than only saying that one exists.
docs/contract.md gave the blankness and duplicate-equality definitions but not the rules they serve, so a reader of only this file got the hard cases and none of the ordinary ones. All eight are now listed, including the one that must be allowed -- a string as an example of one alternative and a counterexample of another -- because an implementation that refused it would break the case the feature exists for, and a list of refusals with no permission in it invites exactly that. Also stated: items go in verbatim, nothing is escaped, and the rendering is not reversible. No SDK parses a rendered rubric back into its parts and none should be written to. The changelog had its notes under a dated 0.1.1 while nothing is published. They move back under Unreleased, where this repo's own release procedure says the release commit picks them up and sets the date.
Items are joined onto one line with "; ", so a newline in one reads as a clause the rubric never declared -- an example ending "x\nNot this option: anything" renders byte-identically to a counterexample clause that was never written. `U+000D` goes with it as hygiene: pasted-text residue that breaks the joined line. Items only. A newline in the rubric text stays legal, because a bare rubric has to stay legal whatever it holds, so refusing it in the clause-bearing case would stop a caller writing ordinary multi-line prose without closing any path -- a runtime rubric built from data still passes its text through verbatim either way. What the rule guarantees is narrower than input trust and is now stated as such: a string rendered from declared parts carries exactly the clauses those parts declared. The test is the literal codepoint, never `str.splitlines()`. That call is what "does this contain a line break?" means in Python and it splits on eight codepoints, where Rust's `lines()` splits on one, so the idiomatic spelling in each language produces two different rules and neither looks wrong read alone. The contract now names the codepoints and forbids the idiom. Blank still wins over this, which strip-first already gave: an item of one newline is empty, not broken.
One sentence decides every normalisation question a third implementer meets, and replaces the per-case rules that were accumulating: `["a; b"]` renders the same as `["a", "b"]` and is still one item, `" a"` is not `"a"`, `"A"` is not `"a"`. Compare the strings the caller declared and nothing else; an implementation that trimmed, split or case-folded first would refuse declarations both SDKs accept. The `; ` against newline asymmetry sits under it as the worked example, since it reads as an inconsistency until the two are separated: `; ` changes how many examples a reader sees, a newline changes which clause they are in, and only the second forges a clause. An item reading "Not this option: x" with no line break in it is the same class as `; ` and stays legal. Also recorded, because the earlier framing overstated it: rubric text is trusted, the SDK does not sanitise it, and the guarantee is rendering integrity rather than input trust.
Takes spec/ at guideme-rust d908ecf, which is the commit that added spec/vectors/rubric.json. Only SOURCE and the new vector differ; the rest of the tree was already identical, and `mise run spec-check` confirms the copy matches main. The golden table was a hand copy of the design document, so three files claimed the rendering was "pinned by" the vector while nothing read it. The test now parametrises over the vector's ten cases and the copy is gone, which makes the claim true: spec-check proves this file matches Rust's, and the test proves the renderer matches this file. Each case is rebuilt through the constructor its `kind` names rather than handed to `render` directly, so the vector exercises what a caller writes. An absent clause arrives as `[]` and is mapped back to no argument: passing it through would be a different declaration, since an empty clause written out is refused. Recorded in docs/contract.md, because a consumer that mapped `[]` to an empty list would fail every case that has one. An empty vector is refused rather than silently parametrising nothing, which would be a test that passes without running. The line-break message now names its code points, matching the Rust SDK: a caller who pasted a stray carriage return cannot see it and needs to know which character to hunt for.
Rust found a sentence in its own contract asserting that a level never renders a counterexample clause, which a pre-rendered string passed to the runtime constructor falsifies. Python does the same thing, so the claim would have been false here too; this file states the level rule as a refusal and so was clean, but two other sentences had the shape. "The examples clause always precedes the counterexamples clause" is true of the renderer and false of the rendered string, since a caller who writes the labels into the rubric text by hand can order them however they like. The renderer is now the subject. "Embedded they cannot produce a clause boundary", about U+2028 and its neighbours, is true of the bytes and unprovable about what a model's tokenizer does with them. The domain is now pinned to the bytes, and the unmeasured half is named as staying out of the contract, the same way the wider line-break set does. The pattern is a claim about the rendered string where the guarantee is only over the rendering function. A sentence whose subject is the output can always be falsified by a caller who builds that output by hand; one whose subject is the renderer cannot.
main takes pull requests only, so the release cannot date this with a direct push; it has to be in the branch that merges. The version and both lock files already carry 0.1.1 because CI resolves them --locked, and this is the half that was deliberately held back until the release was real. Dated the 22nd, which is today. The earlier draft of this heading said the 21st, which was the date the notes were written rather than the date they ship. Unreleased goes back above it, empty, which is how the file sat after 0.1.0 and where the next change belongs.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implements §4 of the cross-SDK rubric-examples design, in Python. Purely additive: 0.1.1.
What this adds
option(rubric, examples=…, counterexamples=…)andlevel(rubric, examples=…)joinfallback(…), which now takes the same keywords. All three return astrsubclass whosestring value is the bare rubric, with the parts carried as attributes;
_Fallbackis gone,folded into that one type as its fallback-marked flavour. All three are exported from
guidemeand__all__is now 35 names.enums.rendercomposes the parts into wire text, and every site that turns a rubric into wiretext goes through it —
Choice.rubric(),Levels.levels(),choose_among(),score_levels()and
NoulQuestion.criteria()— so anoption(…)cannot reach the wire having quietly lost itsexamples.
All three kinds of question take examples. A noul's yes and no are two described alternatives
of one question, as confusable as two options, so
.criteria(…)accepts anoption(…)oneither side;
optionis the one constructor for a described alternative and there is no fourth.Measured live, on a nightly export job with a manual workaround asked "is this urgent?":
Urgent/Not urgentThe load-bearing property
A rubric with no examples and no counterexamples renders byte-identically to today.
tests/test_rubrics.pyasserts it twice: as a Hypothesis property over any non-blank text fora bare string and for all three constructors, and as the first row of the golden table.
Existing users' wire bytes do not move.
Rendering
Clauses joined with
"\n", items within a clause joined with"; ", in the order written, thetext used verbatim — never trimmed, never re-punctuated. The golden cases of the design's §3
are a parametrised test and reproduce exactly:
Order is contract, so a seventh golden case writes both clauses out of alphabetical order: a
renderer that sorted or took a set would fail it.
level(…)deliberately takes nocounterexamples. Anoption(…)carrying them written wherea level belongs is a
ConfigError, on the class statement for aLevelsand on the call forscore_levels, rather than a clause silently dropped.Validation
examplesandcounterexamplesdefault toNone, which is how a clause is left out, so anyempty sequence that arrives was typed by the caller and says nothing.
[]and()are refusedalike, and the distinction is a plain
is None— no reliance on CPython interning the emptytuple, which would be correct today and silently wrong on a runtime that does not intern it.
Each is a
ConfigErrorwhere the rubric is written: a blank or whitespace-only rubric or entry,an empty clause written out, a repeat within one clause, a counterexample on a level, and the
two contradictions — one string as an example of two alternatives of the same question, and one
string as both an example and a counterexample of the same alternative.
The overlap that reads alike and is the whole point stays legal: one string as an example of
one alternative and a counterexample of another. The wire test now sends exactly that (
The dashboard is downis an example oftechnicaland a counterexample ofbilling) and assertsthe rendered bytes for both, so a future tightening cannot take it away unnoticed.
Tests
Four new test functions, taking the ceiling in
AGENTS.mdfrom 43 to 47 with the reasonrecorded there; everything else — eight new refusal cases and the noul wire assertion — went
into a parameter or an assertion of a test that already existed.
Choice, aLevelsand a noul declared with examples produce exactly theexpected
criteriamap, list and{true, false}object on the request the local serverreceived, schema-checked;
covering both a choice and a noul, asserting
described > bareandtold < plain, never afloat.
Run live against
jev-1.13.0on 2026-09-21, three runs, rubric text held constant across eachpair so the examples are the only variable:
In both halves the plain rubric is not merely vaguer but wrong: the choice picks
return_policyat 0.02 confidence, and the noul calls a broken nightly job with a manualworkaround urgent — which is precisely one of the not-urgent examples.
Not in this PR
spec/vectors/rubric.jsonand thespec/SOURCEbump. That vector is generated byguideme-rustand vendored here, and the Rust PR lands first.spec-driftis green on thisbranch only because that vector is not on
guideme-rustmainyet; it will go red the momentthe Rust PR merges, and the vendoring commit on top of this branch is what clears it. The
golden table in the design document is what this PR is tested against instead.
Gate
mise run checkran green through the pre-push hook on the pushed HEAD: fmt-check, gen-check,ruff, pyright, pylint 10.00/10, 184 passed, build, audit. CI is green on every job.