Skip to content

Carry examples and counterexamples on a rubric - #12

Merged
pedro-pscunha merged 18 commits into
mainfrom
feat/rubric-examples
Sep 22, 2026
Merged

pedro-pscunha merged 18 commits into
mainfrom
feat/rubric-examples

Conversation

@pedro-pscunha

@pedro-pscunha pedro-pscunha commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner

Implements §4 of the cross-SDK rubric-examples design, in Python. Purely additive: 0.1.1.

What this adds

option(rubric, examples=…, counterexamples=…) and level(rubric, examples=…) join
fallback(…), which now takes the same keywords. All three return a str subclass whose
string value is the bare rubric, with the parts carried as attributes; _Fallback is gone,
folded into that one type as its fallback-marked flavour. All three are exported from
guideme and __all__ is now 35 names.

enums.render composes the parts into wire text, and every site that turns a rubric into wire
text goes through it — Choice.rubric(), Levels.levels(), choose_among(), score_levels()
and NoulQuestion.criteria() — so an option(…) cannot reach the wire having quietly lost its
examples.

All three kinds of question take examples. A noul's yes and no are two described alternatives
of one question, as confusable as two options, so .criteria(…) accepts an option(…) on
either side; option is the one constructor for a described alternative and there is no fourth.
Measured live, on a nightly export job with a manual workaround asked "is this urgent?":

criteria p(yes) over four runs verdict
plain Urgent / Not urgent 0.75 / 0.75 / 0.74 / 0.76 yes — wrong
carrying examples 0.17 / 0.17 / 0.17 / 0.17 no — right

The load-bearing property

A rubric with no examples and no counterexamples renders byte-identically to today.
tests/test_rubrics.py asserts it twice: as a Hypothesis property over any non-blank text for
a bare string and for all three constructors, and as the first row of the golden table.
Existing users' wire bytes do not move.

Rendering

Clauses joined with "\n", items within a clause joined with "; ", in the order written, the
text used verbatim — never trimmed, never re-punctuated. The golden cases of the design's §3
are a parametrised test and reproduce exactly:

Payments, invoicing, refunds
Examples: My card was charged twice; Where is my refund?
Not this option: The dashboard is down

Order is contract, so a seventh golden case writes both clauses out of alphabetical order: a
renderer that sorted or took a set would fail it.

level(…) deliberately takes no counterexamples. An option(…) carrying them written where
a level belongs is a ConfigError, on the class statement for a Levels and on the call for
score_levels, rather than a clause silently dropped.

Validation

examples and counterexamples default to None, which is how a clause is left out, so any
empty sequence that arrives was typed by the caller and says nothing. [] and () are refused
alike, and the distinction is a plain is None — no reliance on CPython interning the empty
tuple, which would be correct today and silently wrong on a runtime that does not intern it.

Each is a ConfigError where the rubric is written: a blank or whitespace-only rubric or entry,
an empty clause written out, a repeat within one clause, a counterexample on a level, and the
two contradictions — one string as an example of two alternatives of the same question, and one
string as both an example and a counterexample of the same alternative.

The overlap that reads alike and is the whole point stays legal: one string as an example of
one alternative and a counterexample of another. The wire test now sends exactly that (The dashboard is down is an example of technical and a counterexample of billing) and asserts
the rendered bytes for both, so a future tightening cannot take it away unnoticed.

Tests

Four new test functions, taking the ceiling in AGENTS.md from 43 to 47 with the reason
recorded there; everything else — eight new refusal cases and the noul wire assertion — went
into a parameter or an assertion of a test that already existed.

  • golden rendering table (parametrised over §3's cases plus the order case);
  • the passthrough property, under Hypothesis;
  • a wire proof: a Choice, a Levels and a noul declared with examples produce exactly the
    expected criteria map, list and {true, false} object on the request the local server
    received, schema-checked;
  • a live proof, deselected by default, that examples move the distribution — one function
    covering both a choice and a noul, asserting described > bare and told < plain, never a
    float.

Run live against jev-1.13.0 on 2026-09-21, three runs, rubric text held constant across each
pair so the examples are the only variable:

run 1: choice bare=0.50 described=0.87    noul plain=0.75 told=0.17
run 2: choice bare=0.49 described=0.89    noul plain=0.76 told=0.17
run 3: choice bare=0.53 described=0.89    noul plain=0.75 told=0.17

In both halves the plain rubric is not merely vaguer but wrong: the choice picks
return_policy at 0.02 confidence, and the noul calls a broken nightly job with a manual
workaround urgent — which is precisely one of the not-urgent examples.

Not in this PR

spec/vectors/rubric.json and the spec/SOURCE bump. That vector is generated by
guideme-rust and vendored here, and the Rust PR lands first. spec-drift is green on this
branch only because that vector is not on guideme-rust main yet; it will go red the moment
the Rust PR merges, and the vendoring commit on top of this branch is what clears it. The
golden table in the design document is what this PR is tested against instead.

Gate

mise run check ran green through the pre-push hook on the pushed HEAD: fmt-check, gen-check,
ruff, pyright, pylint 10.00/10, 184 passed, build, audit. CI is green on every job.

A choice option and a score level could only say what they cover. Callers want
to show inputs that belong there and, for a choice, inputs that belong to some
other option: measured against live Jev, a deliberately ambiguous ticket goes
from a 0.51/0.49 coin flip on the wrong option to 0.91 on the right one once the
options carry examples.

The parts ride on the rubric string and are composed into it at the four places
a rubric becomes wire text, so an option cannot reach the wire having quietly
lost them. The string value stays the bare text, and a rubric with no parts
renders to that text byte for byte, which is what leaves every 0.1.0 caller's
request unchanged.

They are flattened into the text rather than sent as the API's structured
criteria because a score answer echoes its criteria back in `legend`, which is
`dict[str, str]` here: objects would not deserialise, so a documented API
feature would cost a breaking change to a field nothing reads, for no measured
gain and about 11% more billed input tokens.

`level(...)` takes no counterexamples: "not this option" means nothing on an
ordered scale, where the level above or below is what an input that does not
belong here scores. An `option(...)` written where a level belongs is a
`ConfigError` rather than a clause silently dropped.
The next reader will ask why a documented API feature was ignored, so
docs/design.md records the measurement: object criteria come back in a score
answer's legend and would not parse against a public type nothing reads, for no
gain and about 11% more billed input tokens.

docs/contract.md states the rendered string as a contract item and names the one
deliberate asymmetry with Rust, whose renderer lives in the derive macro.

The README shows option, level and fallback in code, which is also what
tests/test_surface.py asks of every exported name, and AGENTS.md gains the
invariant and the new test ceiling.
Purely additive, so a patch. Both lock files record the version and both are
resolved with --locked, so the root and examples/otlp are relocked here.

The observability capture carried the version on its InstrumentationScope lines
and would now be stale. It was taken once and says that nothing in it is
invented, so the version is elided rather than edited to a number no run
produced.
A yes and a no are two described alternatives of one question, as confusable
as two options of a choice. Measured against live Jev: asked whether a nightly
export job with a manual workaround is urgent, plain `Urgent` / `Not urgent`
answers yes at 0.75 four runs running, and the same criteria carrying examples
answer no at 0.17 four runs running. The plain criteria are wrong, and one of
the no examples is exactly what the state describes. So `.criteria(...)` renders
like every other wire site and `option(...)` serves it, rather than a fourth
constructor for the same idea.

Two contradictions are now refused where the rubric is written: one string as an
example of two alternatives of the same question, which says an input belongs to
both, and one string as both an example and a counterexample of the same
alternative. The overlap that reads alike and is the whole point stays legal --
an example of one alternative and a counterexample of another -- and the wire
test now sends exactly that, so a future tightening cannot take it away
unnoticed.

Rendering order was already declaration order; it is now stated as contract and
pinned by a golden case whose clauses are written out of alphabetical order, so
a renderer that sorted or took a set would fail it.
Telling `examples=[]` from an argument left out used to be identity against the
`()` default, which works only because CPython hands out one empty tuple. That
is an implementation detail rather than a language guarantee, and the failure
mode is silent: on a runtime that does not intern it, the caller mistake this
refuses stops being caught and nothing says so.

`None` is now the clause left out, so the distinction is a plain `is None` and
rests on nothing. Any empty sequence that arrives was typed by the caller, which
makes `examples=()` refused as well -- consistent, because an empty clause
written out says nothing whatever type it is, and now pinned by its own case in
the refusal list rather than left as an accident of the rule.
The noul criteria path was proven to reach the wire by the wire assertion, but
its effect on the answer was only ever measured by hand. It is the same
invariant as the choice half -- examples move the distribution towards the
alternative they describe -- so it folds into the existing live function rather
than becoming a second one, and the entry count does not move.

It is also the stronger half: a vague `Urgent` / `Not urgent` calls a broken
nightly job with a manual workaround urgent at 0.75, and the same criteria
carrying examples call it 0.17, because one of the not-urgent examples is what
the ticket describes. Asserted as `told < plain`, never a float.
Two ways a rubric could be quietly wrong, both found in review.

A `str` is a `Sequence[str]` of its own characters, so `examples="refund"`
type-checked and rendered as six one-letter examples, and a multi-word string
failed with the actively misleading "every entry in examples must say
something, got ' '". This package already guards the identical trap two
functions away, in `score_levels`, so the guard and its message now mirror that
one. One check at the head of the clause validator covers all three
constructors.

`fallback(...)` was accepted and dropped by `choose_among`, `score_levels` and
`noul(...).criteria(...)`, while the declarative `Levels` twin raises. Those
three answer in a `Key`, a `Rank` and a `bool`, none of which has a member to
fall back to, so a caller could believe an unsure answer was handled when it
would raise. All three now refuse and point at `.otherwise(...)`, which is the
rule `Levels` has always stated. Refusing rather than honouring it in
`choose_among` keeps one story: `fallback(...)` marks a Choice member, and
everywhere else the question carries the value.
…sion

docs/contract.md gave the separators and the verbatim rule but never the two
literal labels or the order the clauses come in, and that file is what a third
SDK implements from: as written, its author would pick their own labels and
produce different bytes for the same declaration. The algorithm is now there in
full, with `Examples: ` and `Not this option: ` quoted exactly and the reason
the latter is not a formatting choice -- it was measured against `Not` and
`Counterexamples` and reached the correct option most often.

The observability capture gets its `0.1.0` back. Eliding it avoided inventing a
version no run produced, which was the right instinct and the wrong fix: it also
threw away the provenance. The version is the one the run emitted, the prose now
says so, and the release procedure says to re-take the capture or change
nothing rather than edit the number forward.

Also recorded: that a blank bare rubric stays legal while `option("   ")` is
refused, which is a version boundary rather than an oversight, and guideme-rust
draws the same line from the other side. And that `render` and the `require_*`
checks are internal despite living beside the three exported constructors.
Design 9.7 supersedes 4 here, and this build was on the wrong side of it:
`option("")` and `option("   ")` raised whether or not anything was attached.
That makes a declaration which was legal in 0.1.0 illegal in a patch release,
on a degenerate input that was already meaningless, and it diverges from
guideme-rust, which narrowed the same rule from the same starting point. The
contract says one declaration is legal in both SDKs or in neither.

So the check now fires only when a clause is present, and says what the mistake
actually is: examples were attached to nothing, rather than the description
being blank -- which is not an error this release gets to invent.

The passthrough property drops its non-blank filter as a result. It now runs
over unrestricted text, which is the honest statement of the invariant: a rubric
carrying no examples renders to its own bytes, whatever those bytes are.
Three things a third implementer would otherwise have to guess, and would guess
differently, closed identically in both SDKs per design 9.8.

The vector's `kind` field is named alongside `what`, `examples`,
`counterexamples` and `rendered`, including that the `noul` cases come from the
runtime renderer rather than a derive and are the only cover for that path.

Blankness is Unicode `White_Space`, which is exactly Rust's `str::trim`. Python
strips a superset, and the difference is stated rather than left to be
discovered: measured against this runtime, 29 codepoints to 25, the four extra
being `U+001C`-`U+001F`. They are C0 controls that are never valid rubric text,
so no real declaration reaches the boundary -- but an implementer comparing the
two SDKs would otherwise find it and have no way to tell deliberate from
accidental.

Duplicate detection is exact-string equality. `"a"` and `" a"` may coexist. An
implementation that normalised or trimmed before comparing would refuse
declarations these SDKs accept, which is the same divergence as a different
clause label reached by another route.
The whitespace paragraph already said C0, but only in passing at the end. It now
gives both block ranges, because the whole reason this rule is written down is
that two SDKs must not describe the boundary differently, and naming the wrong
block would be precisely that failure. It also records that nothing goes the
other way: every Unicode White_Space codepoint is one Python strips.

The emptiness check trims and the duplicate check does not, so `" a"` is not
blank yet is a different entry from `"a"`. That reads like a bug until the two
questions are separated: emptiness asks whether the caller wrote anything, where
a leading space changes nothing, and duplication asks whether two entries put
the same bytes in front of the model, where it changes everything, because the
text is rendered verbatim.
Carrying the parts on the string meant `__new__` took four arguments where
`str.__getnewargs__` hands back one, so `copy.copy`, `copy.deepcopy` and
`pickle` raised `TypeError: _Rubric.__new__() missing 2 required positional
arguments`. A bare string did all three in 0.1.0, so this was a regression in a
release whose whole premise is that nothing a 0.1.0 caller does changes.

Enum members hid it: `Enum.__deepcopy__` returns self, so a declared `Choice`
was fine and only bare values broke -- including
`copy.deepcopy({"key": option(...)})`, which is exactly the dict a caller builds
before handing it to `choose_among`.

`__getnewargs_ex__` supplies all four, `is_fallback` included, so a copied
fallback is still the fallback. Covered where the shapes already are: every
golden row now round-trips through all three operations, and the choice test
builds an enum from copied rubrics and checks both the rendering and
`fallback_member()`.

Error vocabulary aligned with the Rust SDK while here: "empty" and "duplicate"
as the nouns, and the duplicate message now names the repeated string rather
than only saying that one exists.
docs/contract.md gave the blankness and duplicate-equality definitions but not
the rules they serve, so a reader of only this file got the hard cases and none
of the ordinary ones. All eight are now listed, including the one that must be
allowed -- a string as an example of one alternative and a counterexample of
another -- because an implementation that refused it would break the case the
feature exists for, and a list of refusals with no permission in it invites
exactly that.

Also stated: items go in verbatim, nothing is escaped, and the rendering is not
reversible. No SDK parses a rendered rubric back into its parts and none should
be written to.

The changelog had its notes under a dated 0.1.1 while nothing is published. They
move back under Unreleased, where this repo's own release procedure says the
release commit picks them up and sets the date.
Items are joined onto one line with "; ", so a newline in one reads as a clause
the rubric never declared -- an example ending
"x\nNot this option: anything" renders byte-identically to a counterexample
clause that was never written. `U+000D` goes with it as hygiene: pasted-text
residue that breaks the joined line.

Items only. A newline in the rubric text stays legal, because a bare rubric has
to stay legal whatever it holds, so refusing it in the clause-bearing case would
stop a caller writing ordinary multi-line prose without closing any path -- a
runtime rubric built from data still passes its text through verbatim either
way. What the rule guarantees is narrower than input trust and is now stated as
such: a string rendered from declared parts carries exactly the clauses those
parts declared.

The test is the literal codepoint, never `str.splitlines()`. That call is what
"does this contain a line break?" means in Python and it splits on eight
codepoints, where Rust's `lines()` splits on one, so the idiomatic spelling in
each language produces two different rules and neither looks wrong read alone.
The contract now names the codepoints and forbids the idiom.

Blank still wins over this, which strip-first already gave: an item of one
newline is empty, not broken.
One sentence decides every normalisation question a third implementer meets,
and replaces the per-case rules that were accumulating: `["a; b"]` renders the
same as `["a", "b"]` and is still one item, `" a"` is not `"a"`, `"A"` is not
`"a"`. Compare the strings the caller declared and nothing else; an
implementation that trimmed, split or case-folded first would refuse
declarations both SDKs accept.

The `; ` against newline asymmetry sits under it as the worked example, since it
reads as an inconsistency until the two are separated: `; ` changes how many
examples a reader sees, a newline changes which clause they are in, and only the
second forges a clause. An item reading "Not this option: x" with no line break
in it is the same class as `; ` and stays legal.

Also recorded, because the earlier framing overstated it: rubric text is
trusted, the SDK does not sanitise it, and the guarantee is rendering integrity
rather than input trust.
Takes spec/ at guideme-rust d908ecf, which is the commit that added
spec/vectors/rubric.json. Only SOURCE and the new vector differ; the rest of the
tree was already identical, and `mise run spec-check` confirms the copy matches
main.

The golden table was a hand copy of the design document, so three files claimed
the rendering was "pinned by" the vector while nothing read it. The test now
parametrises over the vector's ten cases and the copy is gone, which makes the
claim true: spec-check proves this file matches Rust's, and the test proves the
renderer matches this file.

Each case is rebuilt through the constructor its `kind` names rather than handed
to `render` directly, so the vector exercises what a caller writes. An absent
clause arrives as `[]` and is mapped back to no argument: passing it through
would be a different declaration, since an empty clause written out is refused.
Recorded in docs/contract.md, because a consumer that mapped `[]` to an empty
list would fail every case that has one.

An empty vector is refused rather than silently parametrising nothing, which
would be a test that passes without running.

The line-break message now names its code points, matching the Rust SDK: a
caller who pasted a stray carriage return cannot see it and needs to know which
character to hunt for.
Rust found a sentence in its own contract asserting that a level never renders a
counterexample clause, which a pre-rendered string passed to the runtime
constructor falsifies. Python does the same thing, so the claim would have been
false here too; this file states the level rule as a refusal and so was clean,
but two other sentences had the shape.

"The examples clause always precedes the counterexamples clause" is true of the
renderer and false of the rendered string, since a caller who writes the labels
into the rubric text by hand can order them however they like. The renderer is
now the subject.

"Embedded they cannot produce a clause boundary", about U+2028 and its
neighbours, is true of the bytes and unprovable about what a model's tokenizer
does with them. The domain is now pinned to the bytes, and the unmeasured half
is named as staying out of the contract, the same way the wider line-break set
does.

The pattern is a claim about the rendered string where the guarantee is only
over the rendering function. A sentence whose subject is the output can always
be falsified by a caller who builds that output by hand; one whose subject is
the renderer cannot.
main takes pull requests only, so the release cannot date this with a direct
push; it has to be in the branch that merges. The version and both lock files
already carry 0.1.1 because CI resolves them --locked, and this is the half that
was deliberately held back until the release was real.

Dated the 22nd, which is today. The earlier draft of this heading said the 21st,
which was the date the notes were written rather than the date they ship.

Unreleased goes back above it, empty, which is how the file sat after 0.1.0 and
where the next change belongs.
@pedro-pscunha
pedro-pscunha merged commit 59e62a5 into main Sep 22, 2026
7 checks passed
@pedro-pscunha
pedro-pscunha deleted the feat/rubric-examples branch September 22, 2026 11:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant