From 16651bd285be238c9cc3694e9d84c25919313732 Mon Sep 17 00:00:00 2001 From: Pedro Cunha Date: Tue, 22 Sep 2026 19:01:43 -0300 Subject: [PATCH 1/2] Rewrite the README in plain English The README mixed three readers (a new user, a user tuning policy, a contributor) and used long sentences and shifting terms, so a first read was hard. It now follows the heading order and glossary that every guideme SDK shares, in short sentences with one word per meaning, and lists the TypeScript SDK under Other SDKs. Design rationale and measurements leave the README for docs/design.md, where most of them already lived; the one missing reason (why Client sits in guideme.api.client) is added there. The sdist test note moves to CONTRIBUTING.md, whose tier list also drops guideme.question.Question, which AGENTS.md moved to the top list. docs/contract.md points at the section that now carries the timeout divergence. Code blocks are unchanged, so tests/typing/readme.py still mirrors the first example. --- CHANGELOG.md | 7 +- CONTRIBUTING.md | 7 +- README.md | 672 +++++++++++++++++++++-------------------------- docs/contract.md | 2 +- docs/design.md | 4 + 5 files changed, 315 insertions(+), 377 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index bc1972d..5f47869 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,7 +2,12 @@ ## Unreleased -Nothing yet. +### Documentation + +- The README is rewritten in plain English with the headings, the order and the terms that + every guideme SDK now shares. Design rationale and measurements it carried are in + `docs/design.md`, and the contributor notes are in `CONTRIBUTING.md`. The TypeScript SDK is + listed under **Other SDKs**. No behaviour changes. ## 0.2.0 — 2026-09-22 diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index a4e6856..7f1ee27 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -2,8 +2,8 @@ guideme makes a TypeSafe Jev judgment usable as Python control flow. It is one distribution, `guideme`. Its public surface is what `src/guideme/__init__.py` exports, plus a second -supported tier that callers import by their own path: the modules `guideme.api` and -`guideme.policy`, and the single name `guideme.question.Question`. +supported tier that callers import by their own path: the modules `guideme.api`, +`guideme.api.client` and `guideme.policy`. Everything else in the package is private. `AGENTS.md` and the README's **Lower layers** section say what each tier promises. @@ -38,6 +38,9 @@ mise run test # while you work mise run check # the full gate; the pre-push hook runs it too ``` +A source distribution carries the tests but not `mise.toml` or `uv.lock`, so from one rather +than a clone, run the suite with `uv run --group dev pytest`. + The suite is small on purpose and `AGENTS.md` says what a new test may be. A pull request that adds a mock of `policy`, of `Guide` or of the transport will be asked to replace it with a real local server, a property test, or a typing proof. diff --git a/README.md b/README.md index 0c6df9e..b706d37 100644 --- a/README.md +++ b/README.md @@ -1,12 +1,30 @@ -# guideme +# guideme (Python) -Judgments from [TypeSafe Jev](https://docs.typesafe.ai) that read like Python control flow. +[![PyPI](https://img.shields.io/pypi/v/guideme.svg)](https://pypi.org/project/guideme/) +[![Python](https://img.shields.io/pypi/pyversions/guideme.svg)](https://pypi.org/project/guideme/) +[![license](https://img.shields.io/pypi/l/guideme.svg)](https://github.com/pedro-pscunha/guideme-python#license) -A yes/no question is an `if`. A choice is an exhaustive `match` over your own enum. A score is -a comparison against your own ordered levels. Thresholds, unsure bands and fallbacks are -explicit and composable. Every request is one span. The decision logic is pure and its -contract is published under `spec/`, so every guideme SDK, in any language, answers the same -way. +guideme sends a question and your state to [TypeSafe Jev](https://docs.typesafe.ai), the +TypeSafe model that gives judgments. It gives back the answer as a normal Python value: a +`bool`, a member of your own enum, or one of your own ordered levels. Your code then acts on the +answer with an `if`, a `match` or a comparison. Unlike the official +[`typesafe-sdk`](https://docs.typesafe.ai/sdk/python), which mirrors the API, guideme turns an +answer into control flow and calls the API itself. + +## Install + +```sh +uv add guideme # or: pip install guideme +``` + +guideme needs Python 3.12 or newer. The package ships `py.typed`, so your type checker sees +every annotation. Get an API key on the [keys page](https://console.typesafe.ai/keys) of the +TypeSafe console, and set it in the environment as `TYPESAFE_API_KEY`. To give the key in code, +use `Guide.builder().api_key(ApiKey("…")).build()`. + +## Quick start + +This example asks three questions about one support ticket. ```python from guideme import Choice, Guide, Levels, choose, fallback, noul, score @@ -43,74 +61,117 @@ if guide.ask(score(Frustration, "How frustrated is the customer?"), ticket) >= ( prioritise() ``` -The value of each member is the rubric the model reads. The member's name is the wire key. A -docstring on a member is documentation, not a rubric. pyright in strict mode enforces that -every option is handled, so adding a department turns the `match` above into an error until -you handle it. `from_env()` reads `TYPESAFE_API_KEY`, and `sales`, marked with `fallback(…)`, -is also the answer when confidence is below the floor. That floor is `min_confidence`, and it -is `0.0` unless you set it, which is why the choice above asks for `0.6`: with the default a -choice is never unsure and a `fallback(…)` member can never be reached. - -Two members may not share a rubric: Python would make the second an alias of the first, so a -repeat is a `ConfigError` on the class statement rather than a rubric quietly one option -short. `fallback(…)` marks a `Choice` member and only a `Choice` member; a `Levels` is -ordered, so the level to fall back to when a score is unsure is `.otherwise(level)` on the -question. - -That example is `tests/typing/readme.py`, which the gate type-checks with an `assert_type` on -every inferred answer type, so what is on this page cannot drift from what the package infers. -The three asks above are held to it by their use sites instead: the `match` is exhaustive and -the `>=` is between two levels of one scale. +- `Guide.from_env()` reads the key from `TYPESAFE_API_KEY`. +- The value of a member is its *rubric*: the text that tells the model what the option or level + means. The name of the member is its key on the wire. A docstring on a member is not rubric. +- pyright in strict mode makes sure that the `match` handles every option. If you add a + department, the `match` is an error until you handle it. +- `sales` is the fallback: the answer when the choice is unsure. The default `min_confidence` + is `0.0`, and with it a choice is never unsure. That is why this choice sets `0.6`. -## Install +The gate type-checks this example as `tests/typing/readme.py`, so this page cannot drift from +what the package infers. Build one guide and share it: the Configuration section tells why. -```sh -uv add guideme -``` +## Questions -or +There are three kinds of question. A *yes/no question* (TypeSafe calls it a *noul*, and the +wire type is `noul`) gives a `bool`. A *choice* gives one of your options. A *score* gives one +level of your ordered scale. -```sh -pip install guideme +| Constructor | Asks | Plain answer | `.detail()` answer | +|---|---|---|---| +| `noul("…")` | a yes/no question | `bool` | `Verdict`: `verdict` (`"yes"`, `"no"`, `"unsure"`), `p` | +| `choose(C, "…")`, `C` a `Choice` | a choice over 1 to 255 options | a member of `C` | `Ranked[C]`: `choice`, `confidence`, `unsure`, `probabilities` | +| `score(L, "…")`, `L` a `Levels` | a score over 2 to 10 levels, low to high | the most probable member of `L` | `Scored[L]`: the expected `value`, `level`, `confidence`, `unsure`, `distribution` | +| `choose_among("…", options)` | a choice over 1 to 255 `{key: rubric}` pairs given at runtime | `Key`, the key | `Ranked[Key]` | +| `score_levels("…", levels)` | a score over 2 to 10 level texts given at runtime | `Rank`, the index from 0 | `Scored[Rank]` | + +```python +team = guide.ask(choose_among("Which team?", {"billing": "Payments", "technical": "Bugs"}), ticket) +rank = guide.ask(score_levels("How severe?", ["Cosmetic", "Degraded", "Blocking"]), ticket) ``` -Python 3.12 or newer. The package ships `py.typed`, so your checker sees every annotation. +The size limits come from the TypeSafe API. A `Choice` or `Levels` class out of range raises +`ConfigError` on the class statement, and a runtime rubric raises it on the constructor call. +Two members of one class cannot have the same rubric: Python makes the second an alias of the +first, so that is a `ConfigError` too. -Set `TYPESAFE_API_KEY` in the environment, or pass a key to -`Guide.builder().api_key(ApiKey("…")).build()`. Keys come from the TypeSafe console, on its -[keys page](https://console.typesafe.ai/keys). +A yes/no question can say what yes and no mean with +`.criteria("what yes means", "what no means")`. The instructions can be a string or any +JSON-shaped value, so a question can name fields of structured state. -guideme is not the official TypeSafe SDK. That one is -[`typesafe-sdk`](https://docs.typesafe.ai/sdk/python), which mirrors the API: you send -questions and read answers. guideme adds the layer above it, turning an answer into control -flow — your own enums as the option set, thresholds and an unsure ladder as policy, one span -per request — and talks to the API itself rather than wrapping that package. +The *state* is the data you send with the question: a ticket, a message, a record. It can be +anything JSON-shaped: a string, a number, a `dict`, a list, and nestings of them. Convert a +dataclass with `dataclasses.asdict` and a pydantic model with `.model_dump()`. A value that +`json` cannot write, such as `NaN` or `bytes`, raises `ConfigError` before anything is sent. -## Three kinds of question, five constructors +`Question` is what the five constructors return. To annotate a question that you store or pass +on, use the concrete types `NoulQuestion`, `ChoiceQuestion`, `ScoreQuestion`, `DetailedNoul`, +`DetailedChoice` and `DetailedScore`. A `dict` is invariant, so a `dict[str, NoulQuestion]` is +not a `dict[str, Question[bool]]`. -| Constructor | Sends | Plain output | `.detail()` output | -|---|---|---|---| -| `noul("…")` | a yes/no question | `bool` | `Verdict` with the label and the probability | -| `choose(C, "…")` where `C` is a `Choice` | a choice over `C`'s 1 to 255 members | `C` | `Ranked[C]` with confidence and `probabilities`, the whole distribution | -| `score(L, "…")` where `L` is a `Levels` | a score over `L`'s 2 to 10 levels, low to high | `L`, the most probable level | `Scored[L]` with the expected `value`, the level, confidence and `distribution` | -| `choose_among("…", options)` | a choice over 1 to 255 runtime `{key: rubric}` pairs | `Key` | `Ranked[Key]` | -| `score_levels("…", levels)` | a score over 2 to 10 runtime level descriptions | `Rank` | `Scored[Rank]` | +`Probability`, `Confidence`, `Key` and `Rank` are `NewType` brands. Only the wire creates them, +after it validates the value. `Probability(2.0)` in your code is not refused. `Model` names a +model. `ApiKey` holds the key and never prints it. + +`guide.models()` returns a `tuple[ModelInfo, ...]`: the models that your account can use, each +with `name`, `description` and `release_date`. It is one `GET /v1/models` call, with no ask +span. It is retried like an ask, so a `429` while your process starts does not stop the start. + +## When the model is not sure + +An answer is *unsure* when it is not certain enough under the thresholds: + +- Yes/no question: `p >= yes_above` is yes, `p <= no_below` is no, and between the two is + unsure. The defaults are `0.5` and `0.5`, so no answer is unsure. +- Choice and score: `confidence < min_confidence` is unsure. With the default `0.0`, no answer + is unsure. + +A `Policy` is a set of thresholds, and each field is optional. You set it on the question +(`.yes_above(p)`, `.no_below(p)`, `.min_confidence(c)`, `.with_policy(Policy(…))`), on the guide +(`Guide.builder().policy(…)`, or `guide.with_policy(…)` for a copy), or not at all. The question +wins over the guide, and the guide wins over the defaults. A `Policy` settled against the +defaults is a `Thresholds`, the input of `guideme.policy.resolve` and of every golden vector. + +When an answer is unsure, guideme goes down the *unsure ladder*: + +1. The `.otherwise(value)` of the question. +2. The `fallback(…)` member of the `Choice`. A `Levels` class has no fallback member, so for a + score use `.otherwise(level)`. +3. `UnsureError`, which names the question and the threshold that it missed. + +`.detail()` skips the ladder and drops any `.otherwise(…)`. It never raises `UnsureError` and +gives you the full reading to decide yourself. `Policy` is a frozen dataclass, so a house policy +can be a module constant: + +```python +CAUTIOUS = Policy(yes_above=0.7, no_below=0.3) + +guide = Guide.builder().api_key(key).policy(CAUTIOUS).build() +strict = guide.with_policy(Policy(min_confidence=0.8)) + +reading = guide.ask(noul("Is this about billing?").detail(), ticket) +match reading.verdict: + case "yes": + billing() + case "no": + other() + case "unsure": + review(reading.p) -Those size limits are the API's, and guideme checks them where you write the rubric: a `Choice` -or `Levels` class outside the range is a `ConfigError` on the class statement, and a runtime -rubric is one on the constructor call. +picked = strict.ask(choose(Department, "Which team?").detail(), ticket) +``` -A noul can carry `.criteria("what yes means", "what no means")`. Instructions accept a string -or any JSON-shaped value, so a question can reference structured data by field name the way -the TypeSafe docs describe. +`strict` shares the connection pool of `guide` and keeps `CAUTIOUS`, with `min_confidence` set +over it. Thus the choice is unsure below `0.8`, and the yes/no question still uses `0.7 / 0.3`. -### Examples in a rubric +## Examples and counterexamples -Two alternatives that read alike are told apart by showing inputs rather than by describing -harder. `option(…)` takes the inputs that belong to an alternative and the ones that belong -somewhere else, `level(…)` takes the inputs that score at that level, and `fallback(…)` is an -`option(…)` that also marks the unsure member. All three kinds of question take them: a noul's -`.criteria(…)` accepts an `option(…)` for the yes and for the no. +Two options that read alike are easier to tell apart with inputs than with a longer rubric. An +*example* is an input that belongs to an option. A *counterexample* is an input that does not. +`option(…)` takes both. `level(…)` takes examples only, because on an ordered scale an input +that does not belong at one level belongs at another. `fallback(…)` is an `option(…)` that also +marks the fallback member. ```python from guideme import Choice, Levels, fallback, level, option @@ -132,8 +193,8 @@ class Severity(Levels): blocking = level("No workaround exists", examples=["cannot log in", "data loss"]) ``` -The member's value is still the bare rubric; the examples are composed into it only in the -request, as +The value of the member stays the bare rubric. guideme adds the examples only in the request, +as this text: ```text Payments, invoicing, refunds @@ -141,13 +202,12 @@ Examples: My card was charged twice; Where is my refund? Not this option: The dashboard is down ``` -So a rubric with no examples sends exactly what it sent before, and the same strings work in -`choose_among("…", {"billing": option(…)})` and `score_levels("…", [level(…), …])`. Examples -and counterexamples render in the order they are written, always: that order is part of the -published contract. +A rubric with no examples sends the same bytes as before. Examples render in the order that you +write them, and this text is part of the published contract. The same values work at runtime, +in `choose_among("…", {"billing": option(…)})` and `score_levels("…", [level(…), …])`. -A yes and a no are two alternatives of one question, so they take examples too, and this is -where they pay best — a vague pair is the easiest thing to get wrong: +A yes and a no are two options of one question, and a vague pair is the easiest to get wrong. So +`.criteria(…)` also takes an `option(…)` for each side: ```python urgent = noul("Is this ticket urgent?").criteria( @@ -156,89 +216,28 @@ urgent = noul("Is this ticket urgent?").criteria( ) ``` -Asked about a nightly export job that has been failing since Tuesday while the numbers are -pulled by hand, a plain `Urgent` / `Not urgent` answers yes at 0.75. The criteria above answer -no at 0.17, because one of the not-urgent examples is what the ticket describes. - -A string may be an example of one option and a counterexample of another. That is the point -when two options are confusable, and it is the one overlap that stays legal. Offering the same -string as an example of two options, or as both an example and a counterexample of the same -option, says an input belongs where it cannot, so each is refused. - -`level(…)` has no counterexamples, because "not this option" means nothing on an ordered -scale — an input that does not belong at one level scores at another. - -Leave a clause out to say there is none. An empty one written out — `examples=[]` or -`examples=()` — says nothing, so it is refused as the mistake it is, along with a clause given -as one string rather than a list of them (`examples="refund"` would otherwise be six one-letter -examples), a blank entry, a repeat within one clause, a newline or carriage return inside an -entry, a counterexample on a level, a `fallback(…)` given to `choose_among`, `score_levels` or -`.criteria(…)`, and the two contradictions above. Each is a `ConfigError` where the rubric is -written. - -Entries go on one line each, so a newline inside one would read as a clause you never wrote. -`"; "` inside an entry is fine — `"card declined; retry failed"` is ordinary prose, and it -changes how many examples a reader sees rather than which clause they are in. The rubric text -itself may still contain newlines; only the entries are restricted. - -Attaching examples to a blank rubric is refused too, because they describe something that is not -there. A blank rubric on its own is not: it means what it has always meant, and adding examples -to the language does not make an old declaration an error. - -The state is anything JSON-shaped: a text literal, a `dict`, a list of them. A dataclass goes -through `dataclasses.asdict`, a pydantic model through `.model_dump()`. +[`docs/design.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/docs/design.md#decisions) +records a measured case where these examples change a wrong yes into a correct no. -## Policy - -Thresholds decide how a probability or a confidence becomes an answer. They form a patch that -merges from the question, over the guide, over the package defaults. A `Policy` is that patch, -with every field optional; settling one against the defaults gives a `Thresholds`, which is -what `guideme.policy.resolve` takes and what every golden vector is written against. - -| Layer | How to set | Wins over | -|---|---|---| -| question | `.yes_above(p)`, `.no_below(p)` on nouls; `.min_confidence(c)` on choice and score; `.with_policy(Policy(…))` on any | guide | -| guide | `Guide.builder().policy(…)`, or `guide.with_policy(…)` for a scoped copy | defaults | -| defaults | `yes_above 0.5`, `no_below 0.5`, `min_confidence 0.0` | nothing | - -The rules: - -- Noul: `p >= yes_above` is yes, `p <= no_below` is no, strictly between is unsure. With the - defaults there is no unsure band. -- Choice and score: `confidence < min_confidence` is unsure. With the default there is never - an unsure answer. - -When an answer is unsure, resolution goes down a ladder: `.otherwise(value)` on the question, -then the enum's `fallback(…)` member, then `UnsureError` naming the question and the boundary -it missed. `.detail()` skips the ladder and hands you the reading to decide yourself. - -`Policy` is a frozen dataclass, so a house policy is a module constant: - -```python -CAUTIOUS = Policy(yes_above=0.7, no_below=0.3) - -guide = Guide.builder().api_key(key).policy(CAUTIOUS).build() -strict = guide.with_policy(Policy(min_confidence=0.8)) - -reading = guide.ask(noul("Is this about billing?").detail(), ticket) -match reading.verdict: - case "yes": - billing() - case "no": - other() - case "unsure": - review(reading.p) +Each of these raises `ConfigError` where you write the rubric: -picked = strict.ask(choose(Department, "Which team?").detail(), ticket) -``` +- An empty clause, such as `examples=[]`. To say there are none, leave the clause out. +- A clause given as one string, such as `examples="refund"`, in place of a list. +- A blank entry, the same entry twice in one clause, or a newline or carriage return in an entry. +- One string as an example of two options (or of the yes and the no, or of two levels). +- One string as an example and a counterexample of the same option. +- A counterexample on a level, or a `fallback(…)` in `choose_among`, `score_levels` or + `.criteria(…)`. +- Examples on a blank rubric. -`strict` shares the connection pool and inherits `CAUTIOUS`, with `min_confidence` patched over -it, so the choice above is unsure below `0.8` while the noul still reads against `0.7 / 0.3`. +Some overlaps are legal. One string can be an example of one option and a counterexample of +another: that is how you tell two similar options apart. An entry can contain `"; "`, and the +rubric text itself can contain newlines. A blank rubric with no examples is legal too. -## Several judgments, one request +## Several questions in one request -A tuple of questions is a question. So is a list or a dict, and they nest. The answer has the -same shape, from one request and one span. Each question keeps its own policy. +A tuple of questions is also a question. So is a list or a dict, and they nest. The answer has +the same shape and comes from one request and one span. Each question keeps its own policy. ```python urgent, dept, mood, flags = guide.ask( @@ -256,46 +255,19 @@ if urgent or mood.value > 1.5 or flags["vip"]: prioritise() ``` -Question ids are `q0..qN` in encounter order, which for a dict is insertion order. They appear -on the wire, in errors and in events. A batch is atomic: one answer that cannot be resolved -fails the whole call, so put `.otherwise(…)` or `.detail()` on the questions that may come -back unsure. +The question ids are `q0..qN` in encounter order (insertion order for a dict). They appear on +the wire, in errors and in events. A batch is atomic: if one answer cannot be resolved, the +whole call fails. So put `.otherwise(…)` or `.detail()` on each question that can come back +unsure. -Any nesting works at runtime. The forms your checker infers a type for are a single question, -a list, a dict, a tuple of up to eight questions, and a tuple of up to seven followed by one -list or dict, which is the shape above. +Any nesting works at runtime. Your type checker infers a type for one question, a list and a +dict. It also infers a tuple of up to eight questions, or of up to seven followed by one +list or dict (the shape above). -## Sync and async +## The receipt -`Guide` and `AsyncGuide` have the same surface over the same pure core. The difference is the -`await` and the `httpx` client underneath. - -```python -async with AsyncGuide.from_env() as guide: - verdict: Verdict = await guide.ask(noul("Is this about billing?").detail(), ticket) -``` - -The synchronous version is the same two lines with `Guide`, `with` and no `await`. -`Guide.builder()` and `AsyncGuide.builder()` return the same `GuideBuilder`; `.build()` gives -the synchronous guide and `.build_async()` the asynchronous one. Leaving the block closes the -guide, and `guide.close()` does the same thing by hand for a guide that outlives any block. - -A guide holds a connection pool, so build one and share it: both kinds are safe to use from -several threads or several tasks at once, and one guide asking concurrently is what the pool -is for. Building one per request works but opens a pool per request, which is the cost the -pool exists to avoid. `guide.with_policy(…)` returns a second guide over the *same* pool, and -the two are counted, so closing either leaves the other able to ask and the pool closes when -the last of them does. Close each guide once. - -Both guides also answer `models()`, which returns a `tuple[ModelInfo, ...]`: the models the -account may use, each with its `name`, `description` and `release_date`. It is one call to -`GET /v1/models` and gets no ask span of its own. It is retried on exactly the terms an ask -is, so a `429` while your process is starting up does not fail the start. - -## What a request cost - -`ask_with_receipt` is `ask` with the response's own numbers kept. It takes the same shapes and -infers the same types; `ask` is this call followed by `.answer`. +`ask_with_receipt` is `ask` that also keeps the numbers of the response. It takes the same +shapes and infers the same types. `ask` is `ask_with_receipt` followed by `.answer`. ```python receipt: Receipt[bool] = guide.ask_with_receipt(noul("Is this urgent?"), ticket) @@ -305,46 +277,55 @@ if receipt.answer: meter(model=receipt.model, tokens=receipt.usage.input_tokens) ``` -`Receipt` is frozen and carries three things: `answer`, whatever `ask` would have returned; -`model`, the versioned id that actually answered, which is `jev-1.13.0` and not `jev-latest` -even when an alias was asked for; and `usage`, a `Usage` with `input_tokens` and -`output_tokens`. Input tokens are what is billed. Log the model: thresholds are tuned against -one model's numbers, and the alias moves under you. +A `Receipt` is frozen. `answer` is what `ask` returns. `model` is the versioned id of the model +that answered, for example `jev-1.13.0`, also when you asked for the `jev-latest` alias. `usage` +is a `Usage` with `input_tokens` and `output_tokens`, and TypeSafe bills the input tokens. Log +the model: thresholds are tuned against the numbers of one model, and an alias can move. -## Configuration +## Errors -Every setter on `GuideBuilder` returns the builder, and `Guide.builder()` starts one. +Every failure is a `GuidemeError`. Its `.kind` is the error kind: the same string in every +guideme SDK, and the value of `error.type` on the failed span. -| Setter | Default | What it does | +| Class | `.kind` | When | |---|---|---| -| `api_key(ApiKey(…))` | none; required | The key. `from_env()` reads it from `TYPESAFE_API_KEY`. | -| `base_url(…)` | `https://api.typesafe.ai` | The API origin. It may not carry credentials. | -| `model(Model(…))` | `jev-latest` | The model or alias to ask. | -| `policy(Policy(…))` | the defaults | The guide-wide policy patch; a question's own wins over it. | -| `max_retries(n)` | `3` | Resends per call, `0` to never resend. | -| `backoff(…)` | 500 ms | Base of the exponential backoff. | -| `timeout(…)` | 30 s | Per phase of one attempt. Read the paragraph below. | -| `transport(…)` / `async_transport(…)` | `httpx`'s own | Send through your `httpx` transport. | -| `record_state(True)` | off | Put the state JSON on the ask span. It is your users' data. | -| `events(…)` | `"both"` | Whether an answer and a retry go to the span, a log record, or both. | - -**`timeout` is not a deadline for the attempt.** `httpx` gives the whole budget to each phase -separately — connecting, writing, reading, and waiting for a pooled connection — so one attempt -that is slow in more than one phase takes longer than the timeout without breaching anything. -Worst case for a call is `max_retries + 1` attempts of several phases each, plus the backoff -between them. The Rust SDK's `reqwest` deadline covers the attempt as a whole instead; the two -differ because their HTTP clients do, and `docs/contract.md` records it as a divergence rather -than leaving you to find it. - -`transport(…)` and `timeout(…)` refuse each other, in whichever order you write them: a -timeout belongs to the transport that honours it, and `httpx` hands yours this budget as a -request extension it is free to ignore. A silent no-op would be worse than a `ConfigError`. -`transport(…)` is for `build()` and `async_transport(…)` for `build_async()`; using one with -the other build is a `ConfigError` too. +| `AuthError` | `auth` | HTTP 401 | +| `InvalidError` | `invalid` | HTTP 422. `.detail` is the response body. | +| `RateLimitedError` | `rate_limited` | HTTP 429 after the retries, or a `retry-after` that is too long to wait. `.retry_after` holds it. | +| `OverloadedError` | `overloaded` | HTTP 529, on the same terms. `.retry_after` holds it. | +| `TransportError` | `transport` | A connection, TLS or timeout failure. | +| `UnexpectedStatusError` | `unexpected_status` | A status that the contract does not define. | +| `ProtocolError` | `protocol` | The response breaks the contract: a body that does not decode, a wrong answer kind, an option or level that is not in the rubric, a probability outside 0..1. | +| `UnsureError` | `unsure` | The policy said unsure and the ladder had no value. | +| `ConfigError` | `config` | A mistake in your code, raised where you write it: bad thresholds, no key, an empty batch, state that is not JSON, a rubric that breaks a rule, a bad `events(…)`, a timeout that is not positive, negative retries or backoff, a `base_url` that holds credentials. | + +## Retries and timeouts + +guideme retries HTTP 429 and 529, for an ask and for `models()`. The backoff is exponential with +jitter, capped at 30 s. guideme obeys a `retry-after` header in whole seconds. If `retry-after` +is more than 30 s, guideme does not wait: it raises `RateLimitedError` or `OverloadedError`. + +A failed connection is resent in the same budget: a refused or reset connection, or a TLS +handshake that did not complete. That request did not reach a server, so nothing was answered. +A call makes at most `max_retries + 1` attempts, whatever the mix of failures. + +A disconnect during the response is not resent: the API can have answered, and a second request +pays for the same answer twice. No timeout is resent, of any phase, a connect timeout included. +Each of these raises `TransportError` at once. +[`docs/design.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/docs/design.md#decisions) +gives the reasons. + +**`timeout` is not a deadline for the attempt.** `httpx` gives the full value to each phase: +connect, write, read, and the wait for a pooled connection. Thus one slow attempt can take more +than the timeout. The worst case for a call is `max_retries + 1` attempts of several phases each, +plus the backoff between them. The Rust SDK has one deadline for the full attempt, and +[`docs/contract.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/docs/contract.md) +records this divergence. ## Testing your code -Pass a transport and your control flow is testable with no server, no port and no key: +Give the guide an `httpx` transport. Then you can test your control flow with no server, no port +and no key: ```python import httpx @@ -368,25 +349,18 @@ def test_an_urgent_ticket_is_prioritised() -> None: assert guide.ask(noul("Is this urgent?"), "payouts failing") is True ``` -`q0` is the first question in encounter order; a batch of three is answered with `q0`, `q1` -and `q2`. Raise from the handler instead of returning and you get the failure paths: an -`httpx.ConnectError` is resent inside the retry budget, and an `httpx.ReadTimeout`, an -`httpx.ConnectTimeout` or an `httpx.RemoteProtocolError` is not. For -`AsyncGuide`, hand the same `httpx.MockTransport` to `async_transport(…)` and `build_async()`; -it is both kinds of transport at once. +`q0` is the first question in encounter order, and a batch of three uses `q0`, `q1` and `q2`. +To test a failure, raise an exception in the handler. guideme resends an `httpx.ConnectError` +inside the retry budget. It does not resend an `httpx.ReadTimeout`, an `httpx.ConnectTimeout` or +an `httpx.RemoteProtocolError`. For `AsyncGuide`, give the same `httpx.MockTransport` to +`async_transport(…)` and call `build_async()`. A `MockTransport` is both kinds of transport. ## Observability -guideme emits OpenTelemetry spans, span events and OTLP log records through -`opentelemetry-api` and installs nothing: no tracer provider, no logger provider, no exporter, -no logging handler. Install a provider and the data appears. The SDK and an exporter are not -dependencies of this package, so install them alongside it: - -```sh -pip install opentelemetry-sdk opentelemetry-exporter-otlp-proto-grpc -``` - -The smallest provider that leaves the process: +guideme sends OpenTelemetry spans, span events and OTLP log records through `opentelemetry-api`. +It installs no provider, no exporter and no logging handler. When you install a provider, the +data appears. Install the SDK and an exporter next to guideme +(`pip install opentelemetry-sdk opentelemetry-exporter-otlp-proto-grpc`), then: ```python from opentelemetry import trace @@ -399,169 +373,121 @@ provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter())) trace.set_tracer_provider(provider) ``` -One span named `guideme.ask` per request, shaped by the OpenTelemetry GenAI conventions: -`gen_ai.request.model`, `gen_ai.response.model`, `gen_ai.usage.*`, and on failure `error.type` -with an error status. Under it, one HTTP client span per attempt with -`http.response.status_code`, so a retry is visible as sibling spans, plus a `guideme.retry` -event when an attempt is resent. One `guideme.answer` event per question with the outcome, -the probability or confidence, the unsure verdict and the settled thresholds that produced it. -The state is never recorded unless you opt in with `record_state(True)`. The API key never -appears anywhere. - -Every answer and every retry is also an OTLP log record, at `INFO` and at `WARN`, carrying the -trace id and the span id of the span it came from, so a logs backend links one straight back to -the decision it explains. Install a `LoggerProvider` too and they arrive; install neither and -they cost nothing. `events(...)` on the builder picks which signal carries an event when you -export both; **Choosing a signal** in the observability document has the table, the default, and -what happens on an `opentelemetry-api` that has no logs API: asking for one is refused, and the -default falls back to the span event alone. - -Because the shapes are standard, any OTLP backend reads them as is. +Each request is one `guideme.ask` span with the OpenTelemetry GenAI fields, and one HTTP client +span per attempt below it. A retry adds a `guideme.retry` event. Each question adds a +`guideme.answer` event with the outcome, the probability or confidence and the thresholds. Each +answer is also an OTLP log record at `INFO`, and each retry one at `WARN`, with the trace id and +span id. To receive them, install a `LoggerProvider`. The state is not recorded unless you set +`record_state(True)`. The API key is never recorded. + +`events(…)` selects the signal for answers and retries: `"span"`, `"log"` or `"both"` (the +default). If you export traces and logs to one backend, set `"span"` or `"log"` to store each +event once. If your `opentelemetry-api` has no logs API, `"log"` and `"both"` raise +`ConfigError`, and the default sends the span event only. [`docs/observability.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/docs/observability.md) -has the field tables and the environment variables that point the exporter anywhere. +has every field and the exporter settings. [`examples/otlp`](https://github.com/pedro-pscunha/guideme-python/tree/main/examples/otlp) runs -all of it against the live API with a collector that prints what arrives. +it all against the live API, with a collector that prints what arrives. -## Errors +## Configuration -Every failure is a `GuidemeError`. `.kind` is the same string every other guideme SDK reports -and the value of `error.type` on the failed span. +`Guide.builder()` starts a `GuideBuilder`. Each setting returns the builder. -| Class | `.kind` | When | +| Setting | Default | What it does | |---|---|---| -| `AuthError` | `auth` | 401 | -| `InvalidError` | `invalid` | 422; `.detail` is the body | -| `RateLimitedError` | `rate_limited` | 429 after retries, or a `retry-after` too long to wait for; `.retry_after` carries it | -| `OverloadedError` | `overloaded` | 529 on the same terms; `.retry_after` carries it too | -| `TransportError` | `transport` | connection, TLS, timeout | -| `UnexpectedStatusError` | `unexpected_status` | anything the contract does not define | -| `ProtocolError` | `protocol` | the response violates the contract: undecodable body, wrong answer kind, option or level not in the rubric, probability outside 0..1 | -| `UnsureError` | `unsure` | the policy said unsure and nothing caught it | -| `ConfigError` | `config` | raised where the mistake is written: bad thresholds, missing key, empty batch, unserialisable state, a duplicate rubric, a rubric outside 1..255 options or 2..10 levels, a bad `events(...)`, a non-positive timeout, negative retries or backoff, a `base_url` carrying credentials, and so on | - -Retries on 429 and 529 use exponential backoff with jitter, capped at 30 s, and honour an -integer `retry-after`. They apply to `models()` as much as to `ask`. - -A failed connection is resent in the same budget: refused, reset, or a TLS handshake that did -not complete. The request never reached a server, so nothing was judged and nothing is -repeated. A disconnect part-way through a response is **not** resent — the request arrived, -the API may have answered it, and asking again would buy the same judgment twice. - -**No timeout is resent, of any phase.** A connect timeout included, although `httpx` names it -separately: Rust's SDK sets one deadline over the whole attempt and cannot tell a connect -timeout from a read timeout, so retrying one here would make the two SDKs disagree about the -same failure, and a retried timeout multiplies the wall time `timeout(…)` is there to bound. -Everything not resent raises `TransportError` on the first failure. +| `api_key(ApiKey(…))` | none, required | The API key. | +| `base_url(…)` | `https://api.typesafe.ai` | The API origin. It cannot hold credentials. | +| `model(Model(…))` | `jev-latest` | The model or alias to ask. | +| `policy(Policy(…))` | the defaults | The policy of the guide. The policy of a question wins over it. | +| `max_retries(n)` | `3` | Resends per call. `0` never resends. | +| `backoff(…)` | 500 ms | The base of the exponential backoff. | +| `timeout(…)` | 30 s | The limit for each phase of one attempt. | +| `transport(…)` / `async_transport(…)` | the transport of `httpx` | Send through your own `httpx` transport. | +| `record_state(True)` | off | Put the state JSON on the ask span. It is the data of your users. | +| `events(…)` | `"both"` | Send answers and retries as span events, log records, or both. | +| `from_env()` | | Apply the environment variables below. | + +`transport(…)` and `timeout(…)` refuse each other, in either order, with a `ConfigError`: an +injected transport owns its deadlines. `transport(…)` is for `build()`, and `async_transport(…)` +is for `build_async()`. The wrong pair is a `ConfigError` too. + +| Variable | Meaning | +|---|---| +| `TYPESAFE_API_KEY` | The API key. `Guide.from_env()` and `AsyncGuide.from_env()` require it. | +| `TYPESAFE_BASE_URL` | Optional. A different API origin. | +| `GUIDEME_MODEL` | Optional. The model or alias. The default is `jev-latest`. | + +Build one guide per process and share it. A guide holds a connection pool, and it is safe to use +from many threads or tasks at once. A guide per request also works, but it opens a pool for each +request. + +## Sync and async + +`Guide` and `AsyncGuide` have the same methods over the same core. The differences are the +`await` and the `httpx` client below it. + +```python +async with AsyncGuide.from_env() as guide: + verdict: Verdict = await guide.ask(noul("Is this about billing?").detail(), ticket) +``` + +For the synchronous version, use `Guide`, `with`, and no `await`. `Guide.builder()` and +`AsyncGuide.builder()` return the same `GuideBuilder`: `.build()` gives a `Guide`, and +`.build_async()` gives an `AsyncGuide`. A guide closes at the end of its `with` block. Outside a +block, call `guide.close()`. `guide.with_policy(…)` returns a second guide over the same pool. +The pool counts its guides: after you close one, the other can still ask, and the pool closes +with the last guide. If you close one guide twice, the second close does nothing. ## Lower layers -Everything above is re-exported from the `guideme` package, and `guideme.__all__` is that list. +The `guideme` package re-exports everything above, and `guideme.__all__` is that list. Three modules are a second supported tier: `guideme.api`, `guideme.api.client` and -`guideme.policy`. You import those by their own path, they are not re-exported at the top -level, and they are under the same rule as the first tier — nothing in them is removed or -renamed without a major version and a `CHANGELOG.md` entry. Anything else in the package is -private, whatever its name looks like. - -The two bullets after them are not a tier. They say where some of the names above are -declared, which is worth knowing when two of them share a spelling. - -- `guideme.api` is the exact wire mirror of `POST /v1/systemone` and `GET /v1/models`, and - `guideme.api.__all__` is what it offers: the request and response models, its own `Usage`, - and the four adapters between them and the core. That `Usage` is the pydantic model a - response is parsed into, not the `Usage` a receipt carries — a receipt gets the frozen - dataclass of the same name from the top level, copied out of this one, so that nothing - pydantic sits on the surface you import from `guideme`. `guideme.api.client` holds `Client` and - `AsyncClient` for callers who want to build requests themselves. They live one level down - rather than on `guideme.api` because re-exporting them would make `api` and `api.client` - import each other, and the gate fails an import cycle. -- `guideme.policy.resolve(answer, thresholds)` is the pure decision function. `spec/` holds - its JSON Schemas and 42 golden vectors, vendored from - [guideme-rust](https://github.com/pedro-pscunha/guideme-rust), which publishes the contract. - [`docs/contract.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/docs/contract.md) - says what every guideme SDK must satisfy and - [`docs/design.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/docs/design.md) - records the design and its sharp edges. -- `guideme.question` is where the question types are declared, and all of them are re-exported - above: `Question` is what `noul`, `choose`, `choose_among`, `score` and `score_levels` - return, and `NoulQuestion`, `ChoiceQuestion`, `ScoreQuestion`, `DetailedNoul`, - `DetailedChoice` and `DetailedScore` are the concrete ones. Inference covers most uses, so - reach for them when you need to annotate a question you are storing or passing on: a `dict` - is invariant, so a `dict[str, NoulQuestion]` is not a `dict[str, Question[bool]]` and the - annotation has to be written. Everything else in `guideme.question` is private. - `guideme.api` declares its own `Question`, `NoulQuestion`, `ChoiceQuestion` and - `ScoreQuestion`: same names, different classes. Those are the wire shapes the ones above - become on the way out, and you only meet them if you build requests by hand. The import - line says which you have. -- The scalars are validated once and never re-checked: `Probability` and `Confidence` hold the - unit-interval numbers on `Verdict`, `Ranked` and `Scored`, `Key` and `Rank` are what a runtime - rubric answers with, `Model` names the model to ask, and `ApiKey` carries the key without ever - printing it. The first four are `NewType` brands, so the guarantee is that only the wire mints - them, not that `Probability(2.0)` is rejected; it is not. +`guideme.policy`. Import them by their own path. The top level does not re-export them. They +have the same rule as the first tier: nothing in them is removed or renamed without a major +version and a `CHANGELOG.md` entry. Everything else in the package is private, whatever its +name looks like. + +- `guideme.api` is the exact wire mirror of `POST /v1/systemone` and `GET /v1/models`. + `guideme.api.__all__` lists the request and response models, its own `Usage`, and four + adapters between them and the core. +- `guideme.api.client` holds `Client` and `AsyncClient`, to build requests yourself. +- `guideme.policy.resolve(answer, thresholds)` is the pure decision function. `spec/` holds its + JSON Schemas and 42 golden vectors. + +`guideme.api` declares its own `Question`, `NoulQuestion`, `ChoiceQuestion`, `ScoreQuestion` and +`Usage`. These wire shapes are different classes from the ones above, and the import line tells +you which one you have. +[`docs/design.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/docs/design.md#sharp-edges) +explains these pairs and the other sharp edges. ## Other SDKs -Every guideme SDK is written from scratch in its own language and answers the same way, -because they all satisfy one contract: the wire schemas, the 42 golden policy vectors and the -interface shape that [guideme-rust](https://github.com/pedro-pscunha/guideme-rust) publishes -under `spec/` and states in +Each guideme SDK is written from scratch in its own language, and they all answer the same way. +They satisfy one contract: the wire schemas, the 42 golden policy vectors and the interface +shape. [guideme-rust](https://github.com/pedro-pscunha/guideme-rust) publishes it under `spec/` +and states it in [`docs/contract.md`](https://github.com/pedro-pscunha/guideme-rust/blob/main/docs/contract.md). | Language | Package | Repository | |---|---|---| -| Python | `guideme` | this repository | | Rust | [`guideme`](https://crates.io/crates/guideme) | [guideme-rust](https://github.com/pedro-pscunha/guideme-rust) | +| Python | `guideme` | this repository | +| TypeScript | `@guideme/sdk` (not yet on npm) | [guideme-typescript](https://github.com/pedro-pscunha/guideme-typescript) | -This repository vendors that `spec/` and records the commit it came from in `spec/SOURCE`; a -CI job fails when the copy drifts from the Rust repository's `main`. The span, event and -attribute names are shared too, so one dashboard reads both SDKs. - -## Environment - -| Variable | Meaning | -|---|---| -| `TYPESAFE_API_KEY` | required by `Guide.from_env()` and `AsyncGuide.from_env()` | -| `TYPESAFE_BASE_URL` | optional API origin override | -| `GUIDEME_MODEL` | optional model or alias; default `jev-latest` | +This repository copies `spec/` and records the source commit in `spec/SOURCE`. A CI job fails +when the copy drifts from the `main` branch of guideme-rust. The span, event and attribute names +are shared too, so one dashboard reads every SDK. ## Development -Tooling is managed by [mise](https://mise.jdx.dev), which pins `uv` and `gitleaks`; `uv` pins -everything else from `pyproject.toml` and `uv.lock`. - -``` -mise install # fetch the tools -mise run sync # install the locked environment -mise run check # fmt-check, gen-check, ruff, pyright, pylint, pytest, build, audit -mise run test # pytest alone -mise run hooks # point core.hooksPath at the tracked hooks in .githooks -``` - -The hooks are tracked, not generated: `mise run hooks` sets this repository's `core.hooksPath` -to `.githooks` and verifies it took effect. -[`AGENTS.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/AGENTS.md) says what -each stage runs. - -Library code is held to a strict checker set: `pyright` in strict mode, `ruff` with every rule -selected, and `pylint` with every check enabled. Tests are few and high-grade: property tests -for the policy laws, a real local HTTP server for the wire and retry contract, structural -tracing assertions, pyright files that must fail, and a drift guard that re-resolves every -golden vector. - -From a source distribution rather than a clone, `mise.toml` and `uv.lock` are not present, so -the suite runs with `uv run --group dev pytest`. - -Two opt-in tests hit the real API and are deselected by default: - -``` -TYPESAFE_API_KEY=… uv run --locked pytest -m live -``` - -Contributor rules live in -[`AGENTS.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/AGENTS.md). Report a -vulnerability privately, as -[`SECURITY.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/SECURITY.md) -describes, never in a public issue. +To set up a clone, run `mise install`, `mise run sync` and `mise run hooks`. `mise run check` is +the gate. +[`CONTRIBUTING.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/CONTRIBUTING.md) +and [`AGENTS.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/AGENTS.md) have the +rules. Report a vulnerability privately, as +[`SECURITY.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/SECURITY.md) tells, +and never in a public issue. ## License diff --git a/docs/contract.md b/docs/contract.md index da12807..9dfc231 100644 --- a/docs/contract.md +++ b/docs/contract.md @@ -216,7 +216,7 @@ for a pooled connection each get the whole of it — so an attempt that is slow phase outlasts the number written in the builder. Wrapping it to match would need a different wrapper for the synchronous and the asyncio surfaces and would change what cancellation means, which is a worse trade than saying so. It is documented on `GuideBuilder.timeout`, in the -README's configuration section, and here. Nothing on the wire depends on it. +README's **Retries and timeouts** section, and here. Nothing on the wire depends on it. Drift is caught rather than trusted. `mise run spec-check` clones guideme-rust, diffs its `spec/` against this one and fails on any difference except `spec/SOURCE`, which is provenance and has diff --git a/docs/design.md b/docs/design.md index 572cdcf..47f5821 100644 --- a/docs/design.md +++ b/docs/design.md @@ -258,6 +258,10 @@ Each module survives the test. second `Usage` — and it is the right way round: a caller reaching the top-level surface gets the value object, and the name they would otherwise collide with is in a module they only import when they are building requests by hand. +- **`Client` and `AsyncClient` live in `guideme.api.client`, not on `guideme.api`.** + Re-exporting them from `guideme.api` would make `api` and `api.client` import each other, + and the gate fails an import cycle. So a caller who builds requests by hand imports the + wire models from one module and the clients from the other. - **State is JSON-shaped.** Anything `json.dumps` accepts without a default hook. A dataclass goes through `dataclasses.asdict`, a pydantic model through `.model_dump()`. This is the one untyped value in the package, and it is serialised at the boundary. From 0fa6c34fd6db42948659963e020c489700732254 Mon Sep 17 00:00:00 2001 From: Pedro Cunha Date: Tue, 22 Sep 2026 19:17:52 -0300 Subject: [PATCH 2/2] Apply the README review fixes The review found places where a first read still stalled. The quick start used ticket, noul and the fallback before saying what they are, and it did not say that levels are ordered by declaration. Some terms drifted, and the ConfigError and TransportError rows missed cases. The quick start now defines the state, the three constructors, the level order and "unsure". The policy layers list the setter for each kind of question again. The score value states its scale from the TypeSafe API docs. The base_url row says that guideme follows redirects. Typing notes move under Lower layers, and so does the Thresholds sentence. Testing points to the retry rules in place of repeating them. The contributor lines on spec vendoring leave Other SDKs, and same-topic paragraphs merge. No code block changes, so tests/typing/readme.py still mirrors the README. --- README.md | 181 ++++++++++++++++++++++++++++-------------------------- 1 file changed, 95 insertions(+), 86 deletions(-) diff --git a/README.md b/README.md index b706d37..28daf24 100644 --- a/README.md +++ b/README.md @@ -19,8 +19,8 @@ uv add guideme # or: pip install guideme guideme needs Python 3.12 or newer. The package ships `py.typed`, so your type checker sees every annotation. Get an API key on the [keys page](https://console.typesafe.ai/keys) of the -TypeSafe console, and set it in the environment as `TYPESAFE_API_KEY`. To give the key in code, -use `Guide.builder().api_key(ApiKey("…")).build()`. +TypeSafe console. Set it in the environment as `TYPESAFE_API_KEY`. To give the key in code, use +`Guide.builder().api_key(ApiKey("…")).build()`. ## Quick start @@ -61,31 +61,41 @@ if guide.ask(score(Frustration, "How frustrated is the customer?"), ticket) >= ( prioritise() ``` +- `ticket` is the *state*: the data that you send with the question, here a support ticket. + `escalate()` and the `route_…()` functions are your own code. +- `noul(…)` asks a yes/no question (TypeSafe calls it a *noul*). `choose(…)` asks a choice, and + `score(…)` asks a score. - `Guide.from_env()` reads the key from `TYPESAFE_API_KEY`. - The value of a member is its *rubric*: the text that tells the model what the option or level means. The name of the member is its key on the wire. A docstring on a member is not rubric. +- Declare the levels from low to high. The first member is the lowest level, and `>=` compares + by this order. - pyright in strict mode makes sure that the `match` handles every option. If you add a department, the `match` is an error until you handle it. -- `sales` is the fallback: the answer when the choice is unsure. The default `min_confidence` - is `0.0`, and with it a choice is never unsure. That is why this choice sets `0.6`. +- `sales` is the fallback: the answer when the choice is *unsure*, that is, when its confidence + is less than `min_confidence`. The default `min_confidence` is `0.0`, so a choice is never + unsure and `sales` is never used as the fallback. That is why this choice sets `0.6`. -The gate type-checks this example as `tests/typing/readme.py`, so this page cannot drift from -what the package infers. Build one guide and share it: the Configuration section tells why. +Build one guide and share it: the Configuration section tells why. ## Questions -There are three kinds of question. A *yes/no question* (TypeSafe calls it a *noul*, and the -wire type is `noul`) gives a `bool`. A *choice* gives one of your options. A *score* gives one -level of your ordered scale. +There are three kinds of question. A yes/no question gives a `bool`. A choice gives one of your +options. A score gives one level of your ordered scale. Each constructor below gives the plain +answer. Add `.detail()` to a question to get the full reading in its place: the probabilities, +the confidence and the `unsure` flag. | Constructor | Asks | Plain answer | `.detail()` answer | |---|---|---|---| | `noul("…")` | a yes/no question | `bool` | `Verdict`: `verdict` (`"yes"`, `"no"`, `"unsure"`), `p` | | `choose(C, "…")`, `C` a `Choice` | a choice over 1 to 255 options | a member of `C` | `Ranked[C]`: `choice`, `confidence`, `unsure`, `probabilities` | -| `score(L, "…")`, `L` a `Levels` | a score over 2 to 10 levels, low to high | the most probable member of `L` | `Scored[L]`: the expected `value`, `level`, `confidence`, `unsure`, `distribution` | +| `score(L, "…")`, `L` a `Levels` | a score over 2 to 10 levels, low to high | the most probable member of `L` | `Scored[L]`: `value`, `level`, `confidence`, `unsure`, `distribution` | | `choose_among("…", options)` | a choice over 1 to 255 `{key: rubric}` pairs given at runtime | `Key`, the key | `Ranked[Key]` | | `score_levels("…", levels)` | a score over 2 to 10 level texts given at runtime | `Rank`, the index from 0 | `Scored[Rank]` | +The `value` of a score is the probability-weighted level number. The lowest level is 0, and the +value can land between two levels. + ```python team = guide.ask(choose_among("Which team?", {"billing": "Payments", "technical": "Bugs"}), ticket) rank = guide.ask(score_levels("How severe?", ["Cosmetic", "Degraded", "Blocking"]), ticket) @@ -96,23 +106,14 @@ The size limits come from the TypeSafe API. A `Choice` or `Levels` class out of Two members of one class cannot have the same rubric: Python makes the second an alias of the first, so that is a `ConfigError` too. -A yes/no question can say what yes and no mean with -`.criteria("what yes means", "what no means")`. The instructions can be a string or any -JSON-shaped value, so a question can name fields of structured state. - -The *state* is the data you send with the question: a ticket, a message, a record. It can be -anything JSON-shaped: a string, a number, a `dict`, a list, and nestings of them. Convert a -dataclass with `dataclasses.asdict` and a pydantic model with `.model_dump()`. A value that -`json` cannot write, such as `NaN` or `bytes`, raises `ConfigError` before anything is sent. +The question text (the `instructions` argument) can be a string or any JSON-shaped value, so it +can name fields of structured state. A yes/no question can also say what yes and no mean with +`.criteria("what yes means", "what no means")`. -`Question` is what the five constructors return. To annotate a question that you store or pass -on, use the concrete types `NoulQuestion`, `ChoiceQuestion`, `ScoreQuestion`, `DetailedNoul`, -`DetailedChoice` and `DetailedScore`. A `dict` is invariant, so a `dict[str, NoulQuestion]` is -not a `dict[str, Question[bool]]`. - -`Probability`, `Confidence`, `Key` and `Rank` are `NewType` brands. Only the wire creates them, -after it validates the value. `Probability(2.0)` in your code is not refused. `Model` names a -model. `ApiKey` holds the key and never prints it. +The state can be anything JSON-shaped: a string, a number, a `dict`, a list, and nestings of +them. Convert a dataclass with `dataclasses.asdict` and a pydantic model with `.model_dump()`. +A value that `json` cannot write, such as `NaN` or `bytes`, raises `ConfigError` before anything +is sent. `guide.models()` returns a `tuple[ModelInfo, ...]`: the models that your account can use, each with `name`, `description` and `release_date`. It is one `GET /v1/models` call, with no ask @@ -127,13 +128,14 @@ An answer is *unsure* when it is not certain enough under the thresholds: - Choice and score: `confidence < min_confidence` is unsure. With the default `0.0`, no answer is unsure. -A `Policy` is a set of thresholds, and each field is optional. You set it on the question -(`.yes_above(p)`, `.no_below(p)`, `.min_confidence(c)`, `.with_policy(Policy(…))`), on the guide -(`Guide.builder().policy(…)`, or `guide.with_policy(…)` for a copy), or not at all. The question -wins over the guide, and the guide wins over the defaults. A `Policy` settled against the -defaults is a `Thresholds`, the input of `guideme.policy.resolve` and of every golden vector. +A `Policy` is a set of thresholds, and each field is optional. You can set it at two layers: + +- On a question: `.yes_above(p)` and `.no_below(p)` on a yes/no question, `.min_confidence(c)` + on a choice or a score, and `.with_policy(Policy(…))` on any question. +- On a guide: `Guide.builder().policy(…)`, or `guide.with_policy(…)` for a copy. -When an answer is unsure, guideme goes down the *unsure ladder*: +The question wins over the guide, and the guide wins over the defaults. When an answer is unsure, +guideme goes down the *unsure ladder*: 1. The `.otherwise(value)` of the question. 2. The `fallback(…)` member of the `Choice`. A `Levels` class has no fallback member, so for a @@ -141,8 +143,9 @@ When an answer is unsure, guideme goes down the *unsure ladder*: 3. `UnsureError`, which names the question and the threshold that it missed. `.detail()` skips the ladder and drops any `.otherwise(…)`. It never raises `UnsureError` and -gives you the full reading to decide yourself. `Policy` is a frozen dataclass, so a house policy -can be a module constant: +gives you the full reading to decide yourself. + +`Policy` is a frozen dataclass, so you can keep one in a module constant: ```python CAUTIOUS = Policy(yes_above=0.7, no_below=0.3) @@ -202,12 +205,12 @@ Examples: My card was charged twice; Where is my refund? Not this option: The dashboard is down ``` -A rubric with no examples sends the same bytes as before. Examples render in the order that you -write them, and this text is part of the published contract. The same values work at runtime, -in `choose_among("…", {"billing": option(…)})` and `score_levels("…", [level(…), …])`. +A rubric with no examples sends only its text. Examples render in the order that you write them, +and this text is part of the published contract. The same values work at runtime, in +`choose_among("…", {"billing": option(…)})` and `score_levels("…", [level(…), …])`. -A yes and a no are two options of one question, and a vague pair is the easiest to get wrong. So -`.criteria(…)` also takes an `option(…)` for each side: +`.criteria(…)` takes an `option(…)` for the yes and one for the no. A vague yes/no pair is easy +to get wrong, and examples help most there: ```python urgent = noul("Is this ticket urgent?").criteria( @@ -219,25 +222,24 @@ urgent = noul("Is this ticket urgent?").criteria( [`docs/design.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/docs/design.md#decisions) records a measured case where these examples change a wrong yes into a correct no. -Each of these raises `ConfigError` where you write the rubric: - -- An empty clause, such as `examples=[]`. To say there are none, leave the clause out. -- A clause given as one string, such as `examples="refund"`, in place of a list. -- A blank entry, the same entry twice in one clause, or a newline or carriage return in an entry. -- One string as an example of two options (or of the yes and the no, or of two levels). -- One string as an example and a counterexample of the same option. -- A counterexample on a level, or a `fallback(…)` in `choose_among`, `score_levels` or - `.criteria(…)`. -- Examples on a blank rubric. +A rubric obeys these rules. A rubric that breaks one raises `ConfigError` where you write it: -Some overlaps are legal. One string can be an example of one option and a counterexample of -another: that is how you tell two similar options apart. An entry can contain `"; "`, and the -rubric text itself can contain newlines. A blank rubric with no examples is legal too. +- Leave a clause out to say there are none. An empty clause, such as `examples=[]`, is refused. +- Give a clause as a list, not as one string such as `examples="refund"`. +- An entry is not blank, is on one line (no newline or carriage return), and appears once in its + clause. An entry can contain `"; "`. +- One string cannot be an example of two options (or of the yes and the no, or of two levels). +- One string cannot be an example and a counterexample of the same option. It can be an example + of one option and a counterexample of another: that is how you tell two similar options apart. +- A level has no counterexamples, and `choose_among`, `score_levels` and `.criteria(…)` take no + `fallback(…)`. +- Examples or counterexamples on a blank rubric are refused. A blank rubric alone is legal, and + the rubric text itself can contain newlines. ## Several questions in one request -A tuple of questions is also a question. So is a list or a dict, and they nest. The answer has -the same shape and comes from one request and one span. Each question keeps its own policy. +`ask` also takes a tuple, a list or a dict of questions, and they nest. The answer has the same +shape and comes from one request and one span. Each question keeps its own policy. ```python urgent, dept, mood, flags = guide.ask( @@ -261,8 +263,8 @@ whole call fails. So put `.otherwise(…)` or `.detail()` on each question that unsure. Any nesting works at runtime. Your type checker infers a type for one question, a list and a -dict. It also infers a tuple of up to eight questions, or of up to seven followed by one -list or dict (the shape above). +dict. It also infers a tuple of up to eight questions, or of up to seven followed by one list or +dict (the shape above). ## The receipt @@ -293,11 +295,11 @@ guideme SDK, and the value of `error.type` on the failed span. | `InvalidError` | `invalid` | HTTP 422. `.detail` is the response body. | | `RateLimitedError` | `rate_limited` | HTTP 429 after the retries, or a `retry-after` that is too long to wait. `.retry_after` holds it. | | `OverloadedError` | `overloaded` | HTTP 529, on the same terms. `.retry_after` holds it. | -| `TransportError` | `transport` | A connection, TLS or timeout failure. | +| `TransportError` | `transport` | A connection, TLS or timeout failure, or a disconnect during the response. | | `UnexpectedStatusError` | `unexpected_status` | A status that the contract does not define. | | `ProtocolError` | `protocol` | The response breaks the contract: a body that does not decode, a wrong answer kind, an option or level that is not in the rubric, a probability outside 0..1. | | `UnsureError` | `unsure` | The policy said unsure and the ladder had no value. | -| `ConfigError` | `config` | A mistake in your code, raised where you write it: bad thresholds, no key, an empty batch, state that is not JSON, a rubric that breaks a rule, a bad `events(…)`, a timeout that is not positive, negative retries or backoff, a `base_url` that holds credentials. | +| `ConfigError` | `config` | A mistake in your code, raised before anything is sent. For example: bad thresholds, no key, an empty batch, state or instructions that `json` cannot write, a rubric that breaks a rule, two fallback members in one `Choice`, a bad `events(…)`, a timeout that is not positive, negative retries or backoff, a malformed `base_url` or one that holds credentials, a timeout next to a transport, a transport given to the wrong build, `with_policy(…)` on a closed guide. | ## Retries and timeouts @@ -310,17 +312,17 @@ handshake that did not complete. That request did not reach a server, so nothing A call makes at most `max_retries + 1` attempts, whatever the mix of failures. A disconnect during the response is not resent: the API can have answered, and a second request -pays for the same answer twice. No timeout is resent, of any phase, a connect timeout included. -Each of these raises `TransportError` at once. +pays for the same answer twice. guideme does not resend a timeout, in any phase. This includes a +connect timeout. Each of these raises `TransportError` at once. [`docs/design.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/docs/design.md#decisions) gives the reasons. **`timeout` is not a deadline for the attempt.** `httpx` gives the full value to each phase: connect, write, read, and the wait for a pooled connection. Thus one slow attempt can take more than the timeout. The worst case for a call is `max_retries + 1` attempts of several phases each, -plus the backoff between them. The Rust SDK has one deadline for the full attempt, and +plus the backoff between them. The Rust SDK differs here, as [`docs/contract.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/docs/contract.md) -records this divergence. +records. ## Testing your code @@ -350,9 +352,8 @@ def test_an_urgent_ticket_is_prioritised() -> None: ``` `q0` is the first question in encounter order, and a batch of three uses `q0`, `q1` and `q2`. -To test a failure, raise an exception in the handler. guideme resends an `httpx.ConnectError` -inside the retry budget. It does not resend an `httpx.ReadTimeout`, an `httpx.ConnectTimeout` or -an `httpx.RemoteProtocolError`. For `AsyncGuide`, give the same `httpx.MockTransport` to +To test a failure, raise an `httpx` exception in the handler. The Retries and timeouts section +says which ones guideme resends. For `AsyncGuide`, give the same `httpx.MockTransport` to `async_transport(…)` and call `build_async()`. A `MockTransport` is both kinds of transport. ## Observability @@ -373,12 +374,12 @@ provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter())) trace.set_tracer_provider(provider) ``` -Each request is one `guideme.ask` span with the OpenTelemetry GenAI fields, and one HTTP client +Each `ask` is one `guideme.ask` span with the OpenTelemetry GenAI fields, and one HTTP client span per attempt below it. A retry adds a `guideme.retry` event. Each question adds a `guideme.answer` event with the outcome, the probability or confidence and the thresholds. Each answer is also an OTLP log record at `INFO`, and each retry one at `WARN`, with the trace id and span id. To receive them, install a `LoggerProvider`. The state is not recorded unless you set -`record_state(True)`. The API key is never recorded. +`record_state(True)`, and the API key is never recorded. `events(…)` selects the signal for answers and retries: `"span"`, `"log"` or `"both"` (the default). If you export traces and logs to one backend, set `"span"` or `"log"` to store each @@ -396,7 +397,7 @@ it all against the live API, with a collector that prints what arrives. | Setting | Default | What it does | |---|---|---| | `api_key(ApiKey(…))` | none, required | The API key. | -| `base_url(…)` | `https://api.typesafe.ai` | The API origin. It cannot hold credentials. | +| `base_url(…)` | `https://api.typesafe.ai` | The API origin. It cannot hold credentials. Give the final https origin. guideme follows redirects. | | `model(Model(…))` | `jev-latest` | The model or alias to ask. | | `policy(Policy(…))` | the defaults | The policy of the guide. The policy of a question wins over it. | | `max_retries(n)` | `3` | Resends per call. `0` never resends. | @@ -419,7 +420,8 @@ is for `build_async()`. The wrong pair is a `ConfigError` too. Build one guide per process and share it. A guide holds a connection pool, and it is safe to use from many threads or tasks at once. A guide per request also works, but it opens a pool for each -request. +request. `guide.with_policy(…)` returns a second guide over the same pool. The pool counts its +guides: after you close one, the other can still ask, and the pool closes with the last guide. ## Sync and async @@ -434,13 +436,21 @@ async with AsyncGuide.from_env() as guide: For the synchronous version, use `Guide`, `with`, and no `await`. `Guide.builder()` and `AsyncGuide.builder()` return the same `GuideBuilder`: `.build()` gives a `Guide`, and `.build_async()` gives an `AsyncGuide`. A guide closes at the end of its `with` block. Outside a -block, call `guide.close()`. `guide.with_policy(…)` returns a second guide over the same pool. -The pool counts its guides: after you close one, the other can still ask, and the pool closes -with the last guide. If you close one guide twice, the second close does nothing. +block, call `guide.close()`. If you close one guide twice, the second close does nothing. Do not +ask through a closed guide. ## Lower layers -The `guideme` package re-exports everything above, and `guideme.__all__` is that list. +The `guideme` package re-exports everything above, and `guideme.__all__` is that list. Some of +those names need a note: + +- `Question` is what the five constructors return. To annotate a question that you store or + pass on, use the concrete types `NoulQuestion`, `ChoiceQuestion`, `ScoreQuestion`, + `DetailedNoul`, `DetailedChoice` and `DetailedScore`. A `dict` is invariant, so a + `dict[str, NoulQuestion]` is not a `dict[str, Question[bool]]`. +- `Probability`, `Confidence`, `Key` and `Rank` are `NewType` brands. Only the wire creates + them, after it validates the value, so `Probability(2.0)` in your code is not refused. +- `Model` names a model. `ApiKey` holds the key and never prints it. Three modules are a second supported tier: `guideme.api`, `guideme.api.client` and `guideme.policy`. Import them by their own path. The top level does not re-export them. They @@ -452,21 +462,22 @@ name looks like. `guideme.api.__all__` lists the request and response models, its own `Usage`, and four adapters between them and the core. - `guideme.api.client` holds `Client` and `AsyncClient`, to build requests yourself. -- `guideme.policy.resolve(answer, thresholds)` is the pure decision function. `spec/` holds its - JSON Schemas and 42 golden vectors. +- `guideme.policy.resolve(answer, thresholds)` is the pure decision function. It takes a + `Thresholds`: a `Policy` settled against the defaults. `spec/` holds its JSON Schemas and 42 + golden vectors, which use the same `Thresholds`. -`guideme.api` declares its own `Question`, `NoulQuestion`, `ChoiceQuestion`, `ScoreQuestion` and -`Usage`. These wire shapes are different classes from the ones above, and the import line tells -you which one you have. +`guideme.api` also declares a `Question`, `NoulQuestion`, `ChoiceQuestion`, `ScoreQuestion` and +`Usage` of its own: wire shapes, different classes from the ones above. You meet them only when +you build requests by hand, and [`docs/design.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/docs/design.md#sharp-edges) -explains these pairs and the other sharp edges. +explains these pairs. ## Other SDKs -Each guideme SDK is written from scratch in its own language, and they all answer the same way. -They satisfy one contract: the wire schemas, the 42 golden policy vectors and the interface -shape. [guideme-rust](https://github.com/pedro-pscunha/guideme-rust) publishes it under `spec/` -and states it in +Each guideme SDK is written from scratch in its own language, and they all turn one reading into +the same answer. They satisfy one contract: the wire schemas, the 42 golden policy vectors and +the interface shape. [guideme-rust](https://github.com/pedro-pscunha/guideme-rust) publishes it +under `spec/` and states it in [`docs/contract.md`](https://github.com/pedro-pscunha/guideme-rust/blob/main/docs/contract.md). | Language | Package | Repository | @@ -475,9 +486,7 @@ and states it in | Python | `guideme` | this repository | | TypeScript | `@guideme/sdk` (not yet on npm) | [guideme-typescript](https://github.com/pedro-pscunha/guideme-typescript) | -This repository copies `spec/` and records the source commit in `spec/SOURCE`. A CI job fails -when the copy drifts from the `main` branch of guideme-rust. The span, event and attribute names -are shared too, so one dashboard reads every SDK. +The span, event and attribute names are shared too, so one dashboard reads every SDK. ## Development