diff --git a/CHANGELOG.md b/CHANGELOG.md index bc1972d..5f47869 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,7 +2,12 @@ ## Unreleased -Nothing yet. +### Documentation + +- The README is rewritten in plain English with the headings, the order and the terms that + every guideme SDK now shares. Design rationale and measurements it carried are in + `docs/design.md`, and the contributor notes are in `CONTRIBUTING.md`. The TypeScript SDK is + listed under **Other SDKs**. No behaviour changes. ## 0.2.0 — 2026-09-22 diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index a4e6856..7f1ee27 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -2,8 +2,8 @@ guideme makes a TypeSafe Jev judgment usable as Python control flow. It is one distribution, `guideme`. Its public surface is what `src/guideme/__init__.py` exports, plus a second -supported tier that callers import by their own path: the modules `guideme.api` and -`guideme.policy`, and the single name `guideme.question.Question`. +supported tier that callers import by their own path: the modules `guideme.api`, +`guideme.api.client` and `guideme.policy`. Everything else in the package is private. `AGENTS.md` and the README's **Lower layers** section say what each tier promises. @@ -38,6 +38,9 @@ mise run test # while you work mise run check # the full gate; the pre-push hook runs it too ``` +A source distribution carries the tests but not `mise.toml` or `uv.lock`, so from one rather +than a clone, run the suite with `uv run --group dev pytest`. + The suite is small on purpose and `AGENTS.md` says what a new test may be. A pull request that adds a mock of `policy`, of `Guide` or of the transport will be asked to replace it with a real local server, a property test, or a typing proof. diff --git a/README.md b/README.md index 0c6df9e..28daf24 100644 --- a/README.md +++ b/README.md @@ -1,12 +1,30 @@ -# guideme +# guideme (Python) -Judgments from [TypeSafe Jev](https://docs.typesafe.ai) that read like Python control flow. +[![PyPI](https://img.shields.io/pypi/v/guideme.svg)](https://pypi.org/project/guideme/) +[![Python](https://img.shields.io/pypi/pyversions/guideme.svg)](https://pypi.org/project/guideme/) +[![license](https://img.shields.io/pypi/l/guideme.svg)](https://github.com/pedro-pscunha/guideme-python#license) -A yes/no question is an `if`. A choice is an exhaustive `match` over your own enum. A score is -a comparison against your own ordered levels. Thresholds, unsure bands and fallbacks are -explicit and composable. Every request is one span. The decision logic is pure and its -contract is published under `spec/`, so every guideme SDK, in any language, answers the same -way. +guideme sends a question and your state to [TypeSafe Jev](https://docs.typesafe.ai), the +TypeSafe model that gives judgments. It gives back the answer as a normal Python value: a +`bool`, a member of your own enum, or one of your own ordered levels. Your code then acts on the +answer with an `if`, a `match` or a comparison. Unlike the official +[`typesafe-sdk`](https://docs.typesafe.ai/sdk/python), which mirrors the API, guideme turns an +answer into control flow and calls the API itself. + +## Install + +```sh +uv add guideme # or: pip install guideme +``` + +guideme needs Python 3.12 or newer. The package ships `py.typed`, so your type checker sees +every annotation. Get an API key on the [keys page](https://console.typesafe.ai/keys) of the +TypeSafe console. Set it in the environment as `TYPESAFE_API_KEY`. To give the key in code, use +`Guide.builder().api_key(ApiKey("…")).build()`. + +## Quick start + +This example asks three questions about one support ticket. ```python from guideme import Choice, Guide, Levels, choose, fallback, noul, score @@ -43,74 +61,120 @@ if guide.ask(score(Frustration, "How frustrated is the customer?"), ticket) >= ( prioritise() ``` -The value of each member is the rubric the model reads. The member's name is the wire key. A -docstring on a member is documentation, not a rubric. pyright in strict mode enforces that -every option is handled, so adding a department turns the `match` above into an error until -you handle it. `from_env()` reads `TYPESAFE_API_KEY`, and `sales`, marked with `fallback(…)`, -is also the answer when confidence is below the floor. That floor is `min_confidence`, and it -is `0.0` unless you set it, which is why the choice above asks for `0.6`: with the default a -choice is never unsure and a `fallback(…)` member can never be reached. - -Two members may not share a rubric: Python would make the second an alias of the first, so a -repeat is a `ConfigError` on the class statement rather than a rubric quietly one option -short. `fallback(…)` marks a `Choice` member and only a `Choice` member; a `Levels` is -ordered, so the level to fall back to when a score is unsure is `.otherwise(level)` on the -question. - -That example is `tests/typing/readme.py`, which the gate type-checks with an `assert_type` on -every inferred answer type, so what is on this page cannot drift from what the package infers. -The three asks above are held to it by their use sites instead: the `match` is exhaustive and -the `>=` is between two levels of one scale. +- `ticket` is the *state*: the data that you send with the question, here a support ticket. + `escalate()` and the `route_…()` functions are your own code. +- `noul(…)` asks a yes/no question (TypeSafe calls it a *noul*). `choose(…)` asks a choice, and + `score(…)` asks a score. +- `Guide.from_env()` reads the key from `TYPESAFE_API_KEY`. +- The value of a member is its *rubric*: the text that tells the model what the option or level + means. The name of the member is its key on the wire. A docstring on a member is not rubric. +- Declare the levels from low to high. The first member is the lowest level, and `>=` compares + by this order. +- pyright in strict mode makes sure that the `match` handles every option. If you add a + department, the `match` is an error until you handle it. +- `sales` is the fallback: the answer when the choice is *unsure*, that is, when its confidence + is less than `min_confidence`. The default `min_confidence` is `0.0`, so a choice is never + unsure and `sales` is never used as the fallback. That is why this choice sets `0.6`. + +Build one guide and share it: the Configuration section tells why. + +## Questions + +There are three kinds of question. A yes/no question gives a `bool`. A choice gives one of your +options. A score gives one level of your ordered scale. Each constructor below gives the plain +answer. Add `.detail()` to a question to get the full reading in its place: the probabilities, +the confidence and the `unsure` flag. + +| Constructor | Asks | Plain answer | `.detail()` answer | +|---|---|---|---| +| `noul("…")` | a yes/no question | `bool` | `Verdict`: `verdict` (`"yes"`, `"no"`, `"unsure"`), `p` | +| `choose(C, "…")`, `C` a `Choice` | a choice over 1 to 255 options | a member of `C` | `Ranked[C]`: `choice`, `confidence`, `unsure`, `probabilities` | +| `score(L, "…")`, `L` a `Levels` | a score over 2 to 10 levels, low to high | the most probable member of `L` | `Scored[L]`: `value`, `level`, `confidence`, `unsure`, `distribution` | +| `choose_among("…", options)` | a choice over 1 to 255 `{key: rubric}` pairs given at runtime | `Key`, the key | `Ranked[Key]` | +| `score_levels("…", levels)` | a score over 2 to 10 level texts given at runtime | `Rank`, the index from 0 | `Scored[Rank]` | -## Install +The `value` of a score is the probability-weighted level number. The lowest level is 0, and the +value can land between two levels. -```sh -uv add guideme +```python +team = guide.ask(choose_among("Which team?", {"billing": "Payments", "technical": "Bugs"}), ticket) +rank = guide.ask(score_levels("How severe?", ["Cosmetic", "Degraded", "Blocking"]), ticket) ``` -or +The size limits come from the TypeSafe API. A `Choice` or `Levels` class out of range raises +`ConfigError` on the class statement, and a runtime rubric raises it on the constructor call. +Two members of one class cannot have the same rubric: Python makes the second an alias of the +first, so that is a `ConfigError` too. -```sh -pip install guideme -``` +The question text (the `instructions` argument) can be a string or any JSON-shaped value, so it +can name fields of structured state. A yes/no question can also say what yes and no mean with +`.criteria("what yes means", "what no means")`. -Python 3.12 or newer. The package ships `py.typed`, so your checker sees every annotation. +The state can be anything JSON-shaped: a string, a number, a `dict`, a list, and nestings of +them. Convert a dataclass with `dataclasses.asdict` and a pydantic model with `.model_dump()`. +A value that `json` cannot write, such as `NaN` or `bytes`, raises `ConfigError` before anything +is sent. -Set `TYPESAFE_API_KEY` in the environment, or pass a key to -`Guide.builder().api_key(ApiKey("…")).build()`. Keys come from the TypeSafe console, on its -[keys page](https://console.typesafe.ai/keys). +`guide.models()` returns a `tuple[ModelInfo, ...]`: the models that your account can use, each +with `name`, `description` and `release_date`. It is one `GET /v1/models` call, with no ask +span. It is retried like an ask, so a `429` while your process starts does not stop the start. -guideme is not the official TypeSafe SDK. That one is -[`typesafe-sdk`](https://docs.typesafe.ai/sdk/python), which mirrors the API: you send -questions and read answers. guideme adds the layer above it, turning an answer into control -flow — your own enums as the option set, thresholds and an unsure ladder as policy, one span -per request — and talks to the API itself rather than wrapping that package. +## When the model is not sure -## Three kinds of question, five constructors +An answer is *unsure* when it is not certain enough under the thresholds: -| Constructor | Sends | Plain output | `.detail()` output | -|---|---|---|---| -| `noul("…")` | a yes/no question | `bool` | `Verdict` with the label and the probability | -| `choose(C, "…")` where `C` is a `Choice` | a choice over `C`'s 1 to 255 members | `C` | `Ranked[C]` with confidence and `probabilities`, the whole distribution | -| `score(L, "…")` where `L` is a `Levels` | a score over `L`'s 2 to 10 levels, low to high | `L`, the most probable level | `Scored[L]` with the expected `value`, the level, confidence and `distribution` | -| `choose_among("…", options)` | a choice over 1 to 255 runtime `{key: rubric}` pairs | `Key` | `Ranked[Key]` | -| `score_levels("…", levels)` | a score over 2 to 10 runtime level descriptions | `Rank` | `Scored[Rank]` | +- Yes/no question: `p >= yes_above` is yes, `p <= no_below` is no, and between the two is + unsure. The defaults are `0.5` and `0.5`, so no answer is unsure. +- Choice and score: `confidence < min_confidence` is unsure. With the default `0.0`, no answer + is unsure. + +A `Policy` is a set of thresholds, and each field is optional. You can set it at two layers: + +- On a question: `.yes_above(p)` and `.no_below(p)` on a yes/no question, `.min_confidence(c)` + on a choice or a score, and `.with_policy(Policy(…))` on any question. +- On a guide: `Guide.builder().policy(…)`, or `guide.with_policy(…)` for a copy. + +The question wins over the guide, and the guide wins over the defaults. When an answer is unsure, +guideme goes down the *unsure ladder*: + +1. The `.otherwise(value)` of the question. +2. The `fallback(…)` member of the `Choice`. A `Levels` class has no fallback member, so for a + score use `.otherwise(level)`. +3. `UnsureError`, which names the question and the threshold that it missed. + +`.detail()` skips the ladder and drops any `.otherwise(…)`. It never raises `UnsureError` and +gives you the full reading to decide yourself. + +`Policy` is a frozen dataclass, so you can keep one in a module constant: + +```python +CAUTIOUS = Policy(yes_above=0.7, no_below=0.3) + +guide = Guide.builder().api_key(key).policy(CAUTIOUS).build() +strict = guide.with_policy(Policy(min_confidence=0.8)) + +reading = guide.ask(noul("Is this about billing?").detail(), ticket) +match reading.verdict: + case "yes": + billing() + case "no": + other() + case "unsure": + review(reading.p) -Those size limits are the API's, and guideme checks them where you write the rubric: a `Choice` -or `Levels` class outside the range is a `ConfigError` on the class statement, and a runtime -rubric is one on the constructor call. +picked = strict.ask(choose(Department, "Which team?").detail(), ticket) +``` -A noul can carry `.criteria("what yes means", "what no means")`. Instructions accept a string -or any JSON-shaped value, so a question can reference structured data by field name the way -the TypeSafe docs describe. +`strict` shares the connection pool of `guide` and keeps `CAUTIOUS`, with `min_confidence` set +over it. Thus the choice is unsure below `0.8`, and the yes/no question still uses `0.7 / 0.3`. -### Examples in a rubric +## Examples and counterexamples -Two alternatives that read alike are told apart by showing inputs rather than by describing -harder. `option(…)` takes the inputs that belong to an alternative and the ones that belong -somewhere else, `level(…)` takes the inputs that score at that level, and `fallback(…)` is an -`option(…)` that also marks the unsure member. All three kinds of question take them: a noul's -`.criteria(…)` accepts an `option(…)` for the yes and for the no. +Two options that read alike are easier to tell apart with inputs than with a longer rubric. An +*example* is an input that belongs to an option. A *counterexample* is an input that does not. +`option(…)` takes both. `level(…)` takes examples only, because on an ordered scale an input +that does not belong at one level belongs at another. `fallback(…)` is an `option(…)` that also +marks the fallback member. ```python from guideme import Choice, Levels, fallback, level, option @@ -132,8 +196,8 @@ class Severity(Levels): blocking = level("No workaround exists", examples=["cannot log in", "data loss"]) ``` -The member's value is still the bare rubric; the examples are composed into it only in the -request, as +The value of the member stays the bare rubric. guideme adds the examples only in the request, +as this text: ```text Payments, invoicing, refunds @@ -141,13 +205,12 @@ Examples: My card was charged twice; Where is my refund? Not this option: The dashboard is down ``` -So a rubric with no examples sends exactly what it sent before, and the same strings work in -`choose_among("…", {"billing": option(…)})` and `score_levels("…", [level(…), …])`. Examples -and counterexamples render in the order they are written, always: that order is part of the -published contract. +A rubric with no examples sends only its text. Examples render in the order that you write them, +and this text is part of the published contract. The same values work at runtime, in +`choose_among("…", {"billing": option(…)})` and `score_levels("…", [level(…), …])`. -A yes and a no are two alternatives of one question, so they take examples too, and this is -where they pay best — a vague pair is the easiest thing to get wrong: +`.criteria(…)` takes an `option(…)` for the yes and one for the no. A vague yes/no pair is easy +to get wrong, and examples help most there: ```python urgent = noul("Is this ticket urgent?").criteria( @@ -156,89 +219,27 @@ urgent = noul("Is this ticket urgent?").criteria( ) ``` -Asked about a nightly export job that has been failing since Tuesday while the numbers are -pulled by hand, a plain `Urgent` / `Not urgent` answers yes at 0.75. The criteria above answer -no at 0.17, because one of the not-urgent examples is what the ticket describes. - -A string may be an example of one option and a counterexample of another. That is the point -when two options are confusable, and it is the one overlap that stays legal. Offering the same -string as an example of two options, or as both an example and a counterexample of the same -option, says an input belongs where it cannot, so each is refused. +[`docs/design.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/docs/design.md#decisions) +records a measured case where these examples change a wrong yes into a correct no. -`level(…)` has no counterexamples, because "not this option" means nothing on an ordered -scale — an input that does not belong at one level scores at another. - -Leave a clause out to say there is none. An empty one written out — `examples=[]` or -`examples=()` — says nothing, so it is refused as the mistake it is, along with a clause given -as one string rather than a list of them (`examples="refund"` would otherwise be six one-letter -examples), a blank entry, a repeat within one clause, a newline or carriage return inside an -entry, a counterexample on a level, a `fallback(…)` given to `choose_among`, `score_levels` or -`.criteria(…)`, and the two contradictions above. Each is a `ConfigError` where the rubric is -written. - -Entries go on one line each, so a newline inside one would read as a clause you never wrote. -`"; "` inside an entry is fine — `"card declined; retry failed"` is ordinary prose, and it -changes how many examples a reader sees rather than which clause they are in. The rubric text -itself may still contain newlines; only the entries are restricted. - -Attaching examples to a blank rubric is refused too, because they describe something that is not -there. A blank rubric on its own is not: it means what it has always meant, and adding examples -to the language does not make an old declaration an error. - -The state is anything JSON-shaped: a text literal, a `dict`, a list of them. A dataclass goes -through `dataclasses.asdict`, a pydantic model through `.model_dump()`. - -## Policy - -Thresholds decide how a probability or a confidence becomes an answer. They form a patch that -merges from the question, over the guide, over the package defaults. A `Policy` is that patch, -with every field optional; settling one against the defaults gives a `Thresholds`, which is -what `guideme.policy.resolve` takes and what every golden vector is written against. - -| Layer | How to set | Wins over | -|---|---|---| -| question | `.yes_above(p)`, `.no_below(p)` on nouls; `.min_confidence(c)` on choice and score; `.with_policy(Policy(…))` on any | guide | -| guide | `Guide.builder().policy(…)`, or `guide.with_policy(…)` for a scoped copy | defaults | -| defaults | `yes_above 0.5`, `no_below 0.5`, `min_confidence 0.0` | nothing | +A rubric obeys these rules. A rubric that breaks one raises `ConfigError` where you write it: -The rules: +- Leave a clause out to say there are none. An empty clause, such as `examples=[]`, is refused. +- Give a clause as a list, not as one string such as `examples="refund"`. +- An entry is not blank, is on one line (no newline or carriage return), and appears once in its + clause. An entry can contain `"; "`. +- One string cannot be an example of two options (or of the yes and the no, or of two levels). +- One string cannot be an example and a counterexample of the same option. It can be an example + of one option and a counterexample of another: that is how you tell two similar options apart. +- A level has no counterexamples, and `choose_among`, `score_levels` and `.criteria(…)` take no + `fallback(…)`. +- Examples or counterexamples on a blank rubric are refused. A blank rubric alone is legal, and + the rubric text itself can contain newlines. -- Noul: `p >= yes_above` is yes, `p <= no_below` is no, strictly between is unsure. With the - defaults there is no unsure band. -- Choice and score: `confidence < min_confidence` is unsure. With the default there is never - an unsure answer. +## Several questions in one request -When an answer is unsure, resolution goes down a ladder: `.otherwise(value)` on the question, -then the enum's `fallback(…)` member, then `UnsureError` naming the question and the boundary -it missed. `.detail()` skips the ladder and hands you the reading to decide yourself. - -`Policy` is a frozen dataclass, so a house policy is a module constant: - -```python -CAUTIOUS = Policy(yes_above=0.7, no_below=0.3) - -guide = Guide.builder().api_key(key).policy(CAUTIOUS).build() -strict = guide.with_policy(Policy(min_confidence=0.8)) - -reading = guide.ask(noul("Is this about billing?").detail(), ticket) -match reading.verdict: - case "yes": - billing() - case "no": - other() - case "unsure": - review(reading.p) - -picked = strict.ask(choose(Department, "Which team?").detail(), ticket) -``` - -`strict` shares the connection pool and inherits `CAUTIOUS`, with `min_confidence` patched over -it, so the choice above is unsure below `0.8` while the noul still reads against `0.7 / 0.3`. - -## Several judgments, one request - -A tuple of questions is a question. So is a list or a dict, and they nest. The answer has the -same shape, from one request and one span. Each question keeps its own policy. +`ask` also takes a tuple, a list or a dict of questions, and they nest. The answer has the same +shape and comes from one request and one span. Each question keeps its own policy. ```python urgent, dept, mood, flags = guide.ask( @@ -256,46 +257,19 @@ if urgent or mood.value > 1.5 or flags["vip"]: prioritise() ``` -Question ids are `q0..qN` in encounter order, which for a dict is insertion order. They appear -on the wire, in errors and in events. A batch is atomic: one answer that cannot be resolved -fails the whole call, so put `.otherwise(…)` or `.detail()` on the questions that may come -back unsure. +The question ids are `q0..qN` in encounter order (insertion order for a dict). They appear on +the wire, in errors and in events. A batch is atomic: if one answer cannot be resolved, the +whole call fails. So put `.otherwise(…)` or `.detail()` on each question that can come back +unsure. -Any nesting works at runtime. The forms your checker infers a type for are a single question, -a list, a dict, a tuple of up to eight questions, and a tuple of up to seven followed by one -list or dict, which is the shape above. +Any nesting works at runtime. Your type checker infers a type for one question, a list and a +dict. It also infers a tuple of up to eight questions, or of up to seven followed by one list or +dict (the shape above). -## Sync and async +## The receipt -`Guide` and `AsyncGuide` have the same surface over the same pure core. The difference is the -`await` and the `httpx` client underneath. - -```python -async with AsyncGuide.from_env() as guide: - verdict: Verdict = await guide.ask(noul("Is this about billing?").detail(), ticket) -``` - -The synchronous version is the same two lines with `Guide`, `with` and no `await`. -`Guide.builder()` and `AsyncGuide.builder()` return the same `GuideBuilder`; `.build()` gives -the synchronous guide and `.build_async()` the asynchronous one. Leaving the block closes the -guide, and `guide.close()` does the same thing by hand for a guide that outlives any block. - -A guide holds a connection pool, so build one and share it: both kinds are safe to use from -several threads or several tasks at once, and one guide asking concurrently is what the pool -is for. Building one per request works but opens a pool per request, which is the cost the -pool exists to avoid. `guide.with_policy(…)` returns a second guide over the *same* pool, and -the two are counted, so closing either leaves the other able to ask and the pool closes when -the last of them does. Close each guide once. - -Both guides also answer `models()`, which returns a `tuple[ModelInfo, ...]`: the models the -account may use, each with its `name`, `description` and `release_date`. It is one call to -`GET /v1/models` and gets no ask span of its own. It is retried on exactly the terms an ask -is, so a `429` while your process is starting up does not fail the start. - -## What a request cost - -`ask_with_receipt` is `ask` with the response's own numbers kept. It takes the same shapes and -infers the same types; `ask` is this call followed by `.answer`. +`ask_with_receipt` is `ask` that also keeps the numbers of the response. It takes the same +shapes and infers the same types. `ask` is `ask_with_receipt` followed by `.answer`. ```python receipt: Receipt[bool] = guide.ask_with_receipt(noul("Is this urgent?"), ticket) @@ -305,46 +279,55 @@ if receipt.answer: meter(model=receipt.model, tokens=receipt.usage.input_tokens) ``` -`Receipt` is frozen and carries three things: `answer`, whatever `ask` would have returned; -`model`, the versioned id that actually answered, which is `jev-1.13.0` and not `jev-latest` -even when an alias was asked for; and `usage`, a `Usage` with `input_tokens` and -`output_tokens`. Input tokens are what is billed. Log the model: thresholds are tuned against -one model's numbers, and the alias moves under you. +A `Receipt` is frozen. `answer` is what `ask` returns. `model` is the versioned id of the model +that answered, for example `jev-1.13.0`, also when you asked for the `jev-latest` alias. `usage` +is a `Usage` with `input_tokens` and `output_tokens`, and TypeSafe bills the input tokens. Log +the model: thresholds are tuned against the numbers of one model, and an alias can move. -## Configuration +## Errors -Every setter on `GuideBuilder` returns the builder, and `Guide.builder()` starts one. +Every failure is a `GuidemeError`. Its `.kind` is the error kind: the same string in every +guideme SDK, and the value of `error.type` on the failed span. -| Setter | Default | What it does | +| Class | `.kind` | When | |---|---|---| -| `api_key(ApiKey(…))` | none; required | The key. `from_env()` reads it from `TYPESAFE_API_KEY`. | -| `base_url(…)` | `https://api.typesafe.ai` | The API origin. It may not carry credentials. | -| `model(Model(…))` | `jev-latest` | The model or alias to ask. | -| `policy(Policy(…))` | the defaults | The guide-wide policy patch; a question's own wins over it. | -| `max_retries(n)` | `3` | Resends per call, `0` to never resend. | -| `backoff(…)` | 500 ms | Base of the exponential backoff. | -| `timeout(…)` | 30 s | Per phase of one attempt. Read the paragraph below. | -| `transport(…)` / `async_transport(…)` | `httpx`'s own | Send through your `httpx` transport. | -| `record_state(True)` | off | Put the state JSON on the ask span. It is your users' data. | -| `events(…)` | `"both"` | Whether an answer and a retry go to the span, a log record, or both. | - -**`timeout` is not a deadline for the attempt.** `httpx` gives the whole budget to each phase -separately — connecting, writing, reading, and waiting for a pooled connection — so one attempt -that is slow in more than one phase takes longer than the timeout without breaching anything. -Worst case for a call is `max_retries + 1` attempts of several phases each, plus the backoff -between them. The Rust SDK's `reqwest` deadline covers the attempt as a whole instead; the two -differ because their HTTP clients do, and `docs/contract.md` records it as a divergence rather -than leaving you to find it. - -`transport(…)` and `timeout(…)` refuse each other, in whichever order you write them: a -timeout belongs to the transport that honours it, and `httpx` hands yours this budget as a -request extension it is free to ignore. A silent no-op would be worse than a `ConfigError`. -`transport(…)` is for `build()` and `async_transport(…)` for `build_async()`; using one with -the other build is a `ConfigError` too. +| `AuthError` | `auth` | HTTP 401 | +| `InvalidError` | `invalid` | HTTP 422. `.detail` is the response body. | +| `RateLimitedError` | `rate_limited` | HTTP 429 after the retries, or a `retry-after` that is too long to wait. `.retry_after` holds it. | +| `OverloadedError` | `overloaded` | HTTP 529, on the same terms. `.retry_after` holds it. | +| `TransportError` | `transport` | A connection, TLS or timeout failure, or a disconnect during the response. | +| `UnexpectedStatusError` | `unexpected_status` | A status that the contract does not define. | +| `ProtocolError` | `protocol` | The response breaks the contract: a body that does not decode, a wrong answer kind, an option or level that is not in the rubric, a probability outside 0..1. | +| `UnsureError` | `unsure` | The policy said unsure and the ladder had no value. | +| `ConfigError` | `config` | A mistake in your code, raised before anything is sent. For example: bad thresholds, no key, an empty batch, state or instructions that `json` cannot write, a rubric that breaks a rule, two fallback members in one `Choice`, a bad `events(…)`, a timeout that is not positive, negative retries or backoff, a malformed `base_url` or one that holds credentials, a timeout next to a transport, a transport given to the wrong build, `with_policy(…)` on a closed guide. | + +## Retries and timeouts + +guideme retries HTTP 429 and 529, for an ask and for `models()`. The backoff is exponential with +jitter, capped at 30 s. guideme obeys a `retry-after` header in whole seconds. If `retry-after` +is more than 30 s, guideme does not wait: it raises `RateLimitedError` or `OverloadedError`. + +A failed connection is resent in the same budget: a refused or reset connection, or a TLS +handshake that did not complete. That request did not reach a server, so nothing was answered. +A call makes at most `max_retries + 1` attempts, whatever the mix of failures. + +A disconnect during the response is not resent: the API can have answered, and a second request +pays for the same answer twice. guideme does not resend a timeout, in any phase. This includes a +connect timeout. Each of these raises `TransportError` at once. +[`docs/design.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/docs/design.md#decisions) +gives the reasons. + +**`timeout` is not a deadline for the attempt.** `httpx` gives the full value to each phase: +connect, write, read, and the wait for a pooled connection. Thus one slow attempt can take more +than the timeout. The worst case for a call is `max_retries + 1` attempts of several phases each, +plus the backoff between them. The Rust SDK differs here, as +[`docs/contract.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/docs/contract.md) +records. ## Testing your code -Pass a transport and your control flow is testable with no server, no port and no key: +Give the guide an `httpx` transport. Then you can test your control flow with no server, no port +and no key: ```python import httpx @@ -368,25 +351,17 @@ def test_an_urgent_ticket_is_prioritised() -> None: assert guide.ask(noul("Is this urgent?"), "payouts failing") is True ``` -`q0` is the first question in encounter order; a batch of three is answered with `q0`, `q1` -and `q2`. Raise from the handler instead of returning and you get the failure paths: an -`httpx.ConnectError` is resent inside the retry budget, and an `httpx.ReadTimeout`, an -`httpx.ConnectTimeout` or an `httpx.RemoteProtocolError` is not. For -`AsyncGuide`, hand the same `httpx.MockTransport` to `async_transport(…)` and `build_async()`; -it is both kinds of transport at once. +`q0` is the first question in encounter order, and a batch of three uses `q0`, `q1` and `q2`. +To test a failure, raise an `httpx` exception in the handler. The Retries and timeouts section +says which ones guideme resends. For `AsyncGuide`, give the same `httpx.MockTransport` to +`async_transport(…)` and call `build_async()`. A `MockTransport` is both kinds of transport. ## Observability -guideme emits OpenTelemetry spans, span events and OTLP log records through -`opentelemetry-api` and installs nothing: no tracer provider, no logger provider, no exporter, -no logging handler. Install a provider and the data appears. The SDK and an exporter are not -dependencies of this package, so install them alongside it: - -```sh -pip install opentelemetry-sdk opentelemetry-exporter-otlp-proto-grpc -``` - -The smallest provider that leaves the process: +guideme sends OpenTelemetry spans, span events and OTLP log records through `opentelemetry-api`. +It installs no provider, no exporter and no logging handler. When you install a provider, the +data appears. Install the SDK and an exporter next to guideme +(`pip install opentelemetry-sdk opentelemetry-exporter-otlp-proto-grpc`), then: ```python from opentelemetry import trace @@ -399,169 +374,129 @@ provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter())) trace.set_tracer_provider(provider) ``` -One span named `guideme.ask` per request, shaped by the OpenTelemetry GenAI conventions: -`gen_ai.request.model`, `gen_ai.response.model`, `gen_ai.usage.*`, and on failure `error.type` -with an error status. Under it, one HTTP client span per attempt with -`http.response.status_code`, so a retry is visible as sibling spans, plus a `guideme.retry` -event when an attempt is resent. One `guideme.answer` event per question with the outcome, -the probability or confidence, the unsure verdict and the settled thresholds that produced it. -The state is never recorded unless you opt in with `record_state(True)`. The API key never -appears anywhere. - -Every answer and every retry is also an OTLP log record, at `INFO` and at `WARN`, carrying the -trace id and the span id of the span it came from, so a logs backend links one straight back to -the decision it explains. Install a `LoggerProvider` too and they arrive; install neither and -they cost nothing. `events(...)` on the builder picks which signal carries an event when you -export both; **Choosing a signal** in the observability document has the table, the default, and -what happens on an `opentelemetry-api` that has no logs API: asking for one is refused, and the -default falls back to the span event alone. - -Because the shapes are standard, any OTLP backend reads them as is. +Each `ask` is one `guideme.ask` span with the OpenTelemetry GenAI fields, and one HTTP client +span per attempt below it. A retry adds a `guideme.retry` event. Each question adds a +`guideme.answer` event with the outcome, the probability or confidence and the thresholds. Each +answer is also an OTLP log record at `INFO`, and each retry one at `WARN`, with the trace id and +span id. To receive them, install a `LoggerProvider`. The state is not recorded unless you set +`record_state(True)`, and the API key is never recorded. + +`events(…)` selects the signal for answers and retries: `"span"`, `"log"` or `"both"` (the +default). If you export traces and logs to one backend, set `"span"` or `"log"` to store each +event once. If your `opentelemetry-api` has no logs API, `"log"` and `"both"` raise +`ConfigError`, and the default sends the span event only. [`docs/observability.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/docs/observability.md) -has the field tables and the environment variables that point the exporter anywhere. +has every field and the exporter settings. [`examples/otlp`](https://github.com/pedro-pscunha/guideme-python/tree/main/examples/otlp) runs -all of it against the live API with a collector that prints what arrives. +it all against the live API, with a collector that prints what arrives. -## Errors +## Configuration -Every failure is a `GuidemeError`. `.kind` is the same string every other guideme SDK reports -and the value of `error.type` on the failed span. +`Guide.builder()` starts a `GuideBuilder`. Each setting returns the builder. -| Class | `.kind` | When | +| Setting | Default | What it does | |---|---|---| -| `AuthError` | `auth` | 401 | -| `InvalidError` | `invalid` | 422; `.detail` is the body | -| `RateLimitedError` | `rate_limited` | 429 after retries, or a `retry-after` too long to wait for; `.retry_after` carries it | -| `OverloadedError` | `overloaded` | 529 on the same terms; `.retry_after` carries it too | -| `TransportError` | `transport` | connection, TLS, timeout | -| `UnexpectedStatusError` | `unexpected_status` | anything the contract does not define | -| `ProtocolError` | `protocol` | the response violates the contract: undecodable body, wrong answer kind, option or level not in the rubric, probability outside 0..1 | -| `UnsureError` | `unsure` | the policy said unsure and nothing caught it | -| `ConfigError` | `config` | raised where the mistake is written: bad thresholds, missing key, empty batch, unserialisable state, a duplicate rubric, a rubric outside 1..255 options or 2..10 levels, a bad `events(...)`, a non-positive timeout, negative retries or backoff, a `base_url` carrying credentials, and so on | - -Retries on 429 and 529 use exponential backoff with jitter, capped at 30 s, and honour an -integer `retry-after`. They apply to `models()` as much as to `ask`. - -A failed connection is resent in the same budget: refused, reset, or a TLS handshake that did -not complete. The request never reached a server, so nothing was judged and nothing is -repeated. A disconnect part-way through a response is **not** resent — the request arrived, -the API may have answered it, and asking again would buy the same judgment twice. - -**No timeout is resent, of any phase.** A connect timeout included, although `httpx` names it -separately: Rust's SDK sets one deadline over the whole attempt and cannot tell a connect -timeout from a read timeout, so retrying one here would make the two SDKs disagree about the -same failure, and a retried timeout multiplies the wall time `timeout(…)` is there to bound. -Everything not resent raises `TransportError` on the first failure. +| `api_key(ApiKey(…))` | none, required | The API key. | +| `base_url(…)` | `https://api.typesafe.ai` | The API origin. It cannot hold credentials. Give the final https origin. guideme follows redirects. | +| `model(Model(…))` | `jev-latest` | The model or alias to ask. | +| `policy(Policy(…))` | the defaults | The policy of the guide. The policy of a question wins over it. | +| `max_retries(n)` | `3` | Resends per call. `0` never resends. | +| `backoff(…)` | 500 ms | The base of the exponential backoff. | +| `timeout(…)` | 30 s | The limit for each phase of one attempt. | +| `transport(…)` / `async_transport(…)` | the transport of `httpx` | Send through your own `httpx` transport. | +| `record_state(True)` | off | Put the state JSON on the ask span. It is the data of your users. | +| `events(…)` | `"both"` | Send answers and retries as span events, log records, or both. | +| `from_env()` | | Apply the environment variables below. | + +`transport(…)` and `timeout(…)` refuse each other, in either order, with a `ConfigError`: an +injected transport owns its deadlines. `transport(…)` is for `build()`, and `async_transport(…)` +is for `build_async()`. The wrong pair is a `ConfigError` too. + +| Variable | Meaning | +|---|---| +| `TYPESAFE_API_KEY` | The API key. `Guide.from_env()` and `AsyncGuide.from_env()` require it. | +| `TYPESAFE_BASE_URL` | Optional. A different API origin. | +| `GUIDEME_MODEL` | Optional. The model or alias. The default is `jev-latest`. | + +Build one guide per process and share it. A guide holds a connection pool, and it is safe to use +from many threads or tasks at once. A guide per request also works, but it opens a pool for each +request. `guide.with_policy(…)` returns a second guide over the same pool. The pool counts its +guides: after you close one, the other can still ask, and the pool closes with the last guide. + +## Sync and async + +`Guide` and `AsyncGuide` have the same methods over the same core. The differences are the +`await` and the `httpx` client below it. + +```python +async with AsyncGuide.from_env() as guide: + verdict: Verdict = await guide.ask(noul("Is this about billing?").detail(), ticket) +``` + +For the synchronous version, use `Guide`, `with`, and no `await`. `Guide.builder()` and +`AsyncGuide.builder()` return the same `GuideBuilder`: `.build()` gives a `Guide`, and +`.build_async()` gives an `AsyncGuide`. A guide closes at the end of its `with` block. Outside a +block, call `guide.close()`. If you close one guide twice, the second close does nothing. Do not +ask through a closed guide. ## Lower layers -Everything above is re-exported from the `guideme` package, and `guideme.__all__` is that list. +The `guideme` package re-exports everything above, and `guideme.__all__` is that list. Some of +those names need a note: + +- `Question` is what the five constructors return. To annotate a question that you store or + pass on, use the concrete types `NoulQuestion`, `ChoiceQuestion`, `ScoreQuestion`, + `DetailedNoul`, `DetailedChoice` and `DetailedScore`. A `dict` is invariant, so a + `dict[str, NoulQuestion]` is not a `dict[str, Question[bool]]`. +- `Probability`, `Confidence`, `Key` and `Rank` are `NewType` brands. Only the wire creates + them, after it validates the value, so `Probability(2.0)` in your code is not refused. +- `Model` names a model. `ApiKey` holds the key and never prints it. Three modules are a second supported tier: `guideme.api`, `guideme.api.client` and -`guideme.policy`. You import those by their own path, they are not re-exported at the top -level, and they are under the same rule as the first tier — nothing in them is removed or -renamed without a major version and a `CHANGELOG.md` entry. Anything else in the package is -private, whatever its name looks like. - -The two bullets after them are not a tier. They say where some of the names above are -declared, which is worth knowing when two of them share a spelling. - -- `guideme.api` is the exact wire mirror of `POST /v1/systemone` and `GET /v1/models`, and - `guideme.api.__all__` is what it offers: the request and response models, its own `Usage`, - and the four adapters between them and the core. That `Usage` is the pydantic model a - response is parsed into, not the `Usage` a receipt carries — a receipt gets the frozen - dataclass of the same name from the top level, copied out of this one, so that nothing - pydantic sits on the surface you import from `guideme`. `guideme.api.client` holds `Client` and - `AsyncClient` for callers who want to build requests themselves. They live one level down - rather than on `guideme.api` because re-exporting them would make `api` and `api.client` - import each other, and the gate fails an import cycle. -- `guideme.policy.resolve(answer, thresholds)` is the pure decision function. `spec/` holds - its JSON Schemas and 42 golden vectors, vendored from - [guideme-rust](https://github.com/pedro-pscunha/guideme-rust), which publishes the contract. - [`docs/contract.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/docs/contract.md) - says what every guideme SDK must satisfy and - [`docs/design.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/docs/design.md) - records the design and its sharp edges. -- `guideme.question` is where the question types are declared, and all of them are re-exported - above: `Question` is what `noul`, `choose`, `choose_among`, `score` and `score_levels` - return, and `NoulQuestion`, `ChoiceQuestion`, `ScoreQuestion`, `DetailedNoul`, - `DetailedChoice` and `DetailedScore` are the concrete ones. Inference covers most uses, so - reach for them when you need to annotate a question you are storing or passing on: a `dict` - is invariant, so a `dict[str, NoulQuestion]` is not a `dict[str, Question[bool]]` and the - annotation has to be written. Everything else in `guideme.question` is private. - `guideme.api` declares its own `Question`, `NoulQuestion`, `ChoiceQuestion` and - `ScoreQuestion`: same names, different classes. Those are the wire shapes the ones above - become on the way out, and you only meet them if you build requests by hand. The import - line says which you have. -- The scalars are validated once and never re-checked: `Probability` and `Confidence` hold the - unit-interval numbers on `Verdict`, `Ranked` and `Scored`, `Key` and `Rank` are what a runtime - rubric answers with, `Model` names the model to ask, and `ApiKey` carries the key without ever - printing it. The first four are `NewType` brands, so the guarantee is that only the wire mints - them, not that `Probability(2.0)` is rejected; it is not. +`guideme.policy`. Import them by their own path. The top level does not re-export them. They +have the same rule as the first tier: nothing in them is removed or renamed without a major +version and a `CHANGELOG.md` entry. Everything else in the package is private, whatever its +name looks like. + +- `guideme.api` is the exact wire mirror of `POST /v1/systemone` and `GET /v1/models`. + `guideme.api.__all__` lists the request and response models, its own `Usage`, and four + adapters between them and the core. +- `guideme.api.client` holds `Client` and `AsyncClient`, to build requests yourself. +- `guideme.policy.resolve(answer, thresholds)` is the pure decision function. It takes a + `Thresholds`: a `Policy` settled against the defaults. `spec/` holds its JSON Schemas and 42 + golden vectors, which use the same `Thresholds`. + +`guideme.api` also declares a `Question`, `NoulQuestion`, `ChoiceQuestion`, `ScoreQuestion` and +`Usage` of its own: wire shapes, different classes from the ones above. You meet them only when +you build requests by hand, and +[`docs/design.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/docs/design.md#sharp-edges) +explains these pairs. ## Other SDKs -Every guideme SDK is written from scratch in its own language and answers the same way, -because they all satisfy one contract: the wire schemas, the 42 golden policy vectors and the -interface shape that [guideme-rust](https://github.com/pedro-pscunha/guideme-rust) publishes -under `spec/` and states in +Each guideme SDK is written from scratch in its own language, and they all turn one reading into +the same answer. They satisfy one contract: the wire schemas, the 42 golden policy vectors and +the interface shape. [guideme-rust](https://github.com/pedro-pscunha/guideme-rust) publishes it +under `spec/` and states it in [`docs/contract.md`](https://github.com/pedro-pscunha/guideme-rust/blob/main/docs/contract.md). | Language | Package | Repository | |---|---|---| -| Python | `guideme` | this repository | | Rust | [`guideme`](https://crates.io/crates/guideme) | [guideme-rust](https://github.com/pedro-pscunha/guideme-rust) | +| Python | `guideme` | this repository | +| TypeScript | `@guideme/sdk` (not yet on npm) | [guideme-typescript](https://github.com/pedro-pscunha/guideme-typescript) | -This repository vendors that `spec/` and records the commit it came from in `spec/SOURCE`; a -CI job fails when the copy drifts from the Rust repository's `main`. The span, event and -attribute names are shared too, so one dashboard reads both SDKs. - -## Environment - -| Variable | Meaning | -|---|---| -| `TYPESAFE_API_KEY` | required by `Guide.from_env()` and `AsyncGuide.from_env()` | -| `TYPESAFE_BASE_URL` | optional API origin override | -| `GUIDEME_MODEL` | optional model or alias; default `jev-latest` | +The span, event and attribute names are shared too, so one dashboard reads every SDK. ## Development -Tooling is managed by [mise](https://mise.jdx.dev), which pins `uv` and `gitleaks`; `uv` pins -everything else from `pyproject.toml` and `uv.lock`. - -``` -mise install # fetch the tools -mise run sync # install the locked environment -mise run check # fmt-check, gen-check, ruff, pyright, pylint, pytest, build, audit -mise run test # pytest alone -mise run hooks # point core.hooksPath at the tracked hooks in .githooks -``` - -The hooks are tracked, not generated: `mise run hooks` sets this repository's `core.hooksPath` -to `.githooks` and verifies it took effect. -[`AGENTS.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/AGENTS.md) says what -each stage runs. - -Library code is held to a strict checker set: `pyright` in strict mode, `ruff` with every rule -selected, and `pylint` with every check enabled. Tests are few and high-grade: property tests -for the policy laws, a real local HTTP server for the wire and retry contract, structural -tracing assertions, pyright files that must fail, and a drift guard that re-resolves every -golden vector. - -From a source distribution rather than a clone, `mise.toml` and `uv.lock` are not present, so -the suite runs with `uv run --group dev pytest`. - -Two opt-in tests hit the real API and are deselected by default: - -``` -TYPESAFE_API_KEY=… uv run --locked pytest -m live -``` - -Contributor rules live in -[`AGENTS.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/AGENTS.md). Report a -vulnerability privately, as -[`SECURITY.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/SECURITY.md) -describes, never in a public issue. +To set up a clone, run `mise install`, `mise run sync` and `mise run hooks`. `mise run check` is +the gate. +[`CONTRIBUTING.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/CONTRIBUTING.md) +and [`AGENTS.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/AGENTS.md) have the +rules. Report a vulnerability privately, as +[`SECURITY.md`](https://github.com/pedro-pscunha/guideme-python/blob/main/SECURITY.md) tells, +and never in a public issue. ## License diff --git a/docs/contract.md b/docs/contract.md index da12807..9dfc231 100644 --- a/docs/contract.md +++ b/docs/contract.md @@ -216,7 +216,7 @@ for a pooled connection each get the whole of it — so an attempt that is slow phase outlasts the number written in the builder. Wrapping it to match would need a different wrapper for the synchronous and the asyncio surfaces and would change what cancellation means, which is a worse trade than saying so. It is documented on `GuideBuilder.timeout`, in the -README's configuration section, and here. Nothing on the wire depends on it. +README's **Retries and timeouts** section, and here. Nothing on the wire depends on it. Drift is caught rather than trusted. `mise run spec-check` clones guideme-rust, diffs its `spec/` against this one and fails on any difference except `spec/SOURCE`, which is provenance and has diff --git a/docs/design.md b/docs/design.md index 572cdcf..47f5821 100644 --- a/docs/design.md +++ b/docs/design.md @@ -258,6 +258,10 @@ Each module survives the test. second `Usage` — and it is the right way round: a caller reaching the top-level surface gets the value object, and the name they would otherwise collide with is in a module they only import when they are building requests by hand. +- **`Client` and `AsyncClient` live in `guideme.api.client`, not on `guideme.api`.** + Re-exporting them from `guideme.api` would make `api` and `api.client` import each other, + and the gate fails an import cycle. So a caller who builds requests by hand imports the + wire models from one module and the clients from the other. - **State is JSON-shaped.** Anything `json.dumps` accepts without a default hook. A dataclass goes through `dataclasses.asdict`, a pydantic model through `.model_dump()`. This is the one untyped value in the package, and it is serialised at the boundary.