Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -5,4 +5,5 @@ dist/
!.env.example
coverage/
/docs/
/benchmarks/.data/
/benchmarks/results/
134 changes: 75 additions & 59 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,10 +2,37 @@

`choosekit` scores a finite set of choices with a language model and returns a typed decision with a probability distribution. It supports local llama.cpp models and an optional OpenRouter backend.

![SuperGPQA direct-choice benchmark](benchmarks/supergpqa-benchmark.svg)

The chart compares accuracy with a lower-is-better cost-latency product. The green line and confidence band show local Qwen3.8 27B Q4_XL accuracy; it has no cloud cost coordinate. [Method and reproduction](benchmarks/README.md#supergpqa)

| Model | Accuracy | Cost / 1,000 decisions | Decisions/s |
|---|---:|---:|---:|
| Granite 4.0 H Micro | 19.3% | $0.0053 | 2.84 |
| Llama 3.1 8B | 19.0% | $0.0061 | 1.33 |
| GLM 4.7 Flash | 25.7% | $0.0172 | 1.18 |
| Gemma 4 26B | 37.6% | $0.0184 | 2.19 |
| Jev 1.13 | 53.6% | $0.0244 | 2.86 |
| Granite 4.2 8B | 23.8% | $0.0293 | 3.15 |
| DeepSeek V4.1 Flash | 45.6% | $0.0544 | 1.40 |
| DeepSeek V4 Pro | 43.2% | $0.3588 | 0.88 |
| GLM 5.2 | 44.6% | $0.3932 | 0.69 |
| Kimi K3 | 59.3% | $0.6243 | 0.73 |

## Install

Library:

```sh
npm install choosekit
```

MCP server:

```sh
npm install --global choosekit-mcp
```

## Why

Agents often need to choose from known options:
Expand All @@ -18,13 +45,13 @@ Agents often need to choose from known options:

`choosekit` scores choices using the model's conditional log probabilities at the token branches that distinguish them.

The project was inspired by [Jev and the System One model interface](https://typesafe.ai/blog/introducing-system-one-models-and-jev): application state in, typed probabilistic decisions out. Jev is a specialized hosted model. `choosekit` explores the same useful interface with a model you control. The llama.cpp backend keeps application state on infrastructure you choose; OpenRouter is available when a hosted model is more convenient.
The project was inspired by [Jev and the System One model interface](https://typesafe.ai/blog/introducing-system-one-models-and-jev): application state in, typed probabilistic decisions out. Jev is a specialized hosted model. `choosekit` brings the same typed decision interface to general-purpose language models. The llama.cpp backend runs on infrastructure you choose; OpenRouter provides hosted inference.

`choosekit` is an independent project with no affiliation to TypeSafe or Jev.

## MCP server

[`choosekit-mcp`](packages/choosekit-mcp/README.md) exposes choosekit through llama.cpp or OpenRouter as a read-only stdio tool for Claude Code, Codex, and OpenCode. Select the backend and configure it with environment variables when starting the MCP server. Every `choose` call uses this configuration.
[`choosekit-mcp`](packages/choosekit-mcp/README.md) exposes choosekit through llama.cpp or OpenRouter as a read-only stdio tool for Claude Code, Codex, and OpenCode. Select the backend and configure it with environment variables when starting the MCP server.

## llama.cpp

Expand Down Expand Up @@ -65,17 +92,11 @@ const choose = fromOpenRouter({
});
```

The OpenRouter backend supports models and providers that return first-token `top_logprobs`, with up to 20 choices. Unlike llama.cpp, this backend sends the prompt to OpenRouter. It requests reasoning to be disabled. Choices omitted from `top_logprobs` receive zero probability. Returned probabilities are normalized across the supplied choices and are not calibrated correctness estimates.
The OpenRouter backend supports models and providers that return first-token `top_logprobs`, with up to 20 choices. It sends the prompt to OpenRouter and requests reasoning to be disabled.

OpenRouter may route the same model through different providers. Set `provider` to an OpenRouter provider slug to use only that provider and disable fallback:
Choices omitted from `top_logprobs` receive zero probability. Returned probabilities are normalized across the supplied choices and are not calibrated correctness estimates.

```ts
const choose = fromOpenRouter({
apiKey: process.env.OPENROUTER_API_KEY!,
model: "qwen/qwen3.8-27b",
provider: process.env.OPENROUTER_PROVIDER!,
});
```
OpenRouter may route the same model through different providers. Set `provider: "provider-slug"` to use only that provider and disable fallback.

## Scoring modes

Expand All @@ -88,14 +109,53 @@ In `labels` mode, choices are shown to the model as `A`, `B`, `C` instead of the

`minimal-prefix` walks the token tree until every key is distinguishable. For keys such as `watermelon` and `watermelon juice`, the shared token path is handled once and scoring stops when the paths separate.

## Benchmark
## Return value

`choose()` resolves to:

```ts
{
choice, // selected caller key
distribution, // normalized probability for every supplied key
scores, // backend log-probability score for every key
margin, // largest probability minus the second largest
entropy, // Shannon entropy in nats
boundaryTokens, // prompt tokens rolled back at a tokenization boundary
usage, // backend work, when reported
}
```

The result and its nested records are immutable. Each call is stateless. The caller controls action execution, inference retries, and model selection.

## Prompt formatting

`context` is copied unchanged to the start of the scoring prompt. The default formatter then appends the question, choice descriptions, and an answer marker.

Use `formatPrompt` only when you need custom prompt formatting. The result must preserve `context` as an unchanged prefix so an existing server-side prefix cache can still be reused.

## Custom scorer

Use `createChooser` with any backend that can return one comparable conditional log-probability score per candidate:

```ts
import { createChooser, type Scorer } from "choosekit";

const scorer: Scorer = async ({ prompt, candidates, signal }) => ({
logprobs: await scoreCandidateSequences(prompt, candidates, signal),
});

const choose = createChooser(scorer);
```

Scores use natural logarithms and must be at most zero.

## SemIf comparison

The local adapter was compared with `typesafe/jev-1.13` on SemIf's official 144-row `authored144` benchmark, which covers evidence interpretation, rule application, and candidate selection. The local model was **Qwen 3.8 27B Q4_XL** served by llama.cpp on an **NVIDIA RTX 4090**. The Qwen run used the default A/B/C mode. The model was already loaded, and requests were sent one at a time to a llama.cpp server on the same machine.

| Metric | Qwen 3.8 27B Q4_XL + choosekit | Jev 1.13 |
|---|---:|---:|
| Accuracy | 96.53% (139/144) | 96.53% (139/144) |
| Average balanced accuracy across task families | 96.01% | 95.56% |
| Median latency (p50) | 239 ms | 368 ms |
| 95th percentile latency (p95) | 286 ms | 546 ms |
| Throughput | 4.02 decisions/s | 2.43 decisions/s |
Expand All @@ -106,8 +166,6 @@ These results are specific to this 144-case benchmark, and performance can diffe

### Probability examples

Examples from the same benchmark:

[`eafc22c8c40df3932a8e`](benchmarks/data/semif-authored144.jsonl#L112) asks whether the crate is currently in storage. The protocol gives the inventory priority; the current inventory and desk-log entries are missing.

| Choice | Qwen probability | Jev probability |
Expand All @@ -124,7 +182,7 @@ Examples from the same benchmark:
| **Insufficient evidence (selected by both)** | **97.605%** | **99.000%** |
| Contradicted | 0.029% | 1.000% |

The Qwen + llama.cpp probabilities shown here are [uncalibrated](https://proceedings.mlr.press/v70/guo17a.html). [Jev is trained for calibrated decisions](https://typesafe.ai/blog/introducing-system-one-models-and-jev). The distributions look similar in these examples. This benchmark measures accuracy, latency, and distribution similarity.
The Qwen + llama.cpp probabilities shown here are [uncalibrated](https://proceedings.mlr.press/v70/guo17a.html). [Jev is trained for calibrated decisions](https://typesafe.ai/blog/introducing-system-one-models-and-jev). The distributions look similar in these examples.

### Distribution comparison

Expand All @@ -142,51 +200,9 @@ Total variation distance (TVD) compares two complete probability distributions.
| Cases with TVD at or below 10% | 77.08% (111/144) |
| Cases with TVD above 20% | 14.58% (21/144) |

Most distributions are similar. Some differ substantially: the systems select different choices in eight cases, and the largest TVD is 94.23%.

## Prompt formatting

`context` is copied unchanged to the start of the scoring prompt. The default formatter then appends the question, choice descriptions, and an answer marker.

Use `formatPrompt` only when you need custom prompt formatting. The result must preserve `context` as an unchanged prefix so an existing server-side prefix cache can still be reused.

## Custom scorer

Use `createChooser` with any backend that can return one comparable conditional log-probability score per candidate:

```ts
import { createChooser, type Scorer } from "choosekit";

const scorer: Scorer = async ({ prompt, candidates, signal }) => ({
logprobs: await scoreCandidateSequences(prompt, candidates, signal),
});

const choose = createChooser(scorer);
```

Scores use natural logarithms and must be at most zero.

## Result

`choose()` resolves to:

```ts
{
choice, // selected caller key
distribution, // normalized probability for every supplied key
scores, // backend log-probability score for every key
margin, // largest probability minus the second largest
entropy, // Shannon entropy in nats
boundaryTokens, // prompt tokens rolled back at a tokenization boundary
usage, // backend work, when reported
}
```

The result and its nested records are immutable. Each call is stateless. The caller controls action execution, inference retries, and model selection.

## Requirements

- Node.js 20 or newer.
- No runtime dependencies, model downloads, installation hooks, or bundled inference servers.
- The `choosekit` package has no runtime dependencies, model downloads, installation hooks, or bundled inference servers.

[Apache-2.0](LICENSE). Copyright 2026 NotXf1le.
92 changes: 77 additions & 15 deletions benchmarks/README.md
Original file line number Diff line number Diff line change
@@ -1,41 +1,103 @@
# Benchmarks

`data/semif-authored144.jsonl` is an exact copy of SemIf's official [`benchmarks/data/authored144.jsonl`](https://github.com/TheoLeeCJ/SemIf/blob/b9cb32537e78be65f19abfcb1de8fc504b627d84/benchmarks/data/authored144.jsonl) at commit `b9cb32537e78be65f19abfcb1de8fc504b627d84`. Only the local filename differs.
## SuperGPQA

The 144 examples were authored by the SemIf project. The dataset is distributed under SemIf's MIT license, reproduced in `data/SEMIF-LICENSE.txt`.
The chart uses a deterministic 1,000-question sample stratified by discipline and
difficulty. The random-choice baseline is the mean of `1 / number of choices` across
the sample. OpenRouter models are included only when they return `top_logprobs`.
The evaluation sample excludes a deterministic 100-question pilot used to select
working model and provider pairs.

Build the package before running a benchmark:
### Prepare the dataset

```sh
npm ci
hf download m-a-p/SuperGPQA SuperGPQA-all.jsonl \
--repo-type dataset \
--revision 4430d4458112c7d4497fdcf94d7cc223313d6acf \
--local-dir benchmarks/.data/supergpqa/source
node benchmarks/prepare-supergpqa.mjs
npm run build
```

Run the published llama.cpp adapter against a local server:
### Run the benchmark

```sh
OPENROUTER_API_KEY=... node benchmarks/run-supergpqa.mjs \
--backend openrouter \
--model ibm-granite/granite-4.0-h-micro \
--provider cloudflare \
--sample-method proportional \
--sample-size 1000 \
--output benchmarks/results/supergpqa-granite-4.0-h-micro-cloudflare.json
```

The chart uses these OpenRouter model and provider pairs:

| Model | Provider | Result |
|---|---|---|
| `ibm-granite/granite-4.0-h-micro` | `cloudflare` | `supergpqa-granite-4.0-h-micro-cloudflare.json` |
| `meta-llama/llama-3.1-8b-instruct` | `novita` | `supergpqa-llama-3.1-8b-novita.json` |
| `z-ai/glm-4.7-flash` | `cloudflare` | `supergpqa-glm-4.7-flash-cloudflare.json` |
| `z-ai/glm-5.2` | `cloudflare` | `supergpqa-glm-5.2-cloudflare.json` |
| `google/gemma-4-26b-a4b-it` | `dekallm` | `supergpqa-gemma-4-26b-dekallm.json` |
| `ibm-granite/granite-4.2-8b` | `coreweave` | `supergpqa-granite-4.2-8b-coreweave.json` |
| `deepseek/deepseek-v4.1-flash` | `wafer` | `supergpqa-deepseek-v4.1-flash-wafer.json` |
| `deepseek/deepseek-v4-pro-0813` | `cloudflare` | `supergpqa-deepseek-v4-pro-cloudflare.json` |
| `moonshotai/kimi-k3` | `morph` | `supergpqa-kimi-k3-morph.json` |

Run Jev on the same sample:

```sh
OPENROUTER_API_KEY=... node benchmarks/run-supergpqa.mjs \
--backend jev \
--sample-method proportional \
--sample-size 1000 \
--output benchmarks/results/supergpqa-jev-1.13.json
```

Run a local model through llama.cpp:

```sh
node benchmarks/run-supergpqa.mjs \
--backend llama-cpp \
--base-url http://127.0.0.1:8080 \
--model qwen3.8-27b-text-64k \
--sample-method proportional \
--sample-size 1000 \
--output benchmarks/results/supergpqa-qwen3.8-27b-local.json
```

The X axis is average cost per decision multiplied by seconds per decision. The local
model is shown as a horizontal accuracy line.

```sh
node benchmarks/generate-supergpqa-chart.mjs
```

## SemIf

`data/semif-authored144.jsonl` is SemIf's official
[`benchmarks/data/authored144.jsonl`](https://github.com/TheoLeeCJ/SemIf/blob/b9cb32537e78be65f19abfcb1de8fc504b627d84/benchmarks/data/authored144.jsonl)
at commit `b9cb32537e78be65f19abfcb1de8fc504b627d84`. The examples were authored by the
SemIf project and are distributed under its MIT license, reproduced in
`data/SEMIF-LICENSE.txt`.

```sh
npm ci
npm run build

node benchmarks/run-semif.mjs \
--mode labels \
--base-url http://127.0.0.1:11434/ \
--model qwen3.8-27b-text-64k \
--output benchmarks/results/semif-qwen-labels.json
```

Run the Jev comparison with an OpenRouter API key:

```sh
OPENROUTER_API_KEY=... node benchmarks/run-semif-openrouter-jev.mjs \
--model typesafe/jev-1.13 \
--output benchmarks/results/semif-jev-1.13.json
```

Compare the complete distributions:

```sh
node benchmarks/compare-semif.mjs \
--qwen benchmarks/results/semif-qwen-labels.json \
--jev benchmarks/results/semif-jev-1.13.json \
--output benchmarks/results/semif-comparison.json
```

Benchmark result files are ignored because they can contain environment-specific timing and provider metadata.
Loading
Loading