The chart uses a deterministic 1,000-question sample stratified by discipline and
difficulty. The random-choice baseline is the mean of 1 / number of choices across
the sample. OpenRouter models are included only when they return top_logprobs.
The evaluation sample excludes a deterministic 100-question pilot used to select
working model and provider pairs.
hf download m-a-p/SuperGPQA SuperGPQA-all.jsonl \
--repo-type dataset \
--revision 4430d4458112c7d4497fdcf94d7cc223313d6acf \
--local-dir benchmarks/.data/supergpqa/source
node benchmarks/prepare-supergpqa.mjs
npm run buildOPENROUTER_API_KEY=... node benchmarks/run-supergpqa.mjs \
--backend openrouter \
--model ibm-granite/granite-4.0-h-micro \
--provider cloudflare \
--sample-method proportional \
--sample-size 1000 \
--output benchmarks/results/supergpqa-granite-4.0-h-micro-cloudflare.jsonThe chart uses these OpenRouter model and provider pairs:
| Model | Provider | Result |
|---|---|---|
ibm-granite/granite-4.0-h-micro |
cloudflare |
supergpqa-granite-4.0-h-micro-cloudflare.json |
meta-llama/llama-3.1-8b-instruct |
novita |
supergpqa-llama-3.1-8b-novita.json |
z-ai/glm-4.7-flash |
cloudflare |
supergpqa-glm-4.7-flash-cloudflare.json |
z-ai/glm-5.2 |
cloudflare |
supergpqa-glm-5.2-cloudflare.json |
google/gemma-4-26b-a4b-it |
dekallm |
supergpqa-gemma-4-26b-dekallm.json |
ibm-granite/granite-4.2-8b |
coreweave |
supergpqa-granite-4.2-8b-coreweave.json |
deepseek/deepseek-v4.1-flash |
wafer |
supergpqa-deepseek-v4.1-flash-wafer.json |
deepseek/deepseek-v4-pro-0813 |
cloudflare |
supergpqa-deepseek-v4-pro-cloudflare.json |
moonshotai/kimi-k3 |
morph |
supergpqa-kimi-k3-morph.json |
Run Jev on the same sample:
OPENROUTER_API_KEY=... node benchmarks/run-supergpqa.mjs \
--backend jev \
--sample-method proportional \
--sample-size 1000 \
--output benchmarks/results/supergpqa-jev-1.13.jsonRun a local model through llama.cpp:
node benchmarks/run-supergpqa.mjs \
--backend llama-cpp \
--base-url http://127.0.0.1:8080 \
--model qwen3.8-27b-text-64k \
--sample-method proportional \
--sample-size 1000 \
--output benchmarks/results/supergpqa-qwen3.8-27b-local.jsonThe X axis is average cost per decision multiplied by seconds per decision. The local model is shown as a horizontal accuracy line.
node benchmarks/generate-supergpqa-chart.mjsdata/semif-authored144.jsonl is SemIf's official
benchmarks/data/authored144.jsonl
at commit b9cb32537e78be65f19abfcb1de8fc504b627d84. The examples were authored by the
SemIf project and are distributed under its MIT license, reproduced in
data/SEMIF-LICENSE.txt.
npm ci
npm run build
node benchmarks/run-semif.mjs \
--mode labels \
--base-url http://127.0.0.1:11434/ \
--model qwen3.8-27b-text-64k \
--output benchmarks/results/semif-qwen-labels.jsonBefore running, the script requires the selected model ID to appear in GET /v1/models. This prevents
a report from being labelled with a model the server did not expose. Use --skip-model-check to bypass
the check; the report then records modelChecked: false in its runtime block.
Run the Jev comparison with an OpenRouter API key:
OPENROUTER_API_KEY=... node benchmarks/run-semif-openrouter-jev.mjs \
--model typesafe/jev-1.13 \
--output benchmarks/results/semif-jev-1.13.json
node benchmarks/compare-semif.mjs \
--qwen benchmarks/results/semif-qwen-labels.json \
--jev benchmarks/results/semif-jev-1.13.json \
--output benchmarks/results/semif-comparison.json