Skip to content

Latest commit

 

History

History
111 lines (89 loc) · 3.78 KB

File metadata and controls

111 lines (89 loc) · 3.78 KB

Benchmarks

SuperGPQA

The chart uses a deterministic 1,000-question sample stratified by discipline and difficulty. The random-choice baseline is the mean of 1 / number of choices across the sample. OpenRouter models are included only when they return top_logprobs. The evaluation sample excludes a deterministic 100-question pilot used to select working model and provider pairs.

Prepare the dataset

hf download m-a-p/SuperGPQA SuperGPQA-all.jsonl \
  --repo-type dataset \
  --revision 4430d4458112c7d4497fdcf94d7cc223313d6acf \
  --local-dir benchmarks/.data/supergpqa/source
node benchmarks/prepare-supergpqa.mjs
npm run build

Run the benchmark

OPENROUTER_API_KEY=... node benchmarks/run-supergpqa.mjs \
  --backend openrouter \
  --model ibm-granite/granite-4.0-h-micro \
  --provider cloudflare \
  --sample-method proportional \
  --sample-size 1000 \
  --output benchmarks/results/supergpqa-granite-4.0-h-micro-cloudflare.json

The chart uses these OpenRouter model and provider pairs:

Model Provider Result
ibm-granite/granite-4.0-h-micro cloudflare supergpqa-granite-4.0-h-micro-cloudflare.json
meta-llama/llama-3.1-8b-instruct novita supergpqa-llama-3.1-8b-novita.json
z-ai/glm-4.7-flash cloudflare supergpqa-glm-4.7-flash-cloudflare.json
z-ai/glm-5.2 cloudflare supergpqa-glm-5.2-cloudflare.json
google/gemma-4-26b-a4b-it dekallm supergpqa-gemma-4-26b-dekallm.json
ibm-granite/granite-4.2-8b coreweave supergpqa-granite-4.2-8b-coreweave.json
deepseek/deepseek-v4.1-flash wafer supergpqa-deepseek-v4.1-flash-wafer.json
deepseek/deepseek-v4-pro-0813 cloudflare supergpqa-deepseek-v4-pro-cloudflare.json
moonshotai/kimi-k3 morph supergpqa-kimi-k3-morph.json

Run Jev on the same sample:

OPENROUTER_API_KEY=... node benchmarks/run-supergpqa.mjs \
  --backend jev \
  --sample-method proportional \
  --sample-size 1000 \
  --output benchmarks/results/supergpqa-jev-1.13.json

Run a local model through llama.cpp:

node benchmarks/run-supergpqa.mjs \
  --backend llama-cpp \
  --base-url http://127.0.0.1:8080 \
  --model qwen3.8-27b-text-64k \
  --sample-method proportional \
  --sample-size 1000 \
  --output benchmarks/results/supergpqa-qwen3.8-27b-local.json

The X axis is average cost per decision multiplied by seconds per decision. The local model is shown as a horizontal accuracy line.

node benchmarks/generate-supergpqa-chart.mjs

SemIf

data/semif-authored144.jsonl is SemIf's official benchmarks/data/authored144.jsonl at commit b9cb32537e78be65f19abfcb1de8fc504b627d84. The examples were authored by the SemIf project and are distributed under its MIT license, reproduced in data/SEMIF-LICENSE.txt.

npm ci
npm run build

node benchmarks/run-semif.mjs \
  --mode labels \
  --base-url http://127.0.0.1:11434/ \
  --model qwen3.8-27b-text-64k \
  --output benchmarks/results/semif-qwen-labels.json

Before running, the script requires the selected model ID to appear in GET /v1/models. This prevents a report from being labelled with a model the server did not expose. Use --skip-model-check to bypass the check; the report then records modelChecked: false in its runtime block.

Run the Jev comparison with an OpenRouter API key:

OPENROUTER_API_KEY=... node benchmarks/run-semif-openrouter-jev.mjs \
  --model typesafe/jev-1.13 \
  --output benchmarks/results/semif-jev-1.13.json

node benchmarks/compare-semif.mjs \
  --qwen benchmarks/results/semif-qwen-labels.json \
  --jev benchmarks/results/semif-jev-1.13.json \
  --output benchmarks/results/semif-comparison.json