Skip to content

Add SuperGPQA benchmark - #12

Merged
NotXf1le merged 2 commits into
masterfrom
feature/supergpqa-benchmark
Sep 21, 2026
Merged

NotXf1le merged 2 commits into
masterfrom
feature/supergpqa-benchmark

Conversation

@NotXf1le

Copy link
Copy Markdown
Owner

What changed

  • Added a reproducible SuperGPQA benchmark runner for OpenRouter, llama.cpp, and Jev.
  • Added deterministic dataset preparation and a 1,000-question stratified evaluation sample.
  • Kept the evaluation sample separate from the 100-question pilot used to select working model and provider pairs.
  • Added checkpoint/resume support, dataset integrity checks, pinned OpenRouter providers, cost tracking, and latency measurements.
  • Added the benchmark chart and exact results to the README.
  • Added tests for dataset preparation, deterministic sampling, pilot separation, and the pilot-size limit.
  • Cleaned up benchmark names and reorganized the README reference sections.

Why

This provides a reproducible comparison of decision accuracy, cost, and latency across compatible language models and Jev. The dataset is downloaded separately and is not included in the repository.

Checks

  • npm run check
  • npm pack --dry-run

@NotXf1le
NotXf1le merged commit a2fb44c into master Sep 21, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant