Skip to content

[RFC] Move growing data off GitHub: HF as canonical, Pages as cached browse view #682

Description

@junjiezhou1122

Summary

Decouple data growth from the GitHub code repository, stop pushing the
complete dataset through GitHub Pages, and stop relying on "anti-scraping"
as a solution. The current workflow keeps growing data/snapshots/ on
main (~53 MB today, ~780 MB projected in a year) and still serves
radar.json to any visitor who clicks the homepage Download button.

Problem

Three concrete signals, measured in this repository today:

  1. data/snapshots/ contains 58 daily JSON files, ~53 MB total,
    adding ~2 MB per day. At that rate it crosses the GitHub Pages
    1 GB site-size soft limit within a year.
  2. site/data/hf_dataset/ is generated by benchmark-radar export-hf
    and then uploaded as part of the Pages artifact. Even though the
    same data is published to huggingface.co/datasets/ktwu01/benchmark-radar,
    a copy still ships inside the Pages bundle.
  3. site/index.html exposes <a href="/data/radar.json">Download</a>
    and several JS paths still fetch the full radar.json. With ~2 MB
    per request and no cache headers, a small amount of repeated
    traffic puts real pressure on the 100 GB / month Pages soft limit.

These are not separate problems. The fix is the same shape: stop
making the code repository the canonical place for growing data, stop
making Pages the canonical place for the complete dataset, and use
caching + a real download channel for what remains.

Proposed Solution

Split responsibilities along the lines of LiveBench, SWE-bench,
Our World in Data, and OpenAlex:

  • GitHub keeps code, configuration, human-reviewed identity rules,
    small fixtures, and a tiny version-pointer file.
  • Hugging Face keeps the complete daily snapshots, the exported
    catalog, and versioned releases with SHA-256 checksums.
  • The Pages site keeps only what a visitor needs to browse
    (index, per-benchmark shards, first-paint summaries, optionally
    per-day radar slices), and uses long-lived cache headers.

Land this in four small, low-risk PRs, each passing the full CI
sequence in AGENTS.md.

PR 1 — Stop shipping the HF export inside the Pages bundle

  • Change hf_dataset.py's default export directory from
    site/data/hf_dataset to a build-only path such as
    build/hf_dataset.
  • In .github/workflows/pages.yml, rm -rf site/data/hf_dataset
    after export-hf and before the Pages upload step. The HF
    publish job is unchanged.
  • Result: the Pages artifact no longer carries a duplicate of what
    Hugging Face already hosts. Zero behavioral change for visitors.

PR 2 — Stop committing daily snapshots to main

  • In .github/workflows/daily-radar.yml, replace
    git add data/snapshots && git commit && git push with an upload
    step that syncs snapshots/<date>.json to
    huggingface.co/datasets/ktwu01/benchmark-radar.
  • Add a small data/snapshot-pointer.yml that records the current
    snapshot set name and HF revision. snapshots.py reads it on
    startup, with the existing data/snapshots/ directory as the
    offline fallback for tests.
  • Keep test fixtures (tests/test_social.py already uses
    data/snapshots/2026-08-10.json) intact.
  • Result: the code repository stops growing. A fresh checkout still
    builds, because the pointer file is committed and HF is the
    canonical source.

PR 3 — Cache the public JSON and redirect the Download button

  • Serve site/data/benchmark-index.json and site/data/radar.json
    with Cache-Control: public, max-age=2592000 (30 days).
  • On each publish, write the file under a versioned name
    (radar.v20260920.json) and update a tiny
    site/data/manifest.json that points at the active version.
    Update only the manifest, not the old file, when the underlying
    data changes. This is the same pattern OpenAlex uses for its
    snapshots.
  • Change the homepage Download link from /data/radar.json to
    https://huggingface.co/datasets/ktwu01/benchmark-radar, so
    bulk-data users go to HF and the Pages endpoint stops being the
    easiest way to pull the full file.
  • Result: most requests are served by Cloudflare's edge cache. The
    repo gets a real download channel. The dashboard still works
    because applyDashboardData already supports path-based loads.

PR 4 (optional, depends on measured traffic) — Slice radar.json per day

  • Split radar.json into radar-manifest.json (a list of available
    dates, ~10 KB) plus radar/<date>.json (one day, tens of KB).
  • The dashboard already lazy-loads by path; only the data layout
    changes.
  • Result: a visitor looking at one day no longer downloads all of
    history.

Why not "anti-scraping" first

Tempting, but it solves the wrong problem:

  • GNOME's traffic spike in 2026 came from CI tasks re-downloading
    the same artifact, not from hostile bots. Their fix was a Fastly
    cache layer, not IP blocking.
  • FreeCAD's Anubis deployment is useful in their context, but
    Anubis's own README warns that it can break legitimate bots
    (Internet Archive, etc.). Cloudflare's own docs warn that page
    challenges on JSON endpoints break CLI/agent consumers that
    expect JSON.
  • A public Hugging Face dataset is already the easiest way to get
    the full data; if we make that path pleasant, there is less
    reason to scrape the dashboard.
  • Our own domain's WAF does not protect github.com/... releases
    or huggingface.co/... downloads, so a domain-level rule cannot
    actually "shield" every public channel anyway.

Caching + a real download channel is what we have not yet done, and
it is the highest leverage step.

Boundary Conditions

  • HF public storage is best-effort and rate-limited; keep SHA-256
    checksums, retries, and a local fallback in data/snapshots/.
  • HF Datasets are versioned. Removing a file from the latest revision
    does not reclaim space from older revisions. Use HF Storage Buckets
    only for intermediate artifacts; the canonical catalog stays on the
    versioned Dataset.
  • GitHub Releases are NOT subject to the Pages 1 GB / 100 GB limits
    (per-file < 2 GiB, no documented total/bandwidth cap). The CLI ZIP
    flow in publish-cli-data does not need to change.
  • Git LFS is not a solution here: download bandwidth is still billed
    to the repo owner.
  • Old Git history still contains ~53 MB of snapshots even after
    these PRs land. Cleanup of historical blobs must be planned
    separately, never bundled with this migration.
  • The subproject at docs/technical-report/latex is a separate
    repository per AGENTS.md and is out of scope here.

Reference Implementations We Studied

  • LiveBench / SWE-bench: code on GitHub, full data on HF,
    with download_*.py scripts pointing at the dataset.
  • Our World in Data: layered raw → curated → web, with DVC for
    provenance and Cloudflare R2 for hot JSON.
  • OpenAlex: S3 snapshots with manifest written last, supports
    incremental partition-based sync.
  • GNOME (2026-04-17 infra post): caching layer for repeated
    downloads, not IP-level blocking.

Acceptance Checklist

  • data/snapshots/ size in git ls-tree stops growing on main
    (verify after 14 days).
  • Pages artifact size does not contain hf_dataset/ (verified
    by find site -name hf_dataset in the Pages build step).
  • radar.json and benchmark-index.json are served with a
    Cache-Control header (verified by curl -I).
  • Homepage Download button points to
    huggingface.co/datasets/ktwu01/benchmark-radar.
  • benchmark-radar classify and pytest -q both pass on a
    clean checkout that has only the pointer file in
    data/snapshots/.

Related

  • Closes the open piece of [6 points] Switching dashboard tabs waits on the full 70MB radar.json #528 (already-closed PR cut the client
    side; this issue covers the serving side).
  • Related to: HF dataset at
    https://huggingface.co/datasets/ktwu01/benchmark-radar.
  • Status note: we have not yet pulled Cloudflare access logs, so
    quantitative "how close to the limit are we" is a deliberate
    follow-up, not a blocker.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

enhancementNew feature or requestnew featureNew product capability, data workflow, or significant enhancement

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions