You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Decouple data growth from the GitHub code repository, stop pushing the
complete dataset through GitHub Pages, and stop relying on "anti-scraping"
as a solution. The current workflow keeps growing data/snapshots/ on main (~53 MB today, ~780 MB projected in a year) and still serves radar.json to any visitor who clicks the homepage Download button.
Problem
Three concrete signals, measured in this repository today:
data/snapshots/ contains 58 daily JSON files, ~53 MB total,
adding ~2 MB per day. At that rate it crosses the GitHub Pages
1 GB site-size soft limit within a year.
site/data/hf_dataset/ is generated by benchmark-radar export-hf
and then uploaded as part of the Pages artifact. Even though the
same data is published to huggingface.co/datasets/ktwu01/benchmark-radar,
a copy still ships inside the Pages bundle.
site/index.html exposes <a href="/data/radar.json">Download</a>
and several JS paths still fetch the full radar.json. With ~2 MB
per request and no cache headers, a small amount of repeated
traffic puts real pressure on the 100 GB / month Pages soft limit.
These are not separate problems. The fix is the same shape: stop
making the code repository the canonical place for growing data, stop
making Pages the canonical place for the complete dataset, and use
caching + a real download channel for what remains.
Proposed Solution
Split responsibilities along the lines of LiveBench, SWE-bench,
Our World in Data, and OpenAlex:
GitHub keeps code, configuration, human-reviewed identity rules,
small fixtures, and a tiny version-pointer file.
Hugging Face keeps the complete daily snapshots, the exported
catalog, and versioned releases with SHA-256 checksums.
The Pages site keeps only what a visitor needs to browse
(index, per-benchmark shards, first-paint summaries, optionally
per-day radar slices), and uses long-lived cache headers.
Land this in four small, low-risk PRs, each passing the full CI
sequence in AGENTS.md.
PR 1 — Stop shipping the HF export inside the Pages bundle
Change hf_dataset.py's default export directory from site/data/hf_dataset to a build-only path such as build/hf_dataset.
In .github/workflows/pages.yml, rm -rf site/data/hf_dataset
after export-hf and before the Pages upload step. The HF
publish job is unchanged.
Result: the Pages artifact no longer carries a duplicate of what
Hugging Face already hosts. Zero behavioral change for visitors.
PR 2 — Stop committing daily snapshots to main
In .github/workflows/daily-radar.yml, replace git add data/snapshots && git commit && git push with an upload
step that syncs snapshots/<date>.json to huggingface.co/datasets/ktwu01/benchmark-radar.
Add a small data/snapshot-pointer.yml that records the current
snapshot set name and HF revision. snapshots.py reads it on
startup, with the existing data/snapshots/ directory as the
offline fallback for tests.
Keep test fixtures (tests/test_social.py already uses data/snapshots/2026-08-10.json) intact.
Result: the code repository stops growing. A fresh checkout still
builds, because the pointer file is committed and HF is the
canonical source.
PR 3 — Cache the public JSON and redirect the Download button
Serve site/data/benchmark-index.json and site/data/radar.json
with Cache-Control: public, max-age=2592000 (30 days).
On each publish, write the file under a versioned name
(radar.v20260920.json) and update a tiny site/data/manifest.json that points at the active version.
Update only the manifest, not the old file, when the underlying
data changes. This is the same pattern OpenAlex uses for its
snapshots.
Change the homepage Download link from /data/radar.json to https://huggingface.co/datasets/ktwu01/benchmark-radar, so
bulk-data users go to HF and the Pages endpoint stops being the
easiest way to pull the full file.
Result: most requests are served by Cloudflare's edge cache. The
repo gets a real download channel. The dashboard still works
because applyDashboardData already supports path-based loads.
PR 4 (optional, depends on measured traffic) — Slice radar.json per day
Split radar.json into radar-manifest.json (a list of available
dates, ~10 KB) plus radar/<date>.json (one day, tens of KB).
The dashboard already lazy-loads by path; only the data layout
changes.
Result: a visitor looking at one day no longer downloads all of
history.
Why not "anti-scraping" first
Tempting, but it solves the wrong problem:
GNOME's traffic spike in 2026 came from CI tasks re-downloading
the same artifact, not from hostile bots. Their fix was a Fastly
cache layer, not IP blocking.
FreeCAD's Anubis deployment is useful in their context, but
Anubis's own README warns that it can break legitimate bots
(Internet Archive, etc.). Cloudflare's own docs warn that page
challenges on JSON endpoints break CLI/agent consumers that
expect JSON.
A public Hugging Face dataset is already the easiest way to get
the full data; if we make that path pleasant, there is less
reason to scrape the dashboard.
Our own domain's WAF does not protect github.com/... releases
or huggingface.co/... downloads, so a domain-level rule cannot
actually "shield" every public channel anyway.
Caching + a real download channel is what we have not yet done, and
it is the highest leverage step.
Boundary Conditions
HF public storage is best-effort and rate-limited; keep SHA-256
checksums, retries, and a local fallback in data/snapshots/.
HF Datasets are versioned. Removing a file from the latest revision
does not reclaim space from older revisions. Use HF Storage Buckets
only for intermediate artifacts; the canonical catalog stays on the
versioned Dataset.
GitHub Releases are NOT subject to the Pages 1 GB / 100 GB limits
(per-file < 2 GiB, no documented total/bandwidth cap). The CLI ZIP
flow in publish-cli-data does not need to change.
Git LFS is not a solution here: download bandwidth is still billed
to the repo owner.
Old Git history still contains ~53 MB of snapshots even after
these PRs land. Cleanup of historical blobs must be planned
separately, never bundled with this migration.
The subproject at docs/technical-report/latex is a separate
repository per AGENTS.md and is out of scope here.
Reference Implementations We Studied
LiveBench / SWE-bench: code on GitHub, full data on HF,
with download_*.py scripts pointing at the dataset.
Our World in Data: layered raw → curated → web, with DVC for
provenance and Cloudflare R2 for hot JSON.
OpenAlex: S3 snapshots with manifest written last, supports
incremental partition-based sync.
GNOME (2026-04-17 infra post): caching layer for repeated
downloads, not IP-level blocking.
Acceptance Checklist
data/snapshots/ size in git ls-tree stops growing on main
(verify after 14 days).
Pages artifact size does not contain hf_dataset/ (verified
by find site -name hf_dataset in the Pages build step).
radar.json and benchmark-index.json are served with a Cache-Control header (verified by curl -I).
Homepage Download button points to huggingface.co/datasets/ktwu01/benchmark-radar.
benchmark-radar classify and pytest -q both pass on a
clean checkout that has only the pointer file in data/snapshots/.
Summary
Decouple data growth from the GitHub code repository, stop pushing the
complete dataset through GitHub Pages, and stop relying on "anti-scraping"
as a solution. The current workflow keeps growing
data/snapshots/onmain(~53 MB today, ~780 MB projected in a year) and still servesradar.jsonto any visitor who clicks the homepage Download button.Problem
Three concrete signals, measured in this repository today:
data/snapshots/contains 58 daily JSON files, ~53 MB total,adding ~2 MB per day. At that rate it crosses the GitHub Pages
1 GB site-size soft limit within a year.
site/data/hf_dataset/is generated bybenchmark-radar export-hfand then uploaded as part of the Pages artifact. Even though the
same data is published to
huggingface.co/datasets/ktwu01/benchmark-radar,a copy still ships inside the Pages bundle.
site/index.htmlexposes<a href="/data/radar.json">Download</a>and several JS paths still fetch the full
radar.json. With ~2 MBper request and no cache headers, a small amount of repeated
traffic puts real pressure on the 100 GB / month Pages soft limit.
These are not separate problems. The fix is the same shape: stop
making the code repository the canonical place for growing data, stop
making Pages the canonical place for the complete dataset, and use
caching + a real download channel for what remains.
Proposed Solution
Split responsibilities along the lines of LiveBench, SWE-bench,
Our World in Data, and OpenAlex:
small fixtures, and a tiny version-pointer file.
catalog, and versioned releases with SHA-256 checksums.
(index, per-benchmark shards, first-paint summaries, optionally
per-day radar slices), and uses long-lived cache headers.
Land this in four small, low-risk PRs, each passing the full CI
sequence in
AGENTS.md.PR 1 — Stop shipping the HF export inside the Pages bundle
hf_dataset.py's default export directory fromsite/data/hf_datasetto a build-only path such asbuild/hf_dataset..github/workflows/pages.yml,rm -rf site/data/hf_datasetafter
export-hfand before the Pages upload step. The HFpublish job is unchanged.
Hugging Face already hosts. Zero behavioral change for visitors.
PR 2 — Stop committing daily snapshots to
main.github/workflows/daily-radar.yml, replacegit add data/snapshots && git commit && git pushwith an uploadstep that syncs
snapshots/<date>.jsontohuggingface.co/datasets/ktwu01/benchmark-radar.data/snapshot-pointer.ymlthat records the currentsnapshot set name and HF revision.
snapshots.pyreads it onstartup, with the existing
data/snapshots/directory as theoffline fallback for tests.
tests/test_social.pyalready usesdata/snapshots/2026-08-10.json) intact.builds, because the pointer file is committed and HF is the
canonical source.
PR 3 — Cache the public JSON and redirect the Download button
site/data/benchmark-index.jsonandsite/data/radar.jsonwith
Cache-Control: public, max-age=2592000(30 days).(
radar.v20260920.json) and update a tinysite/data/manifest.jsonthat points at the active version.Update only the manifest, not the old file, when the underlying
data changes. This is the same pattern OpenAlex uses for its
snapshots.
/data/radar.jsontohttps://huggingface.co/datasets/ktwu01/benchmark-radar, sobulk-data users go to HF and the Pages endpoint stops being the
easiest way to pull the full file.
repo gets a real download channel. The dashboard still works
because
applyDashboardDataalready supports path-based loads.PR 4 (optional, depends on measured traffic) — Slice
radar.jsonper dayradar.jsonintoradar-manifest.json(a list of availabledates, ~10 KB) plus
radar/<date>.json(one day, tens of KB).changes.
history.
Why not "anti-scraping" first
Tempting, but it solves the wrong problem:
the same artifact, not from hostile bots. Their fix was a Fastly
cache layer, not IP blocking.
Anubis's own README warns that it can break legitimate bots
(Internet Archive, etc.). Cloudflare's own docs warn that page
challenges on JSON endpoints break CLI/agent consumers that
expect JSON.
the full data; if we make that path pleasant, there is less
reason to scrape the dashboard.
github.com/...releasesor
huggingface.co/...downloads, so a domain-level rule cannotactually "shield" every public channel anyway.
Caching + a real download channel is what we have not yet done, and
it is the highest leverage step.
Boundary Conditions
checksums, retries, and a local fallback in
data/snapshots/.does not reclaim space from older revisions. Use HF Storage Buckets
only for intermediate artifacts; the canonical catalog stays on the
versioned Dataset.
(per-file < 2 GiB, no documented total/bandwidth cap). The CLI ZIP
flow in
publish-cli-datadoes not need to change.to the repo owner.
these PRs land. Cleanup of historical blobs must be planned
separately, never bundled with this migration.
docs/technical-report/latexis a separaterepository per
AGENTS.mdand is out of scope here.Reference Implementations We Studied
with
download_*.pyscripts pointing at the dataset.provenance and Cloudflare R2 for hot JSON.
incremental partition-based sync.
downloads, not IP-level blocking.
Acceptance Checklist
data/snapshots/size ingit ls-treestops growing onmain(verify after 14 days).
hf_dataset/(verifiedby
find site -name hf_datasetin the Pages build step).radar.jsonandbenchmark-index.jsonare served with aCache-Controlheader (verified bycurl -I).huggingface.co/datasets/ktwu01/benchmark-radar.benchmark-radar classifyandpytest -qboth pass on aclean checkout that has only the pointer file in
data/snapshots/.Related
side; this issue covers the serving side).
https://huggingface.co/datasets/ktwu01/benchmark-radar.quantitative "how close to the limit are we" is a deliberate
follow-up, not a blocker.