Skip to content

Import Benchmark-Radar.com corpus as a catalog source - #693

Open
Claire1217 wants to merge 14 commits into
mainfrom
data/import-claire-radar-corpus
Open

Claire1217 wants to merge 14 commits into
mainfrom
data/import-claire-radar-corpus

Conversation

@Claire1217

Copy link
Copy Markdown
Collaborator

Closes #692

Summary

Import the complete 1,914-record Benchmark-Radar.com snapshot into Benchmark-Radar.org as the claire_radar catalog source.

The import adds benchmark metadata, categories, release dates, artifact links, provenance, and review status without modifying existing score observations.

Data handling

  • Map compatible fields to the existing catalog schema.
  • Preserve every original source object under source_metadata.claire_radar.
  • Verify the snapshot checksum and each record's SHA-256 hash.
  • Retain accepted, deferred, excluded, and unreviewed records with their original review state.
  • Convert supported arXiv, GitHub, and Hugging Face URLs into catalog identity anchors.
  • Keep one stable source record per original source ID.
  • Do not merge records by name alone.
  • Leave ambiguous identity candidates separate for later review.

The catalog grows from 1,284 to 3,198 source records. These are source records, not necessarily 3,198 distinct benchmark instruments.

Score safety

This import contains metadata only. The existing 12,929 score observations, their protocols, source records, and citations remain unchanged.

Validation

The full clean-checkout CI sequence passes:

  • ruff check .
  • ruff format --check .
  • benchmark-radar normalize-catalog
  • benchmark-radar classify
  • benchmark-radar build-data-release
  • pytest -q — 1,404 passed

A generated Claire Radar benchmark detail page was also checked in the browser for metadata, categories, release date, and artifact links.

Copilot AI lite review requested due to automatic review settings September 27, 2026 08:30
@Claire1217

Copy link
Copy Markdown
Collaborator Author

330226 GPT

@Claire1217
Claire1217 marked this pull request as ready for review September 27, 2026 08:31

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Row-level record hash validation is not enforced during loading or normalization.

Review effort: Lite
Findings: None

What changed in this PR

Imports 1,914 Claire Radar records as a metadata-only catalog source while preserving provenance and existing scores.

Changes:

  • Adds snapshot registration, export tooling, and integrity tests.
  • Normalizes metadata, artifacts, review states, and identity anchors.
  • Updates search relevance fixtures.
File Summary
tests/​test_claire_radar_snapshot.py Tests snapshot export and corpus preservation.
tests/​test_catalog.py Tests catalog integrity behavior.
tests/​fixtures/​search_relevance.yml Updates search relevance expectations.
tests/​fixtures/​search_evaluation.yml Adds relevance judgments.
src/​benchmark_radar/​leaderboard_snapshots.py Loads snapshots and validates registry checksums; row-level digest validation remains unresolved.
src/​benchmark_radar/​catalog.py Normalizes imported metadata and artifacts; per-record digest validation remains unresolved.
scripts/​export_claire_radar_snapshot.py Exports and hashes source records.
data/​leaderboard_snapshots.yml Registers the Claire Radar snapshot.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 1537b27a90

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

SOURCES = {
LLM_STATS_SNAPSHOT_ID: (LLM_STATS_SOURCE, LLM_STATS_KEY_PREFIX),
ARTIFICIAL_ANALYSIS_SNAPSHOT_ID: (ARTIFICIAL_ANALYSIS_SOURCE, ARTIFICIAL_ANALYSIS_KEY_PREFIX),
"claire_radar_2026-09-25": ("claire_radar", "claire-radar"),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Omit score groups for this metadata-only source

Registering this snapshot sends all 1,914 rows through normalize_snapshot, which creates a series even when score_rows is empty. Consequently each Claire shard has a non-empty scores_by_source.claire_radar block with zero rows, whereas catalog_shards._scores_by_source uses {} to represent unscored records. The static page builder then treats every imported record as having “Results” in its title, and the dashboard shows a source-specific empty-score block instead of the normal no-scores state. Skip series creation for scoreless registry sources, or make shard construction require observations before emitting a source key.

Useful? React with 👍 / 👎.

return ""
if host in repo_hosts | paper_hosts | dataset_hosts or host == "huggingface.co":
return ""
return text

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Reject loopback URLs from project links

The checked-in export reaches this fallback for http://localhost:8737 on Frontier Agent Benchmark and http://127.0.0.1:8080/v1\`` on LLMs Benchmark, so they are emitted as Website artifacts. Generated static benchmark pages place artifact URLs directly in href` attributes, meaning the former sends a visitor to a service on their own machine rather than to benchmark documentation. Reject loopback/private hosts (and malformed URLs) before retaining a project link.

Useful? React with 👍 / 👎.

@junjiezhou1122 junjiezhou1122 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed commit 1537b27a90 against issue #692 and the repository's full-corpus rules. The source-row import is lossless and the existing 12,929 score observations remain unchanged, but two generated-product behaviors need changes before merge.

1. Exact source identities are still published as duplicate benchmarks (P1)

src/benchmark_radar/catalog.py:226-240 now gives Claire artifact URLs stable arXiv, GitHub, and Hugging Face identity anchors. However, src/benchmark_radar/catalog.py:467-489 still creates a new claire-radar:* record for every imported row. The candidate path at src/benchmark_radar/catalog_identity.py:118-121 requires at least two shared anchors, and this PR does not add reviewed identity decisions for these records.

A clean normalization produced two current, unambiguous cross-source overlaps:

  • claire-radar:2608.05948 and opencompass:2574 are both GAUGE and share arxiv:2608.05948.
  • claire-radar:2608.09548 and opencompass:2571 are both ELBench and share hf:zeroloss-lab/elbench.

Neither pair reaches the reviewed identity data, so search publishes two GAUGE records and two ELBench records. This conflicts with #692's acceptance criterion that an exact identity match must not duplicate an existing benchmark.

Please add an import identity step that enriches one unambiguous existing target while preserving the full Claire object and provenance. Keep ambiguous names or versions as separate source records. Add integration tests for GAUGE and ELBench that assert one searchable benchmark identity, retained Claire metadata, and unchanged existing scores.

2. Unscored Claire pages advertise "Results" (P2)

src/benchmark_radar/catalog.py:480-489 creates a score series for every Claire row, although this snapshot contains no score observations. src/benchmark_radar/catalog_shards.py:55-65 then emits a source bucket for the empty series, and src/benchmark_radar/site_pages.py:206-208 adds the Results title facet whenever that bucket exists.

For example, the generated GAUGE page title is:

GAUGE Benchmark (claire radar): Paper & Results | Benchmark Radar

The same page reports zero scores and says that no reported scores are on record. This path applies to all 1,914 imported rows.

Please omit empty score buckets, or make the page facet depend on actual score rows. Add a generated-page regression test for an imported record without scores.

Validation

I checked the import in a disposable clean archive:

  • 1,914 unique Claire records reached the catalog, shards, search, and offline release.
  • The build produced 3,198 index rows and 3,198 shards.
  • Existing score observations remained at 12,929 rows with the same canonical SHA-256 as the base revision.
  • 74 focused catalog and search tests passed after normalization.
  • 1,398 tests passed when technical-report tests were excluded. The report submodule was not initialized in the review checkout.
  • git merge-tree found no textual conflict with the current local origin/main; the blockers above are semantic.

AI-assisted review by GPT. Marker: 330226.

@ktwu01

ktwu01 commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Thanks for your effort Claire and Junjie

Please show the importer alongside the original evidence URL and review status. Renaming alone shouldn’t imply that every imported description or date has been verified.

Two concrete overlaps remain unlinked:

  • GAUGE: claire-radar:2608.05948 and opencompass:2574 share the same arXiv paper.
  • ELBench: claire-radar:2608.09548 and opencompass:2571 share the same Hugging Face dataset.

Please reconcile these through reviewed identity links and test that search/detail exposes the relationship. Preserve all imported source records and provenance, with scores attached to their original records. Another GAUGE record cites a different paper, so matching names alone isn’t sufficient.

Please also address the existing findings about empty “Results” sections and localhost project links. One additional date issue: DeepSWE retains a first-public date of May 18 but imports its July 8 paper date as “benchmark release.” LoopsBench and EdgeBench have similar mismatches. Please distinguish those dates and retain their supporting evidence.

The rebuild preserves all 3,198 source records and 12,929 score observations. The remaining concern is how identities and evidence are represented.

Copilot AI review requested due to automatic review settings September 28, 2026 07:26

@junjiezhou1122 junjiezhou1122 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

330226 GPT

Re-reviewed commit 748bf9dcb70d421d3dac766867fef8401488b222. The requested changes are resolved:

  • GAUGE and ELBench are reviewed, bidirectional sibling links. Both source records remain intact, neither inherits identity fields or scores, and the separate GAUGE paper remains unlinked.
  • Claire's 1,914 scoreless records no longer create empty score series or scores_by_source buckets. Legacy zero-observation series remain for sources that use them for coverage auditing.
  • The Claire import boundary rejects malformed, localhost, loopback, Unicode-equivalent, numeric-IPv4, mapped-IPv6, backslash-authority, and empty-port project URLs. The four bad derived CSV cells are cleared while all embedded original objects and record hashes remain unchanged.
  • Static and interactive details show importer, original evidence, and recorded review status. Missing review metadata is omitted rather than described as "not reviewed". Static evidence links accept only HTTP(S).
  • DeepSWE, LoopsBench, and EdgeBench retain separate first-public and paper-first-version dates with explicit evidence. The evidence mapping is keyed by exact source ID in the snapshot registry.
  • Search and detail expose reviewed siblings. Score rows stay attached to their original source records and remain partitioned by source.

Verification:

  • Clean-checkout CI sequence passed in the documented order.
  • 1446 passed.
  • Generated output contains 3,198 records and 3,198 shards.
  • Claire contributes all 1,914 records and zero score observations.
  • The catalog retains all 12,929 score observations.
  • Direct generated-artifact checks covered index, shards, search, detail, static pages, date evidence, provenance, URL sanitation, and source-object hashes.
  • Final independent review reported no findings in scope.

Principles and decisions:

  • Foundational Thinking kept source records, reviewed identity links, score ownership, provenance, and date facts as separate structures.
  • Redesign From First Principles replaced merge/dedup behavior with source-preserving reviewed sibling links.
  • Boundary Discipline put URL validation at export and link-rendering boundaries instead of scattering display-only guards.
  • Fix Root Causes removed empty score buckets and false summaries in normalization/shard/index generation rather than hiding labels in one UI.
  • Test Behavior, Not Implementation verifies query results, generated shards, generated pages, and actual bad URL forms.
  • Sequence Verifiable Units kept the failing repro in b5f1485 before the implementation in 748bf9d.
  • Laziness Protocol kept the CSV change to four derived cells and moved the three reviewed date citations to an exact-key registry mapping, avoiding a 1,914-row rewrite.
  • Prove It Works required the full clean-checkout CI and direct generated-product assertions before approval.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 748bf9dcb7

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

SOURCES = {
LLM_STATS_SNAPSHOT_ID: (LLM_STATS_SOURCE, LLM_STATS_KEY_PREFIX),
ARTIFICIAL_ANALYSIS_SNAPSHOT_ID: (ARTIFICIAL_ANALYSIS_SOURCE, ARTIFICIAL_ANALYSIS_KEY_PREFIX),
"claire_radar_2026-09-25": ("claire_radar", "claire-radar"),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Require Claire Radar in export source validation

Adding this registry entry makes Claire Radar part of the catalog, but REQUIRED_CATALOG_SOURCES in src/benchmark_radar/hf_dataset.py still omits claire_radar. As a result, a build that silently drops the entire 1,914-record Claire snapshot can still pass the Hugging Face export checks because the four older sources and the minimum-count guard remain satisfied, publishing an incomplete corpus without an error. Add the new source to the required-source invariant (and its tests).【AGENTS.md†L3-L6】

Useful? React with 👍 / 👎.

SOURCES = {
LLM_STATS_SNAPSHOT_ID: (LLM_STATS_SOURCE, LLM_STATS_KEY_PREFIX),
ARTIFICIAL_ANALYSIS_SNAPSHOT_ID: (ARTIFICIAL_ANALYSIS_SOURCE, ARTIFICIAL_ANALYSIS_KEY_PREFIX),
"claire_radar_2026-09-25": ("claire_radar", "claire-radar"),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Name Claire Radar in the generated dataset documentation

This source registration causes the generated catalog export to contain Claire Radar records, but generate_dataset_card() still describes the catalog as spanning only LLM Stats, OpenCompass Hub, Artificial Analysis, and model reports, and its source-provider field lists only those four values. Consumers reading the published dataset card therefore cannot tell that the largest newly imported source is present; update the generated overview and source vocabulary alongside this registration.【AGENTS.md†L75-L82】

Useful? React with 👍 / 👎.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unresolved record-hash verification and URL validation issues remain.

Review effort: Lite
Findings: 1 High severity

Open (1)

Comment thread src/benchmark_radar/catalog.py Outdated
Comment on lines +262 to +264
source_metadata = json_object(
row.get("extra_json", ""), label=f"{snapshot_id}:{source_id}:extra_json"
)
Copilot AI review requested due to automatic review settings September 28, 2026 08:07

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

junjiezhou1122
junjiezhou1122 previously approved these changes Sep 28, 2026

@junjiezhou1122 junjiezhou1122 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

330226 GPT

Re-reviewed head b418e4f34a5f1dd84e3c927bb7ebaae294d613b7 after the source-neutral normalization refactor.

The two remaining patch-specific branches are gone:

  • Empty-series behavior is declared by each registered snapshot as preserve_empty or observed_only. Catalog normalization applies the policy without checking a source name. Legacy coverage series remain intact; Claire emits no empty series.
  • Source-private row schemas are translated by registry-selected adapters before catalog normalization. The Claire adapter owns releaseDates, while catalog.py consumes only common released, released_basis, released_source_url, publication_dates, and source_metadata fields.

Additional boundary behavior:

  • Missing or unknown score-series policies fail at registry load.
  • Unknown adapters and malformed adapter options fail visibly.
  • Normalized date bases and ISO dates are validated at the common catalog boundary.
  • Source metadata is retained generically under the actual source key.
  • Exact source IDs select reviewed first-public evidence; benchmark names are never used for evidence lookup.

Verification on this exact head:

  • GitHub Actions CI run 36395499343 passed all six stages.
  • A separate clean detached worktree with initialized submodules passed the same six-stage sequence.
  • 1458 passed in the clean worktree.
  • Generated output remains 3,198 records, 3,198 shards, and 12,929 score observations.
  • All 1,914 Claire records remain, with zero Claire score series/observations.
  • Legacy series counts remain 687 for LLM Stats and 25 for Artificial Analysis.
  • All six Claire private date objects are translated; reviewed evidence for EdgeBench, DeepSWE, and LoopsBench remains exact-ID keyed.
  • The Claire CSV SHA-256 remains cb15e3e585c4234517f4dfe6f93235980b3acf762f093739caea4338e77f166d; all 1,914 embedded source-object hashes still validate.
  • catalog.py contains none of releaseDates, firstPublicAt, paperV1At, or the prior Claire-specific score condition.

Principles and decisions:

  • Model the Domain introduced explicit snapshot policies and an adapter registry instead of growing source-name conditionals.
  • Redesign From First Principles made source translation a pre-normalization boundary, as if multiple private source schemas had existed from day one.
  • Boundary Discipline validates registry policy, adapter selection, adapter options, and normalized dates at their owning boundaries.
  • Test Behavior, Not Implementation uses synthetic non-Claire snapshots to prove both series policies and the common date contract.
  • Sequence Verifiable Units committed failing architecture tests before implementation and kept malformed-date tests in their own red step.
  • Prove It Works checked the real generated corpus, source-object hashes, clean-checkout CI, and GitHub CI on the exact reviewed SHA.

@junjiezhou1122 junjiezhou1122 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

330226 GPT

Re-reviewed head b418e4f34a5f1dd84e3c927bb7ebaae294d613b7 after the source-neutral normalization refactor.

The two remaining patch-specific branches are gone:

  • Empty-series behavior is declared by each registered snapshot as preserve_empty or observed_only. Catalog normalization applies the policy without checking a source name. Legacy coverage series remain intact; Claire emits no empty series.
  • Source-private row schemas are translated by registry-selected adapters before catalog normalization. The Claire adapter owns releaseDates, while catalog.py consumes only common released, released_basis, released_source_url, publication_dates, and source_metadata fields.

Additional boundary behavior:

  • Missing or unknown score-series policies fail at registry load.
  • Unknown adapters and malformed adapter options fail visibly.
  • Normalized date bases and ISO dates are validated at the common catalog boundary.
  • Source metadata is retained generically under the actual source key.
  • Exact source IDs select reviewed first-public evidence; benchmark names are never used for evidence lookup.

Verification on this exact head:

  • GitHub Actions CI run 36395499343 passed all six stages.
  • A separate clean detached worktree with initialized submodules passed the same six-stage sequence.
  • 1458 passed in the clean worktree.
  • Generated output remains 3,198 records, 3,198 shards, and 12,929 score observations.
  • All 1,914 Claire records remain, with zero Claire score series/observations.
  • Legacy series counts remain 687 for LLM Stats and 25 for Artificial Analysis.
  • All six Claire private date objects are translated; reviewed evidence for EdgeBench, DeepSWE, and LoopsBench remains exact-ID keyed.
  • The Claire CSV SHA-256 remains cb15e3e585c4234517f4dfe6f93235980b3acf762f093739caea4338e77f166d; all 1,914 embedded source-object hashes still validate.
  • catalog.py contains none of releaseDates, firstPublicAt, paperV1At, or the prior Claire-specific score condition.

Principles and decisions:

  • Model the Domain introduced explicit snapshot policies and an adapter registry instead of growing source-name conditionals.
  • Redesign From First Principles made source translation a pre-normalization boundary, as if multiple private source schemas had existed from day one.
  • Boundary Discipline validates registry policy, adapter selection, adapter options, and normalized dates at their owning boundaries.
  • Test Behavior, Not Implementation uses synthetic non-Claire snapshots to prove both series policies and the common date contract.
  • Sequence Verifiable Units committed failing architecture tests before implementation and kept malformed-date tests in their own red step.
  • Prove It Works checked the real generated corpus, source-object hashes, clean-checkout CI, and GitHub CI on the exact reviewed SHA.

@junjiezhou1122
junjiezhou1122 dismissed their stale review September 28, 2026 08:12

Duplicate approval created despite a client-side TLS timeout; superseded by review 5335692320 on the same commit.

series = series_by_key.get(key)
rows = observations_by_key.get(key)
if not series and not rows:
if not rows:

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This condition change drops known scale metadata from 8 existing LLM Stats records.

Before, a record with a score series but zero observation rows kept its scores_by_source entry. With if not rows, those records now get {} in their shard and score_summary: null in benchmark-index.json. Affected records (compared against a build of current main):

  • llm-stats:cvtg-2k
  • llm-stats:longtext-bench
  • 6 llm-stats:community-* records (5f95f778…, ed90e889…, 2256e9c9…, fd462fc2…, 07c9946d…, 64d67847…)

Their series carry source-declared evidence: declared_max: 1.0, bounds.basis: aggregator_declared, direction: higher_is_better, direction_basis: source_rank_descending. No numeric score is lost, and the records stay in the catalog. But known values become unknown, which conflicts with principle.md ("Unknown values stay unknown", "Keep score values, units, protocols, dates and citations attached to their source record").

Suggested fix: keep if not series and not rows: so a declared series with zero observations still ships, and handle the Claire case (no series, no rows) through the existing branch. The _score_count change in site_pages.py already keeps the "Results" title facet correct for zero-row series.

minimum_candidates: 0
must_include_candidates: []
kind: topical
must_include: [RealBench, RESCAST-100K]

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This fixture now treats "weather forecasting" as a resolved topical query, but the website search box still returns 0 results for it on this branch.

Checked on a local build of this PR (merged with current main, 3,212 records), Saturation tab:

Query QueryService (this fixture) Website search
weather forecasting RealBench, RESCAST-100K 0 of 0 matches
RealBench n/a 2 matches (both Claire Radar)

Cause: the two target records match only through their description. searchBenchmarkIndex in site/assets/app.js checks a folded substring of name, aliases, publisher, modality and categories. It never reads descriptions. The imported categories don't help either: RealBench, a numerical weather prediction benchmark, is tagged ["General AI", "Language & Knowledge", "cs.LG"].

So "the catalog now covers weather forecasting" holds for the CLI, HTTP and Skill surfaces, not for a reader using the site. The same applies to robot grasp planning → R2HandoverSim.

Two related points on the fixture change:

  • RESCAST-100K is a residential load and indoor-temperature forecasting benchmark that uses weather as an input covariate. Labelling it relevant for "weather forecasting" is a judgement call. docs/query-surfaces.md says a source refresh that changes a judgement requires record review, not a mechanical fixture update.
  • The catalog_gap → topical flip removes the only case that recorded this as a gap. If the website search stays name-based, the gap still exists for site readers.

Suggested resolution, either:

  1. scope the claim: keep these as QueryService-only cases and don't describe weather forecasting as newly searchable on the site; or
  2. make the site's catalog search reach descriptions (or route it through the same index QueryService uses), so both surfaces agree.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Merge Benchmark-Radar.com benchmark data into Benchmark-Radar.org

4 participants