From a4576c81923c6a836cdc622571dfb48617d57c31 Mon Sep 17 00:00:00 2001 From: ktwu01 Date: Sat, 3 Oct 2026 00:33:49 -0500 Subject: [PATCH 01/29] chore: remove implemented superpowers plan and spec catalog_snapshot_adapters.py shipped the plan; the checklists no longer track anything. Co-Authored-By: Claude Opus 5.5 --- ...8-source-neutral-snapshot-normalization.md | 166 ------------------ ...e-neutral-snapshot-normalization-design.md | 164 ----------------- 2 files changed, 330 deletions(-) delete mode 100644 docs/superpowers/plans/2026-09-28-source-neutral-snapshot-normalization.md delete mode 100644 docs/superpowers/specs/2026-09-28-source-neutral-snapshot-normalization-design.md diff --git a/docs/superpowers/plans/2026-09-28-source-neutral-snapshot-normalization.md b/docs/superpowers/plans/2026-09-28-source-neutral-snapshot-normalization.md deleted file mode 100644 index 62427016..00000000 --- a/docs/superpowers/plans/2026-09-28-source-neutral-snapshot-normalization.md +++ /dev/null @@ -1,166 +0,0 @@ -# Source-Neutral Snapshot Normalization Implementation Plan - -> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. - -**Goal:** Remove Claire-specific date and score-series behavior from catalog normalization by translating source rows through validated adapters and applying registry-declared series policies. - -**Architecture:** `leaderboard_snapshots.py` validates declarative snapshot policy and adapter configuration. A new `catalog_snapshot_adapters.py` module translates parsed source rows into a common catalog-row contract. `catalog.py` consumes only that contract and never branches on Claire or reads Claire-private date keys. - -**Tech Stack:** Python 3, PyYAML, pytest, Ruff, committed CSV/YAML snapshot data. - -**Spec:** `docs/superpowers/specs/2026-09-28-source-neutral-snapshot-normalization-design.md` - -## Global Constraints - -- Preserve 3,198 catalog records and shards. -- Preserve all 1,914 Claire Radar records and all 12,929 score observations. -- Do not modify Claire `extra_json` objects or `record_sha256` values. -- Scores remain attached to their original source records and partitioned by source. -- Legacy snapshots retain empty score series; Claire emits series only for observations. -- Reviewed date evidence is selected by exact source ID, never benchmark name. -- `catalog.py` must not interpret Claire-private date keys or branch on Claire for score behavior. -- Do not merge identity-linked source records or change taxonomy. - ---- - -### Task 1: Lock the source-neutral contracts with failing tests - -**Files:** -- Modify: `tests/test_catalog.py` -- Modify: `tests/test_claire_radar_snapshot.py` - -**Interfaces:** -- Consumes: existing `load_snapshots()` and `normalize_snapshot()` entry points. -- Produces: regression requirements for `score_series_policy`, `catalog_adapter`, normalized dates, generic source metadata, and the architecture boundary. - -- [ ] Add registry fixture support for `score_series_policy` and tests showing missing/unknown policies fail. -- [ ] Add tests showing unknown adapters fail. -- [ ] Add synthetic snapshot tests showing `observed_only` omits empty series and `preserve_empty` retains them without checking a source name. -- [ ] Add Claire tests requiring adapter configuration and all six `releaseDates` objects to be translated. -- [ ] Add a structural test proving `catalog.py` does not contain `releaseDates`, `firstPublicAt`, `paperV1At`, or a Claire-specific score condition. -- [ ] Run the focused tests and record the expected failures before production edits. - -Run: - -```bash -uv run --frozen --extra dev pytest -q \ - tests/test_catalog.py \ - tests/test_claire_radar_snapshot.py -``` - -Expected: failures for missing policy validation, missing adapter validation/translation, and source-specific normalizer code. - -- [ ] Commit the failing tests. - -```bash -git add tests/test_catalog.py tests/test_claire_radar_snapshot.py -git commit -m "test: require source-neutral snapshot normalization" -``` - -### Task 2: Validate snapshot policy and adapter configuration - -**Files:** -- Create: `src/benchmark_radar/catalog_snapshot_adapters.py` -- Modify: `src/benchmark_radar/leaderboard_snapshots.py` -- Modify: `data/leaderboard_snapshots.yml` - -**Interfaces:** -- Produces: `SCORE_SERIES_POLICIES`, `CATALOG_ADAPTERS`, and validated snapshot fields `score_series_policy`, `catalog_adapter`, and `adapter_options`. -- `score_series_policy` is exactly `preserve_empty` or `observed_only`. -- `catalog_adapter` defaults to `identity` and must be registered. - -- [ ] Implement the adapter registry names and snapshot configuration validators. -- [ ] Require every registry entry to declare `score_series_policy`. -- [ ] Set existing snapshots to `preserve_empty` and Claire to `observed_only`. -- [ ] Move Claire first-public evidence under `adapter_options.first_public_evidence` and declare `catalog_adapter: claire_radar_v1`. -- [ ] Run loader-focused tests until policy and adapter validation pass. - -Run: - -```bash -uv run --frozen --extra dev pytest -q \ - tests/test_catalog.py -k 'loader or policy or adapter' \ - tests/test_claire_radar_snapshot.py -k 'registered_snapshot' -``` - -Expected: PASS for registry validation; date translation tests remain red until Task 3. - -### Task 3: Translate rows through the adapter boundary - -**Files:** -- Modify: `src/benchmark_radar/catalog_snapshot_adapters.py` -- Modify: `src/benchmark_radar/catalog.py` -- Modify: `tests/test_catalog.py` -- Modify: `tests/test_claire_radar_snapshot.py` - -**Interfaces:** -- Produces: `adapt_catalog_row(row, *, adapter, adapter_options, snapshot_id) -> dict[str, Any]`. -- Common adapted fields are `source_metadata`, `released`, `released_basis`, `released_source_url`, and `publication_dates`. -- Raises `CatalogSnapshotAdapterError` for malformed source objects or adapter options. - -- [ ] Implement the identity adapter, including generic `extra_json` preservation as `source_metadata`. -- [ ] Implement `claire_radar_v1`, translating exact-ID first-public evidence and paper-version dates. -- [ ] Make `normalize_snapshot()` adapt each benchmark row before record/series generation. -- [ ] Simplify `_source_record()` to consume and validate only common fields. -- [ ] Store source metadata under `{source: metadata}`. -- [ ] Apply `score_series_policy` instead of checking the source name. -- [ ] Run all focused catalog and Claire tests until green. - -Run: - -```bash -uv run --frozen --extra dev pytest -q \ - tests/test_catalog.py \ - tests/test_claire_radar_snapshot.py \ - tests/test_artificial_analysis_snapshot.py -``` - -Expected: PASS. - -- [ ] Commit the implementation. - -```bash -git add data/leaderboard_snapshots.yml \ - src/benchmark_radar/catalog_snapshot_adapters.py \ - src/benchmark_radar/leaderboard_snapshots.py \ - src/benchmark_radar/catalog.py \ - tests/test_catalog.py tests/test_claire_radar_snapshot.py -git commit -m "refactor: normalize snapshots through source adapters" -``` - -### Task 4: Verify generated products and full corpus invariants - -**Files:** -- Modify only if a behavior regression is found in code already in scope. - -**Interfaces:** -- Consumes: generated catalog, shards, index, search, detail pages, and release artifacts. -- Produces: evidence that the refactor changed architecture without changing the published corpus. - -- [ ] Run focused generation commands and assert 3,198 records, 1,914 Claire records, 3,198 shards, and 12,929 observations. -- [ ] Verify reviewed GAUGE/ELBench siblings, Claire empty score groups, provenance, date evidence, and source-owned scores. -- [ ] Run `git diff --check` and confirm the Claire CSV checksum remains `cb15e3e585c4234517f4dfe6f93235980b3acf762f093739caea4338e77f166d`. -- [ ] Run the documented six-stage CI in a clean detached worktree: - -```bash -ruff check . -ruff format --check . -benchmark-radar normalize-catalog -benchmark-radar classify -benchmark-radar build-data-release -pytest -q -``` - -Expected: every command exits zero. - -### Task 5: Review and update PR #693 - -**Files:** -- No source changes unless review identifies a bounded defect. - -- [ ] Run `no-comments` and `deslop` checks required by the project workflow. -- [ ] Inspect the final diff and commit history. -- [ ] Push commits to `origin/data/import-claire-radar-corpus`. -- [ ] Wait for GitHub CI on the new head. -- [ ] Re-review the new head and document how registry policy and adapters removed the two source-specific branches. -- [ ] Do not merge the PR. diff --git a/docs/superpowers/specs/2026-09-28-source-neutral-snapshot-normalization-design.md b/docs/superpowers/specs/2026-09-28-source-neutral-snapshot-normalization-design.md deleted file mode 100644 index 00718ec0..00000000 --- a/docs/superpowers/specs/2026-09-28-source-neutral-snapshot-normalization-design.md +++ /dev/null @@ -1,164 +0,0 @@ -# Source-neutral snapshot normalization - -## Goal - -Remove source-specific decisions from catalog normalization. Snapshot configuration and source adapters should translate source data into one catalog input contract before the normalizer runs. - -The change must preserve the current generated product: - -- 3,198 catalog records and shards; -- all 1,914 Claire Radar records; -- all 12,929 score observations; -- source-owned scores and reviewed identity relationships; -- unchanged Claire `extra_json` objects and `record_sha256` values; -- legacy empty series where they are used for source-coverage auditing. - -## Problem - -The current normalizer contains two Claire-specific branches: - -1. It omits empty score series by comparing the source name with `claire_radar`. -2. It reads Claire's private `releaseDates`, `firstPublicAt`, and `paperV1At` fields. - -These branches make a general catalog stage responsible for source policy and source schema. A new scoreless source or a source with equivalent date metadata would require another normalizer edit. - -## Design - -### Snapshot policy - -Every registered snapshot declares a score-series policy: - -```yaml -score_series_policy: preserve_empty -``` - -or: - -```yaml -score_series_policy: observed_only -``` - -`preserve_empty` emits a series even when it has no observations. Existing sources use this policy to preserve coverage-audit behavior. - -`observed_only` emits a series only when at least one observation exists. The Claire snapshot uses this policy because it contains no score data. - -The snapshot loader validates the value and rejects missing or unknown policies. The catalog normalizer receives the validated policy and applies it without inspecting the source name. - -### Source adapters - -A snapshot can name a catalog adapter: - -```yaml -catalog_adapter: claire_radar_v1 -``` - -The adapter runs after the source CSV row is parsed and before `_source_record` receives it. Its output follows the common row contract used by catalog normalization. - -The Claire adapter translates its preserved source object into these common fields: - -- `released`; -- `released_basis`; -- `released_source_url`; -- `publication_dates`. - -`publication_dates` is a list of normalized date facts. Each fact has a date, basis, source URL, and optional note. The common normalizer validates and copies these facts; it does not parse source-private date keys. - -The adapter retains the complete Claire source object as source metadata. It does not mutate `extra_json`, rewrite the CSV, or change `record_sha256`. - -Snapshots without an adapter use the identity adapter, which returns the parsed row unchanged. Unknown adapter names fail during snapshot loading. - -### Date evidence - -The Claire adapter reads the six records that contain `releaseDates` in their preserved source objects. It maps: - -- `firstPublicAt` to the common `released` fact with basis `first_public`; -- `paperV1At` to a `publication_dates` fact with basis `paper_first_version`. - -Reviewed first-public evidence remains keyed by exact source record ID in snapshot configuration. The adapter uses it when present. It never matches evidence by benchmark name. - -When a source object contains a date but no reviewed first-public URL, the adapter retains the date and uses the original record evidence defined by the common provenance contract. It does not imply that the imported date received an independent review. - -Paper evidence comes from the record's paper URL when available. A missing evidence URL does not cause the date fact to be invented or discarded; the common representation records only evidence that exists. - -### Source metadata - -The normalizer stores source metadata under the actual source key: - -```python -{"source_metadata": {source: source_metadata}} -``` - -It does not name Claire directly. Adapters decide which source metadata to return, while the common normalizer only preserves it. - -## Module boundaries - -### `leaderboard_snapshots.py` - -- Parse and validate `score_series_policy`. -- Parse and validate `catalog_adapter`. -- Preserve adapter options as configuration data. -- Reject unsupported policies or adapters before normalization. - -### `catalog_snapshot_adapters.py` - -- Own the adapter registry. -- Provide the identity adapter. -- Translate Claire's private schema into the common row contract. -- Validate adapter-specific options that cannot be validated generically. - -### `catalog.py` - -- Consume only the common row contract. -- Apply the declared score-series policy. -- Validate normalized date facts. -- Preserve returned source metadata under the source key. -- Contain no Claire date-schema branches or Claire score-series branches. - -`catalog.py` may retain the existing snapshot-to-source registration until that separate concern has a source-neutral replacement. This refactor does not broaden into redesigning source identifiers. - -## Validation and failure behavior - -Normalization fails visibly for: - -- a missing or unsupported score-series policy; -- an unsupported adapter name; -- malformed adapter options; -- an invalid normalized date or date basis; -- malformed `publication_dates` output. - -The adapter must not silently infer relationships, evidence, or dates from a benchmark name. - -## Tests - -Tests are written before implementation and must fail for the missing general behavior. - -1. A synthetic source using `observed_only` emits no empty series. -2. A synthetic source using `preserve_empty` retains an empty series. -3. Missing and unknown policies fail during snapshot loading. -4. An unknown adapter fails during snapshot loading. -5. The common normalizer accepts normalized first-public and paper-version facts without source knowledge. -6. The Claire adapter converts all six source objects with `releaseDates`. -7. Reviewed evidence is selected by exact source ID. -8. Claire `extra_json` and `record_sha256` remain byte-for-byte unchanged. -9. A structural test prevents the removed private schema keys and source-specific score condition from returning to `catalog.py`. -10. Generated-product tests retain 3,198 records, 1,914 Claire records, 3,198 shards, and 12,929 observations. -11. Search, detail pages, provenance, reviewed siblings, and source-partitioned scores retain their current behavior. - -## Delivery sequence - -1. Add failing policy, adapter, normalized-date, and architecture-boundary tests. -2. Add policy and adapter configuration validation. -3. Add the adapter registry and Claire adapter. -4. Simplify catalog normalization to consume the common contract. -5. Run focused tests and inspect generated records. -6. Run the documented six-stage CI in a clean checkout. -7. Push the new commits and re-review the new PR head. Do not merge the PR. - -## Non-goals - -- Merging source records across reviewed identity links. -- Moving scores between sources. -- Changing benchmark counts or taxonomy. -- Replacing the existing snapshot-to-source ID registration. -- Introducing a general SSRF policy for static outbound links. -- Rewriting all Claire CSV rows to materialize six normalized date records. From 9124e8ea59f34de658e16cfa96d07b73e007876d Mon Sep 17 00:00:00 2001 From: ktwu01 Date: Sat, 3 Oct 2026 00:33:49 -0500 Subject: [PATCH 02/29] chore: remove one-off OpenReview auth probe The daily radar workflow exercises the same credentials; the probe last ran 2026-08-16. Co-Authored-By: Claude Opus 5.5 --- .github/workflows/test-openreview.yml | 18 ------------- scripts/test_openreview_gh.py | 39 --------------------------- 2 files changed, 57 deletions(-) delete mode 100644 .github/workflows/test-openreview.yml delete mode 100644 scripts/test_openreview_gh.py diff --git a/.github/workflows/test-openreview.yml b/.github/workflows/test-openreview.yml deleted file mode 100644 index 4af04535..00000000 --- a/.github/workflows/test-openreview.yml +++ /dev/null @@ -1,18 +0,0 @@ -name: Test OpenReview Auth - -on: - workflow_dispatch: - -jobs: - test-openreview: - runs-on: ubuntu-latest - env: - OPENREVIEW_USERNAME: ${{ secrets.OPENREVIEW_USERNAME }} - OPENREVIEW_PASSWORD: ${{ secrets.OPENREVIEW_PASSWORD }} - steps: - - uses: actions/checkout@v7 - - uses: actions/setup-python@v7 - with: - python-version: "3.11" - - run: pip install openreview-py - - run: python3 scripts/test_openreview_gh.py diff --git a/scripts/test_openreview_gh.py b/scripts/test_openreview_gh.py deleted file mode 100644 index ff499076..00000000 --- a/scripts/test_openreview_gh.py +++ /dev/null @@ -1,39 +0,0 @@ -#!/usr/bin/env python3 -"""Test OpenReview authentication in GitHub Actions.""" - -import os - -import openreview - - -def main(): - username = os.getenv("OPENREVIEW_USERNAME") - password = os.getenv("OPENREVIEW_PASSWORD") - - if not username or not password: - print("ERROR: OPENREVIEW_USERNAME or OPENREVIEW_PASSWORD not set") - return 1 - - print(f"Username: {username}") - print(f"Password length: {len(password)}") - - client = openreview.api.OpenReviewClient( - baseurl="https://api2.openreview.net", - username=username, - password=password, - ) - - # Test with the working invitation ID - invitation = "ICLR.cc/2026/Conference/-/Submission" - notes = client.get_notes(invitation=invitation, limit=5) - print(f"SUCCESS: Got {len(notes)} notes from {invitation}") - - for n in notes: - title = n.content.get("title", {}).get("value", "N/A") - print(f" {n.id}: {title[:80]}") - - return 0 - - -if __name__ == "__main__": - exit(main()) From 86aad97b2e4a2b13481d7ffa4085d6e621b5a4c6 Mon Sep 17 00:00:00 2001 From: ktwu01 Date: Sat, 3 Oct 2026 00:41:49 -0500 Subject: [PATCH 03/29] Restore the OpenReview auth probe Co-Authored-By: Claude Opus 5.5 --- .github/workflows/test-openreview.yml | 18 +++++++++++++ scripts/test_openreview_gh.py | 39 +++++++++++++++++++++++++++ 2 files changed, 57 insertions(+) create mode 100644 .github/workflows/test-openreview.yml create mode 100644 scripts/test_openreview_gh.py diff --git a/.github/workflows/test-openreview.yml b/.github/workflows/test-openreview.yml new file mode 100644 index 00000000..4af04535 --- /dev/null +++ b/.github/workflows/test-openreview.yml @@ -0,0 +1,18 @@ +name: Test OpenReview Auth + +on: + workflow_dispatch: + +jobs: + test-openreview: + runs-on: ubuntu-latest + env: + OPENREVIEW_USERNAME: ${{ secrets.OPENREVIEW_USERNAME }} + OPENREVIEW_PASSWORD: ${{ secrets.OPENREVIEW_PASSWORD }} + steps: + - uses: actions/checkout@v7 + - uses: actions/setup-python@v7 + with: + python-version: "3.11" + - run: pip install openreview-py + - run: python3 scripts/test_openreview_gh.py diff --git a/scripts/test_openreview_gh.py b/scripts/test_openreview_gh.py new file mode 100644 index 00000000..ff499076 --- /dev/null +++ b/scripts/test_openreview_gh.py @@ -0,0 +1,39 @@ +#!/usr/bin/env python3 +"""Test OpenReview authentication in GitHub Actions.""" + +import os + +import openreview + + +def main(): + username = os.getenv("OPENREVIEW_USERNAME") + password = os.getenv("OPENREVIEW_PASSWORD") + + if not username or not password: + print("ERROR: OPENREVIEW_USERNAME or OPENREVIEW_PASSWORD not set") + return 1 + + print(f"Username: {username}") + print(f"Password length: {len(password)}") + + client = openreview.api.OpenReviewClient( + baseurl="https://api2.openreview.net", + username=username, + password=password, + ) + + # Test with the working invitation ID + invitation = "ICLR.cc/2026/Conference/-/Submission" + notes = client.get_notes(invitation=invitation, limit=5) + print(f"SUCCESS: Got {len(notes)} notes from {invitation}") + + for n in notes: + title = n.content.get("title", {}).get("value", "N/A") + print(f" {n.id}: {title[:80]}") + + return 0 + + +if __name__ == "__main__": + exit(main()) From 5cc366b67d86b86ecfb9b5db1f0d9a66c166c126 Mon Sep 17 00:00:00 2001 From: ktwu01 Date: Mon, 28 Sep 2026 19:33:38 -0500 Subject: [PATCH 04/29] Add related-work drafting to the query surface (ref #549, #650) `benchmark-radar related-work "Label=query" ...` (and GET /api/v1/related-work?q=...) turns topic queries into a cited LaTeX Related Work draft, a matching BibTeX file, and a Markdown comparison table with the #650 axes (paper, repo, dataset, openness). - One QueryService method feeds CLI and HTTP, so both return the same JSON contract; it reads only the local artifacts, with no network. - Each topic keeps full lexical matches from the catalog and from scholarly Radar sources (arXiv, HF Papers, Semantic Scholar, OpenAlex, Crossref); --include-partial widens it. - Every retained work is cited; the Benchmark Radar paper is cited once, for the size of the benchmark landscape. - Authors come only from recorded snapshot metadata. Records without them get a BibTeX `key` field and an `authors_missing` flag instead of a guessed author list. - Output compiles under pdflatex + natbib: Greek letters become math commands, and scripts pdflatex cannot typeset are dropped. - Radar search candidates now carry the recorded `authors` list. Co-Authored-By: Claude Opus 5.5 --- src/benchmark_radar/query.py | 20 ++ src/benchmark_radar/query_cli.py | 62 +++- src/benchmark_radar/query_http.py | 24 +- src/benchmark_radar/related_work.py | 373 +++++++++++++++++++++ src/benchmark_radar/related_work_render.py | 177 ++++++++++ tests/test_related_work.py | 199 +++++++++++ 6 files changed, 852 insertions(+), 3 deletions(-) create mode 100644 src/benchmark_radar/related_work.py create mode 100644 src/benchmark_radar/related_work_render.py create mode 100644 tests/test_related_work.py diff --git a/src/benchmark_radar/query.py b/src/benchmark_radar/query.py index 3080b797..e92d7b4f 100644 --- a/src/benchmark_radar/query.py +++ b/src/benchmark_radar/query.py @@ -494,6 +494,7 @@ def _radar_candidates(self) -> list[dict[str, Any]]: # review BLOCKER). Same function, same output. "science_domains": science_domains_for_record(item), "publisher": " ".join(item.get("organizations") or []), + "authors": [str(name) for name in item.get("authors") or []], "modality": None, "languages": [], "source": source, @@ -663,6 +664,25 @@ def search( "results": results, } + def related_work( + self, + topics: list[str], + *, + per_topic: int = 6, + include_partial: bool = False, + include_radar: bool = True, + ) -> dict[str, Any]: + """Draft a cited related-work section from topic queries (issues #549, #650).""" + from .related_work import build_related_work + + return build_related_work( + self, + topics, + per_topic=per_topic, + include_partial=include_partial, + include_radar=include_radar, + ) + def show(self, identifier: str) -> dict[str, Any]: identifier = str(identifier).strip() if not identifier: diff --git a/src/benchmark_radar/query_cli.py b/src/benchmark_radar/query_cli.py index 26242725..2aea7bcf 100644 --- a/src/benchmark_radar/query_cli.py +++ b/src/benchmark_radar/query_cli.py @@ -21,7 +21,10 @@ ) from .query_http import serve_query_api -QUERY_COMMANDS = frozenset({"init", "sync", "search", "show", "recent", "status", "serve"}) +QUERY_COMMANDS = frozenset( + {"init", "sync", "search", "show", "recent", "status", "serve", "related-work"} +) +RELATED_WORK_FORMATS = ("latex", "bibtex", "markdown") def _data_parent() -> argparse.ArgumentParser: @@ -79,6 +82,29 @@ def _parser() -> argparse.ArgumentParser: recent.add_argument("--recommended", action="store_true") recent.add_argument("--json", action="store_true") + related = subparsers.add_parser( + "related-work", + parents=[data_parent], + help="Draft a cited related-work section and BibTeX from topic queries.", + ) + related.add_argument( + "topics", + nargs="+", + metavar="TOPIC", + help="A short query, or 'Label=query' to name the paragraph it becomes.", + ) + related.add_argument("--per-topic", type=int, default=6) + related.add_argument( + "--include-partial", + action="store_true", + help="Keep candidates that miss some query tokens (noisier).", + ) + related.add_argument("--no-radar", dest="include_radar", action="store_false") + related.add_argument("--format", choices=RELATED_WORK_FORMATS, default="latex") + related.add_argument("--tex", type=Path, help="Write the LaTeX section to this file.") + related.add_argument("--bib", type=Path, help="Write the BibTeX entries to this file.") + related.add_argument("--json", action="store_true") + status = subparsers.add_parser( "status", parents=[data_parent], help="Inspect local catalog and snapshot health." ) @@ -167,6 +193,32 @@ def _print_show(payload: dict[str, Any]) -> None: print(f" {artifact.get('kind')}: {artifact.get('url')}") +def _related_work_printer(args: argparse.Namespace) -> Callable[[dict[str, Any]], None]: + """Write requested files first, then print one format for the terminal.""" + + def printer(payload: dict[str, Any]) -> None: + for path, field in ((args.tex, "latex"), (args.bib, "bibtex")): + if path is not None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(payload[field], encoding="utf-8") + print(f"wrote {field} to {path}", file=sys.stderr) + if args.json: + _print_json(payload) + return + print(payload[args.format], end="") + flagged = [ + entry for entry in payload["entries"] if "authors_missing" in entry["verification"] + ] + if flagged: + print( + f"\n% {len(flagged)} of {payload['count']} entries lack authors in local data; " + "complete them before citing.", + file=sys.stderr, + ) + + return printer + + def _print_status(payload: dict[str, Any]) -> None: print(f"status: {payload['status']}") print(f"catalog: {payload['catalog']['count']} records at {payload['catalog']['path']}") @@ -228,6 +280,14 @@ def run_query_cli(argv: Sequence[str] | None = None) -> int: recommended=args.recommended, ) printer = _print_json if args.json else _print_recent + elif args.command == "related-work": + payload = service.related_work( + args.topics, + per_topic=args.per_topic, + include_partial=args.include_partial, + include_radar=args.include_radar, + ) + printer = _related_work_printer(args) elif args.command == "status": payload = service.status() printer = _print_json if args.json else _print_status diff --git a/src/benchmark_radar/query_http.py b/src/benchmark_radar/query_http.py index 23b6777e..a1525151 100644 --- a/src/benchmark_radar/query_http.py +++ b/src/benchmark_radar/query_http.py @@ -14,7 +14,9 @@ LOGGER = logging.getLogger(__name__) -def _parse_parameters(query: str, *, allowed: set[str]) -> dict[str, list[str]]: +def _parse_parameters( + query: str, *, allowed: set[str], repeatable: frozenset[str] = frozenset() +) -> dict[str, list[str]]: parameters = parse_qs(query, keep_blank_values=True) unknown = sorted(set(parameters) - allowed) if unknown: @@ -23,7 +25,9 @@ def _parse_parameters(query: str, *, allowed: set[str]) -> dict[str, list[str]]: code="invalid_request", status=400, ) - repeated = sorted(key for key, values in parameters.items() if len(values) != 1) + repeated = sorted( + key for key, values in parameters.items() if len(values) != 1 and key not in repeatable + ) if repeated: raise QueryError( f"query parameter(s) must occur once: {', '.join(repeated)}", @@ -138,6 +142,22 @@ def _route_get(self) -> dict[str, Any]: source=_value(parameters, "source"), ) + if path == "/api/v1/related-work": + parameters = _parse_parameters( + request.query, + allowed={"q", "per_topic", "include_partial", "include_radar"}, + repeatable=frozenset({"q"}), + ) + topics = parameters.get("q") or [] + if not topics: + raise QueryError("q is required", code="invalid_request", status=400) + return service.related_work( + topics, + per_topic=_integer(parameters, "per_topic", default=6), + include_partial=_boolean(parameters, "include_partial"), + include_radar=_boolean(parameters, "include_radar", default=True), + ) + if path.startswith("/api/v1/benchmarks/"): _parse_parameters(request.query, allowed=set()) identifier = unquote(path.removeprefix("/api/v1/benchmarks/")) diff --git a/src/benchmark_radar/related_work.py b/src/benchmark_radar/related_work.py new file mode 100644 index 00000000..10eda486 --- /dev/null +++ b/src/benchmark_radar/related_work.py @@ -0,0 +1,373 @@ +"""Related-work drafting over the local query contract (issues #549, #650). + +A paper's related-work section is the job users most often open this tool for +(#522 R2). This module turns a handful of topic queries into a citable draft: +one entry per retained work, each carrying a BibTeX record, the comparison axes +from #650 (paper, repository, dataset, openness), and which topics retrieved it. + +Everything here reads the same local artifacts as ``search`` and ``show``. No +network is touched, so an author the snapshots never recorded stays missing and +is reported under ``verification`` instead of being guessed. Retrieval stays +lexical: a topic keeps only candidates that cover every query token unless the +caller opts into partial matches, and the draft says plainly that absence from +this corpus is not evidence of absence from the literature (#522 R3). +""" + +from __future__ import annotations + +import re +import unicodedata +from dataclasses import dataclass +from typing import TYPE_CHECKING, Any + +from .citation import BIBTEX_KEY, bibtex_citation +from .related_work_render import latex_escape, render_latex, render_markdown + +if TYPE_CHECKING: + from .query import QueryService + +MAX_TOPICS = 12 +MAX_PER_TOPIC = 30 +_SEARCH_WINDOW = 200 +_ARXIV_ID = re.compile(r"(? Topic: + """Accept ``query`` or ``Label=query``; the label names the paragraph.""" + label, separator, query = str(value).partition("=") + if not separator: + query, label = label, label + label, query = " ".join(label.split()), " ".join(query.split()) + if not query: + from .query import QueryError + + raise QueryError(f"topic {value!r} has an empty query", code="invalid_query", status=400) + return Topic(label=label or query, query=query) + + +def _ascii(value: str) -> str: + return unicodedata.normalize("NFKD", value).encode("ascii", "ignore").decode("ascii") + + +def _key_part(value: str) -> str: + return re.sub(r"[^a-z0-9]", "", _ascii(value).casefold()) + + +def _arxiv_id(record: dict[str, Any], extra_urls: list[str]) -> str | None: + if str(record.get("source") or "").casefold() in _ARXIV_SOURCES: + match = _ARXIV_ID.fullmatch(str(record.get("source_id") or "")) + if match: + return match.group(1) + for url in [str(record.get("url") or ""), *extra_urls]: + if "arxiv.org/" in url or "huggingface.co/papers/" in url: + match = _ARXIV_ID.search(url) + if match: + return match.group(1) + return None + + +def _year(record: dict[str, Any], arxiv_id: str | None) -> str | None: + if arxiv_id: + return f"20{arxiv_id[:2]}" + for field in ("published_at", "released"): + match = re.match(r"(\d{4})", str(record.get(field) or "")) + if match: + return match.group(1) + return None + + +def _summary(description: Any) -> str: + if isinstance(description, dict): + description = description.get("en") or next(iter(description.values()), "") + text = " ".join(str(description or "").split()) + text = re.sub(r"\s*See the full description on the dataset page:.*$", "", text) + letters = [char for char in text if char.isalpha()] + if letters and sum(ord(char) > 0x24F for char in letters) > len(letters) / 3: + return "" # Not Latin-script prose; an English draft cannot restate it. + text = _LATEX_COMMAND.sub(r"\1", text) + text = re.sub(r"\s*[\u2014\u2013]\s*", ", ", text).rstrip("\u2026 ").strip() + if not text: + return "" + sentences = re.split(r"(?<=[.!?])\s+(?=[A-Z0-9])", text) + chosen = next((s for s in sentences if _CONTRIBUTION.search(s)), sentences[0]) + stripped = _LEAD_IN.sub("", chosen) + if stripped != chosen: + chosen = stripped[:1].upper() + stripped[1:] + words = chosen.split() + if len(words) > _MAX_SUMMARY_WORDS: + chosen = " ".join(words[:_MAX_SUMMARY_WORDS]).rstrip(",;:") + " ..." + return chosen + + +def _base_key(entry: dict[str, Any]) -> str: + year = entry.get("year") or "" + title_word = next( + ( + _key_part(word) + for word in re.split(r"[\s:/-]+", entry["name"]) + if _key_part(word) and _key_part(word) not in _TITLE_STOPWORDS + ), + "work", + ) + if entry["authors"]: + surname = _key_part(entry["authors"][0].split()[-1]) or "anon" + return f"{surname}{year}{title_word}" + return f"{_key_part(entry['name'])[:24] or 'benchmark'}{year}" + + +def _catalog_entry(service: QueryService, result: dict[str, Any]) -> dict[str, Any]: + record = service.show(result["key"])["benchmark"]["record"] + artifacts = record.get("artifacts") or [] + urls = [str(artifact.get("url") or "") for artifact in artifacts] + arxiv_id = _arxiv_id(result, urls) + paper_url = next( + (u for a, u in zip(artifacts, urls, strict=True) if a.get("kind") == "paper"), None + ) + openness = result.get("openness") + if isinstance(openness, dict): + openness = openness.get("status") + return { + "kind": "catalog", + "key": result["key"], + "name": result["name"], + "authors": [], + "url": paper_url or (urls[0] if urls else result.get("source_url")), + "arxiv_id": arxiv_id, + "year": _year(result, arxiv_id), + "summary": _summary(record.get("description") or result.get("description")), + "axes": { + "has_paper": bool(result.get("has_paper")), + "has_repo": bool(result.get("has_repo")), + "has_dataset": bool(result.get("has_dataset")), + "openness": openness or "unknown", + }, + } + + +def _radar_entry(result: dict[str, Any]) -> dict[str, Any]: + arxiv_id = _arxiv_id(result, []) + return { + "kind": "radar", + "key": result["key"], + "name": result["name"], + "authors": [str(name) for name in result.get("authors") or [] if str(name).strip()], + "url": result.get("url"), + "arxiv_id": arxiv_id, + "year": _year(result, arxiv_id), + "summary": _summary(result.get("description")), + "axes": { + "has_paper": bool(result.get("has_paper")), + "has_repo": bool(result.get("has_repo")), + "has_dataset": bool(result.get("has_dataset")), + "openness": "unknown", + }, + } + + +def _merge(existing: dict[str, Any], incoming: dict[str, Any]) -> None: + """Fold a second record of the same paper into the first one kept.""" + existing["merged_keys"].append(incoming["key"]) + if not existing["authors"] and incoming["authors"]: + existing["authors"] = incoming["authors"] + for field in ("url", "arxiv_id", "year", "summary"): + existing[field] = existing[field] or incoming[field] + for flag in ("has_paper", "has_repo", "has_dataset"): + existing["axes"][flag] = existing["axes"][flag] or incoming["axes"][flag] + if existing["axes"]["openness"] == "unknown": + existing["axes"]["openness"] = incoming["axes"]["openness"] + + +def _bibtex(entry: dict[str, Any]) -> str: + lines = [] + if not entry["authors"]: + lines.append("% Benchmark Radar has no author list for this record; add it before citing.") + lines.append(f"@misc{{{entry['cite_key']},") + lines.append(f" title = {{{{{latex_escape(entry['name'])}}}}},") + if entry["authors"]: + authors = " and ".join(latex_escape(name) for name in entry["authors"]) + lines.append(f" author = {{{authors}}},") + else: + # BibTeX's standard stand-in: styles sort and label by `key` when no + # author exists, so the citation renders without an invented author. + lines.append(f" key = {{{latex_escape(entry['name'])}}},") + if entry["year"]: + lines.append(f" year = {{{entry['year']}}},") + if entry["arxiv_id"]: + lines.append(f" eprint = {{{entry['arxiv_id']}}},") + lines.append(" archivePrefix = {arXiv},") + lines.append(f" url = {{https://arxiv.org/abs/{entry['arxiv_id']}}},") + elif entry["url"]: + lines.append(f" howpublished = {{\\url{{{entry['url']}}}}},") + lines.append("}") + return "\n".join(lines) + + +def _verification(entry: dict[str, Any]) -> list[str]: + issues = [] + if not entry["authors"]: + issues.append("authors_missing") + if not entry["year"]: + issues.append("year_missing") + if not entry["arxiv_id"] and not entry["url"]: + issues.append("locator_missing") + if entry["kind"] == "radar": + issues.append("radar_lead_unverified") + return issues + + +def _assign_keys(entries: list[dict[str, Any]]) -> None: + used = {BIBTEX_KEY} + for entry in entries: + base = _base_key(entry) + key, suffix = base, ord("a") + while key in used: + key = f"{base}{chr(suffix)}" + suffix += 1 + used.add(key) + entry["cite_key"] = key + + +def _coverage(service: QueryService, *, include_radar: bool) -> dict[str, Any]: + catalog_count = service.validated_catalog_index()["count"] + value: dict[str, Any] = {"catalog_count": catalog_count} + if include_radar: + snapshots = service._load_snapshots() + value.update( + { + "radar_first_date": snapshots[0]["date"], + "radar_latest_date": snapshots[-1]["date"], + "snapshot_count": len(snapshots), + } + ) + window = ( + f"Radar observations from {value['radar_first_date']} to {value['radar_latest_date']}" + if include_radar + else "no Radar observations" + ) + value["statement"] = ( + f"Candidates come from {catalog_count} catalog records and {window}. Retrieval is " + "lexical, and a work missing here is evidence about this corpus, not about the " + "literature; older prior art in particular may be absent." + ) + return value + + +def build_related_work( + service: QueryService, + topics: list[str], + *, + per_topic: int = 6, + include_partial: bool = False, + include_radar: bool = True, +) -> dict[str, Any]: + from .query import QUERY_SCHEMA_VERSION, QueryError + + parsed = [parse_topic(value) for value in topics] + if not parsed or len(parsed) > MAX_TOPICS: + raise QueryError( + f"pass between 1 and {MAX_TOPICS} topics", code="invalid_query", status=400 + ) + if per_topic < 1 or per_topic > MAX_PER_TOPIC: + raise QueryError( + f"per_topic must be between 1 and {MAX_PER_TOPIC}", code="invalid_limit", status=400 + ) + + entries: list[dict[str, Any]] = [] + by_identity: dict[str, dict[str, Any]] = {} + topic_rows = [] + scopes = ("catalog", "radar") if include_radar else ("catalog",) + for topic in parsed: + row: dict[str, Any] = {"label": topic.label, "query": topic.query, "search_status": {}} + kept: list[dict[str, Any]] = [] + for scope in scopes: + payload = service.search(topic.query, scope=scope, limit=_SEARCH_WINDOW) + row["search_status"][scope] = payload["search_status"] + taken = 0 + for result in payload["results"]: + if taken >= per_topic: + break + if result["match"]["missing_tokens"] and not include_partial: + continue + if ( + scope == "radar" + and str(result.get("source") or "").casefold() not in SCHOLARLY_RADAR_SOURCES + ): + continue + entry = ( + _catalog_entry(service, result) if scope == "catalog" else _radar_entry(result) + ) + identity = f"arxiv:{entry['arxiv_id']}" if entry["arxiv_id"] else entry["key"] + if identity in by_identity: + target = by_identity[identity] + if entry["key"] not in {target["key"], *target["merged_keys"]}: + _merge(target, entry) + else: + entry.update({"merged_keys": [], "topics": []}) + by_identity[identity] = entry + entries.append(entry) + target = entry + if topic.label not in target["topics"]: + target["topics"].append(topic.label) + if not any(item is target for item in kept): + kept.append(target) + taken += 1 + row["entries"] = kept + topic_rows.append(row) + + _assign_keys(entries) + for entry in entries: + entry["bibtex"] = _bibtex(entry) + entry["verification"] = _verification(entry) + for row in topic_rows: + row["cite_keys"] = [entry["cite_key"] for entry in row.pop("entries")] + + coverage = _coverage(service, include_radar=include_radar) + by_key = {entry["cite_key"]: entry for entry in entries} + latex = render_latex(topic_rows, by_key, coverage=coverage, self_key=BIBTEX_KEY) + bibtex = "\n\n".join([bibtex_citation(), *(entry["bibtex"] for entry in entries)]) + "\n" + return { + "schema_version": QUERY_SCHEMA_VERSION, + "retrieval_mode": "related_work", + "options": { + "per_topic": per_topic, + "include_partial": include_partial, + "include_radar": include_radar, + }, + "topics": topic_rows, + "count": len(entries), + "entries": entries, + "coverage": coverage, + "latex": latex, + "bibtex": bibtex, + "markdown": render_markdown(topic_rows, by_key, coverage=coverage), + "data": service._data_summary(scope="all" if include_radar else "catalog"), + } diff --git a/src/benchmark_radar/related_work_render.py b/src/benchmark_radar/related_work_render.py new file mode 100644 index 00000000..a1ba8859 --- /dev/null +++ b/src/benchmark_radar/related_work_render.py @@ -0,0 +1,177 @@ +"""Render a related-work payload as a LaTeX section and a Markdown table. + +The prose is a draft, and says so in LaTeX comments the compiled paper never +shows: each sentence restates one record summary so an author (or the agent +driving the Skill) can check it against the paper and rewrite it. Every entry +the payload retained is cited, and the Benchmark Radar paper is cited once, as +the source for the size of the benchmark landscape rather than as a search tool. +""" + +from __future__ import annotations + +import re +import unicodedata +from typing import Any + +_LATEX_SPECIALS = { + "\\": r"\textbackslash{}", + "&": r"\&", + "%": r"\%", + "$": r"\$", + "#": r"\#", + "_": r"\_", + "{": r"\{", + "}": r"\}", + "~": r"\textasciitilde{}", + "^": r"\textasciicircum{}", +} +# Capital Greek letters that have their own LaTeX command; the rest are drawn +# identically to a Latin capital, which pdflatex can typeset. +_UPPER_GREEK = frozenset("Gamma Delta Theta Lambda Xi Pi Sigma Upsilon Phi Psi Omega".split()) +_LATIN_LOOKALIKE = { + "Alpha": "A", + "Beta": "B", + "Epsilon": "E", + "Zeta": "Z", + "Eta": "H", + "Iota": "I", + "Kappa": "K", + "Mu": "M", + "Nu": "N", + "Omicron": "O", + "Rho": "P", + "Tau": "T", + "Chi": "X", +} +# Characters pdflatex (utf8 inputenc, T1 fonts) typesets without extra packages. +_TYPESET_PUNCTUATION = frozenset( + "\u2018\u2019\u201c\u201d\u2013\u2014\u2022\u0218\u0219\u021a\u021b" +) +_WE_VERB = re.compile(r"^We (\w+)\s+(.*)$", re.DOTALL | re.IGNORECASE) + + +def _latex_char(char: str) -> str: + if char in _LATEX_SPECIALS: + return _LATEX_SPECIALS[char] + # pdflatex with inputenc rejects Greek text characters, which arXiv titles + # use for names such as tau-bench; math-mode commands work in every engine. + name = unicodedata.name(char, "") + if name.startswith("GREEK SMALL LETTER ") and name.count(" ") == 3: + return f"\\ensuremath{{\\{name.rsplit(' ', 1)[-1].lower()}}}" + if name.startswith("GREEK CAPITAL LETTER ") and name.count(" ") == 3: + letter = name.rsplit(" ", 1)[-1].capitalize() + if letter in _UPPER_GREEK: + return f"\\ensuremath{{\\{letter}}}" + return _LATIN_LOOKALIKE.get(letter, "") + if ord(char) <= 0x17F or char in _TYPESET_PUNCTUATION: + return char + # CJK text, emoji, and other scripts stop a pdflatex run outright; a + # dropped glyph costs a character, an unknown one costs the whole paper. + return "" + + +def latex_escape(value: str) -> str: + value = unicodedata.normalize("NFKC", value).replace("\u2026", "...") + return re.sub(r" {2,}", " ", "".join(_latex_char(char) for char in value)).strip() + + +def _landscape_size(count: int) -> str: + if count >= 100: + return f"more than {count // 100 * 100:,}" + return f"{count}" + + +def _sentence(entry: dict[str, Any]) -> str: + key, name, summary = entry["cite_key"], entry["name"], entry["summary"] + if not summary: + return f"{latex_escape(name)}~\\citep{{{key}}}." + we_verb = _WE_VERB.match(summary) + if we_verb and entry["authors"]: + return f"\\citet{{{key}}} {we_verb.group(1)} {latex_escape(we_verb.group(2))}" + if summary.casefold().startswith(name.casefold()): + rest = summary[len(name) :].lstrip() + if re.match(r"(a|an|the)\s", rest, re.IGNORECASE): + rest = f"is {rest}" + joiner = "" if rest[:1] in {",", ":", ";", "."} else " " + return f"{latex_escape(name)}~\\citep{{{key}}}{joiner}{latex_escape(rest)}" + return f"{latex_escape(name)}~\\citep{{{key}}}: {latex_escape(summary)}" + + +def render_latex( + topics: list[dict[str, Any]], + entries: dict[str, dict[str, Any]], + *, + coverage: dict[str, Any], + self_key: str, +) -> str: + lines = [ + "% Related-work draft generated by `benchmark-radar related-work`.", + "% Each sentence restates a record summary. Check it against the cited paper,", + "% synthesize across works, and rewrite in your own voice before submission.", + f"% Coverage: {coverage['statement']}", + "\\section{Related Work}", + "\\label{sec:related-work}", + "", + "AI evaluation now spans " + f"{_landscape_size(coverage['catalog_count'])} benchmarks tracked across " + f"leaderboards, papers, and repositories~\\citep{{{self_key}}}; below we group " + "the work closest to ours by theme.", + ] + described: set[str] = set() + for topic in topics: + lines.extend(["", f"\\paragraph{{{latex_escape(topic['label'])}.}}"]) + if not topic["cite_keys"]: + lines.append( + f"% No candidates matched every token of {topic['query']!r}; " + "try a narrower or differently worded query." + ) + continue + fresh = [key for key in topic["cite_keys"] if key not in described] + seen = [key for key in topic["cite_keys"] if key in described] + body = [_sentence(entries[key]) for key in fresh] + body = [text if text.endswith((".", "...")) else f"{text}." for text in body] + if seen: + body.append(f"See also~\\citep{{{', '.join(seen)}}}.") + described.update(fresh) + lines.append("\n".join(body)) + return "\n".join(lines) + "\n" + + +def _flag(value: bool) -> str: + return "yes" if value else "no" + + +def render_markdown( + topics: list[dict[str, Any]], + entries: dict[str, dict[str, Any]], + *, + coverage: dict[str, Any], +) -> str: + lines = ["## Related work candidates", "", coverage["statement"], "", "### Topics", ""] + for topic in topics: + statuses = ", ".join( + f"{scope}: {status}" for scope, status in topic["search_status"].items() + ) + kept = len(topic["cite_keys"]) + lines.append(f"- **{topic['label']}** (`{topic['query']}`): {kept} kept; {statuses}") + lines.extend( + [ + "", + "### Comparison", + "", + "| Work | Cite key | Topics | Record | Paper | Repo | Dataset | Openness | Verify |", + "| --- | --- | --- | --- | --- | --- | --- | --- | --- |", + ] + ) + for entry in entries.values(): + axes = entry["axes"] + name = entry["name"].replace("|", "\\|") + work = f"[{name}]({entry['url']})" if entry["url"] else name + record = "catalog" if entry["kind"] == "catalog" else "Radar lead" + lines.append( + f"| {work} | `{entry['cite_key']}` | {'; '.join(entry['topics'])} | {record} | " + f"{_flag(axes['has_paper'])} | {_flag(axes['has_repo'])} | " + f"{_flag(axes['has_dataset'])} | {axes['openness']} | " + f"{', '.join(entry['verification']) or 'none'} |" + ) + return "\n".join(lines) + "\n" diff --git a/tests/test_related_work.py b/tests/test_related_work.py new file mode 100644 index 00000000..e10fae3d --- /dev/null +++ b/tests/test_related_work.py @@ -0,0 +1,199 @@ +from __future__ import annotations + +import json +import re +import threading +import urllib.parse +import urllib.request +from datetime import UTC, datetime, timedelta +from pathlib import Path + +import pytest +from test_query_surfaces import _catalog + +from benchmark_radar.citation import BIBTEX_KEY +from benchmark_radar.models import RadarItem, RadarRun, SourceHealth +from benchmark_radar.query import QueryError, QueryPaths, QueryService +from benchmark_radar.query_cli import run_query_cli +from benchmark_radar.query_http import create_query_server +from benchmark_radar.related_work_render import latex_escape +from benchmark_radar.snapshots import write_snapshot + + +def _paths(tmp_path: Path) -> QueryPaths: + """The shared query fixture plus a later day holding scholarly Radar leads.""" + paths = _catalog(tmp_path) + generated_at = datetime(2026, 8, 30, 8, 0, tzinfo=UTC) + write_snapshot( + RadarRun( + generated_at=generated_at, + since=generated_at - timedelta(hours=48), + items=[ + RadarItem( + source="arXiv", + source_id="2608.01234", + title="Agent Workbench Pro: Long-Horizon Agent Evaluation", + url="https://arxiv.org/abs/2608.01234", + published_at=generated_at - timedelta(hours=2), + summary=( + "Agents are everywhere. To address this gap, we introduce Agent " + "Workbench Pro, a benchmark of 50% harder tasks." + ), + authors=["Ada Lovelace", "Alan Turing"], + categories=["benchmark", "agentic"], + ), + RadarItem( + source="GitHub", + source_id="example/agent-workbench-fork", + title="Agent Workbench Fork", + url="https://github.com/example/agent-workbench-fork", + published_at=generated_at - timedelta(hours=3), + summary="A repository fork of the agent workbench harness.", + categories=["benchmark", "agentic"], + ), + ], + health=[ + SourceHealth(source=source, ok=True, item_count=1, method="API") + for source in ("arxiv", "github", "huggingface") + ], + ), + paths.snapshots, + ) + return paths + + +def _bib_keys(bibtex: str) -> list[str]: + return re.findall(r"@misc\{([^,]+),", bibtex) + + +def _cited_keys(latex: str) -> set[str]: + keys: set[str] = set() + for group in re.findall(r"\\cite[pt]?\{([^}]*)\}", latex): + keys.update(key.strip() for key in group.split(",")) + return keys + + +def test_draft_cites_every_entry_and_the_radar_paper_once(tmp_path: Path) -> None: + payload = QueryService(_paths(tmp_path)).related_work(["Agent benchmarks=agent workbench"]) + + entry_keys = {entry["cite_key"] for entry in payload["entries"]} + assert entry_keys, "fixture must retain at least one work" + assert set(_bib_keys(payload["bibtex"])) == entry_keys | {BIBTEX_KEY} + assert _cited_keys(payload["latex"]) == entry_keys | {BIBTEX_KEY} + # The self-citation supports a factual claim, not a "found with" sentence. + assert payload["latex"].count(BIBTEX_KEY) == 1 + assert "Benchmark Radar" not in payload["latex"].split("\\section", 1)[1] + + +def test_topics_keep_full_matches_unless_partial_is_requested(tmp_path: Path) -> None: + service = QueryService(_paths(tmp_path)) + strict = service.related_work(["agent workbench"]) + loose = service.related_work(["agent workbench"], include_partial=True) + + strict_names = {entry["name"] for entry in strict["entries"]} + assert "Science Discovery Suite" not in strict_names + assert "New Agent Memory Benchmark" not in strict_names + assert len(loose["entries"]) >= len(strict["entries"]) + + +def test_radar_leads_are_scholarly_records_with_recorded_authors(tmp_path: Path) -> None: + payload = QueryService(_paths(tmp_path)).related_work(["agent workbench"]) + by_name = {entry["name"]: entry for entry in payload["entries"]} + + assert "Agent Workbench Fork" not in by_name, "repository leads are not paper citations" + lead = by_name["Agent Workbench Pro: Long-Horizon Agent Evaluation"] + assert lead["cite_key"] == "lovelace2026agent" + assert lead["arxiv_id"] == "2608.01234" + assert "author = {Ada Lovelace and Alan Turing}" in lead["bibtex"] + assert "radar_lead_unverified" in lead["verification"] + # The lead-in clause is dropped and the "we" sentence becomes an author citation. + assert "\\citet{lovelace2026agent} introduce Agent Workbench Pro" in payload["latex"] + assert "50\\% harder" in payload["latex"] + + +def test_catalog_records_without_authors_are_flagged_not_invented(tmp_path: Path) -> None: + payload = QueryService(_paths(tmp_path)).related_work(["agent workbench"], include_radar=False) + entry = next(item for item in payload["entries"] if item["name"] == "Agent Workbench") + + assert entry["authors"] == [] + assert "authors_missing" in entry["verification"] + assert " author " not in entry["bibtex"] + assert " key = {Agent Workbench}," in entry["bibtex"] + assert all(entry["kind"] == "catalog" for entry in payload["entries"]) + assert "radar_first_date" not in payload["coverage"] + + +def test_coverage_statement_names_the_corpus_window(tmp_path: Path) -> None: + payload = QueryService(_paths(tmp_path)).related_work(["agent workbench"]) + + coverage = payload["coverage"] + assert coverage["radar_first_date"] == "2026-08-29" + assert coverage["radar_latest_date"] == "2026-08-30" + assert "not about the literature" in coverage["statement"] + assert coverage["statement"] in payload["latex"] + assert "| Work | Cite key |" in payload["markdown"] + + +def test_invalid_related_work_requests_are_machine_readable(tmp_path: Path) -> None: + service = QueryService(_paths(tmp_path)) + with pytest.raises(QueryError) as empty: + service.related_work(["Label="]) + assert empty.value.code == "invalid_query" + with pytest.raises(QueryError) as limit: + service.related_work(["agent"], per_topic=0) + assert limit.value.code == "invalid_limit" + + +def test_latex_escape_handles_specials_and_greek() -> None: + assert latex_escape("τ-bench 50% & $5") == "\\ensuremath{\\tau}-bench 50\\% \\& \\$5" + assert latex_escape("ΔΑ") == "\\ensuremath{\\Delta}A" + assert latex_escape("评测 bench ✅,ok") == "bench ,ok" + + +def test_cli_and_http_return_the_same_related_work_contract(tmp_path: Path, capsys) -> None: + paths = _paths(tmp_path) + tex_path, bib_path = tmp_path / "out" / "related.tex", tmp_path / "out" / "related.bib" + exit_code = run_query_cli( + [ + "related-work", + "Agent benchmarks=agent workbench", + "science discovery", + "--json", + "--tex", + str(tex_path), + "--bib", + str(bib_path), + "--index", + str(paths.index), + "--shards", + str(paths.shards), + "--snapshots", + str(paths.snapshots), + ] + ) + cli_payload = json.loads(capsys.readouterr().out) + + server = create_query_server(QueryService(paths), host="127.0.0.1", port=0) + thread = threading.Thread(target=server.serve_forever, daemon=True) + thread.start() + try: + query = urllib.parse.urlencode( + [("q", "Agent benchmarks=agent workbench"), ("q", "science discovery")] + ) + with urllib.request.urlopen( + f"http://127.0.0.1:{server.server_port}/api/v1/related-work?{query}", timeout=5 + ) as response: + http_payload = json.load(response) + finally: + server.shutdown() + server.server_close() + thread.join(timeout=5) + + assert exit_code == 0 + assert cli_payload == http_payload + assert [topic["label"] for topic in cli_payload["topics"]] == [ + "Agent benchmarks", + "science discovery", + ] + assert tex_path.read_text(encoding="utf-8") == cli_payload["latex"] + assert bib_path.read_text(encoding="utf-8") == cli_payload["bibtex"] From 0b146f6a2f6eb2702da5b7f459cb5cd00a4038d7 Mon Sep 17 00:00:00 2001 From: ktwu01 Date: Mon, 28 Sep 2026 19:33:38 -0500 Subject: [PATCH 05/29] Route the related-work job through the new command in the Skill The consumer Skill gains a five-step related-work flow: derive labelled topics, run `related-work --json`, read and prune every entry, rewrite the draft while keeping each citation (including the landscape-size citation of Benchmark Radar, never a "found with" sentence), resolve verification flags, and report the coverage window. query-surfaces.md records the command's contract. Co-Authored-By: Claude Opus 5.5 --- docs/query-surfaces.md | 8 ++++++++ skills/benchmark-radar/SKILL.md | 31 +++++++++++++++++++++++++++++++ tests/test_readme.py | 10 ++++++++++ 3 files changed, 49 insertions(+) diff --git a/docs/query-surfaces.md b/docs/query-surfaces.md index ce67a3fe..aef7e0a2 100644 --- a/docs/query-surfaces.md +++ b/docs/query-surfaces.md @@ -27,6 +27,14 @@ health, across the CLI, the HTTP surface, and the public consumer Skill. as adjacent character pairs, and accepts other Unicode letter words. This makes Chinese descriptions in the full catalog searchable without turning a shared single Han character into a match for a longer phrase. +- `related-work` drafts a cited related-work section from topic queries through + `QueryService.related_work`, over the same offline artifacts as `search` and + `show`. It keeps full lexical matches unless partial matches are requested, + admits only scholarly Radar sources, cites every retained entry, and cites the + Benchmark Radar paper once for the size of the benchmark landscape. Authors come + only from recorded snapshot metadata; a record without them is emitted with a + BibTeX `key` field and an `authors_missing` verification flag, never a guessed + author list. Every payload carries a coverage statement naming the corpus window. - Catalog records and daily discovery observations describe different things. Label a discovery observation as evidence of a mention or release, and retain the benchmark record it refers to. Source membership must not establish a diff --git a/skills/benchmark-radar/SKILL.md b/skills/benchmark-radar/SKILL.md index 8827fe15..0ed0cfb4 100644 --- a/skills/benchmark-radar/SKILL.md +++ b/skills/benchmark-radar/SKILL.md @@ -57,6 +57,8 @@ inspecting details with `show`, or tracking recent evidence). `benchmark-radar show "" --json` - Inspect the newest Radar evidence: `benchmark-radar recent --json` +- Draft a cited related-work section for a paper: + `benchmark-radar related-work "