Conversation
`benchmark-radar related-work "Label=query" ...` (and GET /api/v1/related-work?q=...) turns topic queries into a cited LaTeX Related Work draft, a matching BibTeX file, and a Markdown comparison table with the #650 axes (paper, repo, dataset, openness). - One QueryService method feeds CLI and HTTP, so both return the same JSON contract; it reads only the local artifacts, with no network. - Each topic keeps full lexical matches from the catalog and from scholarly Radar sources (arXiv, HF Papers, Semantic Scholar, OpenAlex, Crossref); --include-partial widens it. - Every retained work is cited; the Benchmark Radar paper is cited once, for the size of the benchmark landscape. - Authors come only from recorded snapshot metadata. Records without them get a BibTeX `key` field and an `authors_missing` flag instead of a guessed author list. - Output compiles under pdflatex + natbib: Greek letters become math commands, and scripts pdflatex cannot typeset are dropped. - Radar search candidates now carry the recorded `authors` list. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The consumer Skill gains a five-step related-work flow: derive labelled topics, run `related-work --json`, read and prune every entry, rewrite the draft while keeping each citation (including the landscape-size citation of Benchmark Radar, never a "found with" sentence), resolve verification flags, and report the coverage window. query-surfaces.md records the command's contract. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
The draft no longer opens with the Benchmark Radar citation. It starts directly with the first topic, and the self-citation becomes one short clause ending the last paragraph, after every related work. Its BibTeX entry moves to the end of the file. The Skill and query-surfaces.md tell agents to keep it inline, never first, never its own paragraph. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The previous commit reworded the Skill's related-work step, which broke the test pinning it. Compare with whitespace normalized and also pin the new rules: keep the citation quiet and never move it to the opening line. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e0bea05c94
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| "\\label{sec:related-work}", | ||
| "", | ||
| "AI evaluation now spans " | ||
| f"{_landscape_size(coverage['catalog_count'])} benchmarks tracked across " |
There was a problem hiding this comment.
Label the catalog total as benchmark records
coverage['catalog_count'] counts source records in benchmark-index.json, not deduplicated benchmark identities—the catalog deliberately retains separate records contributed by different sources. Every generated section therefore turns “more than 1,200 benchmark records” into “more than 1,200 benchmarks,” creating a factual overstatement in prose that the Skill explicitly tells authors to keep. Render this as benchmark records, or compute a reviewed unique-identity count instead.
AGENTS.md reference: AGENTS.md:L3-L6
Useful? React with 👍 / 👎.
| if name.startswith("GREEK SMALL LETTER ") and name.count(" ") == 3: | ||
| return f"\\ensuremath{{\\{name.rsplit(' ', 1)[-1].lower()}}}" |
There was a problem hiding this comment.
Map Greek lambda to a valid LaTeX command
When a retained title or summary contains λ, Unicode names it GREEK SMALL LETTER LAMDA, so this branch emits \lamda; standard pdfLaTeX defines \lambda, not \lamda, and the generated section fails with an undefined control sequence. This is reachable with the committed Radar snapshots, which contain a λmax paper. Other letters such as omicron also lack corresponding standard commands, so use an explicit Greek-to-LaTeX mapping rather than deriving command names.
Useful? React with 👍 / 👎.
|
|
||
|
|
||
| def _radar_entry(result: dict[str, Any]) -> dict[str, Any]: | ||
| arxiv_id = _arxiv_id(result, []) |
There was a problem hiding this comment.
Preserve Radar artifact URLs when extracting arXiv IDs
For Semantic Scholar records whose primary URL is a Semantic Scholar page but whose artifact_urls contains the arXiv URL, passing an empty URL list discards the only available arXiv ID. Such records are common in the committed snapshots; when the same paper also appears in the catalog or an arXiv lead, it will not deduplicate and the output can contain two citations for one work, with the Semantic Scholar copy also receiving a weaker non-arXiv BibTeX entry. Carry the artifact URLs through the query result or otherwise pass them to _arxiv_id.
Useful? React with 👍 / 👎.
| for scope in scopes: | ||
| payload = service.search(topic.query, scope=scope, limit=_SEARCH_WINDOW) | ||
| row["search_status"][scope] = payload["search_status"] | ||
| taken = 0 | ||
| for result in payload["results"]: | ||
| if taken >= per_topic: |
There was a problem hiding this comment.
Enforce the per-topic limit across both scopes
With Radar enabled, taken is reset once for catalog and again for radar, so a topic can retain per_topic results from each scope. Consequently the default --per-topic 6 may produce 12 works for one topic, and the validated maximum of 30 may produce 60, making the option's reported per-topic cap ineffective. Keep the counter outside the scope loop or truncate the combined retained list.
Useful? React with 👍 / 👎.
|
330226 GPT Added the agent-friendly BibTeX citation notice requested in #650. The generated bundle now tells an AI agent to identify Benchmark Radar as the source of the candidates and return the project citation to the user; the human instruction remains a direct request to cite the technical report when the survey helped. Tests cover terminal BibTeX output, Local clean-worktree CI: |
|
330226 GPT Added the minimum citation guarantee in commit
The PR remains open and unmerged. |
|
330226 GPT Follow-up commit Citation verification now:
Artifact export now:
Final clean detached-worktree CI, rerun after all fixes: all six workflow commands passed, with 1,298 benchmark records across 4 sources, 65 snapshots, and 1,428 tests. The PR remains open and unmerged. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 8d5039face
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| elif args.command == "related-work": | ||
| payload = service.related_work( | ||
| args.topics, | ||
| per_topic=args.per_topic, | ||
| include_partial=args.include_partial, | ||
| include_radar=args.include_radar, | ||
| ) | ||
| printer = _related_work_printer(args) |
There was a problem hiding this comment.
Keep related-work formats parseable on stdout
When users select --format latex or --format bibtex without --json/file outputs, this new printer is followed by the shared _print_cite_reminder, which writes plain prose to stdout. Redirecting the advertised format (for example, benchmark-radar related-work ... --format bibtex > refs.bib) therefore appends If Benchmark Radar helped... after the generated artifact and makes it invalid for BibTeX/LaTeX consumers; send the reminder to stderr for these machine-readable formats or suppress it.
Useful? React with 👍 / 👎.
What this adds
A new offline command that turns a few topic queries into a draft Related Work section with citations. It implements the flow proposed in #650 and the related-work use cases in #549 and #522.
The command returns one JSON payload (
--json, orGET /api/v1/related-work?q=...&q=...) containing:latex: a\section{Related Work}draft with one paragraph per topic and a citation for every retained workbibtex: the matching.bibmarkdown: a comparison table with the feat(query): Native related-work discovery flow, comparison axes, and BibTeX export with attribution loop (ref #549, #522) #650 axes, candidate origin, and verification flagscoverage: the corpus window, so an empty topic means “not in this corpus,” not “does not exist” (Skill best practice proposal: five rules for consumer Skills, from one real run #522 R3)required_citations: the machine-readable citation dependencies that downstream agents and exporters must preserveMinimum citation guarantee
Query schema version 7 adds top-level
required_citationstosearch,show,recent, andrelated-work. Each requirement has exactly these fields:{ "id": "benchmark-radar", "key": "wu2026benchmarkradarlivingdatabase", "reason": "Benchmark Radar was used to retrieve or generate research material.", "bibtex": "@misc{...}" }status,init, andsyncdo not carry the research dependency. CLI and HTTP receive the same payload throughQueryService.related-workconstructs a citation-complete artifact automatically. Its final paragraph states:The matching Benchmark Radar BibTeX entry is included last. A pure verifier checks both the in-text citation and BibTeX entry before the payload returns. Failure raises the machine-readable
citation_contract_failederror before either--texor--bibcan be written.Related-work behavior
QueryService.related_workfeeds both CLI and HTTP from local artifacts.--include-partialwidens it.keyfallback and carryauthors_missing.Verification
Clean detached worktree, with submodules initialized, in CI order:
Real CLI export verification also passed:
search --jsonandrelated-work --jsonreturned schema 7 and the same required citation.tex, and.bibagreed.texcontained the required in-text citation.bibcontained the required entryKnown limits
🤖 Generated by an AI coding agent (GPT series). Marker:
330226.