Skip to content

Add related-work command: cited LaTeX draft + BibTeX from topic queries (#650, #549) - #701

Open
ktwu01 wants to merge 7 commits into
mainfrom
feat/related-work
Open

ktwu01 wants to merge 7 commits into
mainfrom
feat/related-work

Conversation

@ktwu01

@ktwu01 ktwu01 commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

What this adds

A new offline command that turns a few topic queries into a draft Related Work section with citations. It implements the flow proposed in #650 and the related-work use cases in #549 and #522.

benchmark-radar related-work \
  "Persona agents=persona" "User simulation=simulated users" \
  "Personalized assistants=personalization" "Computer-use agents=computer use" \
  --tex related.tex --bib related.bib

The command returns one JSON payload (--json, or GET /api/v1/related-work?q=...&q=...) containing:

Minimum citation guarantee

Query schema version 7 adds top-level required_citations to search, show, recent, and related-work. Each requirement has exactly these fields:

{
  "id": "benchmark-radar",
  "key": "wu2026benchmarkradarlivingdatabase",
  "reason": "Benchmark Radar was used to retrieve or generate research material.",
  "bibtex": "@misc{...}"
}

status, init, and sync do not carry the research dependency. CLI and HTTP receive the same payload through QueryService.

related-work constructs a citation-complete artifact automatically. Its final paragraph states:

Candidate benchmarks were retrieved using Benchmark Radar~\citep{wu2026benchmarkradarlivingdatabase} and should be verified against their primary sources.

The matching Benchmark Radar BibTeX entry is included last. A pure verifier checks both the in-text citation and BibTeX entry before the payload returns. Failure raises the machine-readable citation_contract_failed error before either --tex or --bib can be written.

Related-work behavior

  • One source of truth, offline. QueryService.related_work feeds both CLI and HTTP from local artifacts.
  • Strict by default. A topic keeps candidates that match every query token; --include-partial widens it.
  • Scholarly Radar sources only. arXiv, Hugging Face Papers, Semantic Scholar, OpenAlex, and Crossref can become citation entries. Repositories, datasets, and model uploads remain comparison metadata rather than papers.
  • No invented authors. Authors come from snapshot metadata. Records without authors use BibTeX's key fallback and carry authors_missing.
  • Deduplicated. A catalog record and Radar lead with the same arXiv ID merge into one entry and one citation.
  • Export-safe. LaTeX escaping handles recorded titles and summaries, and the citation contract is verified before files are written.

Verification

Clean detached worktree, with submodules initialized, in CI order:

ruff check .                         passed
ruff format --check .                passed
benchmark-radar normalize-catalog    passed: 1,298 records across 4 sources
benchmark-radar classify             passed
benchmark-radar build-data-release   passed: 1,298 benchmarks, 65 snapshots
pytest -q                            passed: 1,413 tests

Real CLI export verification also passed:

  • search --json and related-work --json returned schema 7 and the same required citation
  • generated JSON, .tex, and .bib agreed
  • the .tex contained the required in-text citation
  • the .bib contained the required entry
  • Tectonic compiled the generated artifact to PDF with no undefined citations

Known limits

  • Generated sentences are a draft, not semantic synthesis. The consumer Skill tells the agent to read, prune, verify, and rewrite retained works while preserving required citations.
  • Catalog summaries can be caveat notes rather than paper descriptions.
  • Lexical retrieval and the available Radar window can miss older prior art. The payload reports its coverage and must not be used as proof of novelty.

🤖 Generated by an AI coding agent (GPT series). Marker: 330226.

ktwu01 and others added 2 commits September 28, 2026 19:33
`benchmark-radar related-work "Label=query" ...` (and GET
/api/v1/related-work?q=...) turns topic queries into a cited LaTeX
Related Work draft, a matching BibTeX file, and a Markdown comparison
table with the #650 axes (paper, repo, dataset, openness).

- One QueryService method feeds CLI and HTTP, so both return the same
  JSON contract; it reads only the local artifacts, with no network.
- Each topic keeps full lexical matches from the catalog and from
  scholarly Radar sources (arXiv, HF Papers, Semantic Scholar,
  OpenAlex, Crossref); --include-partial widens it.
- Every retained work is cited; the Benchmark Radar paper is cited
  once, for the size of the benchmark landscape.
- Authors come only from recorded snapshot metadata. Records without
  them get a BibTeX `key` field and an `authors_missing` flag instead
  of a guessed author list.
- Output compiles under pdflatex + natbib: Greek letters become math
  commands, and scripts pdflatex cannot typeset are dropped.
- Radar search candidates now carry the recorded `authors` list.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The consumer Skill gains a five-step related-work flow: derive labelled
topics, run `related-work --json`, read and prune every entry, rewrite
the draft while keeping each citation (including the landscape-size
citation of Benchmark Radar, never a "found with" sentence), resolve
verification flags, and report the coverage window. query-surfaces.md
records the command's contract.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI balanced review requested due to automatic review settings September 29, 2026 00:34

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-29T00:38:30.822683Z e0bea05 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

The draft no longer opens with the Benchmark Radar citation. It starts
directly with the first topic, and the self-citation becomes one short
clause ending the last paragraph, after every related work. Its BibTeX
entry moves to the end of the file. The Skill and query-surfaces.md tell
agents to keep it inline, never first, never its own paragraph.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings September 29, 2026 00:36

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

The previous commit reworded the Skill's related-work step, which broke
the test pinning it. Compare with whitespace normalized and also pin the
new rules: keep the citation quiet and never move it to the opening line.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings September 29, 2026 00:36

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e0bea05c94

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

"\\label{sec:related-work}",
"",
"AI evaluation now spans "
f"{_landscape_size(coverage['catalog_count'])} benchmarks tracked across "

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Label the catalog total as benchmark records

coverage['catalog_count'] counts source records in benchmark-index.json, not deduplicated benchmark identities—the catalog deliberately retains separate records contributed by different sources. Every generated section therefore turns “more than 1,200 benchmark records” into “more than 1,200 benchmarks,” creating a factual overstatement in prose that the Skill explicitly tells authors to keep. Render this as benchmark records, or compute a reviewed unique-identity count instead.

AGENTS.md reference: AGENTS.md:L3-L6

Useful? React with 👍 / 👎.

Comment on lines +59 to +60
if name.startswith("GREEK SMALL LETTER ") and name.count(" ") == 3:
return f"\\ensuremath{{\\{name.rsplit(' ', 1)[-1].lower()}}}"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Map Greek lambda to a valid LaTeX command

When a retained title or summary contains λ, Unicode names it GREEK SMALL LETTER LAMDA, so this branch emits \lamda; standard pdfLaTeX defines \lambda, not \lamda, and the generated section fails with an undefined control sequence. This is reachable with the committed Radar snapshots, which contain a λmax paper. Other letters such as omicron also lack corresponding standard commands, so use an explicit Greek-to-LaTeX mapping rather than deriving command names.

Useful? React with 👍 / 👎.



def _radar_entry(result: dict[str, Any]) -> dict[str, Any]:
arxiv_id = _arxiv_id(result, [])

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve Radar artifact URLs when extracting arXiv IDs

For Semantic Scholar records whose primary URL is a Semantic Scholar page but whose artifact_urls contains the arXiv URL, passing an empty URL list discards the only available arXiv ID. Such records are common in the committed snapshots; when the same paper also appears in the catalog or an arXiv lead, it will not deduplicate and the output can contain two citations for one work, with the Semantic Scholar copy also receiving a weaker non-arXiv BibTeX entry. Carry the artifact URLs through the query result or otherwise pass them to _arxiv_id.

Useful? React with 👍 / 👎.

Comment on lines +311 to +316
for scope in scopes:
payload = service.search(topic.query, scope=scope, limit=_SEARCH_WINDOW)
row["search_status"][scope] = payload["search_status"]
taken = 0
for result in payload["results"]:
if taken >= per_topic:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Enforce the per-topic limit across both scopes

With Radar enabled, taken is reset once for catalog and again for radar, so a topic can retain per_topic results from each scope. Consequently the default --per-topic 6 may produce 12 works for one topic, and the validated maximum of 30 may produce 60, making the option's reported per-topic cap ineffective. Keep the counter outside the scope loop or truncate the combined retained list.

Useful? React with 👍 / 👎.

Copilot AI review requested due to automatic review settings September 29, 2026 02:14

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@junjiezhou1122

Copy link
Copy Markdown
Collaborator

330226 GPT

Added the agent-friendly BibTeX citation notice requested in #650. The generated bundle now tells an AI agent to identify Benchmark Radar as the source of the candidates and return the project citation to the user; the human instruction remains a direct request to cite the technical report when the survey helped. Tests cover terminal BibTeX output, --bib file export, and the unchanged final citation key.

Local clean-worktree CI: ruff check ., ruff format --check ., benchmark-radar normalize-catalog, benchmark-radar classify, benchmark-radar build-data-release, and pytest -q all passed. Test result: 1409 passed.

Copilot AI review requested due to automatic review settings September 29, 2026 09:03

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@junjiezhou1122

Copy link
Copy Markdown
Collaborator

330226 GPT

Added the minimum citation guarantee in commit c5710eb.

  • Query schema is now 7.
  • search, show, recent, and related-work expose top-level required_citations; health/data-management commands do not.
  • related-work now emits explicit Benchmark Radar retrieval provenance and verifies both the in-text citation and BibTeX entry before returning or writing files.
  • Citation-contract failures return citation_contract_failed and leave both export paths absent.
  • Clean-worktree CI passed in workflow order: 1,298 records across 4 sources, 65 snapshots, and 1,413 tests.
  • A real CLI export compiled successfully with Tectonic and no undefined citations.

The PR remains open and unmerged.

Copilot AI review requested due to automatic review settings September 29, 2026 10:04

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@junjiezhou1122

Copy link
Copy Markdown
Collaborator

330226 GPT

Follow-up commit 8d5039f addresses the adversarial subagent review findings.

Citation verification now:

  • ignores commented-out LaTeX/BibTeX,
  • requires an active exact citation key,
  • rejects escaped commands and suffixed keys,
  • requires a complete canonical Benchmark Radar BibTeX entry,
  • accepts valid whitespace variants.

Artifact export now:

  • stages every requested output before committing,
  • rejects directory, symlink, duplicate, and aliased destinations before writing,
  • preserves existing permissions,
  • restores prior files if a later commit fails,
  • returns machine-readable artifact_write_failed on write/rollback failure,
  • preserves unrecovered backups and reports incomplete rollback,
  • returns artifact_cleanup_failed if outputs committed but backup cleanup failed.

Final clean detached-worktree CI, rerun after all fixes: all six workflow commands passed, with 1,298 benchmark records across 4 sources, 65 snapshots, and 1,428 tests. The PR remains open and unmerged.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 8d5039face

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +373 to +380
elif args.command == "related-work":
payload = service.related_work(
args.topics,
per_topic=args.per_topic,
include_partial=args.include_partial,
include_radar=args.include_radar,
)
printer = _related_work_printer(args)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Keep related-work formats parseable on stdout

When users select --format latex or --format bibtex without --json/file outputs, this new printer is followed by the shared _print_cite_reminder, which writes plain prose to stdout. Redirecting the advertised format (for example, benchmark-radar related-work ... --format bibtex > refs.bib) therefore appends If Benchmark Radar helped... after the generated artifact and makes it invalid for BibTeX/LaTeX consumers; send the reminder to stderr for these machine-readable formats or suppress it.

Useful? React with 👍 / 👎.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants