Skip to content

Review: Add NVIDIA mirror + task verifiers (site by @KaKituken, verifiers by reviewer) (#58) - #107

Merged
Raibows merged 164 commits into
aiming-lab:mainfrom
TabsPhasers:review/pr-58-nvidia
Sep 14, 2026
Merged

Review: Add NVIDIA mirror + task verifiers (site by @KaKituken, verifiers by reviewer) (#58)#107
Raibows merged 164 commits into
aiming-lab:mainfrom
TabsPhasers:review/pr-58-nvidia

Conversation

@TabsPhasers

Copy link
Copy Markdown
Contributor

Review PR 正文草稿(A4,Draft 状态使用)

Title: Review: Add NVIDIA mirror + task verifiers (site by @KaKituken, verifiers by reviewer) (#58)

Base: aiming-lab/WebHarbor:main ← Head: TabsPhasers:review/pr-58-nvidia(Draft)


Summary

基于当前 main Review 并接管 @KaKituken 的站点贡献(#55)与 @DEM1TASSE 的 verifier/rubric 交付(#58);两位原作者的提交作为祖先保留(请以 regular merge 合并,不要 squash)。NVIDIA 注册为仓库第 24 个站点、容器端口 40023(原评论中的 40016 已被 IKEA 占用)。

Original review resolution

  1. Rebase onto current main / 端口:已在当前 main 上重放;NVIDIA 位于 40023,websyn_start.sh/control_server.py/Dockerfile 三处一致,boot banner 改为动态计数。
  2. HF 素材 immutable pin:素材由独立 HF PR 交付(见 Asset delivery status);合并后 .assets-revision pin 到 merge SHA,并在干净 checkout 验证 fetch_assets.sh
  3. 站点保真(per-series / commerce):新增 40/50 系列 landing 与技术 section;T11 改为“系列技术事实 + 本地购买信息(NVIDIA Marketplace / United States / en-us)并停在本地页面”;T12/T17 改以本地 Wishlist 表达,通用 Cart/Checkout 入口退役。
  4. T5 去捷径:改为 Nano Super 开发套件 vs Orin NX 16GB 生产模块的身份/容量比较;首页正常价签保留,题目不再依赖首页价格。
  5. Nits:boot banner 计数修复;驱动版本为镜像冻结快照并在页面说明,不伪造系列差异。

Review changes

  • verifier 重写为站点自包含、显式权威输入、确定性 stdlib:无 LLM/网络/默认 DB fallback;exit 0/1/2 语义与 JSON 契约固定;修复数字/单位/对象绑定、状态 delta/账号、T15 标题等误判;10 组对抗性加固(负值、否定作用域、按对象绑定、跨单位、证据路由、OS/地区矛盾、纪元等)。
  • T5/T11/T12/T17 显式改题(query/rubric/verifier 映射记录在 PR 描述与站点文档);其余 16 题语义保持。
  • 站点 UI:系列/购买页、Contact Sales、窄屏 header/表格、图片 caption(示意类图片均标注 “not a photograph of this model”)、登录保留 next 回跳。
  • 站点本地测试:test_ui_contract.py(机械 HTML/test-client)与 tests/(verifier 契约/回归)。

Merge note (ports)

This branch is based on main at 129a2742 and registers NVIDIA as site 24 / container port 40023. If main has since taken 40023 (e.g. walmart_careers), move NVIDIA to the next free slot on merge — in websyn_start.sh, control_server.py, Dockerfile (EXPOSE …) and the web port in sites/nvidia/tasks.jsonl. The verifiers are port-agnostic (they match URL paths, not the host port).

Asset delivery status

在该 HF PR 合并、.assets-revision 更新到 merge SHA 且 clean-fetch 通过前,本 GitHub PR 保持 Draft,不得合并。

Validation

  • 判分门禁(冻结输入,最终 verifier):20/20 真实完成 PASS(exit 0)+20/20 默认 no-op FAIL(exit 1),0 INFRA;tasks.jsonl 未改。
  • 最终构建(candidate003 8f74ff8d…:20/20 题 no-op baseline 严格 FAIL、wrapper success=true;受影响题 T8/T16/T17 fresh 原生重跑 PASS(含 DB delta 核验)。
  • 对抗性回归:93 例合成反例全部符合预期;既有 293 例判分套件与 44 项单元测试全过。
  • UI/流程:20 题×1440/390/320 主流程采集完成;1440 独立持久性 20 项通过(匿名 newsletter relogin N/A);主 Agent 逐张亲看 20/20 终态与关键帧,与判分/DB 一致。
  • 源 pin:site 集合 34d43ec8…、UI e50345cb…、grading 63fe5309…、tasks 6de754…;具体文件哈希随 PR 附件提供。

Scope

  • 保留原贡献归属;排除 reviewer 专用 run/证据/本地 overlay 与内部笔记(UI_REVIEW_NOTES.md 等)。
  • 共享框架改动:无。

Blocking checklist

  • HF 素材 PR 合并;.assets-revision pin 到 merge SHA 并在干净 checkout 通过 fetch/build/byte-identity 检查。

Reviews #58.

XuanRui LI and others added 30 commits June 4, 2026 20:44
Addresses review on PR aiming-lab#53:
- Rebase onto current main; register healthline as index 16 -> port 40016
  (append after merriam_webster in websyn_start.sh + control_server.py;
  Dockerfile EXPOSE -> 40000-40016). merriam_webster preserved.
- websyn_start.sh site-count comments reconciled to 17.
- .assets-revision pinned to HF PR aiming-lab#40 (clean tarball, no macOS AppleDouble
  junk, based on current main -> all 17 tarballs). Bump to merged SHA once
  HF PR aiming-lab#40 lands.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Reviewer deliverable for the Healthline mirror (site by @JeremyJC67, PR aiming-lab#53):
one deterministic verifier per task under sites/healthline/verify/, plus
verifier_path + judge_rubric recorded in every tasks.jsonl row. No answer
key in tasks.jsonl — ground truth lives only inside the verifiers.

Deterministic-first: (1) trajectory navigation gate (anti knowledge-shortcut,
important here since several answers are medically recallable), (2) SQLite DB
after-state for the stateful tasks (save / register / password-change, plus
DB cross-checks for saved-count and reading-history), (3) answer vs frozen
ground truth, with the LLM only as an anchored consistency check.

Validated against the official react agent (agent_demo/agent.py): a no-op run
fails all 20 verifiers; a human answer-check confirms every frozen value
matches what the page renders. On the full 20-task run, after fixing one
too-strict nav gate the run itself surfaced, verifier and LLM judge agree
15/20, and all 5 remaining divergences are the deterministic verifier being
correct while the LLM judge false-positives (blank answer / knowledge-shortcut)
or false-negatives (DB-confirmed save / password change).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Search result cards showed drug_class (e.g. 'ACE inhibitor'), letting T16
be answered from the results page without opening the drug pages. Show the
broader category (e.g. 'Heart Medications') instead — same card layout, but
the ACE-inhibitor-vs-statin distinction now requires opening each detail
page. drug_class still shown on the drug detail page; search backend still
matches on drug_class. Addresses DEM1TASSE review note on PR aiming-lab#53.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- web: all 20 tasks pointed at :40015 (Merriam-Webster) after the site was
  rebased to index 16 / :40016; corrected to http://localhost:40016/.
- judge_rubric for Healthline--3 and Healthline--16: spell out the full
  ground-truth answer the page states (T3 symptom list, T16 what each drug
  treats) so the LLM judge grades against complete facts.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Offline Flask mirror of https://doctor.webmd.com/ ("WebMD Care"): deterministic
synthetic directory of 224 doctors across 10 specialties and 8 cities around
Newark, DE 19711, with hospitals, group practices, Choice Awards, real auth,
saved providers, appointment requests and pending reviews.

- Build-generated seed (.build-generated-seed): the Dockerfile regenerates
  instance_seed/webmd_doctor.db plus 224 Pillow initials avatars and 90 video
  poster frames from seed_data.py; no Hugging Face assets, .assets-revision
  untouched.
- One seeded RNG (20260910), literal reference date 2026-09-10, hardcoded
  scrypt hashes for the four benchmark users, whole-function seed gates,
  seed_metadata version + expected-count validation; byte-identical after
  /reset and docker restart.
- Deterministic longest-match query parser (specialty / condition / procedure
  / insurer / gender / virtual), conjunctive filters, Best Match / Distance /
  Average Rating / Number of Ratings ranking, numbered pagination.
- Registered as index 24 in websyn_start.sh and control_server.py; Dockerfile
  EXPOSE raised to 40024, header comment bumped to 25 sites.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GTvmqJcShyy3v3KfEPxMPe
Contributor-side task definitions only (web_name, id, ques, web, upstream_url):
9 lookups (0-7, 15), 4 multi-constraint (9-12), 2 compares (13-14), 5 stateful
(8, 16-19). Every ques routes through a specialty + city, a hub page or an
awards class before naming a doctor (name search is a non-goal, as upstream);
credentials are embedded for login tasks; no answers, verifiers or rubrics.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GTvmqJcShyy3v3KfEPxMPe
- seed: two extra West Chester cardiology slots appended after the grid
  (task 12 now has 8 doctors on the city page, 6 rated 4+ with exactly one
  male); deterministic language top-up for physicians covering three
  offices (task 1 target speaks 3 languages); EXPECTED_COUNTS refreshed
  (226 doctors, 202 in radius)
- pagination: disabled prev/next rendered as spans instead of live links
  to an empty page
- filter bar: empty and default params are dropped from auto-submitted URLs

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GTvmqJcShyy3v3KfEPxMPe
Section-by-section pass against scraped_data/reference at 1440 px:
- profile: booking widget overhangs the hero into the rail, tabs span the
  main column only, colleagues tile grid / hospital affiliation card in the
  stack, Basic rail below the tabs, NPI folded into the certification card,
  nearby + transparency panels full width, star picker, review controls,
  check-circle perspective icons, two-column bulleted top-20 lists
- home: icon tile row for popular specialties, filled uppercase View Profile
  buttons, upstream heading sizes
- filter pills in dark text with navy active state; typeahead no longer
  clipped by the search bar, prefix highlighting
- hospitals / group practices hubs rebuilt as landing pages (state chips,
  name search, top-4 cards, search-all button, care-type chips); state lists
  drop the sliders button and gain a name filter
- hospital / practice details: specialty select, map-on-top locations,
  ratings card, text flag lines
- awards page art panels + larger headings; recipients grouped by state
  with a state filter; guidelines as a centred numbered card; sign-up terms
  line and mm/dd/yyyy date field; login/state/city copy per upstream
- landing "Highest Rated" strip now the 25-mile ring so no task target
  is showcased there

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GTvmqJcShyy3v3KfEPxMPe
Raibows and others added 27 commits September 12, 2026 11:34
The templates' image fallback pointed at avatars/default.png, which is
byte-identical to avatars/alicejdata.png (md5 cba9811627f5a439dae16a273f458a62),
and User.avatar_url defaulted to the same file, so every account without an
avatar - including any newly registered user - rendered Alice's photo.

A new avatar() macro renders the stored image when present and otherwise a letter
tile (initial on the action-colour gradient); User.avatar_url and register() now
default to an empty string, and all nine call sites (nav, profile, account,
search rows, rankings, competition host, dataset owner, notebook author, model
owner) use the macro. Verified: a new account stores '' and shows a letter tile,
seeded users keep their own photos, default.png is no longer referenced, and a
21-route sweep returns 200 everywhere (a missing macro import briefly returned 500
on /rankings and /search and was fixed before this commit).
base.html's nav calls avatar(current_user, 32) but only imported the macro into
child templates, so every page whose own template does not import it returned 500
while signed in (jinja2 UndefinedError on /account/edit, /discussions/new,
/account/password; logged-out renders were unaffected because the nav avatar sits
behind current_user.is_authenticated). Caught by the authenticated route sweep.
base.html now imports the macro itself; the authenticated and logged-out sweeps
both return 200/404 as expected.
…ccount deletion

SQLite leaves foreign keys off unless a connection enables them, so the schema's
declared FKs were decoration and /account/delete removed only the users row: a
verified probe showed an orphan vote insert succeeding and a deleted account
leaving its votes behind (author usernames in discussions/comments dangled too).

A connect-time event now sets PRAGMA foreign_keys=ON (app context reports
[(1,)] and an orphan vote insert raises IntegrityError), and account deletion
removes the account's comments, discussions, submissions, votes, bookmarks,
follows (both directions) and competition entries before deleting the user, with
a rollback plus a flash message on SQLAlchemyError instead of a 500. Verified:
after deletion the user row is gone, every table reports 0 orphan rows and
PRAGMA foreign_key_check is empty; switching the pragma back to OFF reproduces
the orphan vote.
The review description claimed 44 unit tests passing and a boot-banner fix. 44 is
the discovery count when the driver regression class is skipped (unittest reports
'OK (skipped=1)'); with DRIVER_TEST_INPUTS set the same discovery runs 73 tests and
that class fails test_other_information_task_not_relaxed, and the boot banner
already used SITE_COUNT in base commit 129a274. README now lists the exact commands
and current counts (299 mechanical cases, 15 UI-contract tests, 11 CLI-contract
tests, 29 driver tests, 52/73 discovery) and notes that the GitHub PR body itself is
not editable from this repository.

Evidence: _wh_review_tools/pr107-fixes/fixes/L5/after.txt
Twelve links point at nvidia.com, marketplace.nvidia.com and store.nvidia.com. They
are labelled as leaving the mirror, no route fetches them (0 external requests over
114 routes at four widths), and any run that follows one is graded as an off-origin
failure rather than silently passing. README records that scope.

Evidence: _wh_review_tools/pr107-fixes/fixes/L8/after.txt
Records the phase-1 findings, the per-finding disposition with its commit, the ground
truth derivation for all 20 tasks, the phase-2 validation results and the evidence
locations, following the structure of review-reports/PR-87-FINAL-AUDIT.md.
The verifiers read only the URLs, the answer and the DB state, so a trajectory
with its tail cut off still passed: the truncated Kaggle--6 variant (last step
'check' instead of 'done') returned PASS against the pre-fix verifier, and the
same class passed for Kaggle--10/12 in stage 1.

verify_lib.run_complete() now requires the fields agent_demo always writes:
terminated is True, termination_reason is 'agent_done', the last step is the
final 'done' step, and that step carries an answer. All 20 verifiers assert it via
a run_complete check. Verified: truncated variants now fail with reason
run_complete ('the last step is ... not the final done step'), runs that stopped on
max_steps fail too, and the real runs pass ('terminated with agent_done after 15
steps').
Records the stage-1 findings for sites/kaggle (issue-table.md) and their
dispositions after remediation: the pin/registry/verifier-contract blockers that
were fixed, the application and verifier hardening (H1-H9, M1-M10, L1-L9), the
three items that stay blocked on the Hugging Face asset archive or an LLM
endpoint, the ground truth for all 20 tasks derived from the seed database, and
the validation runs (production grading command 20/20, 227-case negative matrix
with zero mismatches, four-width crawl, cold image build with 25/25 sites alive
and byte-identical reset-all, container-generated seed equality, leak re-scan).
Add the reviewed FedEx mirror as site 25 on port 40024.

Offline FedEx mirror with 18 tasks. The seed database is generated from tracked source at image build time, every graded target is derived from the supplied initial database by verify/ground_truth.py, and the answer-isolation tests keep graded values out of tracked source, help copy and static page surfaces. Media is distributed through the pinned Hugging Face archive at revision 68dcbf2c.
…WebMD Doctor branch

Registry: append `webmd_doctor` after `fedex`, so FedEx keeps index 24 / port
40024 and WebMD Doctor moves to index 25 / port 40025 (26 sites, 40000-40025).

Conflict resolutions:
- `.assets-revision`: keep `ad6f424f` (current HF dataset main, which carries
  `webmd_doctor.tar.gz` and the byte-identical `fedex.tar.gz` blob) instead of
  main's FedEx-only pin `68dcbf2c`.
- `scripts/fetch_assets.sh`: take main's registry-scoped fetch implementation.
- `Dockerfile`: keep both site blocks (FedEx asset gate + WebMD generated-asset
  gate and source-built seeds) and widen `EXPOSE` to 40000-40025.
- `README.md`, `AGENTS.md`, `CLAUDE.md`, `CONTRIBUTING.md`, `agent_demo/README.md`,
  `.claude/skills/*`: 26 mirrors, port range 40000-40025, alt ports 41000-41025.
- `sites/walmart_careers/tests/test_integration.py`,
  `sites/rotten_tomatoes/tests/test_environment_quality.py`: take main's
  registry-derived assertions and extend them for `webmd_doctor` (40025) and its
  generated-asset gate.

Follow-on work required by the port move:
- `sites/webmd_doctor/tasks.jsonl`, `README.md`, `verify/verify_lib.py` comment,
  `tests/test_task_breadth.py` (now derives the port from the registry),
  `verify/tests/test_verify_lib.py`: port 40025.
- new `sites/webmd_doctor/tests/conftest.py`: rebuild the build-generated seed
  when it is absent and install the runtime copy the app binds, so the suite runs
  after `scripts/fetch_assets.sh` alone (the Walmart mirror's convention). In that
  state the suite previously failed 7 tests.

Verified after the merge: webmd site tests 19, webmd verify 447, walmart 51,
fedex 91 (+1341 subtests), rotten_tomatoes 56 (+4129 subtests), compass 94.
feat(webmd_doctor): add WebMD Doctor mirror (site 25, port 40024)
…h_assets.sh

(cherry picked from commit c944482a7e91d3dfedd6e65f2baf4eadf7a34cf7)
…hline

Review: Add Healthline mirror + task verifiers (site by @JeremyJC67, verifiers by reviewer) (aiming-lab#59)
…asset-pin rationale

Shared docs (AGENTS.md, CLAUDE.md, CONTRIBUTING.md, README.md and the five
.claude/skills/*/SKILL.md files) still described 27 sites / 40000-40026 after
kaggle was appended as site 28 (index 27, port 40027). Bump the site counts and
port ranges to 28 / 40000-40027 (alt-port mapping 41000-41027) and add Kaggle to
the README site enumeration.

.assets-revision: rewrite the pin rationale to the measured state. The pin
64264d065cdb0b7755ee99dab356be5733d9ddef is the HF dataset `main` head, its
kaggle.tar.gz is 11278106 bytes (sha256 dd6f1ab3...) and passes
scripts/validate_asset_archive.py with 73 managed members and no bare
kaggle/static directory entry; refs/pr/37 covers only 17 of the 28 registered
sites. The pin value itself is unchanged.
…nvidia to index 28 (port 40028)

Merged review/pr-107 (the local phase-2 fix branch for PR aiming-lab#107, head
84bfe58) onto origin/main 145b200 (27 sites + kaggle). Conflicts in
.assets-revision, Dockerfile, control_server.py and websyn_start.sh were
resolved with orch/integration/resolve_conflicts.py nvidia NVIDIA 28 40023
64264d065cdb0b7755ee99dab356be5733d9ddef (40023 measured from the PR head's
tasks.jsonl).

The PR replaces walmart_careers' slot in its own base; current main already
registers walmart_careers, so the merge keeps it and appends nvidia as site 29
(index 28, port 40028). nvidia code, tasks and verifiers come from review/pr-107
unchanged apart from the port re-slot.
…icial NVIDIA renders

Three pairs of product images were byte-identical duplicates, wasting 325831
bytes: RTX 5070 = RTX 5070 Ti (sha256 e74f45250ec86a…), RTX 5060 = RTX 5060 Ti
(sha256 badb1b1e320fc43f…), RTX 4070 SUPER = RTX 4080 SUPER (sha256
2ffee1dde9594067…). Six further images did not show the product they illustrated:
RTX 4060 (lifestyle photo), H100 (abstract particle graphic), H200 (HGX 8-GPU
baseboard), DGX B200 (abstract chip diagram), L40S (isometric office
illustration), RTX PRO 6000 Blackwell (workstation lifestyle composite).

All 12 now carry the official NVIDIA render for that SKU, so
sites/nvidia/static/images/ holds 33 files with 33 distinct sha256 and no
duplicate group.

- sites/nvidia/asset_inventory.json: the 12 entries record the official NVIDIA
  source_url and source_kind.
- sites/nvidia/templates/_product_image.html and
  sites/nvidia/test_ui_contract.py: captions and UI-contract labels follow the new
  renders.
- .assets-revision: the pin moves from 64264d065cdb0b7755ee99dab356be5733d9ddef
  (the PR aiming-lab#75 nvidia bundle, 6706395-byte archive) to
  a9d3ea79513fa36a51498129cb9d031e4cc096cc, the HF dataset main revision after
  PR aiming-lab#84 merged the repacked nvidia.tar.gz (sha256 ee8c6ba9…, 9927312 bytes, 34
  file members, no directory entries). The pin comment block now states the
  measured registry count (29 registered sites, 29 of 29 covered).
…nce gates

The 20-task audit recorded one blocker and several low-severity verifier gaps. Each
change enforces what the judge_rubric already states; tasks.jsonl is untouched.

- T6 (blocker): answers.cuda_compare() now parses comparative claims and count
  ownership instead of a 100-character window around the winner, so the rubric's own
  sentence ("... 5,376 more CUDA cores than ... (21,760 versus 16,384)"), a
  counts-first answer and a trailing delta all pass, while a reversed subject, a wrong
  metric, a wrong delta, an equality claim and swapped absolute counts still fail.
- T1 (medium): measurements() also reads the site's own unit-first spec rows
  ("CUDA Cores: 10,752"); the allowed-unit set and value checks are unchanged, so
  "Tensor Cores: 10,752" and "CUDA Cores: 11,752" still fail.
- T3: the answer must name the model as well as the price (a bare "$299" fails).
- T5: evidence must cover both Jetson products (a comparison of both, or both detail
  pages), matching the rubric's "one-product-only evidence" failure clause.
- T9/T10: driver_evidence() requires the results page to pin the requested product
  series; branch and OS may stay unset. The target driver's own detail page stays
  sufficient.
- T18: the rubric expressly allows a listing or search result, so those routes are
  kept, but a search query must carry a word token that occurs in the article;
  a numeric-only query no longer counts as evidence.
- T19: the new newsletter row must also carry the site's GeForce topic.

tests/test_verifiers.py grows from 299 to 316 cases and pins every change (the three
previously pinned expectations that changed are renamed and documented in place).
verify/README.md documents the extended contract.
…int defects

Site-side fixes for the audit's low-severity findings; no task text changes.

- The RTX 50/40 series models grid no longer orders by price_usd: it is presented by
  specification tier (memory, then CUDA cores, then name), so the tile order does not
  encode the catalog price ranking while the cheapest model stays the last tile.
- The catalog sort control offers only neutral orders (Name, Price: Low to High).
  "Featured" and "Newest" promoted the category flagship into the first tile and
  "Price: High to Low" led with the most expensive card, which for Studio /
  Professional is also the maximum-memory answer.
- The home hero and the two series heroes use dedicated official artwork
  (static/images/heroes/), so no page renders the same image file twice.
- The footer newsletter input is no longer described as "Newsletter email" and the
  register newsletter checkbox no longer starts with "Email", so get_by_label("Email")
  resolves to exactly one field.
- The drivers filter submit button is named "Filter", so /drivers exposes one button
  named "Search".
- main.js hides each "Scroll horizontally ..." hint while its region does not overflow,
  and the drivers results table drops its extra top margin so the bordered scroller has
  symmetric whitespace.
- .stars cannot split its glyph from its value across a line break at 320 px.
- search scoring ignores one-character tokens and matches numeric tokens only as whole
  numbers, so a numeric-heavy query no longer returns the whole catalog.
…es with official renders

Image side of the audit findings. The files themselves live in the HF asset archive
(static/images is gitignored), so this commit carries the contract and the wording.

- geforce-rtx-5080.png and geforce-rtx-5090.png are replaced with NVIDIA's own per-SKU
  og renders (1200x630): the previous pair were two crops of one mirror-bundle strip
  (grey mean absolute difference 22.43 -> 35.47/255), and both are now labelled
  "Official NVIDIA product render" like the other re-sourced assets.
- Three dedicated hero images are added (static/images/heroes/) taken byte-exact from
  NVIDIA's CES 2025 press kit and the Ada/40-series landing page, so the home hero, the
  50-series hero and the 40-series hero each show a different file from every product
  card on their page.
- Captions now describe what each file actually shows: dgx-b200 and h200-tensor-core
  drop the stale "Local mirror illustration" wording, the RTX 4060 asset states that it
  is the RTX 4060 / 4060 Ti family artwork showing the Ti (NVIDIA publishes no
  RTX 4060-only render), and the RTX 5060 asset is described as the family desktop
  system it depicts.
- asset_inventory.json covers 36 files (33 product/news + 3 heroes) with official
  source URLs; test_ui_contract.py pins the new wording and asserts that the hero pages
  contain no duplicate image source.
The audit's aspect-ratio finding named the RTX 50 series grid, where the 5070/5080/5090
renders appeared as shallow letterboxed strips in the 4:3 card box while the Jetson and
B200 renders filled it. The 5080 and 5090 were re-sourced in the previous commit; this
one replaces the last three outliers with official NVIDIA media, so every product image
now sits between 1.77:1 and 1.90:1:

- geforce-rtx-5070.png (1200x361, 3.32:1) -> the CES 2025 press render (1200x675,
  1.78:1), stored as a PNG like every other product image.
- rtx-6000-ada.png (640x235, 2.72:1) -> the Ada Generation announcement press render
  (699x394, 1.77:1); the live RTX 6000 Ada page is retired (HTTP 404).
- shield-tv-pro.png (700x489, 1.43:1) -> NVIDIA's SHIELD TV Pro og image (1200x630,
  1.90:1).

Both newly official files join the caption map and asset_inventory.json, and
test_ui_contract.py pins the new wording (rtx-6000-ada moves out of the generic list).
…NVIDIA review section

The registry has 29 sites (NVIDIA last, index 28, port 40028), and
`scripts/check_site_registry.py` enforces that `websyn_start.sh`,
`control_server.py`, the Dockerfile EXPOSE line and every per-site tasks.jsonl
`web` URL agree. Several documents still described 28 sites and the port range
40000-40027 / 41000-41027, so they now state the measured inventory:

- AGENTS.md: "28 Flask mirror websites" -> 29; port ranges 40000-40028 and
  41000-41028; `seq 41000 41028` in the pre-PR check.
- CLAUDE.md, CONTRIBUTING.md: the working-container and alt-port ranges.
- .claude/skills/{clone-website,evolve-env,harden-env,review-env,seed-database}:
  the alt-port run commands, the "all 28 sites return 200" checks and the
  "the image now runs 28 sites (40000-40027)" note.
- README.md: the NVIDIA candidate section (29 sites, registry index 28, container
  port 40028, host port 48028) and the contribute paragraph (25 -> 29 mirrors).

The README's NVIDIA asset sections are rewritten against this round's
measurements: the pin is `a9d3ea79513fa36a51498129cb9d031e4cc096cc` (the merged
PR aiming-lab#84 squash commit on the dataset main, on top of the merged PR aiming-lab#75), the
pinned `nvidia.tar.gz` is 9,927,312 bytes / sha256 ee8c6ba9... with 34 file
members and no directory members, and `./scripts/fetch_assets.sh` at that pin
extracts 29 sites. The stale "PR aiming-lab#75 is a draft" and "the validator rejects
nvidia.tar.gz" blocker sections and the old re-pack recipe were removed, and the
prepared follow-up archive (37 members, 16,340,955 bytes, sha256 617a3e37...) is
recorded with the note that these numbers are refreshed together with
`.assets-revision` when that PR merges. The validation record now lists the
current test counts (316 verifier cases, 44/73 unit-discovery tests, 29 driver,
15 UI-contract, 23 T7, 11 CLI-contract, check_assets.sh exit 0).

Placeholders and examples (`EXPOSE 8101 40000-N`, `40000-400NN`, `40000 + index`,
`41000NN`) and the historical `review-reports/**` audits are unchanged.
…official images

HF dataset PR aiming-lab#85 ("Upload nvidia.tar.gz with distinct official product and hero
images (aiming-lab#107 follow-up)") is merged, so the pin moves from
a9d3ea79513fa36a51498129cb9d031e4cc096cc to b7e605c0ec5fc47de85b09e7427162cc50e38980,
the merged revision's squash commit on the dataset main.

Refreshed facts (all re-measured this round):
- nvidia.tar.gz at the new pin: 37 file members, 0 directory members,
  16,340,955 bytes, sha256
  617a3e3740ba6706bcab786c8a5c3f9a22ecbb39eff5728ad2c12e4992cb098b,
  and scripts/validate_asset_archive.py prints
  "[fetch] validated 37 managed members for nvidia" (exit 0);
- it carries 36 images with 36 distinct sha256 (33 product/news images plus three
  dedicated hero images under static/images/heroes/), so no byte-identical
  duplicate remains and no page reuses one file twice;
- instance_seed/nvidia.db is byte-identical to the previous pin
  (2143c954def96cc921760ab2bea79fe119de3d73212d1b01daf6c61792c2b38d).

.assets-revision states the new revision and the new archive identity, keeping the
superseded PR aiming-lab#84 archive as history. README.md's asset-delivery section now names
the merged PR aiming-lab#85 revision and lists the new archive first with the previous pin's
archive as the superseded row.
@Raibows

Raibows commented Sep 14, 2026

Copy link
Copy Markdown
Contributor

Thanks for your contribution! @TabsPhasers @KaKituken @DEM1TASSE

@Raibows
Raibows marked this pull request as ready for review September 14, 2026 03:39
@Raibows
Raibows merged commit 7ace1c7 into aiming-lab:main Sep 14, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants