Review B&H Photo: complete assets, repair grading, and generate reset seed - #123
Merged
Merged
Conversation
Merge upstream/main (3600493) into the reviewer branch carrying @Lxr-max's original B&H Photo mirror commit 4618b7a unchanged. The original registered bh_photo as the 16th site on port 40015 against a 15-site base. Current main has 24 sites on 40000-40023, so this merge takes main's version of every shared file and re-registers bh_photo as site 25 on port 40024 in the three required places (websyn_start.sh, control_server.py, Dockerfile EXPOSE), updates the README/AGENTS port ranges and site list, and repoints the 20 task URLs to :40024.
Malformed quantity/store_id values raised ValueError inside the cart and reserve handlers, so a benchmark agent typing text into a number field got an HTTP 500 instead of a visible error. parse_positive_int() now rejects non-numeric input and each handler flashes a validation message and redirects without writing to the database. Every page also requested /favicon.ico and got a 404; the site now ships sites/bh_photo/static/icons/favicon.svg, links it from base.html, and serves it at /favicon.ico. Adds sites/bh_photo/tests/test_functional_contract.py covering both rejected and accepted input for all three handlers plus the favicon (7 tests, red before this change).
At 320px the auto-fit tracks kept their pixel minimum (bundle cards held 320px inside a 300px shell), so the page scrolled horizontally. Wrapping each minimum in min(<px>, 100%) keeps the desktop track and lets the card collapse to the container on small screens. Verified with Chromium over 10 page templates x 1440/768/390/320: no page scrolls horizontally and no element escapes the viewport outside the intentionally scrollable nav row.
The mirror used a dark-green header, a yellow logo tile and a yellow promo banner. None of that matches bhphotovideo.com. Values here are measured on the Internet Archive capture 20260909004043 of the upstream home page, because the live site answers automation with a bot challenge: page background #ffffff department nav #3f9a59, 32px tall body text #000 / #666 nav links white 14px, Used #e0f2aa logo fills #bd2c26 #ffc745 at 98x60 content container 1350px footer #f2f2f2 link accent #007ab8 The header becomes a white band with a red-bordered B&H lockup and a bordered white search field, the department strip becomes the green bar carrying Used and Specials, and the footer becomes light grey. The demo disclosure stays but moves into the slim utility strip upstream has in that position. The logo is drawn from the upstream fills and proportions rather than copied from the upstream SVG path. Checked over 12 page templates x 1440/768/390/320: all 200, no page scrolls horizontally, no broken images.
The contribution shipped an entirely invented catalog: remixed product
names ("Sony Aurora A7X", "Canon Orbit R6 Mark II"), generated specs and
prices, and 167 SVG line drawings. No task fact could be checked against
bhphotovideo.com, and the project's own reviewer checklist asks for real
product imagery rather than placeholders.
Catalog facts now come from dated Internet Archive captures of
bhphotovideo.com, because the live site answers automated requests with a
bot challenge. sources/harvest.py reads the archive index, parses the
JSON-LD product block and the published specification table, and pulls the
product photography the captured page showed. sources/build_manifest.py
selects the shelf, derives the listing facets from published spec rows
only, and writes source_catalog.json with a capture URL and timestamp on
every product. seed_data.py builds the database from that manifest.
128 products, 50+ brands, real B&H SKUs and MPNs, real prices
real availability including backordered and discontinued
4300+ specification rows taken from the published tables
archived product photography where the archive holds it
Accounts, reviews, questions, bundles, carts, orders, reservations and the
used/open-box shelf stay generated, and both the manifest and the module
docstring say so. A field the source does not state is stored empty rather
than invented; products whose photo is not archived show a placeholder
that says so.
Two reproducibility defects are fixed with it. Benchmark accounts hashed
with bcrypt's random salt, so the shipped database could never be
regenerated from source; they now use PBKDF2 over a derived fixed salt
while live registrations keep bcrypt. Index DDL was emitted in set order,
so two builds of identical data produced different bytes; indexes are now
recreated by name and the file is vacuumed. Six consecutive rebuilds are
byte-identical, and a boot on a populated database still leaves the file
untouched, so the reset invariant holds.
Product cards also showed a review count that did not match the reviews on
the page (188 claimed, 2 present); the badge is now the number of rows
actually created.
The store list drops from eight invented locations to the single B&H
counter the upstream footer states (420 9th Ave, New York).
The home page loses the marketing hero and the "146 synthetic products"
card in favour of the department row with product photography that
upstream shows in that position.
Every page title read "· B&H Demo Mirror", the logo lockup said DEMO
MIRROR, and page intros described the site as a benchmark environment
("designed for cart and checkout benchmark tasks", "synthetic mirror
content"). Generated rows carried the same commentary into the data: each
review body began "Benchmark review text for ...", which makes a task that
asks an agent to read a review meaningless.
Titles now follow the upstream pattern ("<page> | B&H", and the home title
bhphotovideo.com serves), page copy reads like the store it mirrors, and
generated reviews carry ordinary short review prose.
What this site is stays disclosed, in the one place upstream keeps a
utility strip: catalog facts from archived B&H pages, accounts and reviews
and orders generated, no live payment or fulfilment. The same statement is
in the footer, the source manifest and the seed module docstring.
Pickup copy now names the single real B&H counter instead of claiming
several locations.
The gallery column stretched to the height of the buy box, so a single product photo sat above a tall empty panel. The main image now uses the 655x500 box upstream renders it in, the column aligns to the top, and the thumbnail strip only appears when a product actually has more than one archived image. Product price adopts the upstream 28px/700/black. Not verified: the upstream Add to Cart button colour. The archived buy box is script-driven and does not hydrate in the capture, so the call to action keeps the sourced brand green rather than a guessed colour.
Kits were three consecutive rows from one department, which produced a drone kit made of three variants of the same camera and, after a first fix, a laptop bundled with a softbox and a lens. Each kit is now built from a named theme that lists the shelves it draws from (a body with storage and support, a microphone with monitoring headphones, a laptop with a display), and no kit may take two products that share a name prefix. Adds sites/bh_photo/verify/verify_lib.py, the deterministic verifier support library for this site: run loading, same-origin navigation checks, click-transition and form-submission checks, negation-aware text matching, snapshot-bound database comparison, and helpers for this schema's cart, wishlist, comparison, reservation, order and account state. Also freezes sites/bh_photo/source_catalog.json at 128 products so task ground truth stays stable, and adds sources/refresh_images.py, which attaches newly archived photography to products already in the manifest without re-running selection.
Replaces the contributor's 20 tasks, every one of which targeted an invented product that no longer exists, and two of which depended on invented store locations. The new set is anchored on this catalogue: four specification lookups whose answers appear only on a product page, five browse, filter, compare and search tasks, three detail-content tasks reading kit contents, a review and a Q&A answer, and seven authenticated tasks covering wishlist, cart quantity, store reservation, checkout, order history, account settings and cart removal. Every answer is a page-specific fact rather than something a model can recall: the exact figure in a specification row, the cheapest item on a filtered listing, an order number, a saved account setting. Listing cards show name, price, condition and rating only, so no lookup answer is readable without opening the product page. Each row carries verifier_path and a judge_rubric of fact checkpoints. There is no answer key in tasks.jsonl; verifiers read ground truth from the shipped seed database at verify time. Adds the first four verifiers. Seed state fixes that came out of writing them: question and answer rows were drawn independently so a question about store pickup could be answered about warranty; some captured pages list a specification row twice; and cart, wishlist and order picks could choose two variants of the same product, which also left an account in a three-product department with a single cart line.
One verifier per task under the site's own verify/ directory, each reading its ground truth from the shipped seed at verify time so no answer key lives in the agent-facing task file. Every verifier checks the run is same-origin, is the right task, ends with a non-empty answer and has real screenshots, then checks the task's own requirements: that the page carrying the answer was actually opened, that the answer states the specific figure rather than the product name, and that the obvious near-miss was not reported instead. The distractor check is explicit per task: the actual megapixel count rather than the effective one, the write speed rather than the read speed, a CFexpress Type A card for a camera that takes Type B, a delivered order number instead of the processing one. Stateful tasks compare the seed against the after-state: the saved row must exist, sibling rows must be untouched, and no table outside the task's allowed set may change. Read tasks allow only search_logs to change, since every search request appends to it. A no-op run (home page opened, nothing done, empty answer) fails all 20 with no false positives and no crashes, and all 20 read their ground truth from the seed. Also attaches 17 more archived product photos to the frozen manifest; 77 of 128 products now carry upstream photography, and the rest render the placeholder that states the image is not archived. Product identity, prices, specifications and account state are byte-identical across that change.
Measured against the archived B&H listing page, the mirror's routes were shallower than the real ones in ways that change how hard a task is. Listing pages now carry what upstream carries: a breadcrumb, an item count, and left-rail facets grouped into collapsible sections named the way upstream names them (Search Within Results, Brand, Price, Customer Rating, Savings & Stock, Product Features). All but the search box start collapsed, so reaching the mount facet takes the same expand step it takes upstream instead of being handed to the agent in a flat form. The search box is wired to a within-results filter. Listing rows now carry the B&H SKU and the manufacturer part number, the four-row Key Features summary with a link into the full specifications, the saving in dollars, and Add to Cart plus Add to Wish List. products gains an mpn column fed from the source manifest; 102 of 128 products state one upstream. Two consequences for the task contract, both resolved by changing the tasks rather than by hiding what upstream shows: Tasks 1 and 3 asked for a figure that the new Key Features summary puts on the listing card. They are re-anchored onto rows the card does not show: the EOS 90D screen resolution and battery model, and the ThinkPad native resolution and weight. Tasks 13 and 14 required the product page, but upstream lets a shopper add to cart or wishlist straight from a listing row, so their verifiers now accept either route. A no-op run still fails all 20 verifiers with no crashes, and a reset still restores the database byte for byte after a search has written to search_logs.
The archived B&H listing splits 189 items across seven pages addressed as /pn/2 through /pn/7, so a listing runs 30 rows to a page. Every listing route in the mirror now does the same: category pages take the /pn/<n> path upstream uses, search, brand, deals and used take a page parameter, and each page keeps the active filters. The listing head reports the window the way upstream does (Showing 1-30 of 57) and the footer carries numbered page links with previous and next. Search copy drops the last reference to a benchmark catalogue.
Re-anchoring task 1 away from the listing card assumed a screen-resolution row that the body-only Canon EOS 90D page does not carry. Its monitor group publishes a dot count instead, so the task and its verifier now ask for the monitor resolution in dots together with the battery model, both outside the four key features a listing card shows. Caught by a new anchor check kept with the case rather than in the repo: it re-reads every task's ground truth from a candidate seed and reports which ones no longer hold, which matters because growing the catalogue can quietly destroy the uniqueness a task depends on.
Upstream shows a sort control and a list/grid switch above the results and renders the list layout by default. The mirror now does the same: sort applies on change and keeps the active filters and page, the view switch is a link pair, and the list layout puts the image beside the details in one wide row per product.
A shelf of sixteen items does not exercise the same behaviour as a real one: filters return everything, pagination never appears, and a task can be answered by reading one screen. Two further archive sweeps took the cache from 395 to 930 captured product pages, and the shelves now hold 296 products from 122 brands priced from 8 to 19,975 dollars. Photography 136 items, five pages Pro Audio 73, three pages Lenses 52, two pages Microphones 60, two pages Tripods and support 41, two pages Computers 33, two pages Two task anchors did not survive the change, both caught by the anchor check before anything was recorded. Task 19 named a product that Alice no longer has in her cart, so it now asks for the dearer of her two cart lines and its verifier derives that from the seed; a future catalogue change moves the answer instead of breaking the task. Task 4's cheapest open-box item moved from a tripod mount to a roll of flight-zone tape, which needed no change because that verifier already reads the answer from the seed. A no-op run still fails all 20 verifiers with no crashes, and all 20 still read their ground truth.
Upstream moved from 24 sites to 26 while this branch was in progress: the FedEx mirror landed through review PR aiming-lab#82 and the WebMD Doctor mirror through aiming-lab#95, which took ports 40024 and 40025. bh_photo was registered at 40024 and now collides. This merge takes main's version of every shared file and re-registers bh_photo as site 27 on port 40026 across the three places that must stay in sync, bumps the exposed range and the site counts, adds the mirror to the README list, and repoints the 20 task URLs. Final slot assignment is the maintainer's call; several review branches are queued for the same range.
Fills image_path for products the later archive passes recovered: 158 of 296 products now carry the photograph their captured page showed, and the rest keep the placeholder that states the image is not archived. Selection is untouched, so no task answer moves.
The archived B&H Photography page opens with grouped subcategory tiles (Digital Cameras, Lenses, Film Cameras and the rest) before any product rows. The mirror's department pages jumped straight to the grid, so the only way to reach a shelf was the top nav, and a task that browses into Lenses from Photography had nowhere to click. Department pages now carry a subcategory row with a representative archived photograph and the item count for each shelf.
Two findings from the recorded runs, in different layers. The listing carried two sort controls after the toolbar was added, one in the toolbar and one in the filter rail. Upstream has one. The duplicate made the page ambiguous and a run that changed the rail control saw no effect, so the rail control is gone. Separately, eight verifiers required the catalogue product name verbatim, which failed a run that answered "Lenovo 16-inch ThinkPad T1g Gen 8" for a product the catalogue calls 'Lenovo 16" ThinkPad T1g Gen 8 Multi-Touch Laptop'. A person reading that answer would call it correct, and the grading contract is supposed to accept a valid rephrasing. They now match on the distinctive brand and model tokens through names_product(). The checks that reject a WRONG product stay strict, and a no-op run still fails all 20 verifiers with no crashes.
A checkout placed on the site produced BH-<date>-<time>-<user> while every order already in the store reads BH-<date>-<user><sequence>, so one account's history showed two different formats. Checkout now issues the same shape, with a uniqueness guard. The checkout verifier only inspected each step's starting url, so it missed the order confirmation, which is reached after the last recorded action. It now inspects the whole url trail including the final page.
The adversarial matrix passed five answers it should have failed, and failed one legitimate run it should have passed. names_product counted single letters and bare digits, so a rival lens cleared the threshold: TTArtisan AF 40mm f/2 Lens (Sony E) and Rokinon 12mm f/2.0 NCS CS Lens (Sony E) share f, 2 and e. Only tokens of three or more non-digit characters identify a product now. Number checks accepted the right digits anywhere in the sentence. An answer giving 5 stars passed a check for 3 by ending "after 3 months"; one giving 40.2 megapixels passed a check for 26.1 by quoting 26.1 in a parenthetical. number_labelled requires the figure to sit within two words of what it measures. states_count does the same for counts, and takes plural nouns when the count is above one, since "f/2 Lens" would otherwise read as a claim that two matched. Task 12 demanded a URL under /search, which is only the header search box. Searching within a listing is the mirror's other keyword search, so browsing to the tripods and filtering there was failed for taking a different route to the same requirement. Either control counts now. Also names the site this verifier library was ported from correctly. Canonical runs 20/20, alternative routes 6/6, adversarial matrix 170/170. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Healthline took index 26 and Kaggle index 27 while this review was in progress, so the slot this branch held (index 27, port 40026) is now Kaggle's neighbour and the port is Healthline's. B&H Photo moves to the end of the list: index 28, port 40028, EXPOSE widened to 40000-40028, and the twenty task rows point at localhost:40028. scripts/check_site_registry.py, which main added in the meantime, reports 29 sites consistent across websyn_start.sh, control_server.py, the Dockerfile EXPOSE range and the per-site task files. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Search took between 24 and 84 seconds. log_search committed on db.session, and a commit expires every instance in the identity map; the search page had just loaded all 296 products eagerly, so the next attribute read reloaded them one row at a time. Profiling put 17 of 20 seconds in filters_for, all of it in _load_expired. The log is append-only and nothing in the request reads it back, so it goes out on its own connection and leaves the session alone. Search now answers in 1.1 to 5.0 seconds. The home page benefited too: eighty-second searches had been starving the server, and the control plane's readiness probe, which allows one second, went from failing 8% of the time to 0 of 25, so /health stops flapping to 503. Verified no rendered output moved: the 96 distinct pages the recorded runs visit were hashed before and after. All 90 that answered before the fix are byte-identical; the other 6 are the search pages that had been timing out at 30 seconds and now serve. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
check_asset_inventory.py rejected the manifest outright: it was a flat map keyed by image id, and the checker wants schema_version 1 with an assets list of path/bytes/sha256/source_url rows. Every already-merged site fails this check locally too, but only because the images live on the dataset and are not downloaded - a different thing from the wrong shape. Rewritten from the files on disk: 337 assets, 88,413,710 bytes, each hashed and carrying both the original B&H URL and the Wayback capture the bytes actually came from. The checker now verifies all 337. Also adds the site .gitignore the other sites carry. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An isolated blind pass over the twenty frozen runs returned 18 PASS and 2 FAIL. Both failures were real and both were mine. Task 6's verifier asked whether the run opened a matching memory card, and separately whether the answer named a catalogue card. Neither question tied to the other, so a run that opened the 512GB + 128GB kit and reported the 512GB + 64GB kit satisfied both. The card named must now be the card opened, and the price binds to that same card. That exposed more: those two cards are word-for-word identical apart from the bundled capacity, and names_product decides by counting shared tokens, so either name matched either card. Measurements now have to match rather than be outvoted by the dozen tokens the two share. Task 19 asked for the cart total. The page shows Subtotal $560.16, Tax $49.71, Total $609.87; the verifier demanded the subtotal. The task now names the field, and the verifier grades the Total and rejects the subtotal given in its place. The run-side fixes for both tasks, and full-page screenshot capture, live in the reviewer's private harness rather than this repository. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Complete source-backed images and product specifications, correct shopping flows and task contracts, harden deterministic verifiers, and generate the reset seed during the Docker build.
Keep Berkeley at index 29, append B&H Photo at index 30, preserve both build-generated seeds, and synchronize the registry, task URLs, and documentation.
Use immutable HF PR aiming-lab#92 commit f09e586e so fresh checkouts can fetch all 31 site archives. Replace this interim pin with the dataset merge commit before merging the code PR.
HF dataset PR aiming-lab#92 is merged. Its complete asset tree matches the tested candidate; the downloaded B&H archive and all 31 registered asset inventories validate.
QianhuiWu
marked this pull request as ready for review
September 16, 2026 05:53
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
B&H Photo in #122 needs integration with the current registry, complete product images, and a reproducible reset database. This PR includes #122 and the review fixes, preserves Berkeley at port 40029, and appends B&H Photo at port 40030 (31 sites total).
instance_seed/bh_photo.dbduring the Docker build from tracked catalog data, with byte-identical reset behavior.HF asset PR #92 is merged.
.assets-revisionpins dataset commitfa1e8a5b9e8e5d0e42764cd658825f4dea088d8f. Its entire asset tree is identical to the tested candidatef09e586eec8bf1bca0bc0881e08b77f3c2a5508e; all 32 pre-existing archives remain unchanged. The B&H bundle was downloaded again from the merged revision and validated.Validated on the integrated branch:
./scripts/build.sh webharbor:pr122-integrated; all 31 sites alive/ready and all homepages HTTP 200.1e40b398ff930e9e37d29dbe220ce2dc; all 24 seed tables match the reviewed data.agent_demo/eval_judge.py --verifier Trueentry point; 137 browser actions and 198 desktop/mobile page audits across 57 unique URLs, with no reported missing images, overflow, clipped controls, or JS/HTTP errors.These task runs are deterministic browser regression replays. The secondary LLM judge was not run. Please use a merge commit to preserve #122's contribution ancestry.