Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
34 commits
Select commit Hold shift + click to select a range
4618b7a
Add B&H Photo WebHarbor mirror
May 28, 2026
3d188fd
chore(bh_photo): integrate PR #38 onto current main as site 25
jackjin1997 Sep 12, 2026
11dca9f
fix(bh_photo): validate form numbers and serve a favicon
jackjin1997 Sep 12, 2026
cc8b0df
fix(bh_photo): stop card grids overflowing narrow viewports
jackjin1997 Sep 12, 2026
f0acfb7
style(bh_photo): match the upstream B&H header, nav and footer
jackjin1997 Sep 12, 2026
aa7352e
feat(bh_photo): rebuild the catalog from sourced B&H data
jackjin1997 Sep 12, 2026
172e05e
style(bh_photo): use retail page copy instead of mirror commentary
jackjin1997 Sep 12, 2026
7bff3ef
style(bh_photo): size the product gallery from the upstream layout
jackjin1997 Sep 12, 2026
946d560
fix(bh_photo): compose kits from coherent themes
jackjin1997 Sep 12, 2026
598d310
feat(bh_photo): add the reviewed task contract
jackjin1997 Sep 12, 2026
f663cb1
feat(bh_photo): add deterministic verifiers for all 20 tasks
jackjin1997 Sep 12, 2026
087b5ee
feat(bh_photo): match the upstream listing depth on task routes
jackjin1997 Sep 13, 2026
c2b6da4
feat(bh_photo): paginate listings at the upstream page size
jackjin1997 Sep 13, 2026
359c010
fix(bh_photo): anchor task 1 on a row the EOS 90D actually publishes
jackjin1997 Sep 13, 2026
cc85552
feat(bh_photo): add the listing toolbar upstream puts above its rows
jackjin1997 Sep 13, 2026
dd1295b
feat(bh_photo): deepen the catalogue to upstream listing scale
jackjin1997 Sep 13, 2026
e0593ae
chore(bh_photo): re-integrate onto current main as site 27
jackjin1997 Sep 13, 2026
16ad17c
chore(bh_photo): attach the rest of the archived photography
jackjin1997 Sep 13, 2026
2b707e9
feat(bh_photo): front department pages with their subcategories
jackjin1997 Sep 13, 2026
b937863
fix(bh_photo): one sort control, and accept a naturally worded answer
jackjin1997 Sep 13, 2026
37d61cf
fix(bh_photo): one order-number format, and check the whole url trail
jackjin1997 Sep 13, 2026
3b1580e
fix(bh_photo): grade the requirement, not the digits or the click path
jackjin1997 Sep 13, 2026
bdfe52c
merge: onto main 145b200; re-slot bh_photo to index 28 (port 40028)
jackjin1997 Sep 13, 2026
6e0dd87
fix(bh_photo): stop the search log from expiring the loaded catalogue
jackjin1997 Sep 13, 2026
519ec2a
chore(bh_photo): put the asset manifest on the repository schema
jackjin1997 Sep 13, 2026
38048ac
fix(bh_photo): two defects an independent review caught
jackjin1997 Sep 14, 2026
382ced6
fix(bh_photo): restore benchmark task paths
jackjin1997 Sep 14, 2026
893a3dd
merge: integrate main and re-slot bh_photo to 40029
jackjin1997 Sep 14, 2026
c4482a3
fix(bh_photo): align photography landing with source
jackjin1997 Sep 15, 2026
32c336c
fix(bh_photo): repair reviewed pages, fixtures, and task grading
QianhuiWu Sep 16, 2026
9ecc96e
merge: integrate main and assign B&H Photo port 40030
QianhuiWu Sep 16, 2026
8003cb7
build(bh_photo): pin the reviewed HF asset candidate
QianhuiWu Sep 16, 2026
cabf211
style(bh_photo): remove trailing blank line in health helper
QianhuiWu Sep 16, 2026
f6aaaaa
build(bh_photo): pin assets to merged HF revision fa1e8a5
QianhuiWu Sep 16, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
88 changes: 17 additions & 71 deletions .assets-revision
Original file line number Diff line number Diff line change
@@ -1,72 +1,18 @@
# Hugging Face dataset holding webharbor's heavy static assets
# (instance_seed/*.db, static/images/, static/external_cache/).
#
# fetch_assets.sh uses the `hf download` CLI. The pin below
# is a git revision (branch name like `main`, a tag, or a specific commit
# sha). Override at runtime with the ASSETS_REVISION env var.
#
# The revision is a commit sha, never a moving branch name.
#
# Current pin = c32018ca3b3d67e7b858b1b85fb101aea5090cd7, the head commit of HF
# dataset `main` (verified 2026-09-15): the squash-merge commit of HF dataset PR
# #91 "berkeley: synthetic imagery bundle (164 files)" (merged 2026-09-15T04:00:54Z),
# which added berkeley.tar.gz (sha256
# ab9d2716ae8d06540a181b5e60c37f613d87b103864b467511da546b1b173789,
# 6951483 bytes, 171 managed members) — the site's first archive, so that
# fetch_assets.sh can install sites/berkeley/static/images/.
#
# The berkeley bundle first shipped through the interim pin `refs/pr/91` while that
# dataset PR was open; the pin was moved to this merged commit once the PR merged,
# because a `refs/pr/<n>` ref moves whenever its branch is updated. Checked at both
# pin changes (2026-09-14 and 2026-09-15): for all 29 sites registered before
# berkeley, the archive size and LFS oid on this commit are identical to the
# b7e605c0 pin below, and all 30 site archives on this commit are byte-identical to
# the ones on `refs/pr/91` (verified by LFS oid and by comparing the downloaded
# berkeley.tar.gz byte for byte), so the pin change alters no site's assets. This
# revision carries an archive for all 30 registered sites.
#
# Previous pin = b7e605c0ec5fc47de85b09e7427162cc50e38980, the commit HF dataset
# `main` pointed at before this pin (verified 2026-09-14): the squash-merge commit
# of HF dataset PR #85 ("Upload nvidia.tar.gz with distinct official product and
# hero images (#107 follow-up)"), which sits on top of the merged PR #84 ("Upload
# repacked nvidia.tar.gz with de-duplicated official product images (#107)"), PR
# #75 ("Add reviewed NVIDIA asset bundle") and PR #74 ("Upload kaggle.tar.gz with
# huggingface_hub").
#
# Its nvidia.tar.gz is the repacked bundle (sha256
# 617a3e3740ba6706bcab786c8a5c3f9a22ecbb39eff5728ad2c12e4992cb098b, 16340955 bytes,
# 37 file members and no directory entries), which carries the official NVIDIA
# renders for 33 product/news images plus three dedicated hero images under
# static/images/heroes/: 36 images with 36 distinct sha256, so
# sites/nvidia/static/images/ has no byte-identical duplicate and no page reuses
# one file twice. It supersedes the 9927312-byte PR #84 archive (sha256
# ee8c6ba966e7a8f7fb5ad2d7ff0134ab98e7b80d6cc77f3328217405b8b34e2f) and the
# original 6706395-byte bundle (blob 0f1d8068af602748d0dec19ef04b52c7559fef28),
# which carried duplicates.
#
# Pin before that = 64264d065cdb0b7755ee99dab356be5733d9ddef, the commit HF dataset
# `main` pointed at (verified 2026-09-13). It carries the merged HF dataset PR
# #74 ("Upload kaggle.tar.gz with huggingface_hub") and the later PR #75 ("Add
# reviewed NVIDIA asset bundle").
#
# Why these pins and not the other candidate (refs/pr/37, resolved
# 600a3e1de158ae56dc82a5fab56c2ca25acb1e27):
# * the pinned revision carries an archive for every registered site (30 of 30 for
# c32018ca and for refs/pr/91: the 29 main-site archives plus berkeley.tar.gz;
# 29 of 29 for b7e605c0), and
# the kaggle seed DB it ships was found identical to the seed the site's own
# code generates (13/13 tables, row by row) in the PR #106 audit;
# * refs/pr/37 ships an older kaggle seed (models table has 10 rows instead of
# 11: densenet121-chestxray is absent) and its archives cover only 17 of the
# 30 registered sites, so pinning there would break fetch_assets.sh.
# * the pinned revision's kaggle.tar.gz (sha256
# dd6f1ab34f99989b300e3366225a8f8b0932d49d033b48f141938d16f1c0607d,
# 11278106 bytes) passes the repository's own
# scripts/validate_asset_archive.py:
# [fetch] validated 73 managed members for kaggle
# That 73-member pack has no directory entries, so the earlier
# `ValueError: unexpected managed path: 'kaggle/static'` rejection (from the
# superseded 11278264-byte pack, sha256 c533c283...) no longer applies and
# `scripts/fetch_assets.sh` installs sites/kaggle's assets from this pin.
# Hugging Face dataset holding WebHarbor's heavy static assets.
# Pin an immutable commit SHA; never a moving branch or refs/pr/<n>.
#
# Merged dataset-main commit from HF asset PR #92.
# https://huggingface.co/datasets/ChilleD/WebHarbor/discussions/92
# Its complete asset tree matches the tested candidate commit
# f09e586eec8bf1bca0bc0881e08b77f3c2a5508e.
#
# Adds bh_photo.tar.gz: 79,658,793 bytes, 511 managed archive members.
# SHA-256: 867363d5484eb114d647e236991017992d5ac91ae3415996ad43bf654d99bd9a
# All 32 pre-existing archives are unchanged from the previous main pin
# c32018ca3b3d67e7b858b1b85fb101aea5090cd7 (Berkeley, HF PR #91).
# This revision has archives for all 31 registered sites. The two additional
# archives (bandcamp and drugs_com) are ignored by fetch_assets.sh.
# B&H's bundle contains images and external cache; Docker generates its seed
# from tracked catalog data (sites/bh_photo/.build-generated-seed).
repo: ChilleD/WebHarbor
revision: c32018ca3b3d67e7b858b1b85fb101aea5090cd7
revision: fa1e8a5b9e8e5d0e42764cd658825f4dea088d8f
12 changes: 6 additions & 6 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ A coding agent (Claude Code, Cursor, Aider, Codex, ...) is reading this. Read on

## What it is

30 Flask mirror websites (Amazon, GitHub, BBC News, ...) packaged into one Docker image, plus a control plane on `:8101` for resetting per-site state. Used as a deterministic offline environment for web-agent benchmarks. ~3 GB image.
31 Flask mirror websites (Amazon, GitHub, BBC News, ...) packaged into one Docker image, plus a control plane on `:8101` for resetting per-site state. Used as a deterministic offline environment for web-agent benchmarks. ~3 GB image.

Two repos:
- **code** (this one) — Flask apps, control plane, scripts.
Expand Down Expand Up @@ -48,17 +48,17 @@ Inside the image, sites live at `/opt/WebSyn/<site>/`. The path predates the ren
# fresh clone
./scripts/fetch_assets.sh # pulls assets from HF
./scripts/build.sh # docker build -t webharbor:dev .
docker run -d -p 8101:8101 -p 40000-40029:40000-40029 webharbor:dev
docker run -d -p 8101:8101 -p 40000-40030:40000-40030 webharbor:dev
```

Or use the published image directly:

```bash
docker run -d -p 8101:8101 -p 40000-40029:40000-40029 \
docker run -d -p 8101:8101 -p 40000-40030:40000-40030 \
battalion7244/webharbor:latest
```

Sites are on `40000`-`40029` in the order declared by `SITES=( ... )` in `websyn_start.sh`. Control plane:
Sites are on `40000`-`40030` in the order declared by `SITES=( ... )` in `websyn_start.sh`. Control plane:

| Method | Path | Purpose |
|--------|---------------------|-------------------------------------------|
Expand Down Expand Up @@ -136,13 +136,13 @@ python3 -m py_compile sites/<site>/app.py

# 3. run on alt ports (don't collide with anything you already have running)
docker run -d --rm --name wh-test \
-p 8201:8101 -p 41000-41029:40000-40029 webharbor:dev
-p 8201:8101 -p 41000-41030:40000-40030 webharbor:dev

# 4. control plane healthy, all sites alive
curl -s http://localhost:8201/health | python3 -m json.tool | head

# 5. every site renders 200
for p in $(seq 41000 41029); do
for p in $(seq 41000 41030); do
curl -so /dev/null -w "$p:%{http_code}\n" http://localhost:$p/
done

Expand Down
2 changes: 1 addition & 1 deletion CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,7 @@ git clone https://github.com/<you>/webharbor && cd webharbor
./scripts/fetch_assets.sh # pull current assets
./scripts/new_site.py mywebsite # OR edit an existing site
./scripts/build.sh && docker run -d --rm \
-p 8101:8101 -p 40000-40029:40000-40029 webharbor:dev
-p 8101:8101 -p 40000-40030:40000-40030 webharbor:dev
# iterate locally...

./scripts/extract_assets.sh ../webharbor-static-pr/ # split assets out
Expand Down
14 changes: 12 additions & 2 deletions Dockerfile
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
# WebHarbor — slim, self-contained image.
# 30 Flask mirror sites + control plane on :8101.
# 31 Flask mirror sites + control plane on :8101.

FROM python:3.12-slim-bookworm

Expand Down Expand Up @@ -103,6 +103,16 @@ os.makedirs('instance_seed', exist_ok=True); \
shutil.copy2('instance/rotten_tomatoes.db', 'instance_seed/rotten_tomatoes.db'); \
print('Rotten Tomatoes seed DB generated at build time.')" && rm -rf /opt/WebSyn/rotten_tomatoes/instance

EXPOSE 8101 40000-40029
# B&H's asset bundle contains images; generate the reset seed from its tracked
# catalog in the image so a fresh checkout needs no locally prepared database.
RUN python3 /opt/check_asset_inventory.py /opt/WebSyn/bh_photo
RUN cd /opt/WebSyn/bh_photo && rm -rf instance instance_seed && PYTHONHASHSEED=0 python3 -c "\
import app; \
import os, shutil; \
os.makedirs('instance_seed', exist_ok=True); \
shutil.copy2('instance/bh_photo.db', 'instance_seed/bh_photo.db'); \
print('B&H Photo seed DB generated at build time.')" && rm -rf instance

EXPOSE 8101 40000-40030

CMD ["/opt/websyn_start.sh"]
63 changes: 29 additions & 34 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,17 +36,17 @@ WebHarbor takes a different approach. We leverage coding agent (e.g., Claude Cod
- **Deep features unlocked** — carts, checkouts, accounts, all fully testable
- **Evolving** — harder tasks drive richer mirrors; the environment grows with agents
- **RL-ready** — sub-second database resets between rollouts
- **Community-driven** — 30 sites today, scaling to 100+ together
- **Community-driven** — 31 sites today, scaling to 100+ together

## 🚀 Quickstart

One command to run all web environments:

```bash
docker run -p 8101:8101 -p 40000-40029:40000-40029 battalion7244/webharbor:latest
docker run -p 8101:8101 -p 40000-40030:40000-40030 battalion7244/webharbor:latest
```

Then point your agent at `http://localhost:40000` through `http://localhost:40029` to explore 30 local mirrors of WebVoyager sites: `Allrecipes, Amazon, Apple, ArXiv, BBC News, Booking, GitHub, Google Flights, Google Maps, Google Search, Hugging Face, Wolfram Alpha, Cambridge Dictionary, Coursera, ESPN, Merriam-Webster, IKEA, Phys.org, Target, TED, Ohio State University, Rotten Tomatoes, Compass, Walmart Careers, FedEx, WebMD Doctor, Healthline, Kaggle, NVIDIA, and UC Berkeley`.
Then point your agent at `http://localhost:40000` through `http://localhost:40030` to explore 31 local mirrors of WebVoyager sites: `Allrecipes, Amazon, Apple, ArXiv, BBC News, Booking, GitHub, Google Flights, Google Maps, Google Search, Hugging Face, Wolfram Alpha, Cambridge Dictionary, Coursera, ESPN, Merriam-Webster, IKEA, Phys.org, Target, TED, Ohio State University, Rotten Tomatoes, Compass, Walmart Careers, FedEx, WebMD Doctor, Healthline, Kaggle, NVIDIA, UC Berkeley, and B&H Photo`.

For sub-second reset between rollouts, expose the control plane and call `/reset/<site>`:

Expand All @@ -63,29 +63,29 @@ git clone https://github.com/aiming-lab/WebHarbor && cd WebHarbor
./scripts/build.sh # docker build -t webharbor:dev .
```

### Local NVIDIA review candidate
### Site registry

This branch registers **30 sites**, the 30 entries listed above; NVIDIA took index 28
when #107 merged, so UC Berkeley (the site under review here) is the last entry,
registry index 29, container port 40029 (local review host port 48029). The
published-image quickstart above is not a claim that this review candidate has been
published or accepted.
This checkout registers **31 sites**. NVIDIA remains at index 28, UC Berkeley
remains at index 29, and B&H Photo is appended at index 30. Build the image from
this checkout to use this registry; publishing source does not update the
published Docker image automatically.

| Site | Registry position | Container port | Local review host port |
| Site | Registry position | Container port | Example local review host port |
| --- | --- | --- | --- |
| NVIDIA | 28 | 40028 | 48028 |
| UC Berkeley | 29 | 40029 | 48029 |
| B&H Photo | 30 | 40030 | 48030 |

`websyn_start.sh`, `control_server.py`, the `Dockerfile` `EXPOSE` line and every
site's `tasks.jsonl` `web` URL agree on 30 sites and `40000-40029`;
site's `tasks.jsonl` `web` URL agree on 31 sites and `40000-40030`;
`scripts/check_site_registry.py` (run by `scripts/check_assets.sh`) fails when they
drift.

After preparing the candidate assets and building `webharbor:dev`, the local
review deployment uses:

```bash
docker run -p 127.0.0.1:48080:8101 -p 127.0.0.1:48000-48029:40000-40029 webharbor:dev
docker run -p 127.0.0.1:48080:8101 -p 127.0.0.1:48000-48030:40000-40030 webharbor:dev
```

NVIDIA inherits the site contribution from @KaKituken
Expand All @@ -97,28 +97,22 @@ passed.

### Asset delivery status

`.assets-revision` is pinned to `c32018ca3b3d67e7b858b1b85fb101aea5090cd7`, the
head commit of HF dataset `main` and the squash-merge commit of HF dataset PR
[#91](https://huggingface.co/datasets/ChilleD/WebHarbor/discussions/91)
("berkeley: synthetic imagery bundle (164 files)", merged 2026-09-15T04:00:54Z).
Every other registered site's archive on that commit has the same size and LFS oid
as on the previous pin `b7e605c0ec5fc47de85b09e7427162cc50e38980`, and all 30 site
archives are byte-identical to the ones the interim `refs/pr/91` pin served, so the
pin change adds the UC Berkeley bundle without altering any other site's assets:

- the pinned revision carries 32 `*.tar.gz` (30 registered sites plus
`bandcamp.tar.gz` and `drugs_com.tar.gz`, which `fetch_assets.sh` ignores for
sites this checkout does not register);
- `nvidia.tar.gz` at that revision is the same 37-member archive as at the previous
pin, so `scripts/validate_asset_archive.py nvidia.tar.gz nvidia` prints
`[fetch] validated 37 managed members for nvidia` and exits 0;
- `berkeley.tar.gz` at that revision has 171 managed members, so
`scripts/validate_asset_archive.py berkeley.tar.gz berkeley` prints
`[fetch] validated 171 managed members for berkeley`;
- `./scripts/fetch_assets.sh` at this pin extracts all 30 registered sites
(`[fetch] done — 30 site(s) extracted into sites/`).

The previous pin `b7e605c0ec5fc47de85b09e7427162cc50e38980` is the squash-merge
`.assets-revision` pins the merged dataset commit `fa1e8a5b9e8e5d0e42764cd658825f4dea088d8f`
from [HF asset PR #92](https://huggingface.co/datasets/ChilleD/WebHarbor/discussions/92).
Its complete asset tree matches the tested candidate commit
`f09e586eec8bf1bca0bc0881e08b77f3c2a5508e`.

This revision adds `bh_photo.tar.gz` and preserves all 32 existing archives from
`c32018ca3b3d67e7b858b1b85fb101aea5090cd7` byte-for-byte, including Berkeley and
NVIDIA. It carries 33 archives for 31 registered sites plus the unregistered
Bandcamp and Drugs.com archives, which `fetch_assets.sh` ignores. A clean asset
fetch downloaded and extracted all 31 registered sites successfully.

B&H's archive contains images and external cache. The Docker build validates
its 508 declared assets and generates `instance_seed/bh_photo.db` from the tracked
catalog. No manually prepared B&H database is required for a fresh build.

The earlier pin `b7e605c0ec5fc47de85b09e7427162cc50e38980` is the squash-merge
commit of HF dataset PR
[#85](https://huggingface.co/datasets/ChilleD/WebHarbor/discussions/85) on the
dataset's `main`. It sits on top of PR
Expand All @@ -130,6 +124,7 @@ added the first reviewed NVIDIA bundle.
| --- | --- | --- | --- |
| `nvidia.tar.gz` at the current pin | 37 | 16,340,955 | `617a3e3740ba6706bcab786c8a5c3f9a22ecbb39eff5728ad2c12e4992cb098b` |
| `berkeley.tar.gz` at the current pin (HF PR #91) | 171 | 6,951,483 | `ab9d2716ae8d06540a181b5e60c37f613d87b103864b467511da546b1b173789` |
| `bh_photo.tar.gz` at the current pin (HF PR #92) | 511 | 79,658,793 | `867363d5484eb114d647e236991017992d5ac91ae3415996ad43bf654d99bd9a` |
| previous pin's `nvidia.tar.gz` (HF PR #84, superseded) | 34 | 9,927,312 | `ee8c6ba966e7a8f7fb5ad2d7ff0134ab98e7b80d6cc77f3328217405b8b34e2f` |

PR #85 replaces five product images and adds three dedicated hero images (see
Expand Down
2 changes: 1 addition & 1 deletion control_server.py
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,7 @@
'ikea', 'phys_org', 'target', 'ted',
'osu', 'rotten_tomatoes', 'compass', 'walmart_careers',
'fedex', 'webmd_doctor', 'healthline', 'kaggle',
'nvidia', 'berkeley',
'nvidia', 'berkeley', 'bh_photo',
]
BASE_PORT = 40000
WEBSYN_DIR = '/opt/WebSyn'
Expand Down
1 change: 1 addition & 0 deletions sites/bh_photo/.build-generated-seed
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
The Dockerfile generates instance_seed/bh_photo.db from the tracked source_catalog.json and seed_data.py.
8 changes: 8 additions & 0 deletions sites/bh_photo/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
# Harvested archive HTML and the manifests built from it are build-time
# intermediates; the runtime catalogue lives in instance_seed/bh_photo.db.
sources/cache/

# Recreated from instance_seed/ on every boot; not committed.
instance/
__pycache__/
*.pyc
1 change: 1 addition & 0 deletions sites/bh_photo/.requires-images
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
This site requires static/images from the pinned Hugging Face asset bundle.
3 changes: 3 additions & 0 deletions sites/bh_photo/_health.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
"""Per-site health probe (optional, called by control_server)."""
def health():
return {"ok": True, "site": "bh_photo"}
Loading