diff --git a/.claude/agents/parser-fixture-writer.md b/.claude/agents/parser-fixture-writer.md new file mode 100644 index 0000000..02860cc --- /dev/null +++ b/.claude/agents/parser-fixture-writer.md @@ -0,0 +1,75 @@ +--- +name: parser-fixture-writer +description: >- + Write fixture-based regression tests for one open-wine-map parser + (FR cahier extractor, ES national-pliego parser, a country's + national-spec parser, fiche_technique, etc.). Use during Phase 5 of the + improvements plan, or after any parser regression ("add fixture tests + for the RO parser", "lock in the uppercase-title guard with a test"). + Produces tests/fixtures/_*.txt|.html + tests/test__parser.py. + The agent reads cached raw/ documents to harvest real excerpts but must + redact them to minimal sections and never commit anything from raw/ + itself. +tools: Read, Grep, Glob, Bash, Write +--- + +You write regression tests for ONE parser at a time in open-wine-map, a +pipeline that extracts wine-appellation facts from public regulator +documents. Parsers here are regex/keyword-routing functions that have +historically regressed when a tweak for country N broke country M — your +tests exist to make that impossible. + +## Inputs you are given + +- The parser module path (e.g. `scripts/_lib/es/national_pliego.py`, + `scripts/02_extract_cahiers.py`, `scripts/_lib/fiche_technique.py`). +- Optionally: known-tricky cases (a slug, a git commit that fixed a bug, + a CURATOR_TODO note). + +## Method + +1. **Read the parser.** List its public entry points and the template / + heading variants it claims to handle (docstrings and comment tables + enumerate them — e.g. national_pliego.py names the JCCM / INCAVI / + AGACAL / ITACyL… heading variants; 02_extract_cahiers handles + uppercase-title guards, ordinal/bullet grape prefixes, cidre + `1) DENOMINATION` variants). +2. **Mine git history for past regressions:** + `git log --oneline -- ` — every "fix"-flavoured commit is + a mandatory test case. +3. **Harvest fixtures from raw/ caches** (these exist locally but are + gitignored): find a document exercising each variant, extract the + MINIMAL section the parser routes on (target < 200 lines), and save as + `tests/fixtures/__.txt` (pdftotext output) or + `.html` (OJ pages). These are excerpts of public, licence-clear + regulator documents — fine to commit. If raw/ lacks a document for a + variant, synthesize a minimal fixture from the parser's own expected + shape and mark it `# synthetic` in a header comment. +4. **Write `tests/test__parser.py`.** Import the parser the same way + existing tests do (see tests/test_content_block.py: + `sys.path.insert(0, …/scripts)` then import). Assert on: + - section routing (the right text lands in the right semantic role), + - grape parsing: principal vs accessory split, threshold percentages, + prefix stripping, + - style/colour detection where the parser does it, + - one regression test per mined bug-fix commit, named + `test_regression_`. + Keep assertions on STRUCTURE (keys, counts, specific slugs), not on + full-output snapshots — snapshots rot. +5. **Run:** + `uv run python -m pytest tests/test__parser.py -v` → all pass, and + the full suite `uv run python -m pytest -q` stays green. + `uv run ruff check tests/` → clean. + +## Hard rules + +- Never commit whole documents or anything under `raw/`; fixtures are + short redacted excerpts only. +- Never modify the parser to make a test pass. If the parser's actual + behaviour disagrees with what its docs/comments claim, write the test + against ACTUAL behaviour and flag the discrepancy in your final report. +- No network access needed or allowed — everything comes from the local + raw/ cache or is synthetic. +- One country/parser per run; report at the end: fixtures added, cases + covered, variants NOT covered (missing raw documents), discrepancies + found. diff --git a/.claude/skills/ruff-zero/SKILL.md b/.claude/skills/ruff-zero/SKILL.md new file mode 100644 index 0000000..c9a6ae3 --- /dev/null +++ b/.claude/skills/ruff-zero/SKILL.md @@ -0,0 +1,90 @@ +--- +name: ruff-zero +description: >- + Drive open-wine-map's ruff error count to zero WITHOUT changing behaviour. + Use when executing Phase 1 of the improvements plan, or whenever + `uv run ruff check scripts/ tests/` reports errors. Covers the safe + procedure for each rule class present in this repo (F401, F541, F601, + E741, E702, E731, F841, E402), with a mandatory special procedure for + F601 duplicate keys in scripts/_lib/grape_lexicon.py — those are DATA, + not style, and blind deletion can silently change grape-synonym folding. +--- + +# ruff-zero + +Goal state: `uv run ruff check scripts/ tests/` → `All checks passed!` and +`uv run python -m pytest -q` → all pass, after EVERY commit. + +## Order of operations + +1. `uv run ruff check scripts/ tests/ --fix` (auto-fixes F401 unused + imports, F541 empty f-strings). Run pytest. Commit: + `fix: ruff auto-fixes (unused imports, empty f-strings)` +2. F601 in `scripts/_lib/grape_lexicon.py` — see special procedure below. + Own commit. +3. Remaining classes, ONE RULE CLASS PER COMMIT, pytest between each: + - **E741** (`l`, `I`, `O` as names): rename to something contextual + (`l` → `line`, `lst`, `layer`, `lon`…). Rename ONLY within the + function scope shown by ruff; use exact-match find within that + function, never file-wide sed. + - **E702** (semicolon-joined statements): split onto separate lines, + same indentation. + - **E731** (lambda assignment): convert to a `def` with the same name. + - **F841** (unused variable): if the right-hand side has side effects + (function call), keep the call and drop the assignment; if pure, + delete the line. + - **E402** (import not at top): only move the import if nothing between + file top and the import mutates `sys.path` or env that the import + needs. In this repo several scripts do + `sys.path.insert(0, …/scripts)` BEFORE importing `_lib` — those + imports must stay put; silence with `# noqa: E402` instead. + +## F601 special procedure (grape_lexicon.py) + +Each finding = the same dict key literal appears twice. Python keeps the +LAST one. This file maps grape-name slugs to canonical slugs — a wrong +deletion changes which VIVC identity a grape folds into. + +Per finding: + +1. Find BOTH occurrences: `grep -n '""' scripts/_lib/grape_lexicon.py` +2. Compare the two VALUES. + - **Identical values** → delete the LATER occurrence (the one ruff + points at). If the later line carries a more informative comment + (e.g. a VIVC number), move that comment to the surviving line. + - **Different values** → DO NOT TOUCH. Append the key + both + line numbers + both values to a report list and continue. These are + real data conflicts requiring VIVC research (hand off to the + `grape-colour-researcher` agent or a human). +3. After all findings: + ``` + uv run ruff check scripts/_lib/grape_lexicon.py --select F601 + uv run python -m pytest -q # tests/test_no_duplicate_keys.py guards semantics + ``` +4. Commit: `fix: dedupe repeated grape_lexicon keys (identical-value F601s)` + — include the skipped-conflict report (if any) in the commit body. + +## Freezing at zero + +After all classes are clean, add to `pyproject.toml`: + +```toml +[tool.ruff.lint] +select = ["E", "F", "W"] +``` + +Run the check again — adding `W` may surface new findings; fix them the +same way (one class per commit). Do NOT add rule families beyond E/F/W +without being asked. + +## Pitfalls + +- Never run file-wide search-replace for renames; ruff gives exact + line/col — edit surgically. +- `scripts/04_build_maps.py` and `scripts/_lib/map_template.py` are + 5.9k/4k lines; load only the relevant region (read_file with offset), + not the whole file. +- If pytest fails after a change, `git checkout -- ` and redo that + single finding; do not stack fixes on a broken tree. +- `done2.json` / `todo.json` in the worktree are translation round-trip + artifacts — never commit them. diff --git a/.claude/skills/stage04-extraction/SKILL.md b/.claude/skills/stage04-extraction/SKILL.md new file mode 100644 index 0000000..6032d80 --- /dev/null +++ b/.claude/skills/stage04-extraction/SKILL.md @@ -0,0 +1,99 @@ +--- +name: stage04-extraction +description: >- + Move-only refactor procedure for carving functions out of + scripts/04_build_maps.py into scripts/_lib/ modules without changing any + build output. Use when executing Phase 6 of the improvements plan + ("extract augmenters", "modularize stage 04") or any time a function + must move out of 04_build_maps.py / map_template.py. The danger this + skill defends against: stage 04 failures are often SILENT (a country or + feature just disappears from the map), so verification relies on the + golden-output comparator, not on "it didn't crash". +--- + +# stage04-extraction + +## Iron rules + +1. **Move, never edit.** The function body is copied verbatim. No renames, + no cleanups, no type-hint additions, no f-string fixes — those are + separate commits OUTSIDE this skill. +2. **One module per commit.** +3. **The build output is the test.** A move is proven correct only by + `scripts/compare_build_output.py` reporting `identical` after a full + stage-04 rebuild (Task "final proof" below). py_compile + pytest are + intermediate gates only. + +## Prerequisite + +A golden snapshot must exist (plan Phase 0): +`tmp/golden-before/` populated from the CURRENT `wiki/` output, and +`scripts/compare_build_output.py` present. If missing, create them first — +never extract without a golden snapshot. + +## Procedure (per module) + +1. **Map the block.** Identify the function(s) to move and EVERYTHING they + reference: module-level constants, helper functions, imports. Use: + ``` + grep -n "" scripts/04_build_maps.py + ``` + for every symbol used inside the body. A symbol used by BOTH the moved + block and remaining code stays in a shared location (move it to the new + module and import it back, or leave it and import it into the new + module — prefer whichever produces fewer cross-imports). +2. **Create the module** under `scripts/_lib/` (e.g. + `scripts/_lib/augment/es.py`). Top of file: only the imports the moved + code actually needs. Note: `_lib` modules use relative imports + (`from .grape_lexicon import …`) — follow the pattern of existing + `_lib` files; check one (e.g. `scripts/_lib/lieu_dit.py`) first. +3. **Replace in 04_build_maps.py** the moved block with an import near the + other `_lib` imports: + ```python + from _lib.augment.es import augment_es_records_with_national_pliegos + ``` + (Match how 04_build_maps.py already imports from `_lib` — copy an + existing import line's style exactly.) +4. **Intermediate gates:** + ``` + uv run python -m py_compile scripts/04_build_maps.py scripts/_lib/augment/*.py + uv run python -m pytest -q + uv run ruff check scripts/ + ``` + All three must pass. (04_build_maps.py cannot be imported by module + name — it starts with a digit — hence py_compile.) +5. **Smoke the entry point** without a full build: + ``` + uv run scripts/04_build_maps.py --help 2>&1 | head -5 + ``` + If the script has no argparse, run it for ~20 seconds and Ctrl-C after + the record-loading phase prints; an ImportError/NameError appears + immediately. +6. **Commit:** `refactor: extract to _lib/augment/.py (no-op)` + +## Final proof (once per phase, after the LAST module move) + +``` +uv run scripts/04_build_maps.py +uv run scripts/compare_build_output.py tmp/golden-before wiki +``` + +Expected: `identical`, exit 0. ANY diff = a move changed behaviour → +bisect by `git stash` / re-applying module moves one at a time. Do not +rationalize a diff away; stage 04 output is deterministic by design (the +codebase sorts set-derived structures specifically to guarantee this). + +## Known coupling traps in 04_build_maps.py + +- Several `augment_*` functions read module-level path constants + (`ROOT`, `RAW`, …) defined near the top of the file — they must be + imported or re-derived (`Path(__file__).resolve()` depth changes when + the file moves from `scripts/` to `scripts/_lib/augment/`! Re-derive + carefully: `parents[2]` from `scripts/_lib/augment/x.py` = repo root, + vs `parents[1]` from `scripts/04_build_maps.py`). +- `_backfill_it_nonstub_from_masaf` is a private helper of the IT + augmenter — moves with it. +- `synthesize_it_sottozone_records` is called between other IT steps in + `main()`; preserve call ORDER in main() exactly. +- tqdm/print progress lines inside moved functions stay verbatim (curators + grep build logs for them). diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml new file mode 100644 index 0000000..66dece0 --- /dev/null +++ b/.github/workflows/ci.yml @@ -0,0 +1,35 @@ +name: ci +on: + push: + branches: [main, seo-geo-prerender] + pull_request: + +jobs: + checks: + runs-on: ubuntu-latest + steps: + - uses: actions/checkout@v4 + - uses: astral-sh/setup-uv@v5 + with: + enable-cache: true + - name: Sync dev dependencies + run: uv sync --group dev + - name: Ruff (lint, pinned at zero) + run: uv run ruff check scripts/ tests/ + - name: Pytest + run: uv run python -m pytest -q + + eslint: + # Advisory only — the map app JS lives in scripts/_lib/assets/app.js + # (extracted from the Python template in Phase 4). Build correctness is + # guarded by the byte-identity golden check, not by lint, so this job is + # non-blocking (continue-on-error). + runs-on: ubuntu-latest + continue-on-error: true + steps: + - uses: actions/checkout@v4 + - uses: actions/setup-node@v4 + with: + node-version: "20" + - name: ESLint (map app JS) + run: npx --yes eslint@9 scripts/_lib/assets/app.js diff --git a/.gitignore b/.gitignore index 97ae1e7..dfdaf7d 100644 --- a/.gitignore +++ b/.gitignore @@ -6,6 +6,11 @@ __pycache__/ .DS_Store *.log +# Bunny CDN access logs pulled via scripts/fetch_bunny_logs.py (v2 logging API) +logs/ + +.hermes/* + # secrets — PISTE / Légifrance OAuth credentials, etc. .env .env.* @@ -37,6 +42,9 @@ raw/wikipedia/* *.pmtiles *.zip *.html +# …except test fixtures (short, redacted, licence-clear OJ-page HTML +# excerpts) — these ARE committed; see tests/fixtures/README.md. +!tests/fixtures/*.html # stage 02c round-trip work files (todo.json emitted by --emit-todo, # done*.json returned by the translator) — ephemeral diff --git a/CLAUDE.md b/CLAUDE.md index 292e412..140aed8 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -95,7 +95,10 @@ details and "Hard rules" for invariants that apply to every country. in the tooltip source-block alongside the Wikipedia attribution. JKI publishes no explicit data licence — the codebase therefore ships VIVC IDs + prime names (factual citation) and does **not** republish - verbatim synonym strings in the UI pending JKI confirmation. Citation: + verbatim synonym strings in the UI pending JKI confirmation (a draft + licence-query email to JKI is staged at + [docs/vivc-jki-licence-query.md](docs/vivc-jki-licence-query.md) for a + human to send; record the reply there). Citation: Röckel et al., Vitis International Variety Catalogue — www.vivc.de. Ambiguous slugs (multiple candidate VIVC entries) get pinned via `raw/vivc/slug_overrides.json` (template at `slug_overrides.example.json`). @@ -321,6 +324,26 @@ correcting the upstream commune list / resolver. Thresholds are CLI-configurable (`--sliver-max`, `--max-sliver-km2`, …); `--strict` exits non-zero on unreviewed slivers. +## Bétard-snapshot delta audit + +[scripts/audit_betard_delta.py](scripts/audit_betard_delta.py) flags +appellations whose geometry comes from Bétard 2022 (`geom_source` +`figshare-pdo` / `figshare-pdo-alias`) but whose GI was **registered or +amended after the dataset's data snapshot** (Nov-2021 cutoff) — those +polygons may not reflect a post-snapshot boundary change, and a brand-new GI +may only be on the map via an incidental file-number match. It reads the +compact `wiki/data/aocs.en.*.js` startup blob (geom_source is a startup +field) cross-referenced against each `raw//eambrosia/index.json`'s +`eu_protection_date` / `modification_date`, and buckets each record +**FLAGGED** / **REVIEWED** / **OK** / **NO-DATE** (the sibling-audit pattern). +Curator-confirmed-unchanged boundaries go in +[scripts/_lib/betard_delta_overrides.json](scripts/_lib/betard_delta_overrides.json) +(slug → `{reason, source}`) and report as REVIEWED. Read-only detector — +changes no geometry; a real FLAGGED finding is fixed upstream (a newer Bétard +release, a regional zone layer, or a commune-list resolver). `--strict` exits +non-zero on any unreviewed FLAGGED finding; `--cutoff` overrides the snapshot +date. + ## Page format (per-AOC pages) ``` @@ -381,8 +404,9 @@ twice with no changes upstream must be a no-op (cache hits). | 02d_extract_terroir_facts.py | raw/inao/cahier-extracted/*.json + raw/wikipedia/aocs/fr/ | raw/terroir-facts/*.json + manifest.json | | 02e_translate_terroir_facts.py | raw/terroir-facts/*.json | raw/translations/terroir-facts//*.json | | 02g_fetch_vivc.py | raw/inao/cahier-extracted/*.json + raw/es/pliegos-extracted/ + raw/pt/cadernos-extracted/ + raw/vivc/slug_overrides.json | raw/vivc/{search,passport,by-slug}/*.html\|json + manifest.json + slug_overrides.example.json | +| 02i_fetch_wikidata_qids.py | raw/*/*-extracted/*.json (slug + id_eambrosia) + raw/wikipedia/aocs// + raw/wikidata/slug_overrides.json | raw/wikidata/qids-by-slug.json + p9854.json + manifest.json + slug_overrides.example.json | | 03_generate_wiki.py | raw/inao/cahier-extracted/*.json + raw/terroir-facts/ | wiki/*.md, wiki/_index.json | -| 04_build_maps.py | raw/inao/cahier-extracted/*.json + raw/wikipedia/grapes/ + raw/translations/grapes/ + raw/vivc/by-slug/ + raw/wikipedia/styles/ + raw/translations/styles/ + raw/wikipedia/aocs/ + raw/translations/summaries/ + raw/translations/terroir-facts/ + raw/terroir-facts/ + raw/ign/communes.geojson + raw/inao/parcellaire/ + raw/cadastre/lieux-dits/ | wiki/index.html (EN canonical = homepage), wiki/{fr,es,nl}/index.html, wiki/map-data/*.pmtiles, wiki/robots.txt, wiki/sitemap.xml (homepage × 4 locales) | +| 04_build_maps.py | raw/inao/cahier-extracted/*.json + raw/wikipedia/grapes/ + raw/translations/grapes/ + raw/vivc/by-slug/ + raw/wikidata/qids-by-slug.json + raw/wikipedia/styles/ + raw/translations/styles/ + raw/wikipedia/aocs/ + raw/translations/summaries/ + raw/translations/terroir-facts/ + raw/terroir-facts/ + raw/ign/communes.geojson + raw/inao/parcellaire/ + raw/cadastre/lieux-dits/ | wiki/index.html (EN canonical = homepage), wiki/{fr,es,nl}/index.html, wiki/{en,fr,es,nl}//index.html (per-appellation entity pages), wiki/{en,fr,es,nl}/appellations/index.html (browse-index hub linking every indexable slug, grouped by country — fixes the entity-page link-graph orphan problem), wiki/map-data/*.pmtiles, wiki/robots.txt, wiki/sitemap.xml (4 home + 4 browse + 1,638 index slugs × 4 locales), wiki/llms.txt (AI-crawler index), wiki/404.html | ## Spain pipeline (`scripts/es/`) @@ -915,13 +939,16 @@ Per IT record, in priority order (each step records the chosen source in `geom_source` so the panel can attribute correctly): 1. **`geoportal-zone:`** — official regional-geoportal - production-zone polygon, matched by appellation name. Five regions - are harvested (Piemonte, Veneto, Lazio, Lombardia, Toscana — all - CC-BY 4.0 / IODL 2.0); ~218 of 531 IT wines resolve here, including - every flagship (Barolo, Soave, Valpolicella, Chianti, Brunello, - Bolgheri, Franciacorta, Frascati). An appellation spanning regions - is the union of its per-region pieces. Umbria + Puglia are tracked - to-dos (see CURATOR_TODO.md). + production-zone polygon, matched by appellation name. Six regions + are harvested (Piemonte, Veneto, Lazio, Lombardia, Toscana, Umbria — + all CC-BY 4.0 / IODL 2.0); ~237 of 531 IT wines resolve here, + including every flagship (Barolo, Soave, Valpolicella, Chianti, + Brunello, Bolgheri, Franciacorta, Frascati, Sagrantino di + Montefalco). An appellation spanning regions is the union of its + per-region pieces (e.g. Orvieto = `lazio+umbria`). Umbria is harvested + via a bespoke CKAN per-appellation `.zip`/`.7z` shapefile fetch + (`fetch_type: ckan_shapefiles` in `zone_sources.py`); Puglia is the + remaining tracked to-do (see CURATOR_TODO.md). 2. **`parent-appellation`** — sottozone (sub-denominations) inherit the parent's polygon. 3. **`figshare-pdo`** — exact `file_number` (`PDO-IT-A*` / @@ -2711,20 +2738,44 @@ CH-specific notes: - The per-canton règlement parser at [scripts/_lib/ch/reglement.py](scripts/_lib/ch/reglement.py) uses a **whole-document grape-lexicon scan** rather than section-scoped - extraction. The shared `_lib.grape_entity.match_variety` is robust - enough (lexicon-based + per-token rejection) to scan ~50 KB of - règlement text without false positives, and cantonal règlements - frequently bury the variety list in an annex or refer to an - external annex by article — section-scoped extraction missed most - of the recall. Commune extraction stays section-scoped because + extraction, because cantonal règlements frequently bury the variety + list in an annex or refer to an external annex by article — + section-scoped extraction missed most of the recall. Scanning ~50 KB + of regulatory prose is noise-prone, though, so `extract_varieties` + guards the shared `match_variety` three ways (2026-06): a + **commune-name guard** (a chunk that is a Swiss commune — Genève, + Sion, Cortaillod — is area prose, not a grape), a **fuzzy floor** + (`_CH_FUZZY_FLOOR`=95: weak fuzzy hits matched common words / place + fragments to obscure non-Swiss grapes — `vigne`→viognier, + `Nein`→durif), and a small **`_CH_STOP_SURFACES`** set for prose + words that are exact grape aliases out of context (`canton`→chenin, + `Säure`→calitor, `weisse`→valente). Candidate cleaning strips + leading FR/IT/DE determiners (`il Merlot`, `la Bondola`), `a)`/`°` + ordinal markers, and footnote tails, and adds a name-before-number + variant to recover must-weight (Mostgewicht) lines + (`a) Blauburgunder 19,4 °Brix`). The Swiss Agroscope crossings + + natives the scan surfaces (Garanoir, Mara, Galotta, Doral, Divico, + Carminoir, Diolinoir, Bondola, Completer, Gouais→heunisch) live in + the shared `grape_lexicon.py`; Cornalin / Humagne are deliberately + NOT folded yet (Valais Cornalin = Rouge du Pays vs Humagne Rouge = + Cornalin d'Aoste is an identity tangle that needs a VIVC pass). + Commune extraction stays section-scoped because whole-document commune scans generate huge false-positive lists. - Variety extraction recall: 20 of 26 cantons return non-zero - varieties; AG (29), GE (47), VD (16), VS (25), FR (66), NE (18), - TI (14), JU (4), LU (7), BL (8), GR (5), SG (5), OW (4), TG (4), - ZH (3), SH (2), SO (1), BE (1), GL (1), ZG (1). Cantons with 0 - varieties (AI, AR, BS, NW, SZ, UR) defer to federal OVin without - cataloguing varieties locally (BS defers to BL via inter-cantonal - Vereinbarung; the others are tiny corpora ~5 ha each). + Variety extraction recall (post-2026-06 FP cleanup — counts dropped + because the prior figures double-counted prose/place false positives + that the new guards remove, and rose where Mostgewicht recall recovers + real lists): 11 cantons return non-zero cantonale-tier varieties; + AG (65), GE (51), VS (44), VD (38), LU (25), TI (22), NE (15), BL (7), + SG (4), OW (4), GR (2). The other cantons (AI, AR, BE, BS, GL, JU, NW, + SH, SO, SZ, TG, UR, ZG, ZH) now return 0 at the cantonale tier: + some defer to federal OVin without cataloguing varieties locally + (AI/AR/BS/NW/SZ/UR; BS defers to BL via inter-cantonal Vereinbarung; + tiny corpora ~5 ha each), and the rest (BE, GL, JU, SH, SO, TG, ZG, + ZH) DO grow wine but either publish a règlement too sparse to parse + (ZH/SO are ~3 KB with no variety list) or enumerate varieties in a + Mostgewicht / annex form the scan doesn't yet recover — a known + recall gap, not a regression (their prior non-zero counts were + entirely prose/place false positives now removed). - For the 5 multi-AOC cantons (VD has 10 AOCs, GE has 23, TI has 4, BE has 3, FR has 2), the canton-wide règlement body is the v1 default for variety lists. Per-AOC commune-list carving (Phase 2.5) @@ -4131,10 +4182,148 @@ uv run pybabel init -i locale/messages.pot -d locale -l # to add a ne After editing a `.po`, just rerun `uv run scripts/04_build_maps.py`. +## Data bundle: startup blob + lazy panel detail + +The map's per-appellation data ships in two tiers so the front page is light +(the corpus is ~2,900 records). The render-blocking startup bundle +`wiki/data/aocs...js` (`window.__OWM_DATA`) carries only +`STARTUP_AOCS_FIELDS` ([scripts/_lib/map_template.py](scripts/_lib/map_template.py)) +— the fields every startup path reads: the sidebar appellation list, search, +facet filtering (`matchesClient`/`matchesExceptFacets`), fly-to +(`fitToFiltered`/`localityRank`), the active-filter chips +(`renderActiveFilters`→`grapeName`→`grapes_info`), and the map-click stack +dedup (`geom_source`). It is **the contract between the Python emitter and the +JS app**: a field read by any of those paths MUST be in the set, or the +sidebar/map silently breaks (there is no golden-diff guard here — Phase 3 +deliberately changes this bundle). + +Everything else on a record (summary, terroir facts, sources, grape +display-names, dűlők, menzioni, notes, attribution/geometry-provenance fields) +is the **panel payload**, emitted per slug per locale as +`wiki/data/d//.json` (the complement of `STARTUP_AOCS_FIELDS`; +write-if-changed + stale-prune in `emit_html`). The JS fetches it on first +panel open (`hydratePanel` → `Object.assign` into `AOCS[slug]`, deduped while +in flight, graceful degrade on failure), showing a skeleton placeholder +meanwhile; repeat opens are instant. This cut the startup bundle from ~13.2 MB +raw / 1.9 MB gz to ~3.2 MB / ~0.4 MB. `grapes_info` stays inline — it is a +startup dependency via the active-filter chips. The full record is still used +for the server-rendered entity cards (SSR is unchanged), so crawlers see the +same content; only the JS map lazy-loads. The per-slug JSON is not +content-hashed (stable path, fetched at runtime, not referenced from cacheable +HTML) and never enters the sitemap. + +### Map app JS source + +The map application JS lives in +[scripts/_lib/assets/app.js](scripts/_lib/assets/app.js) — a real, lint-able +`.js` file, not an escaped Python string. `_render_app_js` in +[scripts/_lib/map_template.py](scripts/_lib/map_template.py) injects the +per-locale build values by replacing `__OWM___` tokens (one per former +`{slot}` format field) and emits `wiki/assets/app...js` exactly +as before — byte-for-byte equivalent to the old `_APP_JS.format(**kwargs)` +(verified by the golden comparator). `eslint.config.js` (flat config, advisory) +lints it: `npx --yes eslint@9 scripts/_lib/assets/app.js`; the OWM token +identifiers are declared as globals there, and a non-blocking `eslint` CI job +runs it. Edit app.js directly; do not move the JS back into the template. The +CSS still ships inline in `_TEMPLATE` and is lifted to the shared +`style..css` by `_split_template`. + +## Structured data (JSON-LD) on entity pages + +Each **indexable** per-appellation entity page (`//`) carries a +single schema.org `@graph` — `WebSite → WebPage → Place(AdministrativeArea) → +BreadcrumbList`, with stable fragment `@id`s cross-linked via +`mainEntity` / `breadcrumb` / `isPartOf`. Built by `_build_entity_jsonld()` +in [scripts/_lib/map_template.py](scripts/_lib/map_template.py); **folded** +pages (sub-denominations, stubs, no-geometry, thin records) emit none, by +design. Honest modelling only — no `Article` markup (would need fabricated +author / editorial dates for a generated page). + +- **`sameAs`** (entity reconciliation — the highest-value SEO/GEO signal): + Wikidata QID → per-locale Wikipedia article → official regulator / producer + body. Source PDFs and legal acts are *not* identity pages, so they go in the + WebPage's **`isBasedOn`** instead. The eAmbrosia register is excluded — it's + a hash-route SPA (`…/#/detail/EUGI…`), one server shell for every GI, so it's + a poor `sameAs` target. +- **Wikidata QIDs** come from stage **02i** + ([scripts/02i_fetch_wikidata_qids.py](scripts/02i_fetch_wikidata_qids.py)), + a cached, incremental network stage (run between 02g and 04). Two resolution + paths: (1) Wikidata property **P9854 "eAmbrosia ID"** joined on each record's + `id_eambrosia` (the eAmbrosia-sourced countries), and (2) **Wikipedia + sitelink → QID** via the MediaWiki `pageprops.wikibase_item` API, keyed on the + 02b/aocs validated article title (covers FR, which is INAO-sourced and carries + no `id_eambrosia`, plus the P9854 gaps). Coverage is partial by design + (~1.2 k of the corpus at v1; the rest fall back to Wikipedia/regulator + `sameAs` or none). Pin/suppress a QID via `raw/wikidata/slug_overrides.json`. + Stage 04 joins `qids-by-slug.json` onto each record as `wikidata_qid`. +- **BreadcrumbList**: `Open Wine Map → [parent →] appellation`. The country + level is deliberately omitted — there is no per-country landing page, and a + non-final `ListItem` without an `item` URL is invalid for Google's + BreadcrumbList rich result. Every emitted crumb carries an `item`. +- `description` is the localized summary → first terroir-fact bullets → the + 160-char meta description; `inLanguage` is the page locale. +- Contract: the builder returns a pre-serialised opaque string filling the + `{jsonld_html}` `str.format` slot — its JSON braces are data, not format + fields, so it must not be double-braced or `esc()`-ed. + +## Internal linking & crawlability (browse index, cross-links, llms.txt, 404) + +The map homepage is a JS app shell with no crawlable `` links to entity +pages (the sidebar list is client-rendered), so the per-appellation pages were +link-graph orphans discoverable only via the sitemap. Stage 04 now emits a +static link layer (all in [scripts/_lib/map_template.py](scripts/_lib/map_template.py) ++ [scripts/04_build_maps.py](scripts/04_build_maps.py)): + +- **Browse-index pages** — `wiki//appellations/index.html` per locale, + a self-contained page (`_render_browse_page` + `_BROWSE_TEMPLATE`, NOT the + map shell) listing every **index**-classified appellation as a real ``, + grouped by country and sorted by localized name. The static crawl hub. + A namespace `assert "appellations" not in aocs` guards the path collision. +- **Inbound links** to the browse hub: the sidebar footer (`{browse_path}` + slot, on the homepage AND every entity page) + a paragraph in the About + dialog (`about_browse_html`). +- **Entity cross-link nav** (`_entity_nav_html` → `