Skip to content

Gi terms and analytics - #26

Merged
bdevloed merged 42 commits into
mainfrom
gi-terms-and-analytics
Sep 20, 2026
Merged

bdevloed merged 42 commits into
mainfrom
gi-terms-and-analytics

Conversation

@bdevloed

Copy link
Copy Markdown
Contributor

No description provided.

bdevloed and others added 30 commits September 11, 2026 18:55
…llages guard

- app.js: the page-load open of a /<lang>/<slug> landing no longer fires
  Appellation Viewed — the pageview already records the slug, and the
  custom event made every entity landing a non-bounce by construction and
  let sessions start with a custom event. Appellation Viewed gains `via`
  (map / cycle / facet / omnisearch / panel-link) and `stack_size`; new
  `Feedback Clicked` {channel: github|email} on the sidebar and
  About-dialog links (data-feedback attributes in map_template.py).
- stage 02 extract_aire: recognise degree-less section-IV sub-block
  headers ("1 - Aire géographique") and sentence-form aires ("sur le
  territoire de la commune de Mâcon du département de Saône-et-Loire"),
  dropping lowercase asides. Pouilly-Loché's aire was the 366-commune
  proximité list (drawn across all of Burgundy in simple mode); it is now
  Mâcon, and 58 other single-commune AOCs gained a previously empty aire.
  A partial --only run no longer stubs the unselected records or shrinks
  _index.json.
- stage 04: [villages-guard] narrows a cahier-text commune union spanning
  more than 20x the parcellaire bbox to parcel-bearing communes (dormant
  for correct records).
- docs/analytics.md: event/prop reference, goal recipe, reading
  artefacts; CLAUDE.md pointer + scripts-contract note; CURATOR_TODO
  entries (Pouilly-Loché aire, canonical grape-slug ranking fragility);
  six parser regression tests.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…rator notes

Handoff from the 2026-09-11 review of 1,000 terroir-fact bullets:
docs/plan-terroir-facts-quality.md (fixes W1–W8), the two helper modules
the plan introduces (grounding coverage for 02d quotes; intra-record
de-duplication of restated facts) — not yet wired into a stage — and the
CURATOR_TODO entries (three FR parents bound to the wrong BO Agri PDF,
the W1–W8 queue).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Every record now carries `eu_scheme` (pdo / pgi / spirit-gi / uk-pdo /
uk-pgi / none) and `national_term` (the EU-registered traditional term the
regulator attaches GI-wide: AOC, DOCG / DOC / IGT, DOCa / DOQ / DO / Vino
de Pago / Vino de Calidad / Vino de la Tierra, DOC / Vinho Regional, DOC /
IG, DAC / Landwein, DOK / IĠT), derived at stage 04 by _lib/gi_terms.py and
rendered as "TERM (SCHEME)" in the panel meta line, docTitleFor, the SSR
card, entity title / description, browse list and children nav, with a
tooltip (definition + regulator source) on both tokens.

Sources, sha-pinned and joined on the EU file number: IT from the MASAF
elenchi DOP / IGP scraped by it/00 (national_term.py + overrides for Cirò
Classico, Valtènesi, Casauria); ES from the MAPA listado fetched by es/00
(national_term.py + overrides: Priorat DOQ, Tharsys, Urbezo); everyone
else from the checked-in ruling table traditional_terms.json (18 Austrian
DAC pins, Marc d'Alsace, per-scheme / per-term tooltip texts in
en/fr/es/nl). `class_key` feeds a new advanced-mode "Appellation type"
facet (scheme rows over term rows, parents-only wine counts).

audit_gi_terms.py asserts child == parent, term scheme == record scheme,
no empty class_label, and reports the per-country distribution. The
gettext .po sources are now committed (only .mo is ignored) so a fresh
checkout renders EN/ES/NL. Plan in docs/plan-eu-scheme-and-national-tier.md;
cross-check in VERIFICATION.md; curator pin passes (GR, CZ, CH, NL, SI, HU,
BG, CY) in CURATOR_TODO.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…he 02d caches

Two no-LLM post-pass scripts from the terroir-fact quality work
(docs/plan-terroir-facts-quality.md): dedupe_terroir_facts.py collapses
facts restated across the four sub-section calls of one record
(_lib.terroir_dedupe), and recompute_terroir_provenance.py re-grades every
fact's quotes with the ellipsis-aware coverage rule (_lib.terroir_coverage)
and rewrites cahier_coverage / wiki_coverage / provenance in place.
Committed as a snapshot; the terroir-facts session may still be iterating.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…eedback, shared prompts

Every terroir-fact write now goes through terroir_cache.write_source_cache /
write_translation_cache: the first write to a slug in a run snapshots its
source cache and its four translation caches to raw/terroir-facts-backup/<run>/,
and rollback_terroir_facts.py restores or deletes every file a run touched
(source + translations as one unit; a rollback is itself snapshotted).

The 21 × 02d scripts share one grounding test (terroir_coverage: ellipsis-
and block-aware, 0.6 threshold unchanged), dedupe their four sub-section
calls (terroir_dedupe), append the shared English STYLE_RULES block
(terroir_prompts) and the per-record review-feedback constraints
(terroir_feedback, from raw/terroir-facts-feedback/<slug>.json) to their
system prompt; shared cahiers are graded per chapter (terroir_chapters).
The 21 × 02e scripts build their prompt through translation_system_prompt
(names verbatim, common nouns translated, exonyms — exonyms.py).

Post-passes over the caches without an LLM call: normalize (colour codes,
VT/SGN, Latinisation), filter_terroir_boilerplate (tautology quotes shared
by ≥ 3 records), dedupe, recompute_terroir_provenance; build_terroir_feedback
merges a review's evidence into the sidecars; detect_untranslated lists the
(slug, locale) pairs still carrying source-form text.

Per-stage model defaults (providers.STAGE_DEFAULTS): sonnet-5 / thinking off
for 02d, opus-5 / adaptive for the gate and audit, sonnet-4-6 for 02e;
batch.run_two_pass threads `thinking` to the Anthropic batch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…audit, scoped re-run

02d_verify_terroir_facts.py (terroir_gate) grades, per record in one request,
every bullet's claim against the exact source text 02d graded against, the
Wikipedia hints and the record's feedback sidecar; verdicts apply
deterministically — supported / rewrite (guarded) / drop / subsection — with
`support` per fact, a `gate` block per cache, index-aligned prune or
`pending:` re-key of the translation caches, and a feedback `history` entry.

02e_verify_terroir_facts.py (terroir_backcheck) compares each translated
bullet with its source bullet — numbers, hedges, entities, watch-list false
friends, exonyms — and applies the corrected translation under the same
guards (`check` per fact, `backcheck` block per cache).

audit_terroir_facts_llm.py grades a record sample's rendered EN bullets
against the full source with an adversarial Opus-5 verifier (Wilson
interval; --from-backup / --compare for paired before-after; --emit-feedback
to feed the residue back). audit_terroir_facts.py gains gate_pending,
rewrite_rejected, feedback_recurrence, wiki_binding, foreign_name and the
strict FR name guard.

rerun_terroir_facts.py runs one scoped re-run under one run id: mark the
scoped caches stale → 02d --batch per country → 02d_verify → 02e --batch
→ 02e_verify → audit, logs under /tmp/owm-<run>/.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ia binding pins

MASAF consolidated disciplinari append one sub-disciplinare per sottozona,
each restarting at Art. 1; the last occurrence of each article used to win,
so 20 parents took their summary, roster, area and Art. 9 lien from their
last sottozona (Montepulciano d'Abruzzo → San Martino sulla Marrucina).
extract_article_runs now splits the header sequence at every restart, drops
a table of contents and a decree preamble, keeps the parent's own run as
article_bodies and the later runs as `annexes`; stage 04's sottozona
detector finds 77 sottozone in 17 parents (was 38 in 10).

BO Agri serves another appellation's cahier for Pierrevert, L'Étoile and
Grands-Echezeaux; register_overrides.json pins them prefer_cahier so stage
01 binds the eAmbrosia register attachment ahead of BO Agri.

aoc_overrides.json pins the Wikipedia articles the new wiki_binding audit
check flagged; 02b_fetch_aoc_lexicon gains --only / --refresh.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…UDE.md invariants

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…r the panel

derive_terroir's 4,000-character cap (a panel length) was the only text
02d, the gate and the audits ever saw: 469 of 522 disciplinari carry a
longer "Legame con l'ambiente geografico" — 3.66 M characters of regulator
terroir text — and Italy is the largest and worst-scoring country.

pick_terroir_article takes max_chars (None = whole article); stage 02f
emits `link_to_terroir_full` (uncapped) and `terroir_article` next to the
byte-identical `link_to_terroir` (cap_at_sentence, TERROIR_BRIEF_CHARS),
parser_template v3. IT 02d prefers the full text, so terroir_sources, the
gate and the audits follow; the changed source sha marks every IT record
stale for the next run. Sidecars regenerated (522 ok).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The grounding test failed verbatim quotes on typography alone: a curly
apostrophe against a straight one, a low-9 „quote" or « guillemets »
against "quotes", an en dash against a hyphen, the space a line-break
hyphenation leaves behind ("gradi- giorno"), a ligature. normalize() now
NFKC-decomposes, folds every quote / apostrophe / dash variant, drops soft
hyphens (a bullet glyph in this corpus) and zero-width characters, closes
hyphenated line breaks and the spaces inside guillemets — on both sides,
threshold unchanged at 0.6.

Measured on every kept quote of the r1 corpus against the exact source
text (terroir_sources): FR 1,628 of 3,165 quotes score higher and 1,387
go from block-rescued to a single contiguous match; the other countries
484 of 5,679 and 419; no quote crosses the threshold downwards (the few
lower scores are ≤ 0.011 shifts near 1.0 from "…" becoming three
characters).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…` is earned

STYLE_RULES now opens with the claim-support rule, spelled out as the
gate's own over-claim catalogue (a causal wrapper on a co-occurrence, a
narrowed en-bloc attribution, an invented qualifier, a sibling's
statement, a strengthened hedge) — the gate rewrote 27 % of the r1
bullets, 44 % of them light edits.

Gate (gate-v2): a `rewrite` verdict must carry a rewrite that changes
meaning. A rewrite that came back empty keeps the original as supported
(`rewrite_missing`, note kept — 124 did in r1); a cosmetic one (ratio
≥ 95 and no differing word of 4+ letters, so an added hedge is never
cosmetic) keeps the original as supported (`cosmetic_rewrite`) instead of
re-keying four translations. The report and the audit list both.

`interactions` (R3): measured on r1, 69 % of the 1,221 interactions
quotes carry an explicit connective, 8 % only in the bullet, 23 % none.
terroir_interactions.py carries a per-language connective table (15
source languages) and a deterministic test on the grounding quote: the
21 × 02d scripts drop an interactions fact whose quote states no link
(earn_interactions, cap 2 — the fourth call stays, the cahier's X.3 and
the single document's 8.4 are where the regulator states the links); the
gate demotes such a fact to the natural factors after its verdicts. No
bullet is ever promoted into the sub-section.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…t write; recurrence resolves on rewrites

needs_gate (shared by 02d_verify and the audit's gate_pending) now asks
whether every fact carries a gate verdict and whether the gate block
matches the record's source and GATE_VERSION — no exact sha of the
bullets. The normalise, dedupe and boilerplate post-passes change bullets
or remove facts without invalidating the verdicts on the rest, and used to
re-fire the gate corpus-wide after every post-pass.

The 21 × 02d scripts apply normalize_facts at write time (colour codes,
VT / SGN, terminal period), so the normalise post-pass is a no-op on a
fresh cache.

feedback_recurrence is lexical and matched a corrected bullet on its own
misleading original (5 of the 6 residual hits of the r1 audit): an entry
is now resolved when the matched fact's support.original_bullet matches
the claim at least as well and the rewrite was not cosmetic.
is_cosmetic_rewrite moves to terroir_dedupe so the gate and the feedback
module share it without an import cycle.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every fetched result carries the provider's usage (input / output / cache
tokens, Anthropic and Mistral alike); run_batch sums them, prices them at
the Batch-API rate (BATCH_PRICES_USD_PER_M — list price × 0.5, cache reads
at 10 %, cache writes at 125 %), prints the line and appends one row per
batch to raw/.batch/costs.jsonl (stage, model, thinking, batch id,
OWM_TERROIR_RUN). run_two_pass returns the summary; the gate and the
back-check put it in their report under `batch`; the orchestrator logs
the run's spend per stage at the end. Until now a run's cost had to be
re-read from the API afterwards.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
--scoped-02d passes the scope down to each 02d (--slug for FR, --only
elsewhere) so a country whose sources all changed — Italy after the MASAF
full-text change — re-extracts only the scoped records; --scoped-gate
passes it to 02d_verify as --only-file, so a smoke run right after a
GATE_VERSION bump does not re-gate the corpus. Defaults unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
….4, 3.7, 3.10) in CLAUDE.md and the hand-off

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…gs; rewrite_missing note

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… back empty (backcheck-v2)

The prompt requires a non-empty fix for a `fix` verdict; an empty one —
512 in r1 — keeps the translation as ok with check.fix_missing and the
issue note, listed as missing_fixes in the report, instead of
fix-rejected. Same treatment as the gate's rewrite_missing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…a citations caught and trailing ones stripped

The audit's masaf_sidecar_stale check still pinned it-masaf-disciplinare-v2,
so every regenerated v3 sidecar would have reported as stale. META_RE was
English-only: the smoke's Barolo bullet ended in "secondo il disciplinare"
unflagged. It now carries the Italian / French / Spanish / German / Dutch /
Portuguese citation forms, and the normaliser drops such a clause when it
trails the sentence (the fact stays; a mid-sentence citation is left to the
audit).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ike the other 20 scripts

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ation' and the EN/FR/ES/DE single-document forms

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The smoke run's back-check step submitted the whole corpus (5,895
requests): BACKCHECK_VERSION had been bumped to v2 and the step is
corpus-wide on whatever is unchecked. Cancelled before any request was
processed. --scoped-backcheck passes the scope to 02e_verify as
--only-file, the way --scoped-gate does for the gate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…migration

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…nctions, cedilla folding, NL→EN fallback

The smoke's five unearned interactions drops included three false
negatives ("glavni čimbenik", "fördert", "και έτσι"); samples of the
unmatched FR / NL / RO quotes showed the same pattern ("déterminent",
"contribuant", "hetgeen zich vertaalt in", "dă vinuri", cedilla ţ/ş,
Ambt Delden's English source under source_lang nl). The tables now carry
factor / role / influence nouns, word-initial verb stems, the plain
"because" conjunctions and "result of / in combination with" forms;
Romanian text is folded to ț / ș and a Dutch record is also tried
against the English table. Interactions quotes matched on the r1 corpus:
69 % → 85 %; the negatives in the tests still hold.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…tputs; restate the 3.4 acceptance

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ated source sha

Two such caches (LU Moselle, AT Tirol) sat unnoticed until the smoke's
corpus-wide 02e step happened to redo them; the audit now lists them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…, static system prompts in the gate / back-check / audit / 02e

_lib/prompt_cache.py: cached_system(shared, rest) puts the text a record
sends more than once first in the system prompt with cache_control, so
requests after the first read it at 0.1× the input price; mark_cached
wraps a static system prompt as one cached block; system_text flattens
for Mistral / Ollama, the manual round-trip files and the batch request
hash. OWM_CACHE_TTL = 5m (default) / 1h / off.

The 20 non-FR 02d scripts: the four sub-section calls of a record each
resent the whole lien (Santorini ≈ 29 K tokens per call); the lien is now
the leading cached block and the user turn carries only the sub-section
request (split_user_lead for the four USER_LEAD scripts). FR slices
section X per call and shares nothing, so it is unchanged. The gate
(1,365 tokens on Opus 5), the LLM audit (945) and the 21 × 02e scripts
(≈ 3 K tokens per locale, shared by every record of a batch) cache their
static system prompt; the back-check's 860-token prompt is under Sonnet
4.6's 1,024-token minimum, so its marker is a no-op until the prompt or
the model changes.

Batch hits are best-effort (concurrent processing; Anthropic quotes
30–98 %); a record's four calls are submitted adjacently and the ledger's
cache_creation / cache_read tokens show the achieved rate per batch —
with the 5-minute TTL the four-call pattern breaks even at a 29 % hit
rate. Request shapes validated with count_tokens.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…orini, 02d cost halved; revised migration projection

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… cahier defects

Pouilly-Vinzelles' "1°" is OCR'd as "l°"; Menetou-Salon's section X opens
straight at "a) - Description des facteurs naturels" with no "1°"
heading; Floc de Gascogne's a) heading lost its letter. Each lost the
largest slice (Pouilly-Vinzelles extracted 1 fact from 9.9 K chars, now
0). TOP_RE accepts l / I and normalises them; a lien with no "1°" treats
the text before the first numbered heading as section 1; a section 1
whose first sub-marker is not a) keeps the text before it as the natural
factors. 460 / 460 FR jobs now carry the slice.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…udit, caching-in-batches finding

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
bdevloed and others added 12 commits September 20, 2026 14:40
They took the generic Anthropic default, which only coincided with
STAGE_DEFAULTS["02e"]; a change to the 02e default would have changed
nothing. Wiring test added.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…the cached lien

Inside one large batch a record's four sub-section calls are processed
concurrently, so most of them write the lien instead of reading it
(13–47 % hit rates on cfg-2026-09-14 — break-even, not a saving). The 20
non-FR 02d scripts tag each call with its sub-section (cache_phase);
run_two_pass groups the collected requests by phase and run_phased
submits one batch per phase in order, each with its own resumable
sidecar, merging the results for one replay. The lien block carries the
1-hour TTL in phased mode (a read refreshes the timer; each phase only
has to finish within an hour) and the default TTL otherwise, where the 2×
write would be a loss. OWM_BATCH_PHASED=0 disables (one batch, as
before). Trade-off: four sequential batches per country.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…s thinking=disabled

The Claude 5 family runs adaptive thinking when the parameter is omitted;
the 4.x models do not. Every stage's max_tokens is sized for the JSON reply
alone, so Sonnet 5 on 02e (thinking mode None) spent the 2,000-token budget
thinking and 64 % of the replies came back truncated and were rejected by
the length check. effective_thinking() sends "disabled" explicitly for a
Claude 5 model when the stage sets no mode; explicit modes and the 4.x
models are unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ased probe, Claude 5 thinking rule

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…appellation" on records with sub-denominations

A parent's bullets are inherited by its sub-denominations' pages, where
"the northernmost part of the appellation, straddling Rioja Alavesa and
Rioja Alta" (Rioja's own sentence) reads as Alavesa's. For a record with
sub-denominations (terroir_roster — the roster the dedupe post-pass used,
now shared), the 02e user message and the 02d per-record block ask for the
appellation's name wherever the source refers to it generically, keeping
every sub-denomination's own name untouched. Per record, so the cached
per-locale system prompts stay identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…rs; a reverted move leaves no trace

The Rioja pliego states the sub-zones' wine styles (versatility, ageing
aptitude, blending) inside its climate paragraph and 02d filed the bullet
under natural factors; the gate's misfiling examples did not name the
case and Opus left it. Now explicit; re-gated Rioja → produit. When a
verdict moves a bullet into interactions and the earned rule sends it
straight back, the fact no longer carries moved_from / unearned_interaction.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… the 'Marnedallei' blend (glossary + back-check watch-list)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…records with sub-denominations

The 02e context rule produced "the Rioja appellation" for "la denominación";
the back-check, without the context, reverted it as an added entity.
The per-record note now reaches the checker's user message too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…alleys (2026-09-15)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…JSON stayed cached across a #hash navigation

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…' (normaliser, glossary, back-check watch-list)

The NL UI says 'appellatie' throughout; the NL bullets carried the French
loanword 261 times ('De appellation is gelegen in de streek Revermont').
The normaliser (post-pass + render time) rewrites the common noun and
leaves the registered term 'appellation d'origine contrôlée / protégée';
02e's glossary and the back-check's watch-list carry the same rule.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@bdevloed
bdevloed merged commit 2c7c876 into main Sep 20, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant