Skip to content

Tolerate OWID population column renames - #21

Merged
setupelz merged 1 commit into
mainfrom
fix/owid-population-column-tolerance
Aug 18, 2026
Merged

setupelz merged 1 commit into
mainfrom
fix/owid-population-column-tolerance

Conversation

@setupelz

Copy link
Copy Markdown
Owner

Summary

  • OWID rewrites its population export in place with no version identifier. The value column name has already changed once ("Population (historical estimates)" -> "Population"), which broke both hardcoded pins at data-refresh time -- the notebook path failed via a silent rename no-op that surfaced as a bare KeyError two lines later.
  • The column now resolves through a new owid_population_column() helper (newest spelling first, diagnostic DataProcessingError on an unrecognised column) instead of two hardcoded string pins.
  • Notebook 103's ingestion filter is now an explicit ISO3-or-world-key allow-list that prints what it drops, replacing a bare Code.notna() check that would have silently admitted any future OWID_* synthetic aggregate code into the intermediate artifact.
  • This matters most for first-run breakage risk: a fresh fetch-data user following the README hits this exact failure path, which is also the likely route a JOSS reviewer takes.
  • natscen already carries the consumer-side equivalent of this allow-list pattern; this brings the ingestion side in line.
  • The durable fix remains the pending Zenodo population vintage, which converts this source from unversioned to a pinned sha256 -- this PR is the interim tolerance layer.

Lint & Test Results

  • Lint: ruff check on the three changed files reports 2 pre-existing findings (UP038, isinstance tuple style) on lines untouched by this diff; no new findings introduced.
  • Tests: 704 unit + 42 integration tests pass, including 4 new regression tests covering both known OWID spellings, the unknown-column failure mode, and newest-spelling-wins precedence.

Test Plan

  • pytest tests/unit/iamc_historical/test_owid_population_column.py -v -- 4/4 passed
  • Full unit + integration suite -- 704 + 42 passed
  • Ruff check on changed files -- no new findings

OWID rewrites the population export in place and has renamed the value
column once already ("Population (historical estimates)" -> "Population"),
which broke both hardcoded pins at data-refresh -- the notebook path via
a silent rename no-op surfacing as a bare KeyError two lines later. The
column now resolves through owid_population_column() (newest spelling
first, diagnostic error on unknown), and notebook 103's ingestion filter
is an explicit ISO3-or-world-key allow-list that reports what it drops,
instead of relying on downstream source intersections to catch synthetic
aggregate codes. Regression tests cover both spellings and the failure
mode. The durable fix remains the pending Zenodo population vintage,
which converts this source from unversioned to a pinned sha256.
@setupelz
setupelz merged commit 44daa38 into main Aug 18, 2026
3 checks passed
@setupelz
setupelz deleted the fix/owid-population-column-tolerance branch August 18, 2026 06:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant