diff --git a/CATALOG.md b/CATALOG.md
index 5bb399a..7ebd110 100644
--- a/CATALOG.md
+++ b/CATALOG.md
@@ -6,7 +6,7 @@
The migrated-dataset registry, **auto-generated** from the sidecar manifests (`lectures/*.yml`). Do not edit by hand — run `python scripts/build_catalog.py`. A dataset appears here once it has a manifest; files not yet migrated are tracked in [PLAN.md](PLAN.md) Phase 9.
-**9 datasets migrated** · 2.4 MB total · 8 permitted / 1 restricted redistribution
+**10 datasets migrated** · 2.4 MB total · 9 permitted / 1 restricted redistribution
| Dataset | Class | Source | Licence | Redist. | Integrity | Builder | Size | Used by |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
@@ -16,6 +16,7 @@ The migrated-dataset registry, **auto-generated** from the sidecar manifests (`l
| [**epl_match_goals.csv**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/epl_match_goals.csv)
English Premier League — full-time scores, 2015-16 to 2024-25 | constructed | [openfootball / football.json](https://github.com/openfootball/football.json) | Public domain | ✅ permitted | ✅ verified | ✅ committed | 203.2 KB | [lecture-python-intro · fitting_distributions.md](https://github.com/QuantEcon/lecture-python-intro/blob/main/lectures/fitting_distributions.md) |
| [**japan_deaths_by_age.csv**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/japan_deaths_by_age.csv)
Japan — deaths by single year of age, 2023 | constructed | [United Nations, Department of Economic and Social Affairs, Population Division — World Population Prospects 2024](https://population.un.org/wpp/downloads) | CC BY 3.0 IGO | ✅ permitted | ✅ verified | ✅ committed | 1.7 KB | [lecture-python-intro · observed_distributions.md](https://github.com/QuantEcon/lecture-python-intro/blob/main/lectures/observed_distributions.md)
[lecture-python-intro · fitting_distributions.md](https://github.com/QuantEcon/lecture-python-intro/blob/main/lectures/fitting_distributions.md) |
| [**japan_earthquakes.csv**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/japan_earthquakes.csv)
Japan region — earthquakes of magnitude 5 and above, 2000-2024 | constructed | [Advanced National Seismic System (ANSS) Comprehensive Earthquake Catalog (ComCat), US Geological Survey](https://earthquake.usgs.gov/earthquakes/search/) | US Government work — public domain | ✅ permitted | ✅ verified | ✅ committed | 172.8 KB | [lecture-python-intro · fitting_distributions.md](https://github.com/QuantEcon/lecture-python-intro/blob/main/lectures/fitting_distributions.md) |
+| [**japan_population_by_age.csv**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/japan_population_by_age.csv)
Japan — population by single year of age, 2024 | constructed | [Population Estimates, Statistics Bureau of Japan, Ministry of Internal Affairs and Communications](https://www.stat.go.jp/english/data/jinsui/index.html) | Japan Statistics Bureau terms of use | ✅ permitted | ✅ verified | ✅ committed | 1.3 KB | [lecture-python-intro · prob_dist.md](https://github.com/QuantEcon/lecture-python-intro/blob/main/lectures/prob_dist.md) |
| [**lingcod_msy_recovery.csv**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/lingcod_msy_recovery.csv)
Pacific Coast lingcod — biomass and fishing pressure relative to MSY | constructed | [RAM Legacy Stock Assessment Database](https://www.ramlegacy.org/) | CC BY 4.0 | ✅ permitted | ⚠️ unverifiable | ⚠️ unrecovered | 2.3 KB | [lecture-python-intro · msy_fishery.md](https://github.com/QuantEcon/lecture-python-intro/blob/main/lectures/msy_fishery.md) |
| [**realwage.csv**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/realwage.csv)
OECD real minimum wages — 32 countries, 2006–2016 | constructed | [OECD — Real minimum wages (RMW)](https://stats.oecd.org/Index.aspx?DataSetCode=RMW) | CC BY 4.0 | ✅ permitted | ⚠️ unverifiable | ⚠️ unrecovered | 118.7 KB | [lecture-python-programming · pandas_panel.md](https://github.com/QuantEcon/lecture-python-programming/blob/main/lectures/pandas_panel.md)
[lecture-python.myst · pandas_panel.md](https://github.com/QuantEcon/lecture-python.myst/blob/main/lectures/pandas_panel.md) |
| [**us_adult_heights.csv**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/us_adult_heights.csv)
United States — adult standing height by sex, NHANES 2015-2018 | constructed | [National Health and Nutrition Examination Survey (NHANES), US Centers for Disease Control and Prevention, National Center for Health Statistics](https://www.cdc.gov/nchs/nhanes/index.htm) | US Government work — public domain | ✅ permitted | ✅ verified | ✅ committed | 123.1 KB | [lecture-python-intro · prob_dist.md](https://github.com/QuantEcon/lecture-python-intro/blob/main/lectures/prob_dist.md)
[lecture-python-intro · observed_distributions.md](https://github.com/QuantEcon/lecture-python-intro/blob/main/lectures/observed_distributions.md)
[lecture-python-intro · fitting_distributions.md](https://github.com/QuantEcon/lecture-python-intro/blob/main/lectures/fitting_distributions.md) |
diff --git a/lectures/japan_population_by_age.csv b/lectures/japan_population_by_age.csv
new file mode 100644
index 0000000..1bb379a
--- /dev/null
+++ b/lectures/japan_population_by_age.csv
@@ -0,0 +1,102 @@
+age,total_population,japanese_population
+0,716,695
+1,757,738
+2,797,780
+3,828,810
+4,836,818
+5,872,853
+6,915,897
+7,940,921
+8,980,961
+9,1004,987
+10,1003,985
+11,1028,1011
+12,1032,1015
+13,1057,1041
+14,1066,1051
+15,1071,1057
+16,1091,1076
+17,1085,1070
+18,1089,1062
+19,1121,1061
+20,1189,1102
+21,1207,1109
+22,1261,1148
+23,1297,1163
+24,1307,1169
+25,1295,1168
+26,1312,1184
+27,1304,1180
+28,1298,1181
+29,1309,1196
+30,1300,1196
+31,1268,1167
+32,1275,1185
+33,1265,1180
+34,1288,1201
+35,1308,1229
+36,1345,1274
+37,1378,1311
+38,1396,1335
+39,1445,1388
+40,1493,1438
+41,1512,1457
+42,1512,1456
+43,1524,1475
+44,1590,1543
+45,1629,1583
+46,1689,1646
+47,1732,1692
+48,1812,1773
+49,1885,1848
+50,1990,1954
+51,2023,1987
+52,1985,1949
+53,1926,1890
+54,1864,1828
+55,1833,1799
+56,1787,1753
+57,1778,1750
+58,1387,1359
+59,1705,1678
+60,1595,1568
+61,1551,1525
+62,1497,1475
+63,1463,1443
+64,1467,1448
+65,1487,1470
+66,1440,1424
+67,1393,1378
+68,1456,1441
+69,1492,1478
+70,1482,1470
+71,1560,1548
+72,1633,1622
+73,1710,1700
+74,1813,1804
+75,1962,1952
+76,1911,1902
+77,1784,1775
+78,1088,1081
+79,1141,1135
+80,1360,1354
+81,1279,1273
+82,1273,1268
+83,1191,1186
+84,1035,1031
+85,852,849
+86,856,853
+87,821,818
+88,758,755
+89,657,655
+90,558,556
+91,494,493
+92,419,418
+93,344,344
+94,273,273
+95,218,218
+96,171,171
+97,112,112
+98,81,81
+99,52,52
+100,87,87
diff --git a/lectures/japan_population_by_age.csv.yml b/lectures/japan_population_by_age.csv.yml
new file mode 100644
index 0000000..789514e
--- /dev/null
+++ b/lectures/japan_population_by_age.csv.yml
@@ -0,0 +1,89 @@
+# Manifest for japan_population_by_age.csv.
+
+filename: japan_population_by_age.csv
+title: Japan — population by single year of age, 2024
+description: >
+ Population of Japan by single year of age from 0 to "100 and over", as of
+ 1 October 2024, for the total population and for Japanese nationals. Counts
+ are in thousands, as published. Used as a first concrete example of a
+ probability mass function.
+
+# Constructed: the upstream workbook is Table 1 of the Population Estimates,
+# laid out for printing -- ages run down a left-hand block and continue in a
+# second block to its right, under two levels of merged headers. Flattening
+# that into one row per age is our processing.
+class: constructed
+
+source:
+ name: >
+ Population Estimates, Statistics Bureau of Japan, Ministry of Internal
+ Affairs and Communications
+ series: >
+ Table 1 — Population by age (single years) and sex, and sex ratio; total
+ population and Japanese population, as of 1 October 2024
+ url: https://www.stat.go.jp/english/data/jinsui/index.html
+ file_url: https://www.stat.go.jp/data/jinsui/2024np/zuhyou/05k2024-1.xlsx
+ citation: >
+ Statistics Bureau, Ministry of Internal Affairs and Communications, Japan.
+ Population Estimates, 1 October 2024.
+
+license:
+ name: Japan Statistics Bureau terms of use
+ url: https://www.stat.go.jp/english/index/copyright.html
+ redistribution: permitted
+ verified: 2026-08-04
+ # The Statistics Bureau permits reproduction of its published statistics with
+ # the source indicated, which the citation above does. Its terms follow the
+ # Japanese government standard open-data licence, compatible with CC BY.
+
+retrieved: 2026-08-04
+maintainer: QuantEcon
+
+# ---------------------------------------------------------------------------
+# Integrity
+# ---------------------------------------------------------------------------
+# Migration check (not a manifest field — a one-time repoint gate): the
+# japanese_population column of this file equals, value for value, the column
+# that lecture-python-intro's prob_dist.md previously read out of an e-Stat
+# export committed at
+# lectures/_static/lecture_specific/prob_dist/japan_population_by_age.xlsx,
+# so the repoint cannot change lecture output.
+#
+# That export was an on-demand e-Stat download with no recorded provenance --
+# the `builder_status: unrecovered` case AGENTS.md describes. Rather than
+# migrate it in that state, the same table was located at its publisher, the
+# Statistics Bureau, where it has a stable URL, and rebuilt from there. The
+# published figures are identical, so the lecture is unaffected and the data
+# gains a source, a citation and a builder.
+
+integrity:
+ sha256: 5bf07638cb75fa99e9dffb7ed2e272982a938384d49c9e1c4065117d696b81d8
+ upstream:
+ status: verified
+ date: 2026-08-04
+ against: 05k2024-1.xlsx
+ note: >
+ Written by re-running the committed builder against the Statistics
+ Bureau workbook on the retrieval date. The 2024 estimates are a closed
+ release, so re-running should reproduce this hash.
+
+schema:
+ format: csv
+ columns:
+ - {name: age, dtype: int64, description: single year of age; 100 means "100 and over"}
+ - {name: total_population, dtype: int64, description: total population at this age, in thousands}
+ - {name: japanese_population, dtype: int64, description: Japanese-national population at this age, in thousands}
+ row_count_floor: 101
+ date_range: {start: 2024, end: 2024}
+ known_nulls: {}
+ # Every figure is rounded to the nearest thousand independently, so the
+ # per-age counts sum to slightly more than the published totals: by 5
+ # thousand for the total population and 3 thousand for the Japanese
+ # population. The builder allows for this and asserts the sums are within 60.
+
+consumers:
+ - repo: QuantEcon/lecture-python-intro
+ file: lectures/prob_dist.md
+
+builder: scripts/japan_population_by_age.py
+builder_status: committed
diff --git a/migration.yml b/migration.yml
index 52d2b6e..6562adf 100644
--- a/migration.yml
+++ b/migration.yml
@@ -155,6 +155,20 @@ datasets:
date: 2026-08-04
cutover: null
+ japan_population_by_age.csv:
+ pilot: P5
+ status: landed
+ # Migrated out of lecture-python-intro, which held an e-Stat export with no
+ # recorded provenance. Rebuilt from the same table at its publisher, the
+ # Statistics Bureau, whose copy has a stable URL; the figures are identical.
+ # Flip to `repointed` when the lecture PR merges.
+ prior_pattern: local-path
+ landed:
+ pr: QuantEcon/data-lectures#33
+ date: 2026-08-04
+ repoints: []
+ cutover: null
+
# Planned waves that have not landed anything here yet. `datasets` names the
# files as the audit sees them today, so the dashboard can join the two views.
# `title` is the reader-facing milestone name (the dashboard is read by people
diff --git a/scripts/audit_annotations.yml b/scripts/audit_annotations.yml
index 73e4dfd..f682c01 100644
--- a/scripts/audit_annotations.yml
+++ b/scripts/audit_annotations.yml
@@ -115,10 +115,14 @@ datasets:
provenance: constructed-committed
note: make_data.py + README beside the data
japan_population_by_age.xlsx:
- description: Japan population by age group (prob_dist bimodality example)
+ description: Japan population by age (prob_dist PMF example)
provenance: author-assembled
- flags: [new-since-2026-07-15]
- note: added by lecture-python-intro#790; upstream source not yet recorded
+ note: >
+ an e-Stat export with no recorded provenance, added by
+ lecture-python-intro#790. Superseded by japan_population_by_age.csv,
+ which is rebuilt from the Statistics Bureau original and manifested
+ here; this entry goes when the lecture repoints and the xlsx leaves the
+ lecture repo
life-expectancy-vs-gdp-per-capita.csv:
description: Our World in Data, life expectancy vs GDP per capita
provenance: verbatim
diff --git a/scripts/japan_population_by_age.py b/scripts/japan_population_by_age.py
new file mode 100644
index 0000000..563e538
--- /dev/null
+++ b/scripts/japan_population_by_age.py
@@ -0,0 +1,102 @@
+"""
+Builder for lectures/japan_population_by_age.csv.
+
+Extracts the population of Japan by single year of age from Table 1 of the
+Population Estimates published by the Statistics Bureau of Japan, as of
+1 October 2024.
+
+The upstream workbook is laid out for printing: ages run down a left-hand block
+and continue in a second block to its right, under two levels of merged
+headers. This builder flattens that into one row per age.
+
+Counts are published in thousands and are left in those units, since that is
+what the source states. The final age, 100, means "100 and over".
+
+Stages: fetch -> pre-process -> validate -> write.
+"""
+
+import io
+import os
+import re
+import urllib.request
+
+import openpyxl
+import pandas as pd
+
+CURRENT_FILE_DIR = os.path.dirname(os.path.abspath(__file__))
+REPO_ROOT = os.path.dirname(CURRENT_FILE_DIR)
+PUBLISHED_DIR = os.path.join(REPO_ROOT, 'lectures')
+
+SOURCE_URL = 'https://www.stat.go.jp/data/jinsui/2024np/zuhyou/05k2024-1.xlsx'
+
+OUT_FILE = 'japan_population_by_age.csv'
+
+SHEET = '第1表'
+
+# (age label column, first data column) for the left and right blocks of the
+# printed table. Data columns run: total both sexes, male, female, sex ratio,
+# Japanese both sexes, male, female.
+BLOCKS = ((1, 2), (10, 11))
+TOTAL_OFFSET = 0 # total population, both sexes
+JAPANESE_OFFSET = 4 # Japanese-national population, both sexes
+
+# The "total" row of the source table, used to check that we read the columns
+# we think we did.
+PUBLISHED_TOTALS = {'total_population': 123802, 'japanese_population': 120296}
+
+# Every figure, including the total, is rounded to the nearest thousand
+# independently, so the per-age figures need not sum exactly to the published
+# total. 101 roundings of up to half a thousand each bound the gap at ~50;
+# the tolerance below sits inside that and far below any column shift.
+ROUNDING_TOLERANCE = 60
+
+
+def fetch():
+ with urllib.request.urlopen(SOURCE_URL) as response:
+ return openpyxl.load_workbook(io.BytesIO(response.read()))[SHEET]
+
+
+def pre_process(ws):
+ rows = []
+ for age_col, data_col in BLOCKS:
+ for r in range(1, ws.max_row + 1):
+ label = ws.cell(r, age_col).value
+ if label is None:
+ continue
+ # age labels look like "0 歳", "1", "100 歳以上"; the totals row
+ # and any notes carry no leading digits
+ match = re.match(r'^(\d+)', str(label).replace(' ', '').replace(' ', ''))
+ if not match:
+ continue
+ rows.append({
+ 'age': int(match.group(1)),
+ 'total_population': ws.cell(r, data_col + TOTAL_OFFSET).value,
+ 'japanese_population': ws.cell(r, data_col + JAPANESE_OFFSET).value,
+ })
+ return pd.DataFrame(rows).sort_values('age').reset_index(drop=True)
+
+
+def validate(df):
+ """Refuse to write anything that is not the shape we expect."""
+ assert list(df.columns) == ['age', 'total_population', 'japanese_population']
+ assert df['age'].tolist() == list(range(101)), 'ages are not 0 to 100'
+ assert not df.isnull().values.any()
+ # Column check: our per-age figures must sum to the published totals, up to
+ # the rounding slack described above. This catches a shifted column, which
+ # is the way this parse would fail.
+ for column, published in PUBLISHED_TOTALS.items():
+ gap = abs(df[column].sum() - published)
+ assert gap <= ROUNDING_TOLERANCE, \
+ f'{column} sums to {df[column].sum()}, published total is {published}'
+
+
+def run():
+ df = pre_process(fetch())
+ validate(df)
+ df.to_csv(os.path.join(PUBLISHED_DIR, OUT_FILE), index=False)
+ print(f'wrote {OUT_FILE}: {len(df)} rows, '
+ f'{df["total_population"].sum():,} thousand people')
+
+
+if __name__ == '__main__':
+ run()
diff --git a/scripts/render_audit.py b/scripts/render_audit.py
index 119f5b0..ee0c677 100644
--- a/scripts/render_audit.py
+++ b/scripts/render_audit.py
@@ -402,18 +402,24 @@ def pattern_barlist(audit: dict) -> str:
"""
-def what_changed() -> str:
- # The narrative diff vs the hand-built 2026-07-15 artifact. Static prose:
- # it describes a fixed historical comparison, not live state.
+def what_changed(audit: dict) -> str:
+ # The narrative diff vs the hand-built 2026-07-15 artifact. The prose is
+ # static -- it describes a fixed historical comparison -- but any count it
+ # quotes is taken from live state, so a landing wave cannot leave a stale
+ # number on the front page.
+ recs = (audit["migration"] or {}).get("datasets") or {}
+ n_re = sum(1 for r in recs.values() if r.get("status") in ("repointed", "final"))
return f"""
-
The audit was first taken by hand on {PREV_SNAPSHOT}. Regenerating from live -repo state two days later, the picture has already changed — which is why this dashboard is generated.
-The first two migration waves moved lingcod_msy_recovery.csv (was a
+
The audit was first taken by hand on {PREV_SNAPSHOT}. Everything below is +regenerated from live repo state, so this section follows the programme rather than freezing +a moment in it — which is why this dashboard is generated.
+The first waves moved lingcod_msy_recovery.csv (was a
Colab-breaking local path) and the pandas_panel trio realwage.csv /
countries.csv / employ.csv (5 of their 6 reads pointed at the retired
-legacy repo). Every consuming lecture now reads the central copy. Details on the
+legacy repo). Later waves landed the data behind the probability lectures. Every consuming
+lecture reads the central copy. Details on the
migration tracker.
The pre-MyST QuantEcon/lecture-python repo was renamed to
@@ -428,11 +434,14 @@ def what_changed() -> str:
intro's copies by URL), and its short_path fetches intro's committed
graph.txt over the network. That file was "shadowed dead weight" in the prior audit;
it is now load-bearing.
us_adult_heights.csv and japan_population_by_age.xlsx entered
-prob_dist as local-path reads — the pattern P1 just migrated away from. New data
-keeps arriving under the old conventions until the styleguide
-({issue_link("QuantEcon/QuantEcon.manual#108")}) lands.
prob_dist (lecture-python-intro#790) as local-path reads — the pattern P1 had just
+migrated away from. Both are now here, and in both cases the missing builder was recovered
+rather than recorded as a gap: the heights extract is reproduced byte for byte from NHANES,
+and the population table was rebuilt from its publisher, the Statistics Bureau, replacing an
+e-Stat export that carried no provenance. New data will keep arriving under the old
+conventions until the styleguide ({issue_link("QuantEcon/QuantEcon.manual#108")}) lands.