From 1971fcce901b8d5ee969f88898e19ecbea865a89 Mon Sep 17 00:00:00 2001 From: Matt McKay Date: Mon, 17 Aug 2026 08:45:12 +1000 Subject: [PATCH] Correct two manifest records and pin the builder environment MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Three defects the #84 validation found in records written during waves B1' and B2', plus the schema/practice gap bundled with them. The 1983 manifest described consumption_per_capita as "Observed 2.14e-4 to 2.66e-4". Re-derived from the committed bytes with the stdlib csv module: the minimum is 2.1254904619e-4 at 1960-12-31, which rounds to 2.13e-4, and the second-smallest value rounds there too — so 2.14e-4 matched no observed value. Both bounds now carry the date they occur on. Both hansen manifests quoted the Ken French zip header for the T-bill leg's lineage but stopped mid-sentence, inside quotation marks. Re-fetched 2026-08-17, unchanged at the 202606 CRSP vintage; the quote is now complete and each manifest says plainly that the ICE BofA clause applies to no row in a sample ending 1978. The issue also expected this quote in a builder docstring — it is not there; both builders reference the zip without quoting its header, so the manifests were the only place to fix. requirements.txt pinned wbgapi, PyYAML and scipy but not pandas, which 9 of the 10 builders import. Both hansen manifests carry integrity.upstream.status: verified, which was therefore measured against an unrecorded ambient environment. Pinned at 2.3.3 — the version they were verified against — with a comment recording that they also reproduce byte-identically under 3.0.5, so the pin documents the measurement rather than guarding a known sensitivity. manifest-schema.yml omitted source.doi (16 of 33 manifests), source.version (10), source.note (20) and license.note (16). All four are documented, and the header now states what the file is: a record of the fields that recur, not a closed schema, with extensions expected and documented back once they earn their place. That gives a field-by-field conformance check a defined answer and names the two existing extensions (schema.sheets, schema.read_as). Verified: all 33 manifests and the schema parse; CATALOG.md regenerates unchanged, so no consumer of these fields exists yet. Closes #85 Co-Authored-By: Claude Opus 5 (1M context) --- lectures/hansen_singleton_1982_data.csv.yml | 9 ++++--- lectures/hansen_singleton_1983_data.csv.yml | 13 +++++---- manifest-schema.yml | 30 +++++++++++++++++++++ requirements.txt | 9 +++++++ 4 files changed, 53 insertions(+), 8 deletions(-) diff --git a/lectures/hansen_singleton_1982_data.csv.yml b/lectures/hansen_singleton_1982_data.csv.yml index 970215c..61a607f 100644 --- a/lectures/hansen_singleton_1982_data.csv.yml +++ b/lectures/hansen_singleton_1982_data.csv.yml @@ -40,9 +40,12 @@ source: version: > Ken French research factors, zip header "This file was created using the 202606 CRSP database. The 1-month TBill rate data until 202405 are from - Ibbotson Associates" — read 2026-08-13. FRED series are unversioned and - revisable; the 1959-1978 window used here has been stable across every - check so far (see integrity.upstream). + Ibbotson Associates. Starting from 202406, the 1-month TBill rate is from + ICE BofA US 1-Month Treasury Bill Index." — read 2026-08-13, re-read + unchanged 2026-08-17. The ICE BofA clause is quoted for completeness and + applies to no row here; this file's sample ends in 1978. FRED series are + unversioned and revisable; the 1959-1978 window used here has been stable + across every check so far (see integrity.upstream). series: > FRED CNP16OV (civilian noninstitutional population 16+, thousands of persons, BLS Employment Situation); diff --git a/lectures/hansen_singleton_1983_data.csv.yml b/lectures/hansen_singleton_1983_data.csv.yml index 1f5e5c8..84e99d7 100644 --- a/lectures/hansen_singleton_1983_data.csv.yml +++ b/lectures/hansen_singleton_1983_data.csv.yml @@ -41,10 +41,13 @@ source: version: > Ken French research factors, zip header "This file was created using the 202606 CRSP database. The 1-month TBill rate data until 202405 are from - Ibbotson Associates" — read 2026-08-13. That Ibbotson clause covers the - whole of this file's sample, since it ends in 1978. FRED series are - unversioned and revisable; the 1959-1978 window used here has been stable - across every check so far (see integrity.upstream). + Ibbotson Associates. Starting from 202406, the 1-month TBill rate is from + ICE BofA US 1-Month Treasury Bill Index." — read 2026-08-13, re-read + unchanged 2026-08-17. The Ibbotson clause covers the whole of this file's + sample, since it ends in 1978; the ICE BofA clause is quoted for + completeness and applies to no row here. FRED series are unversioned and + revisable; the 1959-1978 window used here has been stable across every + check so far (see integrity.upstream). series: > FRED CNP16OV (civilian noninstitutional population 16+, thousands of persons, BLS Employment Situation); @@ -158,7 +161,7 @@ schema: - {name: gross_real_return, dtype: float64, description: "gross real monthly return on the aggregate stock market: (1 + (Mkt-RF + RF)/100) divided by gross_inflation_cons. A ratio near 1 (observed 0.865-1.159), not a percentage and not a net return"} - {name: gross_cons_growth, dtype: float64, description: "gross monthly growth of real per-capita nondurable consumption — the ratio of consecutive months of consumption_per_capita. Observed 0.967-1.033"} - {name: gross_inflation_cons, dtype: float64, description: "gross month-over-month inflation of the PCE nondurables price deflator (DNDGRG3M086SBEA); the deflator applied to both nominal return legs. Observed 0.995-1.024"} - - {name: consumption_per_capita, dtype: float64, description: "real per-capita nondurable consumption, DNDGRA3M086SBEA / CNP16OV. A LEVEL, and the one column with no interpretable unit: it divides a chain-type index (2017=100) by a population in thousands, so it lands near 2.4e-4 and is meaningful only through its growth rate. Observed 2.14e-4 to 2.66e-4"} + - {name: consumption_per_capita, dtype: float64, description: "real per-capita nondurable consumption, DNDGRA3M086SBEA / CNP16OV. A LEVEL, and the one column with no interpretable unit: it divides a chain-type index (2017=100) by a population in thousands, so it lands near 2.4e-4 and is meaningful only through its growth rate. Observed 2.13e-4 (1960-12-31) to 2.66e-4 (1973-02-28)"} - {name: gross_real_tbill, dtype: float64, description: "gross real monthly return on one-month Treasury bills: (1 + RF/100) divided by gross_inflation_cons. Observed 0.984-1.008. Exceeded by gross_real_return in only 54% of months, though the mean excess is positive"} row_count_floor: 239 # exact by design, not a floor with headroom: # the sample is the paper's 1959:2-1978:12 and diff --git a/manifest-schema.yml b/manifest-schema.yml index cc69f42..0b7e84f 100644 --- a/manifest-schema.yml +++ b/manifest-schema.yml @@ -18,6 +18,16 @@ # # The manifests are the source for the generated catalog page (PLAN Phase 2), # which doubles as the public dataset registry. +# +# COMPLETENESS: this file documents the fields that recur across the corpus. It +# is not a closed schema and nothing validates against it — a dataset whose +# provenance needs a field not shown here should add one rather than distort +# itself to fit, and a field that earns its place in several manifests should +# then be documented back here. Two such extensions are already in wide use and +# are described where they apply: `schema.sheets` for multi-sheet workbooks, and +# `schema.read_as` for the pandas read-kwargs a positional read needs. Adding +# the four `source`/`license` fields below closed the last known gap between +# this file and practice (QuantEcon/data-lectures#85). # --------------------------------------------------------------------------- # Identity @@ -43,9 +53,23 @@ source: name: World Bank national accounts data, and OECD National Accounts data files series: NY.GDP.MKTP.KD.ZG url: https://data.worldbank.org/indicator/NY.GDP.MKTP.KD.ZG + # Dataset DOI where the source issues one, else null. Prefer a DOI that + # resolves to the DATA deposit; an article DOI is better than nothing, but say + # which it is. In use by 16 of 33 manifests. + doi: null + # The upstream's own version identifier, quoted as the source states it — an + # edition ("Maddison Project Database 2020"), a vintage stamp, or a file + # header. Null where the source is genuinely unversioned; say so rather than + # inventing a version. In use by 10. + version: null citation: > World Bank, World Development Indicators, series NY.GDP.MKTP.KD.ZG (GDP growth, annual %). + # Anything a reader must know before trusting these bytes that the fields + # above cannot carry — most often that the file is NOT what its name or its + # cited paper implies. Where that is the case, say it here AND at the top of + # the manifest. In use by 20. + note: null license: name: CC BY-4.0 @@ -61,6 +85,12 @@ license: # date is not evidence. P1 found the # authority is often not the project # homepage but its Zenodo/DOI record. + # What `redistribution` rests on, whenever that is anything less obvious than + # a named public licence: a split answer across providers, an inherited + # exposure, a term found only in prose, or a reasoned judgement about what is + # actually published here. A bare `permitted` beside a `name: null` is the + # case that most needs this field. In use by 16. + note: null retrieved: 2024-04-10 # ISO date the bytes were obtained, or # null for inherited-undated files — diff --git a/requirements.txt b/requirements.txt index 481fb1f..528fb73 100644 --- a/requirements.txt +++ b/requirements.txt @@ -1,4 +1,13 @@ wbgapi==1.0.12 +pandas==2.3.3 # imported by 9 of the 10 builders — every one except + # builders/business_cycle.py. Pinned so an + # integrity.upstream.status of `verified` names a + # reproducible environment rather than whatever python is + # ambient (QuantEcon/data-lectures#85). This is the + # version both hansen builders were verified against on + # 2026-08-13; they also reproduce byte-identically under + # 3.0.5, so the pin records the measurement rather than + # guarding a known sensitivity. PyYAML==6.0.3 # scripts/build_catalog.py — parse the sidecar manifests scipy==1.16.3 # builders/NEWQDATA.py — loadmat, the only way to read a # MATLAB 5.0 .MAT; pandas cannot