Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -209,13 +209,15 @@ notebooks/404_*
/data/emissions/primap-202503/Guetschow_et_al_2025-PRIMAP-hist_v2.6.1_final_no_rounding_13-Mar-2025.nc
/data/emissions/cmip7-historical-2025.12.07/country-history.feather
/data/emissions/cmip7-historical-2025.12.07/global-workflow-history.csv
/data/emissions/gcb-2025/GCB2025v15_MtCO2_flat.csv
/data/gdp/wdi-2025/API_NY.GDP.MKTP.KD_DS2_en_csv_v2_213435.csv
/data/gdp/wdi-2025/API_NY.GDP.MKTP.PP.KD_DS2_en_csv_v2_1004.csv
/data/gini/unu-wider-2025/WIID-29APR2025.xlsx
/data/gini/wdi-2025/API_SI.POV.GINI_DS2_en_csv_v2.csv
/data/population/un-owid-2025/population.csv
/data/population/un-owid-2025/UN_PPP2024_Output_PopTot.xlsx
/data/lulucf/melo-2026/timeseries_NGHGI_v3.1.csv
/data/lulucf/melo-2026-v4/timeseries_NGHGI_gap-filled_4.0.0.csv
/data/bunkers/gcb-2024/National_Fossil_Carbon_Emissions_2024v1.0.xlsx
/data/scenarios/ipcc_ar6_gidden/ar6_gidden.xlsx
/data/scenarios/ipcc_ar6_gidden/ar6_gidden.zip
Expand Down
17 changes: 17 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,12 @@ Format: [Keep a Changelog](https://keepachangelog.com/en/1.1.0/). Versions: [Sem
- `fair-shares fetch-data`: downloads every input dataset from source, verifies pinned checksums, records provenance in `data/PROVENANCE.md`. Missing files also fetch automatically on first use.
- Python API reference (`docs/api/python-api.md`) for pip-only users.
- Root `CONTRIBUTING.md`. This changelog.
- LULUCF Data Hub v4.0.0 (gap-filled NGHGI, 187 countries, 2000-2024) as the opt-in LULUCF source `melo-2026-v4`. `melo-2026` (v3.1.1) stays the default.
- Global Carbon Budget 2025 national fossil CO2 (1850-2024) as the opt-in emissions source `gcb-2025`, `co2-ffi` only. Its world row excludes international bunkers and the 1991 Kuwaiti oil fires, and its runs take bunkers from the same file. The pipeline asks an emissions source only for the categories it declares. `primap-202503` stays the default.
- RCB sources `forster_2026` (IGCC 2025, from 2026) and `ar6_wg1_2021` (AR6 WGI Table SPM.2, from 2020), each with 1.5, 1.7 and 2°C budgets. Their deductions use AR6 scenarios in a peak-warming band (`scenario_selection: peak-warming-band` in `rcbs.yaml`). The band keeps only scenarios that reach net-zero CO2 by 2100, because a remaining carbon budget runs to net-zero CO2 (14, 111 and 37 of 14, 131 and 69 scenarios at 1.5, 1.7 and 2°C). A budget label without a scenario set and an empty band raise an error. A rebase with missing emission years raises an error, and the pipeline skips that source with a warning. The band rule leaves the budgets of the existing sources unchanged; the bunker rebase under "Fixed" changes them.
- Opt-in coverage rule for analysis countries: a `coverage` block on an emissions source (`emissions_recorded_before`, `population_from`). `gcb-2025` sets it to 1990 and 1850, so a country needs an emissions record before 1990 and population from 1850. Eight countries with their first GCB record in 1990 or later (AND, FSM, LSO, MHL, NAM, PLW, TLS, TUV) and Macao (population from 1950) join rest-of-world, which leaves 168 analysis countries in a `co2-ffi` run. `country_data_coverage_summary.csv` names the failed test in `coverage_rule_failed`, and notebook 101 saves `emiss_co2-ffi_first_recorded_year.csv`. Sources without the block, including `primap-202503`, keep their country set and outputs.
- `rebase_fill_max_years` in `rcbs.yaml` (default 1). When a world series of the rebase (fossil CO2, bunkers, LULUCF) ends before the year before the budget baseline, each later year takes the last observed value, up to this number of years, with a warning that names the source, the series, the years and the value. The fill is a placeholder until observed data are published. With `gcb-2025` (observed to 2024) `forster_2026` rebases with 2025 filled. With PRIMAP (observed to 2023) it needs two years and stays skipped. 0 turns the fill off.
- Emission category `co2` with a `co2-ffi`-only emissions source and an active LULUCF source, for example `gcb-2025` with `melo-2026-v4` and target `rcbs`. Notebook 107 derives `co2` as `co2-ffi` plus national-inventory LULUCF and skips the non-CO2 categories that such a source cannot supply. Without a LULUCF source the configuration still raises an error.
- World Bank WDI Gini index (`SI.POV.GINI`) as a Gini source, with `analysis/gini_source_comparison.py` reporting the coverage and value differences against WIID.

### Changed
Expand All @@ -19,6 +25,7 @@ Format: [Keep a Changelog](https://keepachangelog.com/en/1.1.0/). Versions: [Sem
- **Breaking:** the default Gini source is now World Bank WDI (`wdi-2025`), which is CC-BY-4.0. WIID stays available as `active_gini_source=unu-wider-2025` but is opt-in, and outputs built on it cannot be redistributed under CC BY 4.0. Output directory names change, so existing WIID runs are not overwritten. The two sources give materially different capability-based allocations — WDI/PIP is consumption-based for most low- and middle-income countries.
- **Breaking:** analysis-country membership no longer depends on Gini availability. A country with complete emissions, GDP and population is now in the analysis even without a Gini value, and receives the analysis-country mean (`general.gini_missing_policy: fallback-mean`, or `strict` to refuse). The country set grows by 9 on the standard sources; `country_data_coverage_summary.csv` gains a `gini_imputed` column.
- Gini source config replaces the unused `world_key` and `gini_year` keys with `selection` and `year_window`, both of which the notebooks read.
- Allocation years start at 1850 (was 1900). With `capability_reference_year` set, the adjusted budget approaches need GDP at the reference year only, so `allocation_year` can precede the GDP series. For `co2` and `all-ghg`, the parameter grid checks `pre_allocation_responsibility_year` against the first year of the responsibility emissions frame it receives.

### Deprecated

Expand All @@ -30,8 +37,18 @@ Format: [Keep a Changelog](https://keepachangelog.com/en/1.1.0/). Versions: [Sem

### Fixed

- `compose_config` did not declare the packaged `data_sources_unified.yaml` as an input, so an existing output folder reused its stale `config.yaml` after a source changed (for example a new `coverage` block) until the run used `--forcerun compose_config`. The rule now re-runs when that file changes.
- Global Carbon Budget citation pointed at a DOI that does not resolve; now the correct paper DOI (10.5194/essd-17-965-2025) plus data-product DOI (10.18160/GCP-2024).
- Licence statements corrected: WIID is CC BY-NC-SA 3.0 IGO, UN/OWID population is mixed-terms, CMIP7 is CC-BY-SA-4.0 (author-confirmed).
- The Python API's preprocessing path wrote `country_gini_stationary.csv` without a Rest-of-World row, unlike the notebook path. Both now use the same code.
- Notebook `100_data_preprocess_rcbs` passed the removed `project_root=` argument to `load_and_process_rcbs`, so the RCB pipeline could not run.
- Adjusted budgets of RCB sources with a baseline after 2020 were too low. The rebase to 2020 added world fossil emissions, which exclude international bunkers, and the bunker deduction covered 2020 to net zero, so the bunkers of 2020 to the year before the baseline were deducted and never added. The rebase now adds them (new output column `rebase_bunkers_mt`). Every adjusted budget of `lamboll_2023` (baseline 2023) rises by 2,822 Mt CO2 and every adjusted budget of `forster_2024` (baseline 2024) by 3,959 Mt CO2, for `co2-ffi` and `co2`, with PRIMAP emissions and `gcb-2024` bunkers. `ar6_2020` (baseline 2020) is unchanged. `process_rcb_to_2020_baseline` takes the bunker series as `world_bunker_emissions`.
- Adjusted `co2` budgets of RCB sources with a baseline after 2020 were too low. The rebase to 2020 adds observed LULUCF in the national-inventory (NGHGI) convention for 2020 to the year before the baseline, and the BM-to-NGHGI convention gap also started in 2020, so the conversion for those years was applied twice. The convention gap now runs from the baseline year of the source to net zero (Weber et al. 2026, Eqs. 2-3). Notebook 104 writes one median gap per baseline year (`convention_gap_median_from` in `rcb_scenario_adjustments.yaml`, which replaces `convention_gap_median`), so existing outputs need a re-run of notebook 104. With PRIMAP emissions and `melo-2026` LULUCF, the `1.5p50` budget of `lamboll_2023` rises from 222,434 to 241,477 Mt CO2 and that of `forster_2024` from 206,034 to 230,905 Mt CO2. `ar6_2020` (baseline 2020) and every `co2-ffi` budget are unchanged. The argument `actual_bm_lulucf_emissions` of `load_and_process_rcbs` and `process_rcb_to_2020_baseline` is now `world_nghgi_lulucf_emissions`: the series is the NGHGI world row, and the pipeline holds no observed bookkeeping-model LULUCF series.
- `build_nghgi_world_co2_timeseries` subtracted bunkers from a world fossil series that already excludes them (871 Mt CO2 in 2020; 20,770 Mt over 2000-2019). World `co2` is now fossil plus LULUCF, as in notebook `100_data_preprocess_rcbs`, and the function no longer takes `bunker_ts`. The error reached `run_rcb_preprocessing` and notebook `106_generate_pathways_from_rcbs`; budget files written by notebook 100 did not contain it.
- `compute_bunker_deduction` treated 2023 as the last observed bunker year for every source. It now reads the last year from the bunker data, so `gcb-2025` runs use the observed 2024 value and extrapolate the 2024 rate. Runs with `gcb-2024` bunkers give the same deduction as before.
- `calculate_budget_from_rcb` summed only the years present when the allocation year preceded the first year of the world emissions series (for example `co2`, which starts in 2000, with allocation year 1990). It now raises an error that names the missing years and the first available year.
- `calculate_budget_from_rcb` raised no error when the allocation year was after 2020 and the world emissions ended earlier, and summed only the years present. It now raises an error that names the missing years and the last available year. A year with a missing value counts as missing in this check, in the pre-2020 check and in the rebase to 2020, which earlier summed NaN-padded years as zero.
- For `co2`, the pipeline could run the scenario notebook (104) before the LULUCF notebook (107) and then write a zero convention gap. The scenario rule now depends on notebook 107. `process_rcb_to_2020_baseline` raises an error for a `co2` budget with a baseline after 2020 when it receives no inventory LULUCF series; it used a zero LULUCF rebase term before. A budget with baseline 2020 needs no series.
- `import fair_shares.library.config` and `import fair_shares.library.config.models` failed with a circular import when they were the first import in a process.
- A budget label without a numeric temperature under the peak-warming band rule raised a bare `ValueError`; it now raises a `ConfigurationError` that names the label. `rebase_fill_max_years` must be a non-negative integer.
- The Gini-adjusted approaches ignored Gini entirely when `capability_reference_year` named a year before `allocation_year`: the capability snapshot was read from the unfiltered inputs without the adjustment, on both the budget and pathway side. Results on that parameterisation change — on the standard sources, China's share of a `per-capita-adjusted-gini-budget` allocation moves from 10.8% to 5.3% and India's from 23.2% to 28.4%.
34 changes: 28 additions & 6 deletions Snakefile
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,9 @@
# Import shared logic from Python modules (single source of truth)
from fair_shares.library.utils.data.config import (
build_source_id,
get_bunkers_source,
get_emission_preprocessing_categories,
get_emissions_data_parameters,
get_final_categories,
get_co2_component,
is_composite_category,
Expand Down Expand Up @@ -95,8 +97,13 @@ NOTEBOOK_DIR = "notebooks"
# Two category lists drive the pipeline:
# EMISSION_CATEGORIES — what PRIMAP extraction (notebook 101) produces
# FINAL_CATEGORIES — what the allocation loop iterates over
# The emissions source is asked only for categories it declares.
_target = active_target_source or ""
EMISSION_CATEGORIES = get_emission_preprocessing_categories(_target, emission_category)
EMISSION_CATEGORIES = get_emission_preprocessing_categories(
_target,
emission_category,
get_emissions_data_parameters(active_emissions_source).get("available_categories"),
)
FINAL_CATEGORIES = get_final_categories(_target, emission_category)
is_multi_category = needs_decomposition(_target, emission_category)

Expand Down Expand Up @@ -143,6 +150,8 @@ _needs_lulucf = emission_category in ("co2", "co2-lulucf", "all-ghg")
# Bunker data is needed for all non-pathway targets (RCBs must subtract
# international bunker emissions before country allocation).
_needs_bunkers = _allocation_mode != "pathway"
# The emissions source names its bunker data; the default is gcb-2024.
_bunkers_source = get_bunkers_source(active_emissions_source)


if _needs_lulucf and active_lulucf_source is None:
Expand All @@ -153,6 +162,13 @@ if _needs_lulucf and active_lulucf_source is None:
f"Example: --config ... active_lulucf_source=melo-2026"
)

# Historical series that the scenario rule declares as input. Notebook 107
# derives co2 when the emissions source declares only co2-ffi.
if emission_category == "co2" and "co2" not in EMISSION_CATEGORIES:
_scenario_emissions_input = f"{OUTPUT_DIR}/intermediate/emissions/emiss_co2_nghgi_timeseries.csv"
else:
_scenario_emissions_input = f"{OUTPUT_DIR}/intermediate/emissions/emiss_{emission_category}_timeseries.csv"

# Resolve scenario source: per-target override → global default
_scenario_source_key = (
_target_yaml.get("scenario_source")
Expand Down Expand Up @@ -274,6 +290,9 @@ rule compose_config:
All validation logic is in config/models.py (Pydantic).
The Snakefile only does minimal checks — Pydantic does comprehensive validation.
"""
input:
# build_data_config composes the output config from this packaged file
sources_yaml=str(packaged_config("data_sources/data_sources_unified.yaml")),
output:
config=f"{OUTPUT_DIR}/config.yaml",
params:
Expand Down Expand Up @@ -405,6 +424,7 @@ if _needs_lulucf:
notebook=f"{OUTPUT_DIR}/notebooks/107_derive_nghgi_categories_{active_lulucf_source}.ipynb",
nghgi_world=f"{OUTPUT_DIR}/intermediate/emissions/world_co2-lulucf_timeseries.csv",
nghgi_metadata=f"{OUTPUT_DIR}/intermediate/emissions/lulucf_metadata.yaml",
nghgi_co2=f"{OUTPUT_DIR}/intermediate/emissions/emiss_co2_nghgi_timeseries.csv",
shell:
notebook_cmd("{input.notebook}", "{output.notebook}")

Expand All @@ -418,10 +438,10 @@ if _needs_bunkers:
allocation. Independent of LULUCF — uses GCB fossil emissions data.
"""
input:
notebook=f"{NOTEBOOK_DIR}/108_data_preprocess_bunkers_gcb-2024.ipynb",
notebook=f"{NOTEBOOK_DIR}/108_data_preprocess_bunkers_{_bunkers_source}.ipynb",
config=f"{OUTPUT_DIR}/config.yaml",
output:
notebook=f"{OUTPUT_DIR}/notebooks/108_data_preprocess_bunkers_gcb-2024.ipynb",
notebook=f"{OUTPUT_DIR}/notebooks/108_data_preprocess_bunkers_{_bunkers_source}.ipynb",
bunker_csv=f"{OUTPUT_DIR}/intermediate/emissions/bunker_timeseries.csv",
shell:
notebook_cmd("{input.notebook}", "{output.notebook}")
Expand All @@ -446,7 +466,7 @@ if uses_scenarios:
input:
notebook=f"{NOTEBOOK_DIR}/{_scenario_nb_stem}.ipynb",
config=f"{OUTPUT_DIR}/config.yaml",
emissions_data=f"{OUTPUT_DIR}/intermediate/emissions/emiss_{emission_category}_timeseries.csv",
emissions_data=_scenario_emissions_input,
lulucf_notebook=(f"{OUTPUT_DIR}/notebooks/107_derive_nghgi_categories_{active_lulucf_source}.ipynb" if _needs_lulucf else []),
output:
notebook=f"{OUTPUT_DIR}/notebooks/{_scenario_nb_stem}.ipynb",
Expand All @@ -459,7 +479,7 @@ if uses_scenarios:
input:
notebook=scenario_notebook,
config=f"{OUTPUT_DIR}/config.yaml",
emissions_data=f"{OUTPUT_DIR}/intermediate/emissions/emiss_{emission_category}_timeseries.csv",
emissions_data=_scenario_emissions_input,
lulucf_notebook=(f"{OUTPUT_DIR}/notebooks/107_derive_nghgi_categories_{active_lulucf_source}.ipynb" if _needs_lulucf else []),
bunker_csv=(f"{OUTPUT_DIR}/intermediate/emissions/bunker_timeseries.csv" if _needs_bunkers else []),
scenario_adjustments=f"{OUTPUT_DIR}/intermediate/scenarios/rcb_scenario_adjustments.yaml",
Expand All @@ -476,7 +496,9 @@ if uses_scenarios:
input:
notebook=scenario_notebook,
config=f"{OUTPUT_DIR}/config.yaml",
emissions_data=f"{OUTPUT_DIR}/intermediate/emissions/emiss_{emission_category}_timeseries.csv",
emissions_data=_scenario_emissions_input,
# Notebook 107 runs first: it writes the NGHGI series that 104 reads.
lulucf_notebook=(f"{OUTPUT_DIR}/notebooks/107_derive_nghgi_categories_{active_lulucf_source}.ipynb" if _needs_lulucf else []),
output:
notebook=scenario_nb_out,
scenarios=f"{OUTPUT_DIR}/intermediate/scenarios/scenarios_{emission_category}_timeseries.csv",
Expand Down
21 changes: 21 additions & 0 deletions data/emissions/gcb-2025/CITATION.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
# Citation

## Global Carbon Budget 2025: National Fossil CO2 Emissions

**Data product (the file in this directory):**

Andrew, R. M., & Peters, G. P. (2025). *The Global Carbon Project's fossil CO2 emissions dataset* (2025v15) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.17417124

**Paper (methods and results):**

Friedlingstein, P., et al. (2026). Global Carbon Budget 2025. *Earth System Science Data*, 18, 3211-3288. https://doi.org/10.5194/essd-18-3211-2026

Cite both when reporting values derived from this file.

## Licence

CC BY 4.0 for both the data product and the paper.

## Files

- `GCB2025v15_MtCO2_flat.csv`: territorial fossil CO2 emissions by country, 1750-2024, long format. Units: MtCO2/yr. Also holds international shipping (`XIS`), international aviation (`XIA`), the 1991 Kuwaiti oil fires and the global total (`WLD`).
Loading
Loading