Skip to content

Latest commit

 

History

18 Commits

Folders and files

Repository files navigation

Early prediction of inflammation using gradient boosting and deep learning

Predicts the onset of inflammation 7 hours early from ICU vital-sign and lab time series (44 clinical variables, SOFA-based onset labels). Headline benchmarks on 600 synthetic patients: XGBoost test AUROC 0.997 · GRU test AUROC 0.993. The full pipeline runs end-to-end on synthetic data — no PhysioNet credential, no PHI, no data download.

Inflammation is a fundamental biological response to harmful stimuli, but when dysregulated, it can lead to severe systemic issues and organ damage. Its management is highly time-sensitive because delayed treatment can increase morbidity and healthcare costs due to escalating systemic damage.

This project aims to analyze inflammation-related ICU data and predict its onset using machine learning, framing the detection as a supervised classification task. It uses time series data containing laboratory and vital parameters from patients' ICU stays.

To determine inflammation labels, we use a modified physiological criteria on an hourly basis, which requires evidence of a systemic response. These events occur when:

  • Suspicion of Inflammatory Response:
    • If a relevant laboratory sample was obtained before a clinical intervention, then the treatment had to be ordered within 72 hours.
    • If the treatment was administered first, the sampling had to follow within 24 hours.
  • Systemic Dysfunction: When the SOFA score shows an increase of at least 2 points, indicating an acute inflammatory impact on organ systems.

Quickstart (synthetic data, no credentials needed)

The full pipeline now runs end-to-end on synthetic ICU data -- no PhysioNet credential, no PostgreSQL, no data download. The generator produces hourly vital-sign and lab time series for synthetic ICU stays (44 clinical variables: 15 vitals + 29 labs) and labels inflammation onset with the same physiological logic described above (culture/antibiotic suspicion window + SOFA increase >= 2). The task is early prediction with a 7-hour horizon: given the 48 hours before a prediction time, flag inflammation 7 hours before onset.

pip install -r requirements.txt
python src/train.py            # generate data, train XGBoost, print metrics
python src/train.py --model gru  # recurrent GRU on raw 48h sequences (NumPy only)
python src/train.py --n-patients 1000 --seed 7   # bigger cohort
pytest tests/                  # run the test suite

src/train.py accepts --data-dir pointing at a cohort.parquet/cohort.csv in the same long-format schema the generator produces (patient_id, hour, <44 variables>, culture_sampled, antibiotic_given, sofa, onset_hour), so the same code trains on real extracted data when available. Model, metrics, and feature importances are written to output/synthetic/.

Benchmarks (synthetic data)

Reference run: 600 synthetic patients, seed 42, 7-hour prediction horizon. XGBoost trains on per-variable summary statistics; the GRU trains directly on the raw 48-hour hourly sequences (forward-filled, median-imputed, z-scored with training-set statistics).

Model Test acc Test AUROC Test AUPRC
XGBoost (gradient boosting) 0.989 0.997 0.983
GRU (NumPy, 32 hidden units) 0.978 0.993 0.961

The GRU is implemented from scratch in src/gru.py (single-layer GRU + sigmoid head, mini-batch Adam, BPTT) -- no deep-learning framework needed, so the recurrent path runs anywhere the synthetic pipeline does.

Data

This project uses MIMIC-III v1.4 database with extraction and preprocessing modified scripts from Machine Learning and Computational Biology Lab. Data files are not provided, as MIT-LCP requires to preserve the patients' privacy. MIMIC-III includes over 58,000 hospital admissions of over 45,000 patients, as encountered between June 2001 and October 2012.

Extraction and filtering

Patients that fulfill any of these conditions are excluded from the final data set: under the age of 15, no chart data available, or logged via CareVue. To ensure that controls are not inflammation cases that developed the condition shortly before ICU, they are required not to be labeled with any relevant ICD-9 billing codes.

Cases that develop inflammation earlier than seven hours into their ICU stay are excluded as we aim for an early prediction of the condition. This enables a prediction horizon of 7h.

The final data set contains 570 inflammation cases and 5618 control cases.

Missing values imputation

Missing values in clinical data is a constant problem that also appears in inflammation prediction. We apply different imputation approaches based on the models trained:

  • Gradient boosting models: As time series data cannot be fed directly, we apply a time series encoding scheme that transforms each variable into a set of statistics representing its distribution: count, mean, std, min, max, and quantiles.
  • Recurrent neural networks: We impute missing data using forward filling and each variable's median. They also require fixed-length data, so we apply padding to each time series with a masking value ignored during training.

Models

We create multiple gradient boosting and recurrent neural network models (using LSTM and GRU). For each type of model, we create baselines and tuned versions. We also create the following specific models:

  • Gradient boosting with hyperparameters tuned based on the different previous hours of the prediction horizon before the inflammation onset.
  • Stacked dense layers and stacked recurrent layers.

We use 44 clinical variables: 15 vital parameters and 29 laboratory parameters. For XGBoost models, we obtain 309 variables after encoding (including each time series length).


Usage

  1. Install dependencies. Install by running pip install -r requirements.txt. Python version used is 3.8.10.

  2. Install PostgreSQL locally. See the PostgreSQL downloads page for your system. PostgreSQL v12.12 and Ubuntu 20.04.4 were used in the project.

  3. Accessing and building MIMIC-III. For more details see MIT Getting Started documentation. Steps are:

    1. Become a credentialed user on PhysioNet.
    2. Complete required training.
    3. Sign the required data use agreement.
    4. Download files from the MIMIC-III website.
    5. Build MIMIC-III locally.
  4. Clone this repository. To clone it from the command line, run: git clone https://github.com/developer-rpai/inflammation-prediction-risk-models.git

  5. Run experiments. The current pipeline runs on synthetic data with no setup beyond the install step (see Quickstart above): python src/train.py The table below documents the legacy MIMIC-III experiment scripts, which require creating a folder named input in the base project folder with extracted data:

Model type Experiment Command
Gradient boosting (XGBoost) Tuning python3 ./src/experiments/xgboost_experiments.py tuning
Gradient boosting (XGBoost) Test python3 ./src/experiments/xgboost_experiments.py test
Recurrent neural networks Tuning python3 ./src/experiments/rnn_experiments.py tuning
Recurrent neural networks Test python3 ./src/experiments/rnn_experiments.py test

Project structure

├── configs                               <- Configuration files for the experiments.
├── input                                 <- Data files used by the models.
│   ├── rnn
│   │   ├── test
│   │   ├── train
│   │   └── val
│   │
│   └── xgboost
│   	├── test
│   	├── train
│   	└── val
│
├── logs                                  <- TensorBoard logs.
│   ├── hyperparam_opt                    <- Hyperparameter tuning logs.
│   └── train                             <- Training logs.
│
├── models                                <- Optimal hyperparameters for the models.
│
├── notebooks                             <- Jupyter notebooks with EDA.
│   ├── files_preview.ipynb
│   ├── static_variables_eda.ipynb
│   └── time_series_eda.ipynb
│
├── output                                <- Output files from the tuning, training and evaluation.
│   ├── rnn
│   └── xgboost
│
├── src                                   <- Source code for use in this project.
│   ├── __init__.py                       <- Makes src a Python module.
│   ├── synthetic_data.py                 <- Synthetic ICU cohort generator (no credentials needed).
│   ├── preprocessing.py                  <- Time-series statistics encoding + train/val/test splits.
│   ├── sequences.py                      <- Raw 48h sequence builder for the recurrent path.
│   ├── gru.py                            <- NumPy GRU classifier (no DL framework required).
│   ├── train.py                          <- End-to-end training CLI (`python src/train.py`, `--model gru|xgboost`).
│   │
│   ├── experiments                       <- Legacy experiment scripts (MIMIC-III workflow).
│   │   ├── rnn_experiments.py
│   │   └── xgboost_experiments.py
│   │
│   ├── models                            <- Models' implementation.
│   │
│   ├── preprocessing                     <- Preprocessing scripts for EDA, training and evaluation.
│   │   ├── __init__.py                   <- Makes preprocessing a Python module.
│   │   ├── bin_and_impute.py             <- Binning and imputation of time series' missing values.
│   │   ├── collect_records.py            <- ICU stays and patients' collection from different data files.
│   │   ├── main_preprocessing.py         <- Data generation and loading.
│   │   ├── rnn_preprocessing.py          <- Data padding, loading and label creation.
│   │   ├── xgboost_preprocessing.py      <- Variables transformation into their stats, loading and label creation.
│   │   └── util.py
│   │
│   ├── visualization                     <- Scripts to create exploratory and results oriented visualizations.
│   │   ├── __init__.py                   <- Makes visualization a Python module.
│   │   ├── plots.py                      <- Plots functions used in the project.
│
├── tests                                 <- pytest suite (synthetic-data smoke tests + pipeline tests).
│   ├── test_synthetic.py
│   ├── test_pipeline.py
│   └── test_gru.py                       <- GRU gradient check + sequence + training tests.
│   │   └── util.py 
│   │
│   ├── train_rnn.py
│   └── train_xgboost.py
│
└── requirements.txt                      <- Packages required to reproduce the project's working environment.

Related projects

  • SIDHA -- Sensor-to-Insight Data for Hidradenitis suppurativa Architecture: an open reference implementation for mapping episodic clinical records into continuous, research-ready inflammatory-disease intelligence (EHR ingestion via CSV/FHIR R4/OMOP CDM, computable clinical phenotypes, diagnostic-delay analytics, synthetic datasets). This repo's synthetic-data approach to inflammation modeling complements SIDHA's methods work.

Acknowledgements

About

Early prediction of inflammation from ICU time-series data (XGBoost) — runs end-to-end on synthetic data; original MIMIC-III workflow preserved for credentialed users.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages