In this article we are going to bring a complete ML pipeline that produces a trained Decision Tree model about Heart Failure Prediction, using a set of tools that allow a better management of the data flow (MLFlow, Hydra, anaconda and Wandb).
Heart Failure Prediction: Analysis and modeling. There are some factors that affects Death Event. This dataset contains person's information like age, sex, blood pressure, smoke, diabetes, ejection fraction, creatinine phosphokinase, serum creatinine, serum sodium, time and we have to predict their DEATH EVENT.
The starter kit contains all the steps we have previously completed, only slightly modified to work better together.
A few notes and instructions:
-
When chaining together the steps, the output artifact of a step should be the input artifact of the next one (when applicable). Also use the
artifact_typeoptions so that the final visualization of the pipeline highlights the different steps. For example, you can useraw_datafor the artifact containing the downloaded data,preprocessed_datafor the artifact containing the data after the preprocessing, and so on. -
For testing, set the
project_nametomlops-final-project. Once you are done developing, do a production run by changing theproject_nametoheart_failure_classification_prod. This way the visualization of the pipeline will not contain all your trials and errors. Remember to tag the produced model export asprod(we are going to use it in the next exercise) -
When developing, you can override the parameter
main.execute_stepsto only execute one or more steps of the pipeline, instead of the entire pipeline. This is useful for debugging. For example, this only executes thedecision_treestep:mlflow run . -P hydra_options="main.execute_steps='decision_tree'"
and this executes
downloadandpreprocess:mlflow run . -P hydra_options="main.execute_steps='download,preprocess'"