This is a reference application for using DVC for data versioning, experiment tracking and model versioning for small to medium datasets. The scope of this project is to introduce the usage of DVC, show how to use DVC for your own project and give an example implementation for comparison and testing.
For more information see the official DVC documentation: https://dvc.org/doc/start.
python -m venv .venv && source .venv/bin/activatepip install -r requirements.txtgit init && echo ".venv" > .gitignore && dvc init && git commit -m "Initialize git and DVCdvc add "path/to/file.parquet" ; git add . ; git commit -m "Add initial dataset to DVC" # Add your datapath herepython update_stock_data.py--> you are now set to train a model with (locally) versioned data
with Live() as live:
live.log_param("epochs", NUM_EPOCHS)
for epoch in range(NUM_EPOCHS):
train_model(...)
metric_name = evaluate_model(...) # your training code here
live.log_metric("metric_name", metric_name)
live.next_step()
torch.save(model.state_dict(), "models/model.pth")
live.log_artifact("model.pth", type="model")Example: Train transformer model from downloaded stock data (note: this is just an example and not a good performing model)
python train_transformer_model.pyNote: Git commit also handles dvc commit for dvc added data, which is done automatically with the usage of dvclive
--> These steps enable you to do manual experiments using the examples update_stock_data.py and train_transformer_model.py, utilizing DVC to control data versioning, experiment tracking and model versioning
- Download the DVC extension for VSCode
- With this extension, you can see and compare the DVC Experiments, show plots, restore models, datasets and parameters to previous experiments
Note: DVC stores the models in the DVC cache. This can also be used with an external bucket for storing the files (see below). When the experiment is commited and pushed, it is tracked in the bucket and can be utilized by all users (and after cleaning the cache). The data, models and experiments that are not commited and pushed can only be used locally, which means they are lost when the cached is cleaned.
- parameterize the pipeline by using a params.yaml file
- modularize pipeline steps in different files
- adapt the code to utilize the params.yaml contents
dvc stage add -n stagename ...adds a pipeline stage to the dvc.yaml file including the stage "train" with its parameters from the params.yaml file, dependencies and ouptuts as well as a command to run the pipeline
Example with the modified train.py:
dvc stage add -n train \
--params base,train \
--deps train.py --deps data/raw \
--outs models/model.pth \
python src/train.pydvc dagvisualizes the pipeline in the terminal
dvc exp runruns the pipeline from the dvc.yaml file and captures the state of the workspace as DVC experimentdvc exp run --name "batch-size_8" --set-param "train.batch_size=8"changes the batch_size to 8 and runs the experimentdvc exp run --name "batch-size-tryout" --queue -S "train.batch_size=8,16,32"queues 3 experiments with different batch sizes, to run:dvc exp run --run-all
dvc remote add -d remotename url:port/pathadds config for remote datastore locationdvc remote add -d remote_name ssh://login@serverip/path/to/datastore#Note: This needs to be the absolute path
- Use existing SSH key if possible or generate SSH Key (
ssh-keygen) ssh-copy-id name@i.p.ad.dre.ssto copy your SSH key to the serverdvc pushauthenticates with the ssh key from the ssh-configuration (on ubuntu~/.ssh)
dvc remote modify remote_name ask_password trueenables password auth at connectiondvc remote modify --local remote_name password yourpasswordadds config.local file with the password stored and adds this file to gitignore
dvc add path/data.xmladds data to DVC (if not done earlier)git add path/data.xml.dvc data/.gitignoreadded den Bezug zugit commit -m "Add raw data"tracked changes for dvcdvc pushuploads the data to the remote store
- DVC stores data and experiments in the .dvc/cache/files/md5 directory
- This means it copies (or links) the data and hashes it
- Only f BTRFS data management is available on your system, the data does not need to be copied (hardlink and symlink are not recommended due to the risk of data corruption)
- For experiment tracking, the dvc.yaml is copied to .dvc/cache/runs directory
dvc pushcopies the cache to the remotedvc gccleans the local cache (deletes all files from cache other than the current working ones)dvc checkout <branch-or-commit>checks out a dedicated experiment