This guide provides an extremely in-depth, step-by-step walkthrough for configuring an Ubuntu system for deep learning and machine learning tasks using NVIDIA hardware. All installations are performed system-wide (i.e., no virtual environments), so please be aware that these changes affect your entire system. It is recommended to back up your system or create a restore point before proceeding.
Disclaimer: Installing packages system-wide may lead to conflicts or stability issues over time. This guide is intended for users who prefer or require a global setup.
- Prerequisites
- System Update and Upgrade
- NVIDIA Drivers and CUDA Toolkit
- Global Python Installation and Essential Packages
- Machine Learning and Data Processing Libraries
- Deep Learning Frameworks
- cuDNN Installation
- TensorRT Installation
- Profiling and Debugging Tools
- Verification and Testing
- Troubleshooting and Further Resources
- Ubuntu Versions Tested: Ubuntu 20.04 and Ubuntu 22.04.
- Hardware: An NVIDIA GPU with supported drivers.
- Internet Connection: Required for downloading packages.
- Administrator Access: You will need
sudoprivileges to install system-wide packages.
Official Resources:
Ensure your system is up-to-date. This minimizes conflicts and ensures all dependencies are current.
sudo apt update && sudo apt upgrade -yInstall the latest NVIDIA drivers using Ubuntu’s driver management tool. This ensures compatibility with CUDA and your GPU hardware.
sudo ubuntu-drivers autoinstallAfter installation, reboot your system:
sudo rebootThe CUDA toolkit provides the necessary libraries and tools to run GPU-accelerated applications.
Option A: Install via Ubuntu Repositories
sudo apt install nvidia-cuda-toolkitOption B: Download from NVIDIA
For the latest version and more control over the installation, visit the NVIDIA CUDA Toolkit Downloads page and follow the instructions for your Ubuntu version.
Verify Installation:
nvcc --version
nvidia-smiExpected output:
nvcc --versionshould display the CUDA compiler version.nvidia-smishould list your GPU details along with driver and CUDA versions.
Install Python 3 and its package manager globally:
sudo apt install python3 python3-pipKeep pip updated to avoid compatibility issues:
sudo -H pip3 install --upgrade pipInstall some utilities to improve your Python experience (e.g., enhanced error messages):
sudo -H pip3 install pretty_errorsTest the installation:
python3 -m pretty_errorsInstall common libraries used in data processing and machine learning:
sudo -H pip3 install numpy scipy pandas matplotlib seaborn scikit-image pillow opencv-python-headlessNote: We install OpenCV via pip in its headless version to avoid GUI dependencies on servers.
Install TensorFlow with GPU support:
sudo -H pip3 install tensorflowFor the most recent instructions and compatibility details, refer to the TensorFlow Installation Guide.
Install PyTorch using the recommended installation command for your CUDA version. For example, for CUDA 11.7, visit the PyTorch Get Started page. A typical command might be:
sudo -H pip3 install torch torchvisionEnable integration between TensorFlow and TensorRT for optimized inference:
sudo -H pip3 install tensorflow-tensorrtInstall PyTorch-TensorRT for accelerated PyTorch inference:
sudo -H pip3 install torch-tensorrtcuDNN is essential for deep learning performance and must match your CUDA version.
- Download cuDNN:
- Visit the NVIDIA cuDNN Download page (login required).
- Select the version that matches your installed CUDA toolkit.
- Installation Instructions:
- Follow the official guide provided with the download package. Typically, you will extract the files and copy them to the appropriate CUDA directories (e.g.,
/usr/local/cuda/includeand/usr/local/cuda/lib64).
- Follow the official guide provided with the download package. Typically, you will extract the files and copy them to the appropriate CUDA directories (e.g.,
Refer to the cuDNN Installation Guide for detailed steps.
TensorRT accelerates inference on NVIDIA GPUs.
Download the TensorRT package for your Ubuntu version from the NVIDIA TensorRT Download page. For Ubuntu 20.04, an example download command is:
sudo wget https://developer.download.nvidia.com/compute/machine-learning/tensorrt/8.6.1.6/ubuntu2004/x86_64/tensorrt-repo-ubuntu2004-8.6.1.6-ga-cuda11.7_1-1_amd64.debsudo dpkg -i tensorrt-repo-ubuntu2004-8.6.1.6-ga-cuda11.7_1-1_amd64.deb
sudo apt-key adv --fetch-keys http://developer.download.nvidia.com/compute/machine-learning/tensorrt/8.6.1.6/ubuntu2004/x86_64/7fa2af80.pub
sudo apt updatesudo apt install libnvinfer8 libnvinfer-dev libnvinfer-plugin8sudo -H pip3 install nvidia-pyindex tensorrtVerify TensorRT Installation:
python3 -c "import tensorrt as trt; print(trt.__version__)"Install NVIDIA’s profiling and debugging tools to optimize and troubleshoot GPU applications:
sudo apt install nsight-systems nsight-computeThese tools provide detailed insights into GPU performance and can help diagnose issues in deep learning workloads.
After installing all components, verify that your system is correctly set up.
nvidia-smi
nvcc --versionpython3 -c "import tensorflow as tf; print(tf.config.list_physical_devices('GPU'))"Expected output: A list showing your GPU(s).
python3 -c "import torch; print(torch.cuda.is_available())"Expected output: True
python3 -c "import tensorrt as trt; print(trt.__version__)"
python3 -c "import tensorflow as tf; from tensorflow.python.compiler.tensorrt import trt_convert; print(trt_convert)"If you require ONNX for model interoperability, install and verify as follows:
sudo -H pip3 install onnx onnx-tf onnx-torch
python3 -c "import onnx; print(onnx.__version__)"
python3 -c "import onnx_tf; print(onnx_tf.__version__)"
python3 -c "import onnx_torch; print(onnx_torch.__version__)"-
Driver or CUDA Mismatch:
Ensure that the installed NVIDIA driver is compatible with your CUDA toolkit. Consult the CUDA Compatibility Guide for details. -
cuDNN Version Mismatch:
The cuDNN version must align with your CUDA version. Double-check the compatibility matrix on the cuDNN Release Notes. -
System-wide Python Conflicts:
Since this guide avoids virtual environments, conflicts might arise from global package installations. Regularly update and manage packages carefully. -
TensorRT Issues:
Refer to the TensorRT Documentation for troubleshooting integration issues and performance tuning.
- CUDA Toolkit: NVIDIA CUDA Documentation
- cuDNN: cuDNN Installation Guide
- TensorFlow: TensorFlow Installation Guide
- PyTorch: PyTorch Get Started
- TensorRT: TensorRT Documentation
- Nsight Tools: Nsight Systems Documentation and Nsight Compute Documentation
This guide is designed to provide a robust, system-wide setup for deep learning and machine learning on Ubuntu with NVIDIA hardware. For additional support, refer to the official documentation linked above or open an issue in this repository.