Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

75 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

LibDDLA — Distributed Dense Linear Algebra on GPUs

LibDDLA 0.0.4 is a C++17 template library for distributed dense linear algebra on GPU devices. It provides ScaLAPACK-style APIs with 2D block-cyclic data distribution over an MPI process grid, with a unified CUDA or HIP backend selected at build time. All functions live in the ddla namespace.

Key facts and caveats

  • One backend only. Exactly one of DDLA_USE_CUDA or DDLA_USE_HIP must be enabled. There is no CPU-only backend.
  • Supported scalar types are float, double, std::complex<float>, and std::complex<double>. Not every routine is instantiated for all four types; notably the standard distributed Cholesky family (ppotrf / ppotrs / pposv) is complex-only.
  • Matrix storage is caller-owned. LibDDLA routines operate on device pointers supplied by the caller. Individual routines may allocate and release temporary device workspaces internally.
  • There is no installed CMake package config (find_package(LibDDLA) is not supported). Set CMAKE_INSTALL_PREFIX, build the install target, and link your application against the installed shared library and headers.

Prerequisites

  • CMake ≥ 3.13
  • C++17 compiler
  • MPI (Open MPI or MPICH)
  • CUDA backend: CUDA Toolkit with cuBLAS, cuSOLVER, cuRAND
  • HIP backend: ROCm / DTK with hipBLAS, hipSOLVER, hipRAND, and a CMake version with first-class HIP language support
  • CCL mode: NCCL or RCCL for direct inter-GPU collectives (enable -DDDLA_USE_CCL=ON)

The current HIP link configuration includes RCCL in its backend library list, so RCCL must be available to HIP builds even when another communication mode is selected.

The build system does not enforce a minimum CUDA/ROCm/NCCL version; use toolchains that are compatible with C++17, the GPU hardware, and the required math/solver libraries.

Quick build

Architecture values (CMAKE_CUDA_ARCHITECTURES, CMAKE_HIP_ARCHITECTURES) are hardware-specific — use the SM number for your NVIDIA GPU (e.g. "80" for A100, "70" for V100) or the gfx target for your AMD GPU (e.g. "gfx90a" for MI200 series).

CUDA

cmake -S . -B build-cuda                               \
  -DDDLA_USE_CUDA=ON                                   \
  -DCMAKE_CUDA_ARCHITECTURES="80"                      \
  -DBUILD_TESTS=ON                                     \
  -DDDLA_USE_CCL=ON                                    \
  -DCMAKE_INSTALL_PREFIX="$PWD/install"
cmake --build build-cuda -j
cmake --build build-cuda --target install

HIP / ROCm

cmake -S . -B build-hip                                \
  -DDDLA_USE_HIP=ON                                    \
  -DCMAKE_HIP_ARCHITECTURES="gfx90a"                   \
  -DBUILD_TESTS=ON                                     \
  -DDDLA_USE_CCL=ON                                    \
  -DCMAKE_INSTALL_PREFIX="$PWD/install"
cmake --build build-hip -j
cmake --build build-hip --target install

The installed layout uses include/ddla/*.h and the platform library directory (normally lib/libddla.so on Linux). There is no CMake package config installed, so downstream builds must add the install prefix to their header and library search paths manually, or locate the library with CMake's find_library. Downstream translation units must also define the same backend macro used to build LibDDLA (DDLA_USE_CUDA or DDLA_USE_HIP).

CMake options

Option Default Description
BUILD_TESTS OFF Build test executables
DDLA_USE_CUDA OFF Build for NVIDIA CUDA GPUs
DDLA_USE_HIP OFF Build for AMD HIP/ROCm GPUs
DDLA_USE_DEBUG OFF Enable DDLA_USE_DEBUG preprocessor macro
DDLA_USE_CCL OFF Use NCCL (CUDA) or RCCL (HIP) for device collectives
DDLA_USE_GPU_CPU_TUNNEL OFF Route communication through host staging buffers (D2H → MPI → H2D)

When both DDLA_USE_CCL and DDLA_USE_GPU_CPU_TUNNEL are enabled, the GPU-CPU tunnel path takes precedence for communication. NCCL/RCCL libraries are still linked, and CMake emits a warning noting that CCL will not be used.

Communication modes

The library selects one of three communication paths at compile time:

Path Preprocessor Behaviour
NCCL / RCCL DDLA_USE_CCL Direct device collectives on GPU-resident buffers
GPU-CPU tunnel DDLA_USE_GPU_CPU_TUNNEL D2H copy → MPI collective → H2D copy
Synchronized MPI on device ptrs (neither) deviceStreamSynchronize then GPU-aware MPI on device pointers

All ranks must execute collectives in the same order regardless of the chosen path.

Data model and lifecycle

Handle

#include <ddla/ddla.h>
#include <ddla/ddla_connector.h>

ddla::DdlaHandle_t handle;          // opaque pointer (DdlaStream*)
ddla::ddla_init(handle);            // create BLAS/Solver handles and device streams
ddla::ddla_set(handle, MPI_COMM_WORLD, 'R');   // auto process grid (row-major)
// or: ddla::ddla_set(handle, MPI_COMM_WORLD, nprows, npcols, 'R');
  • ddla_set stores the process-grid dimensions and communicator on the handle. The automatic form ('R') derives a 2D grid from the number of MPI ranks.
  • Destroy with ddla::ddla_destroy(handle) when done.

Descriptor (DdlaDesc)

A DdlaDesc captures the 2D block-cyclic layout of a distributed matrix: global rows/cols (m, n), block sizes (mb, nb), process-grid coordinates (nprows, npcols, myprow, mypcol), source-process offsets (irsrc, icsrc), and local dimensions (m_loc, n_loc, lld).

Important: most factorization and solve routines require square blocks (mb == nb). Use init_square_blk to enforce this automatically.

int m = 4096, n = 64;
ddla::DdlaDesc desc(handle);
desc.init_square_blk(m, n, 0, 0);   // derives one square block size from the grid

Free index-mapping helpers are in <ddla/ddla_desc.h>: indxg2p, indxg2l, indxl2g, num_loc. They match the ScaLAPACK convention exactly.

Local device matrices

Each MPI rank owns a contiguous local submatrix of size m_loc × n_loc stored in GPU device memory. The caller allocates and frees this memory.

#include <ddla/ddla.h>

int main(int argc, char** argv) {
    MPI_Init(&argc, &argv);

    ddla::DdlaHandle_t handle;
    ddla::ddla_init(handle);
    ddla::ddla_set(handle, MPI_COMM_WORLD, 'R');

    int n = 4096, nrhs = 64;
    ddla::DdlaDesc descA(handle), descB(handle);
    descA.init_square_blk(n, n, 0, 0);
    descB.init_square_blk(n, nrhs, 0, 0);

    double *d_A, *d_B;
    ddla::deviceMalloc(&d_A, descA.m_loc() * descA.n_loc() * sizeof(double));
    ddla::deviceMalloc(&d_B, descB.m_loc() * descB.n_loc() * sizeof(double));

    // ... fill matrices, call ddla routines ...

    ddla::deviceFree(d_A);
    ddla::deviceFree(d_B);
    ddla::ddla_destroy(handle);
    MPI_Finalize();
    return 0;
}

The ddla::deviceMalloc / ddla::deviceFree helpers declared in <ddla/ddla_connector.h> are convenience wrappers over cudaMalloc / hipMalloc that safely handle zero-byte requests.

API overview

Handle and descriptor

Symbol Description
DdlaHandle_t Opaque handle type (DdlaStream*)
ddla_init Allocate handle (streams, BLAS/Solver handles, optional CCL comms)
ddla_set Configure process grid (auto 'R' or explicit nprows×npcols)
ddla_destroy Tear down handle
ddla_get_stream Return the default device stream from a handle
DdlaDesc 2D block-cyclic matrix descriptor
DdlaDesc::init_square_blk Initialize mb = nb from global row/col counts
DdlaDesc::init Initialize with explicit block sizes
indxg2p / indxg2l / indxl2g / num_loc Free index-mapping functions

Distributed BLAS and data movement

Function Description
pgemm C = α·op(A)·op(B) + β·C (SUMMA algorithm)
pgeadd C = α·op(A) + β·op(B)
pdam Add scalar to diagonal of distributed square matrix
ptran Out-of-place distributed transpose (with optional conjugate)
transport_block Extract/transpose a contiguous block from a distributed matrix into a local buffer

LU factorization, solve, and drivers

Function Description
pgetf2 Unblocked panel LU (inner kernel, host pivot array)
pgetf2_panel Alternative panel LU (rank-revealing variant)
pgetrf LU with partial (row) pivoting
pgetrf_bpiv Block LU with partial pivoting per block row (device pivot array)
pgetrf_nopiv Multi-process LU without pivoting
getrf_nopiv Local (single-process) LU without pivoting
pgetrs Solve using pivoted LU factors
pgetrs_nopiv Solve using no-pivot LU factors
pgesv Driver: LU + solve with pivoting
pgesv_nopiv Driver: LU + solve without pivoting
ptrtrs Distributed triangular solve
plapiv Apply row-pivot permutation
pswap Swap rows or columns between two distributed matrices

Cholesky

Function Description
ppotrf Standard Cholesky factorization (complex-only)
ppotrs Solve using Cholesky factor (complex-only)
pposv Driver: Cholesky + solve (complex-only)
potrf_bottom_right Local bottom-right Cholesky (all four scalar types)
ppotrf_bottom_right Distributed bottom-right Cholesky (all four scalar types)

Auxiliary and advanced

Function Description
gemmVbatched Batch of GEMMs with variable dimensions (device-resident dim arrays)
gemmVbatched2s Two-stage variable-batch GEMM with reusable temporary
random_generate Fill device buffer with uniform random values

random_generate<T> is declared in <ddla/ddla_connector.h> as a template, implemented in src/ddla_utils.cpp, and explicitly instantiated for float, double, std::complex<float>, and std::complex<double>.

Testing

Build with -DBUILD_TESTS=ON. Test executables are MPI programs and must be launched with mpirun. The total number of ranks must equal the product of process-grid rows and columns (nprows × npcols), which defaults automatically. You can override the grid with --grid:

# Run a single test on 4 ranks (auto grid)
mpirun -np 4 build-cuda/tests/test_random_generate

# Explicit 2×2 grid for a test that supports --grid
mpirun -np 4 build-cuda/tests/test_api_grid_ptrtrs --grid 2x2

The files tests/test_cuda.sh and tests/test_hip.sh are cluster-specific Slurm batch scripts. They load modules, build, and run the test suite on particular machines; adapt their module/partition settings to your own cluster.

Repository layout

LibDDLA/
├── include/ddla/         Public headers
│   ├── ddla.h            Main API declarations (all Doxygen comments)
│   ├── ddla_handle_t.h   Handle type and init/set/destroy
│   ├── ddla_desc.h       DdlaDesc descriptor and index mapping helpers
│   ├── ddla_comm.h       Communication primitives (bcast, send/recv, allreduce)
│   ├── ddla_connector.h  CUDA/HIP type aliases, macros, CHECK utilities, malloc wrappers, random_generate
│   ├── ddla_stream.h     DdlaStream (internal)
│   ├── transport_block.h Distributed-to-local block extraction
│   ├── ptran.h           Distributed matrix transpose
│   ├── gemmVbatched.h    Variable-size batched GEMM
│   └── Backend wrappers  gemm.h, trsm.h, scal.h, axpy.h, swap.h, geru.h,
│                         iamax.h, geam.h, herk.h, syrk.h, trmm.h,
│                         gemmBatched.h, getrf.h, potrf.h, laswp.h
├── src/                  Implementation (one routine per .cpp)
├── tests/                Integration tests (MPI executables)
├── benchmarks/           Benchmark data and supporting artifacts
├── install_scripts/      Backend-specific build/install scripts
├── cmake/                CMake helper modules
├── CMakeLists.txt        Top-level build
└── LICENSE               LGPL-3.0 license text

Version

0.0.4 — defined in src/version.h. The shared library SONAME tracks the major version.

License

LGPL v3. See LICENSE.

About

LibDDLA(the library of device distributed linear algorithms)

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages