Skip to content
View DanieleSavino's full-sized avatar
  • Rome
  • 23:38 (UTC +02:00)

Organizations

@PlumJuice-HPC-Team

Block or report DanieleSavino

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
DanieleSavino/README.md

Daniele Savino

3rd year CS student at Sapienza. I like low-level code, parallel systems, and understanding why things are slow.

Currently working on my thesis under Daniele De Sensi — implementing BiNE collective algorithms in NCCL, targeting the multi-GPU memory hierarchy of modern HPC clusters. Competing at SCC connect in November 2026 with @PlumJuice-HPC-Team.

Skills

C · C++ · CUDA · MPI · OpenMP · Python
Profiling: Score-P / Scalasca / Cube · perf / PAPI · Nvidia nsights / nsys
Clusters: Leonardo (CINECA) · MareNostrum 5 (BSC) · MeluXina

Pinned Loading

  1. nccl nccl Public

    Forked from HLC-Lab/nccl

    NCCL fork implementing BiNE collective algorithms

    C++

  2. unitccl unitccl Public

    CLI for testing and profiling Bine collectives in NCCL — Slurm rank sweeps with live dashboard, nsys profiling, plotting

    Python

  3. ImageFlow ImageFlow Public

    Modular C image processing pipeline with pluggable devices, operations, and schedulers — CUDA/HIP/OpenMP backends currently registered.

    C++ 1

  4. fastest fastest Public

    C testing framework with Python orchestration — pool comparisons, nanosecond timing, matplotlib plotting, no dependencies in core

    Python 1

  5. CollBench CollBench Public

    C instrumentation library for MPI custom collectives — nanosecond timing, zero overhead without -DCB_PROFILE

    Python 1

  6. CollAlgo CollAlgo Public

    BiNE collective algorithms for MPI — CUDA-aware, profiling via CollBench

    C 1