Skip to content

feat: expand CUDA runtime, memory, profiling, and documentation support - #1

Merged
muhammad-fiaz merged 5 commits into
mainfrom
dev
Aug 8, 2026
Merged

muhammad-fiaz merged 5 commits into
mainfrom
dev

Conversation

@muhammad-fiaz

@muhammad-fiaz muhammad-fiaz commented Aug 7, 2026 •

Copy link
Copy Markdown
Owner

Summary

This PR expands cuda.zig with a broader CUDA runtime API, improved memory and stream management, new tensor capabilities, kernel occupancy utilities, profiling support, additional examples, and a substantial documentation refresh.

The update also prepares the project for the v0.0.2 release and improves the overall developer experience when working with CUDA through Zig.

Core and Runtime

The CUDA runtime and driver integration has been expanded with improved dynamic library loading and version compatibility.

Device management and device property queries have been extended, including corrected CUDA device property mappings for compute capability, SM count, clock rate, cache information, and other device properties.

Runtime memory operations have been expanded, along with additional stream and event functionality for synchronization, querying, event recording, waiting, and timing.

Peer to peer device memory operations and related device helpers have also been added.

Memory

Stream ordered CUDA memory pool support has been added using asynchronous allocation and deallocation.

PoolBuffer provides typed allocations backed by CUDA memory pools.

Pitched 2D memory allocation and copy support has been added, including asynchronous 2D memory transfers.

The memory API now provides broader support for device memory, pinned host memory, managed memory, peer to peer transfers, and CUDA memory IPC.

Memory alignment and error handling have also been improved.

Kernel and Occupancy

A dedicated occupancy module has been added.

The new API provides active blocks per multiprocessor calculations as well as block and grid configuration suggestions.

The kernel API documentation has also been updated to cover occupancy alongside kernel modules, functions, launches, CUDA Graphs, stream capture, and NVRTC.

Streams and Events

The Stream API now provides synchronization and non blocking status queries.

Stream and event dependency helpers have been added.

The Event API supports recording, synchronization, querying, and elapsed time measurement.

Stream capture workflows are also documented and supported.

Tensor Operations

The high level tensor API has been expanded with tensor shape and data type abstractions.

Elementwise and reduction operations have been expanded.

Tensor transformation functionality has been added, including reshape, transpose, and broadcasting.

Matrix oriented and batched tensor operations have also been expanded.

A new N D tensor operations example demonstrates these capabilities.

Profiling

A high precision profiler utility has been added for CPU side timing.

The implementation uses platform native timing facilities on Windows and POSIX systems.

Profiler session markers are available for CUDA profiling tools, along with a scoped profiler guard for convenient instrumentation.

Timing helpers provide nanosecond timestamps and elapsed millisecond measurements.

The profiler API is exposed through the top level CUDA namespace.

Examples

Four new examples have been added covering memory pools and pitched memory, occupancy and profiling, CPU versus GPU benchmarking, and N D tensor operations.

The examples demonstrate the new APIs through practical CUDA workloads and are integrated into the Zig build system.

Documentation

The documentation has been reorganized and expanded to reflect the updated API surface.

The core, device, memory, kernel, stream, and tensor API references have been refreshed.

Documentation for memory pools, occupancy, profiling, tensor operations, benchmarking, and the new examples has been added.

The getting started and installation guides have been updated for v0.0.2.

Navigation, metadata, example listings, and VitePress configuration have also been updated.

Build and Project Configuration

The new examples have been registered with build.zig.

The package version has been updated to 0.0.2.

Windows build artifacts such as .exe and .pdb have been added to .gitignore.

The public CUDA namespace now exposes the memory pool, occupancy, profiler, and tensor functionality.

Release Preparation

This PR brings the project from v0.0.1 toward v0.0.2, focusing on broader CUDA runtime coverage, improved memory management, more capable stream and event APIs, higher level tensor functionality, kernel occupancy support, profiling utilities, practical examples, and improved documentation.

Overall, this release significantly broadens cuda.zig into a more complete Zig interface for CUDA runtime functionality, GPU memory management, asynchronous execution, tensor operations, profiling, and performance oriented workflows.

…vements, and comptime loader refactor

- Add memory pool API (src/memory/pool.zig) with stream-ordered allocation support and
  pitched 2D memory helpers; expose via public cuda.zig re-exports
- Add occupancy profiler (src/kernel/occupancy.zig, src/utils/profiler.zig) that queries
  cuOccupancyMaxActiveBlocksPerMultiprocessor and surfaces achieved vs theoretical
  occupancy percentages alongside active warp counts
- Add N-body gravitational simulation benchmark (examples/12_benchmark_matrix_ops.zig)
  comparing sequential CPU O(N^2) against CUDA parallel execution with high-precision
  wall-clock timing via QueryPerformanceCounter on Windows / clock_gettime on POSIX;
  consistently demonstrates 40x+ GPU speedup on consumer hardware
- Refactor src/core/loader.zig to replace static DLL/SO name arrays with comptime loops
  over major (11,12,13) x minor (0-9) version pairs, eliminating hardcoded lists while
  preserving full cross-version compatibility and zero link-time dependencies
- Fix runtime FFI structs in src/runtime/ffi.zig to correctly map cudaDeviceProp fields
  (compute capability, multiprocessor count, clock rate, L2 cache, memory bandwidth)
  resolving corrupted device info output seen in earlier runs
- Update NVRTC loader to share the same comptime version probe list as the runtime
  loader; fixes 'NVRTC not available' false-negatives on systems where the library
  exists under a versioned name
- Add examples 10 (memory pools / pitched memory) and 11 (occupancy profiler) with
  corresponding VitePress docs pages and updated docs navigation config
- Extend error.zig with pool and occupancy error variants; propagate through public API
- Update README.md with benchmark results table, new example descriptions, and revised
  feature matrix
- Add *.exe and *.pdb to .gitignore to prevent Windows build artifacts from staging
@muhammad-fiaz muhammad-fiaz added documentation Improvements or additions to documentation enhancement New feature or request labels Aug 7, 2026
- Add docs/examples/12-benchmark-matrix-ops.md covering the N-body CPU vs CUDA GPU
  benchmark with full source, sample output (RTX 4070 SUPER: 44x speedup), timing
  methodology, and links to related examples
- Register example 12 in sidebar (docs/.vitepress/config.ts) and examples index table
- Add Memory Pools and Benchmarking entries to guide sidebar navigation
- Expand docs/guide/memory-buffers.md: updated buffer type table to include PoolBuffer,
  added Memory Pools section (cudaMallocAsync/cudaFreeAsync, CUDA 11.2+ note) and
  2D Pitched Memory section (cudaMallocPitch, memcpy2DAsync, coalescing rationale)
- Update docs/guide/getting-started.md: bump zig fetch URL and build.zig.zon snippet
  from v0.0.1 to v0.0.2; expand Next Steps with Tensor Ops and Benchmark links
- Update docs/guide/installation.md: bump stable release URL from v0.0.1 to v0.0.2
- Expand docs/api/core.md: add PoolError and OccupancyError variants to CudaError;
  add Profiler section documenting start/stop, ProfilerGuard, nowNs, and elapsedMs
- Rewrite docs/index.md homepage: add Memory Pools, Occupancy/Profiler, and 40x GPU
  Speedup feature cards; update tagline and description for v0.0.2 capabilities
- Expand SEO keywords in config.ts: add memory pools, occupancy, n-body benchmark,
  gpu speedup terms
Expand Tensor(T) to support up to 8D shapes with row-major strides.

- Implement N-D shape transforms: reshape, flatten, squeeze, unsqueeze, N-D axis transpose, 2D transpose (T2), slice, and concat
- Implement NumPy-style right-aligned broadcast elementwise ops (broadcastAdd, broadcastSub, broadcastMul, broadcastDiv)
- Add global reductions (sum, mean, max, min) and axis reductions (sumAxis, maxAxis)
- Add 3D and 4D batched matrix multiplication (batchedMatmul)
- Add example 13 and update API references and guides
@muhammad-fiaz
muhammad-fiaz merged commit a1091fc into main Aug 8, 2026
15 checks passed
@muhammad-fiaz
muhammad-fiaz deleted the dev branch August 8, 2026 00:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant