Repository navigation
feat: expand CUDA runtime, memory, profiling, and documentation support - #1
Merged
Merged
Conversation
…vements, and comptime loader refactor - Add memory pool API (src/memory/pool.zig) with stream-ordered allocation support and pitched 2D memory helpers; expose via public cuda.zig re-exports - Add occupancy profiler (src/kernel/occupancy.zig, src/utils/profiler.zig) that queries cuOccupancyMaxActiveBlocksPerMultiprocessor and surfaces achieved vs theoretical occupancy percentages alongside active warp counts - Add N-body gravitational simulation benchmark (examples/12_benchmark_matrix_ops.zig) comparing sequential CPU O(N^2) against CUDA parallel execution with high-precision wall-clock timing via QueryPerformanceCounter on Windows / clock_gettime on POSIX; consistently demonstrates 40x+ GPU speedup on consumer hardware - Refactor src/core/loader.zig to replace static DLL/SO name arrays with comptime loops over major (11,12,13) x minor (0-9) version pairs, eliminating hardcoded lists while preserving full cross-version compatibility and zero link-time dependencies - Fix runtime FFI structs in src/runtime/ffi.zig to correctly map cudaDeviceProp fields (compute capability, multiprocessor count, clock rate, L2 cache, memory bandwidth) resolving corrupted device info output seen in earlier runs - Update NVRTC loader to share the same comptime version probe list as the runtime loader; fixes 'NVRTC not available' false-negatives on systems where the library exists under a versioned name - Add examples 10 (memory pools / pitched memory) and 11 (occupancy profiler) with corresponding VitePress docs pages and updated docs navigation config - Extend error.zig with pool and occupancy error variants; propagate through public API - Update README.md with benchmark results table, new example descriptions, and revised feature matrix - Add *.exe and *.pdb to .gitignore to prevent Windows build artifacts from staging
- Add docs/examples/12-benchmark-matrix-ops.md covering the N-body CPU vs CUDA GPU benchmark with full source, sample output (RTX 4070 SUPER: 44x speedup), timing methodology, and links to related examples - Register example 12 in sidebar (docs/.vitepress/config.ts) and examples index table - Add Memory Pools and Benchmarking entries to guide sidebar navigation - Expand docs/guide/memory-buffers.md: updated buffer type table to include PoolBuffer, added Memory Pools section (cudaMallocAsync/cudaFreeAsync, CUDA 11.2+ note) and 2D Pitched Memory section (cudaMallocPitch, memcpy2DAsync, coalescing rationale) - Update docs/guide/getting-started.md: bump zig fetch URL and build.zig.zon snippet from v0.0.1 to v0.0.2; expand Next Steps with Tensor Ops and Benchmark links - Update docs/guide/installation.md: bump stable release URL from v0.0.1 to v0.0.2 - Expand docs/api/core.md: add PoolError and OccupancyError variants to CudaError; add Profiler section documenting start/stop, ProfilerGuard, nowNs, and elapsedMs - Rewrite docs/index.md homepage: add Memory Pools, Occupancy/Profiler, and 40x GPU Speedup feature cards; update tagline and description for v0.0.2 capabilities - Expand SEO keywords in config.ts: add memory pools, occupancy, n-body benchmark, gpu speedup terms
Expand Tensor(T) to support up to 8D shapes with row-major strides. - Implement N-D shape transforms: reshape, flatten, squeeze, unsqueeze, N-D axis transpose, 2D transpose (T2), slice, and concat - Implement NumPy-style right-aligned broadcast elementwise ops (broadcastAdd, broadcastSub, broadcastMul, broadcastDiv) - Add global reductions (sum, mean, max, min) and axis reductions (sumAxis, maxAxis) - Add 3D and 4D batched matrix multiplication (batchedMatmul) - Add example 13 and update API references and guides
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This PR expands
cuda.zigwith a broader CUDA runtime API, improved memory and stream management, new tensor capabilities, kernel occupancy utilities, profiling support, additional examples, and a substantial documentation refresh.The update also prepares the project for the
v0.0.2release and improves the overall developer experience when working with CUDA through Zig.Core and Runtime
The CUDA runtime and driver integration has been expanded with improved dynamic library loading and version compatibility.
Device management and device property queries have been extended, including corrected CUDA device property mappings for compute capability, SM count, clock rate, cache information, and other device properties.
Runtime memory operations have been expanded, along with additional stream and event functionality for synchronization, querying, event recording, waiting, and timing.
Peer to peer device memory operations and related device helpers have also been added.
Memory
Stream ordered CUDA memory pool support has been added using asynchronous allocation and deallocation.
PoolBufferprovides typed allocations backed by CUDA memory pools.Pitched 2D memory allocation and copy support has been added, including asynchronous 2D memory transfers.
The memory API now provides broader support for device memory, pinned host memory, managed memory, peer to peer transfers, and CUDA memory IPC.
Memory alignment and error handling have also been improved.
Kernel and Occupancy
A dedicated occupancy module has been added.
The new API provides active blocks per multiprocessor calculations as well as block and grid configuration suggestions.
The kernel API documentation has also been updated to cover occupancy alongside kernel modules, functions, launches, CUDA Graphs, stream capture, and NVRTC.
Streams and Events
The
StreamAPI now provides synchronization and non blocking status queries.Stream and event dependency helpers have been added.
The
EventAPI supports recording, synchronization, querying, and elapsed time measurement.Stream capture workflows are also documented and supported.
Tensor Operations
The high level tensor API has been expanded with tensor shape and data type abstractions.
Elementwise and reduction operations have been expanded.
Tensor transformation functionality has been added, including reshape, transpose, and broadcasting.
Matrix oriented and batched tensor operations have also been expanded.
A new N D tensor operations example demonstrates these capabilities.
Profiling
A high precision profiler utility has been added for CPU side timing.
The implementation uses platform native timing facilities on Windows and POSIX systems.
Profiler session markers are available for CUDA profiling tools, along with a scoped profiler guard for convenient instrumentation.
Timing helpers provide nanosecond timestamps and elapsed millisecond measurements.
The profiler API is exposed through the top level CUDA namespace.
Examples
Four new examples have been added covering memory pools and pitched memory, occupancy and profiling, CPU versus GPU benchmarking, and N D tensor operations.
The examples demonstrate the new APIs through practical CUDA workloads and are integrated into the Zig build system.
Documentation
The documentation has been reorganized and expanded to reflect the updated API surface.
The core, device, memory, kernel, stream, and tensor API references have been refreshed.
Documentation for memory pools, occupancy, profiling, tensor operations, benchmarking, and the new examples has been added.
The getting started and installation guides have been updated for
v0.0.2.Navigation, metadata, example listings, and VitePress configuration have also been updated.
Build and Project Configuration
The new examples have been registered with
build.zig.The package version has been updated to
0.0.2.Windows build artifacts such as
.exeand.pdbhave been added to.gitignore.The public CUDA namespace now exposes the memory pool, occupancy, profiler, and tensor functionality.
Release Preparation
This PR brings the project from
v0.0.1towardv0.0.2, focusing on broader CUDA runtime coverage, improved memory management, more capable stream and event APIs, higher level tensor functionality, kernel occupancy support, profiling utilities, practical examples, and improved documentation.Overall, this release significantly broadens
cuda.ziginto a more complete Zig interface for CUDA runtime functionality, GPU memory management, asynchronous execution, tensor operations, profiling, and performance oriented workflows.