Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,8 @@ zig-cache/
*.o
*.obj
*.ptx
*.exe
*.pdb

# VitePress & Node build artifacts
docs/.vitepress/dist/
Expand Down
46 changes: 28 additions & 18 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -55,6 +55,8 @@
- For **data validation and serialization** support, check out **[zigantic](https://github.com/muhammad-fiaz/zigantic)**.
- For **build tooling** support, check out **[buildx.zig](https://github.com/muhammad-fiaz/buildx.zig)**.
- For **CUDA/GPU computing** support, check out **[cuda.zig](https://github.com/muhammad-fiaz/cuda.zig)**.
- For **Sqlite** support, check out **[sqlite.zig](https://github.com/muhammad-fiaz/sqlite.zig)**.
- For **Simplified Build.zig** support, check out **[build.zig](https://github.com/muhammad-fiaz/buildx.zig)**.

---

Expand All @@ -64,16 +66,19 @@
| Feature | Description |
|---------|-------------|
| **Dynamic Library Loader** | Runtime resolution of CUDA Driver (`nvcuda.dll` / `libcuda.so`), Runtime (`cudart`), and NVRTC with zero link-time dependencies. |
| **Toolkit Version Compatibility** | Native support for CUDA 12.0 through 13.3 Update 1 with automatic ABI detection. |
| **Toolkit Version Compatibility** | Native support for CUDA 12.0 through 13.4 Developer Preview with automatic ABI detection. |
| **Transparent CPU Fallback** | Automatic fallback to pure Zig CPU implementations for memory allocations and tensor operations when no CUDA GPU is detected. |
| **Device Selection & Properties** | Enumeration of all visible CUDA devices, compute capability queries, memory size reporting, and threadlocal device context management. |
| **Typed Memory Buffers** | High-level `DeviceBuffer(T)`, `PinnedBuffer(T)` (page-locked DMA host memory), and `UnifiedBuffer(T)` (managed memory with prefetch/advise). |
| **Typed Memory Buffers** | High-level `DeviceBuffer(T)`, `PinnedBuffer(T)` (page-locked DMA host memory), `UnifiedBuffer(T)` (managed memory with prefetch/advise), and `PoolBuffer(T)` (stream-ordered memory pools). |
| **Pitched & 2-D Memory** | Hardware-optimal 2-D pitched allocation (`mallocPitch`) and 2-D transfers (`memcpy2D`). |
| **Synchronous & Asynchronous Copies** | Typed H2D, D2H, and D2D memory transfers (sync and stream-ordered async). |
| **Streams & Events** | High-level wrappers for `Stream` and `Event` with elapsed time calculation and stream synchronization. |
| **Streams & Events** | High-level wrappers for `Stream` and `Event` with stream priority ranges (`getStreamPriorityRange`) and elapsed time calculation. |
| **Kernel Launch & Modules** | Arbitrary POD argument marshaling for kernel launches, module loading (`PTX` / `cubin`), and `Function` lookup. |
| **Occupancy Calculator** | Calculate optimal SM block utilization with `maxActiveBlocksPerMultiprocessor` and `maxPotentialBlockSize`. |
| **NVRTC Compilation** | Runtime compilation of CUDA C++ source strings to PTX assembly. |
| **Profiler Integration** | Scoped session tracking via `profiler.start()`, `profiler.stop()`, and `ProfilerGuard`. |
| **Multi-GPU & Peer Access** | `canAccessPeer`, `enablePeerAccess`, `disablePeerAccess`, and cross-device transfers. |
| **Tensor Abstraction** | Generic `Tensor(T)` struct supporting shape manipulation, elementwise ops (add, sub, mul, div, relu), reductions (sum, mean), and matrix multiplication. |
| **Tensor Abstraction** | Generic `Tensor(T)` struct supporting up to **8-D shapes**, elementwise ops, broadcast add/sub/mul/div, reductions (sum, mean, min, max, sumAxis, maxAxis), 2-D/3-D/4-D batched matmul, reshape, transpose, slice, and concat. |
| **CudaAllocator** | `std.mem.Allocator` vtable implementation backed by GPU global device memory. |

</details>
Expand All @@ -93,7 +98,7 @@ Before using `cuda.zig`, ensure you have the following:
|-------------|---------|-------|
| **Zig** | 0.16.0+ | Download from [ziglang.org](https://ziglang.org/download/) |
| **Operating System** | Windows 10+, Linux, macOS | Cross-platform GPU computing |
| **CUDA Driver (Optional)** | 12.0 - 13.3 | Optional runtime dependency; falls back to CPU if absent |
| **CUDA Driver (Optional)** | 12.0 - 13.4 | Optional runtime dependency; falls back to CPU if absent |

---

Expand Down Expand Up @@ -127,10 +132,10 @@ zig build -Dtarget=x86_64-windows

### Method 1: Zig Fetch (Recommended)

**Latest Stable Release (v0.0.1)**
**Latest Release (v0.0.2)**

```bash
zig fetch --save https://github.com/muhammad-fiaz/cuda.zig/archive/refs/tags/0.0.1.tar.gz
zig fetch --save https://github.com/muhammad-fiaz/cuda.zig/archive/refs/tags/0.0.2.tar.gz
```

### Method 2: Zig Fetch (Development / Nightly)
Expand All @@ -148,7 +153,7 @@ Add the dependency to your `build.zig.zon` file.
```zig
.dependencies = .{
.cuda = .{
.url = "https://github.com/muhammad-fiaz/cuda.zig/archive/refs/tags/0.0.1.tar.gz",
.url = "https://github.com/muhammad-fiaz/cuda.zig/archive/refs/tags/0.0.2.tar.gz",
.hash = "...", // Run `zig fetch --save <url>` to generate the hash.
},
},
Expand Down Expand Up @@ -230,28 +235,25 @@ const cuda = @import("cuda");
pub fn main() !void {
const allocator = std.heap.page_allocator;

const a_data = [_]f32{ 1.0, 2.0, 3.0, 4.0 };
const b_data = [_]f32{ 5.0, 6.0, 7.0, 8.0 };

var a = try cuda.Tensor(f32).fromSlice(&a_data, &.{ 2, 2 });
// N-Dimensional Tensor Broadcasting
var a = try cuda.Tensor(f32).fromSlice(&.{ 1, 2, 3, 4, 5, 6 }, &.{ 2, 3 });
defer a.deinit();
var bias = try cuda.Tensor(f32).fromSlice(&.{ 10, 20, 30 }, &.{3});
defer bias.deinit();

var b = try cuda.Tensor(f32).fromSlice(&b_data, &.{ 2, 2 });
defer b.deinit();

var c = try a.matmul(b);
var c = try a.broadcastAdd(bias);
defer c.deinit();

const result = try c.toHost(allocator);
defer allocator.free(result);

std.debug.print("Matmul output: {any}\n", .{result});
std.debug.print("Broadcast Output: {any}\n", .{result});
}
```

## Examples

The `examples/` directory contains **9 runnable examples**:
The `examples/` directory contains **13 runnable examples**:

- [`01_device_info`](examples/01_device_info.zig) - Device enumeration, compute capability, and memory specs
- [`02_memory_transfer`](examples/02_memory_transfer.zig) - Host-to-Device and Device-to-Host transfers
Expand All @@ -262,6 +264,10 @@ The `examples/` directory contains **9 runnable examples**:
- [`07_cpu_fallback`](examples/07_cpu_fallback.zig) - Demonstrating automatic CPU fallback execution
- [`08_managed_memory`](examples/08_managed_memory.zig) - Advanced Unified Memory allocations, prefetching, and advice
- [`09_nvrtc_compilation`](examples/09_nvrtc_compilation.zig) - Dynamic CUDA C++ source compilation to PTX via NVRTC
- [`10_memory_pools_pitched`](examples/10_memory_pools_pitched.zig) - Stream-ordered memory pool allocations and 2D pitched memory
- [`11_occupancy_profiler`](examples/11_occupancy_profiler.zig) - Kernel occupancy calculations, stream priorities, and profiler markers
- [`12_benchmark_matrix_ops`](examples/12_benchmark_matrix_ops.zig) - Parallel N-Body compute benchmark comparing Single-Threaded CPU vs CUDA GPU
- [`13_ndarray_tensor_ops`](examples/13_ndarray_tensor_ops.zig) - Reshape, transpose, broadcast, axis reductions, batched matmul

To run any example:
```bash
Expand All @@ -274,6 +280,10 @@ zig build example-multi-gpu
zig build example-cpu-fallback
zig build example-managed-memory
zig build example-nvrtc-compilation
zig build example-memory-pools-pitched
zig build example-occupancy-profiler
zig build example-benchmark-matrix-ops
zig build example-ndarray-tensor-ops
```

## Validation & Testing
Expand Down
4 changes: 4 additions & 0 deletions build.zig
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,10 @@ pub fn build(b: *std.Build) void {
.{ .name = "example-cpu-fallback", .path = "examples/07_cpu_fallback.zig" },
.{ .name = "example-managed-memory", .path = "examples/08_managed_memory.zig" },
.{ .name = "example-nvrtc-compilation", .path = "examples/09_nvrtc_compilation.zig" },
.{ .name = "example-memory-pools-pitched", .path = "examples/10_memory_pools_pitched.zig" },
.{ .name = "example-occupancy-profiler", .path = "examples/11_occupancy_profiler.zig" },
.{ .name = "example-benchmark-matrix-ops", .path = "examples/12_benchmark_matrix_ops.zig" },
.{ .name = "example-ndarray-tensor-ops", .path = "examples/13_ndarray_tensor_ops.zig" },
};

const run_all_step = b.step("run-all-examples", "Run all example executables");
Expand Down
2 changes: 1 addition & 1 deletion build.zig.zon
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
.{
.name = .cuda,
.version = "0.0.1",
.version = "0.0.2",
.fingerprint = 0x61c9d2274fd46b30,
.minimum_zig_version = "0.16.0",
.dependencies = .{},
Expand Down
10 changes: 8 additions & 2 deletions docs/.vitepress/config.ts
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@ export const GTM_ID = "GTM-P4M9T8ZR";
export const ADSENSE_CLIENT_ID = "ca-pub-2040560600290490";

export const KEYWORDS =
"zig, cuda, gpu, nvidia, nvrtc, device memory, streams, events, tensors, matmul, fallback, parallel computing, cuda runtime, driver api";
"zig, cuda, gpu, nvidia, nvrtc, device memory, streams, events, tensors, matmul, fallback, parallel computing, cuda runtime, driver api, memory pools, occupancy, n-body benchmark, gpu speedup";

export default defineConfig({
lang: "en-US",
Expand Down Expand Up @@ -175,7 +175,7 @@ export default defineConfig({
programmingLanguage: "Zig",
offers: { "@type": "Offer", price: "0", priceCurrency: "USD" },
downloadUrl: "https://github.com/muhammad-fiaz/cuda.zig",
softwareVersion: "0.0.1",
softwareVersion: "0.0.2",
license: "https://opensource.org/licenses/MIT",
});
} else {
Expand Down Expand Up @@ -252,6 +252,7 @@ export default defineConfig({
items: [
{ text: "Device Management", link: "/guide/device-management" },
{ text: "Memory Buffers", link: "/guide/memory-buffers" },
{ text: "Memory Pools", link: "/guide/memory-buffers#memory-pools" },
{ text: "Streams & Events", link: "/guide/streams-events" },
{ text: "Kernel Launch", link: "/guide/kernel-launch" },
{ text: "NVRTC Compilation", link: "/guide/nvrtc" },
Expand All @@ -260,6 +261,7 @@ export default defineConfig({
{ text: "Tensor Operations", link: "/guide/tensor-ops" },
{ text: "CUDA Allocator", link: "/guide/allocator" },
{ text: "Version Compatibility", link: "/guide/version-compat" },
{ text: "Benchmarking", link: "/examples/12-benchmark-matrix-ops" },
{ text: "Related Projects", link: "/guide/related-projects" },
],
},
Expand Down Expand Up @@ -294,6 +296,10 @@ export default defineConfig({
{ text: "07 — CPU Fallback", link: "/examples/07-cpu-fallback" },
{ text: "08 — Managed Memory", link: "/examples/08-managed-memory" },
{ text: "09 — NVRTC Compilation", link: "/examples/09-nvrtc-compilation" },
{ text: "10 — Memory Pools & 2D Pitched", link: "/examples/10-memory-pools-pitched" },
{ text: "11 — Occupancy & Profiler", link: "/examples/11-occupancy-profiler" },
{ text: "12 — CPU vs GPU Benchmark", link: "/examples/12-benchmark-matrix-ops" },
{ text: "13 — N-D Tensor Ops", link: "/examples/13-ndarray-tensor-ops" },
],
},
],
Expand Down
36 changes: 36 additions & 0 deletions docs/api/core.md
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,8 @@ pub const CudaError = error{
DriverError,
RuntimeError,
NvrtcError,
PoolError, // stream-ordered pool allocation failed
OccupancyError, // occupancy query failed
Unknown,
};
```
Expand Down Expand Up @@ -82,3 +84,37 @@ fn checkDriver(code: c_int) !void
```

All runtime and driver calls in cuda.zig pass through these helpers. On failure they emit a `std.log.debug` message and return the appropriate `CudaError`.

## Profiler (`cuda.profiler`)

High-precision wall-clock profiler for measuring CPU-side execution time. Uses `QueryPerformanceCounter` on Windows and `clock_gettime(CLOCK_MONOTONIC)` on POSIX — no dependency on `std.time`.

```zig
// Session markers (interacts with Nsight / nvprof)
cuda.profiler.start();
defer cuda.profiler.stop();

// Scoped guard — automatically calls stop() on scope exit
var guard = cuda.profiler.ProfilerGuard.begin();
defer guard.end();
```

### Timing Functions

```zig
/// Returns current wall-clock time in nanoseconds (platform-native precision).
pub fn nowNs() u64

/// Returns elapsed milliseconds between two nowNs() readings.
pub fn elapsedMs(start_ns: u64, end_ns: u64) f64
```

Example:

```zig
const t0 = cuda.profiler.nowNs();
// ... work ...
const elapsed = cuda.profiler.elapsedMs(t0, cuda.profiler.nowNs());
std.debug.print("Elapsed: {d:.3} ms\n", .{elapsed});
```

101 changes: 42 additions & 59 deletions docs/api/device.md
Original file line number Diff line number Diff line change
@@ -1,75 +1,58 @@
---
title: Device API
description: CUDA device enumeration, property queries, selection, and reset in cuda.zig.
title: Device API Documentation
description: Reference for Device, DeviceProperties, peer access (P2P), and device management in cuda.zig.
---

# Device API
# Device API Reference

## `cuda.device`
The `cuda.device` namespace provides GPU device enumeration, property queries, selection, and peer-to-peer (P2P) memory access.

```zig
/// Return the number of CUDA-capable devices on this host.
/// Returns 0 in CPU-fallback mode.
pub fn count() !u32

/// Set the active device for the calling thread.
pub fn set(index: u32) !void

/// Return the index of the currently active device.
pub fn current() !u32
## Device Management

/// Reset the current device, destroying all resources.
/// Equivalent to cudaDeviceReset(). Use only at shutdown.
pub fn reset() !void
```zig
const count = try cuda.deviceCount();
try cuda.setDevice(0);
const current = try cuda.currentDevice();
try cuda.synchronize();
```

## `cuda.Device`
## `Device` Struct

```zig
pub const Device = struct {
index: u32,

/// Select a device by index and set it as current.
pub fn select(index: u32) !Device

/// Return the device name string (null-terminated, max 256 bytes).
pub fn name(self: Device) []const u8
const dev = try cuda.Device.init(0);

/// Return the full cudaDeviceProp structure for this device.
pub fn properties(self: Device) !DeviceProperties
};
const name = try dev.name();
const cap = try dev.computeCapability();
const total_mem = try dev.totalMemory();
const free_mem = try dev.freeMemory();
const props = try dev.propertiesRaw();
```

## `DeviceProperties`

Direct mapping of `cudaDeviceProp`. Key fields:

| Field | Type | Description |
|---|---|---|
| `name` | `[256]u8` | Device name |
| `totalGlobalMem` | `usize` | Total VRAM in bytes |
| `sharedMemPerBlock` | `usize` | Max shared memory per block |
| `regsPerBlock` | `i32` | Max 32-bit registers per block |
| `warpSize` | `i32` | Warp size in threads |
| `maxThreadsPerBlock` | `i32` | Max threads per block |
| `maxGridSize` | `[3]i32` | Max grid dimensions |
| `clockRate` | `i32` | Clock frequency in kHz |
| `multiProcessorCount` | `i32` | Number of SMs |
| `major` / `minor` | `i32` | Compute capability |
| `totalConstMem` | `usize` | Constant memory size |
| `l2CacheSize` | `i32` | L2 cache size in bytes |
| `maxThreadsPerMultiProcessor` | `i32` | Max threads per SM |
| `isMultiGpuBoard` | `i32` | Non-zero if NVLink present |

## Peer Access
### Peer-to-Peer Access (P2P)

```zig
/// Check if device `src` can directly access memory on device `dst`.
pub fn canAccess(src: u32, dst: u32) !bool

/// Enable peer access from the current device to `target`.
pub fn enable(target: u32) !void

/// Disable peer access from the current device to `target`.
pub fn disable(target: u32) !void
if (try cuda.runtime.device.canAccessPeer(0, 1)) {
try cuda.setDevice(0);
try cuda.runtime.device.enablePeerAccess(1, 0);
// Direct cross-GPU memory transfers are now enabled
}
```

### Device Functions

| Function | Description |
|----------|-------------|
| `cuda.deviceCount()` | Number of visible CUDA GPUs |
| `cuda.setDevice(index)` | Set active device for current thread |
| `cuda.currentDevice()` | Get active device index |
| `cuda.synchronize()` | Synchronize current device context |
| `cuda.allDevices(allocator)` | Return array of all `Device` handles |
| `dev.name()` | Device name string |
| `dev.computeCapability()` | `{ major, minor }` capability struct |
| `dev.totalMemory()` | Total VRAM in bytes |
| `dev.freeMemory()` | Free VRAM in bytes |
| `dev.propertiesRaw()` | Full `DeviceProperties` struct |
| `canAccessPeer(dev, peer)` | Check P2P support between GPUs |
| `enablePeerAccess(peer, flags)` | Enable P2P access to peer GPU |
| `disablePeerAccess(peer)` | Disable P2P access to peer GPU |
| `getAttribute(attr, dev)` | Raw `cuDeviceGetAttribute` query |
Loading
Loading