Skip to content

F3 — CUDA backend #8

Description

@kkokosa

A GPU backend for the same .dfb checkpoints and the same semantics as the CPU backends.

Deliverables

  • PTX kernels for phase A (LIF update, threshold, refractory countdown) and the blocked-CSR delivery; driver-API binding (no CUDA toolkit dependency at run time beyond the driver).
  • CUDA graphs for the 18-step batch; lanes (batched trials) on the GPU.
  • Cross-backend validation: spike-identical to the CPU reference on the golden fixtures and statistically equivalent (per-neuron spike counts) on the full v630 sugar run; determinism for a fixed seed.
  • dotfly bench --backend cuda numbers for the three standard scenarios.

Acceptance: ≥ 10× real time for the full MaleCNS on one consumer GPU; Backend.Cuda selectable from SimulationOptions; the room demo runs unchanged on it.

Context: PLAN.md §11 (F3), §8 (performance design).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    featureA roadmap feature (PLAN.md §11)performanceBackends, throughput, packaging

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions