Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
24 commits
Select commit Hold shift + click to select a range
245481d
Add experimental CUDA reference backend and mandatory hardware valida…
freeman-1984-coder Sep 10, 2026
2844c12
Merge branch 'main' into feat/cuda-reference
freeman-1984-coder Sep 10, 2026
237e7e2
Integrate 0.4.0a4 and extend CUDA gate to game feedback and restore
freeman-1984-coder Sep 10, 2026
d025232
Add reward-trained external reflex readout and experiment replay
freeman-1984-coder Sep 12, 2026
35628f3
Use NVRTC built-in math without requiring host math headers
freeman-1984-coder Sep 12, 2026
fcbc5ca
Document verified A16 conformance and honest small-circuit timings
freeman-1984-coder Sep 12, 2026
b6824af
Keep replay trial index valid on first animation frame
freeman-1984-coder Sep 12, 2026
a0cc31c
Add reproducible full FlyWire CPU CUDA validation runner
freeman-1984-coder Sep 12, 2026
c7fa117
Document complete graph validation protocol and modeling assumptions
freeman-1984-coder Sep 12, 2026
5cb8255
Preserve exact 64-bit FlyWire IDs across signed and unsigned source c…
freeman-1984-coder Sep 12, 2026
4b75aed
Keep full brain runner importable on Windows and document source audit
freeman-1984-coder Sep 12, 2026
7cf9cd1
Record complete FlyWire A16 validation and measured CPU CUDA performance
freeman-1984-coder Sep 12, 2026
357e654
Add sourced bilateral olfactory input and full-brain CUDA probe
freeman-1984-coder Sep 13, 2026
930b5f9
Format olfactory map source metadata
freeman-1984-coder Sep 13, 2026
bc0cf78
Record full-brain odor result and diagnose subthreshold benchmark wei…
freeman-1984-coder Sep 13, 2026
cfa5a0e
Add synaptic LIF CPU reference checked against Brian2
freeman-1984-coder Sep 13, 2026
d970731
Merge published full-brain reports into CUDA development
freeman-1984-coder Sep 13, 2026
7c96a49
Add experimental synaptic CUDA engine and full-graph odor protocol
freeman-1984-coder Sep 13, 2026
592738a
Consume bounded bridge bodies before rejecting browser origins
freeman-1984-coder Sep 13, 2026
130765c
Add voxel odor environment with an explicit neural control boundary
freeman-1984-coder Sep 13, 2026
e89b8fe
Add complete-brain GPU voxel recording with a silenced control
freeman-1984-coder Sep 13, 2026
716dc7f
Publish verified full-brain synaptic CUDA propagation results
freeman-1984-coder Sep 13, 2026
1c53c0f
Publish actual full-brain voxel trajectories and browser replay
freeman-1984-coder Sep 13, 2026
0a58f5e
Add crawlable experiment summary and article metadata
freeman-1984-coder Sep 13, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 16 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,22 @@ jobs:
working-directory: packages/js
- run: pytest tests/test_experiment.py
- run: git diff --exit-code -- site/js/runtime.js site/js/dodge-runtime.js site/toy-bundle.json site/license.txt
synaptic-reference:
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- run: python -m pip install -e . 'Brian2==2.9.0' 'numpy==1.26.4'
- name: Independent synaptic dynamics oracle
run: python scripts/validate_synaptic_reference.py --output synaptic-reference.json
- uses: actions/upload-artifact@v4
if: always()
with:
name: synaptic-reference
path: synaptic-reference.json
godot:
runs-on: ubuntu-latest
timeout-minutes: 10
Expand Down
12 changes: 9 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,9 +1,15 @@
# flybrain-sdk

> Development branch: [experimental CUDA backend](docs/cuda.md). [Full FlyWire benchmark](docs/validation/flywire-full-a16.md): 139,255 neurons, all 16.85M source rows, CPU/CUDA parity and 10 seconds of continuous simulation on A16-8Q. CUDA median 11.55 seconds per simulated second (not real time). [Train an external reflex readout](examples/REFLEX_TRAINING.md). Released v0.4.0a4 is CPU-only.


**No CUDA required.** A Python SDK for connecting small connectome simulations to games and experiments, starting with a working NumPy CPU backend.

**Full-brain CUDA benchmark (experimental branch):** [139,255 FlyWire neurons on a real NVIDIA GPU](https://freeman-1984-coder.github.io/flybrain-sdk/fullbrain.html). Complete published proofread graph, CPU/CUDA checks, reproducible timing and evidence. Released v0.4.0a4 remains CPU-only.
**Full-brain sensory result:** [The GPU odor probe activates the input cells but exposes a propagation limit in the benchmark weight preset](docs/validation/flywire-odor-a16.md). Numerical parity does not establish biological behavior. [Published report](https://freeman-1984-coder.github.io/flybrain-sdk/fullbrain.html#olfaction).

**Full-brain voxel recording:** [Watch the complete GPU brain drive a ground body](https://freeman-1984-coder.github.io/flybrain-sdk/foraging.html). The fixed readout curves away from food; the input-silenced control stays still. [Evidence, assumptions and reproduction](docs/validation/flywire-voxel-a16.md).

**Synaptic dynamics verified on GPU:** [The separate mV engine matches Brian2 and passes actual CUDA tests](docs/synaptic-dynamics.md). [A complete FlyWire sensory probe now reaches downstream populations](docs/validation/flywire-synaptic-odor-a16.md); the candidate DNa02 readout does not simply encode stimulus side. Navigation remains unproven.

**Live game:** [Open the browser sandbox](https://freeman-1984-coder.github.io/flybrain-sdk/live.html) — add obstacles, tune mappings, inspect neural control, save and verify recordings. [Adapter guide](docs/live-sandbox.md).

Expand All @@ -23,7 +29,7 @@ brain.step(100)
print(brain.action().to_dict())
```

**0.4 alpha:** the bundled offline demo is a hand-authored 12-neuron circuit. A separate 3.8 MB MaleCNS model now runs 313 real source neurons and 20,607 anatomical edges with explicitly assumed LIF parameters. [Model card and reproducible recipe](models/male-cns-escape-v1/README.md). CUDA and WASM are reserved interfaces, not implemented runtimes. No GPU, credentials, or network access are needed to run the toy demo after installation.
**0.4 alpha:** the bundled offline demo is a hand-authored 12-neuron circuit. A separate 3.8 MB MaleCNS model now runs 313 real source neurons and 20,607 anatomical edges with explicitly assumed LIF parameters. [Model card and reproducible recipe](models/male-cns-escape-v1/README.md). This development branch adds an experimental CUDA runtime; WASM remains unimplemented. No GPU, credentials, or network access are needed to run the toy demo after installation.

## Make your own demo

Expand Down Expand Up @@ -105,7 +111,7 @@ Game movement and rendering stay in your application; see [game_loop.py](example
| --- | --- |
| `cpu` | Working reference implementation, NumPy edge lists, float64 |
| `wasm` | Reserved name; raises `BackendUnavailableError` |
| `cuda` | Reserved name; raises `BackendUnavailableError`; no CUDA dependencies |
| `cuda` | Experimental optional CuPy runtime; 12/12 A16 hardware cases passed; unavailable without NVIDIA/CuPy |

See [dynamics and backend contract](docs/architecture.md) for equations, spike timing and limitations.

Expand Down
6 changes: 4 additions & 2 deletions docs/api.md
Original file line number Diff line number Diff line change
Expand Up @@ -138,7 +138,9 @@ Use `flybrain.experiment.replay_experiment(recording, download=False)` to verify
accepts an observation with `offer(seq, observation)`, then commits the actual
engine control with `acknowledge(seq, applied)`. Only one action may be pending.
Identical pending offers reuse their result without reintegrating the brain.
`snapshot()` is allowed at acknowledged boundaries; `from_snapshot()` restores
the built-in linear encoder and rate readout. The engine must checkpoint its own
`snapshot()` is allowed at acknowledged boundaries; `from_snapshot(data, backend=None)` restores
the built-in linear encoder and rate readout. In this CUDA development branch,
pass `backend="cpu"` or `backend="cuda"` to select the restore device explicitly.
The engine must checkpoint its own
world at the matching sequence. See the [Godot example](../examples/godot/README.md)
for transport, numeric validation, timing limits and recovery rules.
95 changes: 95 additions & 0 deletions docs/cuda.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,95 @@
# Experimental CUDA backend — A16 hardware validation passed

This development branch implements a CuPy/CUDA reference backend. **All 12 required hardware cases passed on a Vultr NVIDIA A16-2Q on 2026-09-12 UTC.** The published v0.4.0a4 remains CPU-only. This is experimental compatibility support, not a speedup claim: the 313-cell model took 2.64 seconds per simulated second on this GPU, versus 0.108 seconds on the same host CPU. See [the measured report](validation/a16-20260912.md).

## Optional installation

The basic `pip install -e .` path still installs only NumPy. Importing `flybrain` or creating a CPU brain does not import CuPy or initialize CUDA. On an NVIDIA host with a suitable driver, use a separate Python 3.12 environment:

```sh
pip install -e '.[dev]' 'cupy-cuda12x[ctk]==14.2.0'
python scripts/validate_cuda.py --output cuda-validation.json
```

CuPy 14 requires Python 3.10+; the CPU SDK still supports Python 3.9. Match the CuPy wheel to the host driver/toolkit and install only one CuPy distribution. The `ctk` extra installs compatible CUDA components; it cannot supply the host NVIDIA driver. See [CuPy installation](https://docs.cupy.dev/en/stable/install.html).

After hardware validation:

```python
from flybrain import FlyBrain, available_backends

print(available_backends()) # dependency/device probe, not a conformance certificate
brain = FlyBrain.load("male-cns-escape-v1", backend="cuda", download=True)
brain.stimulate("looming_left", duration_ms=100)
brain.advance(duration_ms=100)
print(brain.action().to_dict())
brain.save("gpu.checkpoint.json")
cpu = FlyBrain.restore("gpu.checkpoint.json", backend="cpu")
```

Unavailable dependencies, devices or kernel compilation raise `BackendUnavailableError`. There is no silent CPU fallback. Python-only `backend="cuda"` does not add CUDA to the browser runtime; WASM is still unimplemented.

## What the implementation does

- One CUDA thread owns one postsynaptic row. Stable ordering keeps original parallel-edge accumulation order. It uses float64 and disables fused multiply-add to preserve separate reference operations; the reported hardware suite measures conformance.
- The graph, voltage, spikes, refractory counters, filtered rates and silencing mask reside on one GPU. Each instance owns a stream so external CuPy stream contexts do not reorder its operations.
- One kernel advances one tick using previous-tick spikes. It writes separate output buffers. A scalar error flag is checked before committing the tick, preserving state on nonfinite voltage errors.
- Selected observations transfer only requested cells/fields. Full observations and checkpoints explicitly transfer full state. Restore uses the existing CPU validation rules, with temporary CPU graph/state allocation.
- Current vectors are currently prepared on the host and uploaded every tick. Each tick also synchronizes an error flag. This first implementation prioritizes compatibility; it is **not** the future device-resident batched schedule and may be slower than CPU on small models.

The [CuPy RawKernel API](https://docs.cupy.dev/en/stable/reference/generated/cupy.RawKernel.html) provides runtime compilation. Larger graphs, batching, mixed precision and trainable dynamics need separate work and evidence.

## The hardware gate

`validate_cuda.py` requires an actual CuPy-visible NVIDIA device, runs the hardware suite with `FLYBRAIN_REQUIRE_CUDA=1`, rejects skipped tests, then times toy and real circuits on CPU and GPU. Without a device it writes a failed report and exits nonzero. Ordinary CPU CI skips those hardware tests explicitly.

The suite has 12 required cases. It compares every cell over 180 ticks for toy, real and signed/parallel-edge synthetic graphs at two timesteps (1,080 comparison ticks), plus cross-device checkpoints with pending stimuli, custom readout, interventions, selected observations, invalid-state handling and stream isolation. Spikes/clock must match exactly; voltage/rates use absolute tolerance `1e-10`. A report includes environment versions, device, source hashes, passed test counts, configuration, warmup and per-run timings. This covers the reported hardware and cases only.

Acceptance before merging/advertising CUDA:

1. No skipped or failed hardware tests on at least one NVIDIA GPU.
2. Checkpoint continuation and direct/named current semantics match CPU.
3. Honest CPU/GPU wall times with compilation/loading separated and synchronization included.
4. Store the report with the exact tested source identity and inspect any failures.
5. Re-run normal CPU packaging/CI; missing CUDA never breaks CPU installation.

## Optional rented T4 run

A prepared runner is available at `scripts/validate_cuda_modal.py`. Its configuration is one T4, one CPU core, 4 GiB RAM, at most one container, a 600-second function timeout, no retries, and scale-to-zero. There is no persistent volume, web endpoint or schedule. Only source code, the hardware test, validator and the small model are uploaded; no credentials or local home directory are included.

After the account and spending authorization are resolved, a maintainer can run:

```sh
# Separate Python 3.10+ environment on the controlling computer.
pip install 'modal==1.5.5'
modal setup
modal run scripts/validate_cuda_modal.py --output cuda-validation.json
```

This command starts remote compute and can incur charges. As checked on 2026-09-10, [Modal's published pricing](https://modal.com/pricing) lists T4 at $0.000164/second, CPU at $0.0000131/core/second, and RAM at $0.00000222/GiB/second. Ten minutes at the configured allocations is about $0.112 in execution charges, before build/startup/other charges. This estimate is not an enforced account-wide dollar cap. Review current pricing and available account credits before running; no credits are assumed. Use a $1 trial budget and verify the ephemeral app stops in the provider console when the run finishes or is canceled.

The runner definition was checked against the local Modal SDK without invoking any remote function. It remains untested remotely. [Modal GPU documentation](https://modal.com/docs/guide/gpu) describes device selection. Recheck container termination and actual billed usage before considering the rental step finished.

## Game integration (development branch only)

The draft now includes the v0.4.0a4 Godot and project-generation changes.
`make_demo(..., backend="cuda")` selects CUDA for Python sessions.
`ExternalController.from_snapshot(data, backend="cpu")` explicitly restores a
GPU checkpoint onto CPU, or vice versa with `backend="cuda"`. The new hardware
cases compare both session and acknowledged external feedback against CPU for toy
and real circuits, including CUDA → CPU → CUDA continuation. Both cases passed on the reported A16 device.

On a configured NVIDIA host, the experimental Godot bridge accepts
`python examples/godot/bridge.py --backend cuda`. Its chosen backend applies to
fresh sessions and imported checkpoints. The default remains CPU; a saved backend
label cannot change the bridge's explicitly selected device. Actual Godot with
CUDA has not been verified.

## Complete FlyWire brain benchmark

The [full FlyWire v783 report](validation/flywire-full-a16.md) validates all 139,255
proofread neurons and 16,847,997 source rows on A16-8Q. CPU/CUDA parity, backend
checkpoint replay, 10 seconds of continuous simulated time, and 171 regression
tests passed. CUDA median 11.55 seconds per simulated second versus CPU 110.96
seconds on that host. This is faster than the SDK CPU reference, but not real time.
[Reproduce the full-graph test](fullbrain-validation.md).
71 changes: 71 additions & 0 deletions docs/fullbrain-validation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,71 @@
# Full FlyWire CUDA validation

**Functional limitation:** the [full-brain odor probe](validation/flywire-odor-a16.md) found that this benchmark’s normalized weights cannot drive an initially resting, externally unstimulated neuron to threshold. Use these results for numerical conformance, not sensory-to-motor behavior.

This opt-in runner loads all **139,255 proofread FlyWire v783 neurons** and every
row in the official proofread connections table. It uses the SDK's actual
`Connectome`, `CPUBackend`, and `CUDABackend`, without replacing their execution
paths. No neuron selection, minimum-count pruning, or synthetic graph is used.

The pinned source contains **16,847,997 neuron-pair/neuropil rows**, representing
**54,492,922 synaptic contacts**. Multiple rows may connect the same pair in
different neuropils. All endpoints match the official 139,255-neuron list.

## Reproduce on an NVIDIA host

Use the experimental CUDA branch. A 32 GB system-memory host gives headroom for
this initial Python object-based importer. GPU memory requirements are measured
by the runner; the raw file size is not the required GPU memory.

```sh
python -m venv .venv
. .venv/bin/activate
pip install -e '.[datasets]' 'cupy-cuda12x[ctk]==14.2.0' pytest
mkdir -p data results
curl -fL 'https://zenodo.org/records/10676866/files/proofread_root_ids_783.npy?download=1' -o data/proofread_root_ids_783.npy
curl -fL 'https://zenodo.org/records/10676866/files/proofread_connections_783.feather?download=1' -o data/proofread_connections_783.feather
python scripts/validate_fullbrain.py --data-dir data --output results/fullbrain.json
```

Download source: [FlyWire Consortium, v783.0, Zenodo](https://zenodo.org/records/10676866).
The runner checks both published MD5 values, then records SHA-256 values and byte
sizes. A missing or different file fails the test. The raw data are not bundled
in the package or redistributed by this repository.

The neuron list stores IDs as `uint64`, while the connection table uses `int64`.
Normalize both to the same integer type before lookup: mixed signed/unsigned
NumPy searches can promote large IDs to `float64` and lose their identity. The
runner rejects floating-point IDs and includes regression cases above 2^53.

## What is checked

- All official neuron IDs are unique; all source edge endpoints resolve.
- Every published proofread connection row is retained, including weak connections,
self-connections, and parallel rows for different neuropils.
- The complete network rests without input, then generates spikes under seeded input.
- CPU and CUDA states across every neuron are compared every 25 ticks over 200
active ticks: spikes and refractory counters exactly, voltage/rates within 1e-10.
- A JSON checkpoint reproduces 30 subsequent CUDA ticks exactly.
- Three 1,000 ms runs per backend follow a 100-tick warmup. Compilation/loading are
reported separately; timed steps include host input upload and error synchronization.
- 10,000 additional CUDA ticks check finite, bounded states and record population
activity, device allocations, and peak host memory. These observation-heavy
timings are separate from the throughput benchmark.

## What this means biologically

The whole published proofread **brain** graph is used; this is not the whole
nervous system or the raw EM volume, and not every unproofread segment is a neuron
in this graph. Each source row aggregates anatomical contacts between a neuron
pair in a neuropil. A connection row is not one anatomical synapse.

Dynamics are dimensionless, simplified LIF assumptions. Contact counts are scaled
to incoming absolute weight sum 2.5. Each neuron's sign uses its outgoing
contact-weighted transmitter scores: GABA/glutamate negative, other labels positive.
Unknown transmitter evidence uses positive sign and is counted. This deliberately
simple policy does not model receptor-specific effects or neuromodulation.

The benchmark uses artificial currents, not reconstructed vision, behavior,
physiological calibration, or biological learning. Passing establishes full-graph
software execution and numerical checks for the reported configuration. It does
not establish that a living fly's complete brain function has been reproduced.
Loading
Loading