Skip to content

Repository files navigation

FlyCube

A measured fruit-fly (MaleCNS) connectome is wired into the decision loop of a Rubik's cube solver, and you can watch its real neural activity choose every turn.

FlyCube solving a 3-move scramble

(A ~30-second browser recording of the same flow is in docs/media/demo-recording.webm.)

15-second demo

scripts/flycube setup      # user-local Python 3.12 + npm deps
scripts/flycube prepare    # download + verify MaleCNS v1.0, build the graphs
scripts/flycube demo       # build the viewer, start the backend
# open http://localhost:8766/  -> Random scramble -> Solve

On the single screen: the cube scrambles, the brain point cloud lights up with the activity actually used, the 18 action probabilities and the value appear, the chosen turn is committed, and the cube animates. Search expansions, budgets, timeouts, state hashes and compute time are shown honestly; failures are shown as failures.

What FlyCube actually is

  • A real cube environment. 54-facelet truth, all 18 half-turn-metric moves, a uniform legal-state generator (twist/flip/parity invariants), SHA-256 state hashes, seeded scrambles and scramble-text input. Rendering is downstream of the state, never the reverse.
  • A real connectome in the loop. MaleCNS v1.0 neurons, soma voxels, contact counts and transmitter-predicted signs are imported with exact 64-bit IDs. A reproducible 8,192-neuron subgraph (original IDs and geometry) feeds a sparse recurrent rate network; input (visual_projection) and readout (descending + central complex) groups are disjoint and communicate only through measured edges.
  • A trained policy/value model, not a scripted solver. Training uses a trusted conventional teacher offline only (Kociemba / pycuber / exact BFS); inference and the live server cannot import the teacher (enforced by test).
  • Two honest inference modes. Policy only takes the policy argmax. Brain + search runs bounded best-first search using the learned value and policy and reports expansions, budgets and timeouts, distinguishing the search-selected move from the policy top-1.
  • One event stream. Episode ID, step, evaluated state hash, checkpoint hash, sampled neuron activity, policy/value, selected move and resulting cube state — all from the backend. The renderer never invents activity or moves.
  • A documented correction. The first training run propagated connectivity backwards (r @ W.T instead of r @ W). The code now defaults to the correct direction; the original checkpoints are preserved under checkpoints/legacy-reversed-connectivity/, and every published number is labelled as legacy. See docs/methodology.md.

FlyCube is a fork of DOOMFLY (MIT) that replaces the Doom environment with a cube. It does not require Doom, Mario, a GPU, or root access.

Architecture

cube state ──▶ engineered sticker encoder ──▶ sparse recurrent rate network
                                                    on measured MaleCNS edges
                                                    (fixed reservoir or
                                                     trainable gain+bias)
                                                          │
                                          policy (18 moves) + value
                                                          │
                       ┌──────────────────────────────────┴───────────────┐
                       │ Policy only            brain + bounded search   │
                       └──────────────────────────────────┬───────────────┘
                                                          ▼
                   one SSE event stream ──▶ Vite + Three.js single screen

Full detail: docs/architecture.md.

Quickstart

Requirements: Linux/macOS (or WSL2), Python ≥ 3.10, Node ≥ 20, ~3 GB disk for the MaleCNS download. No sudo is required: scripts/flycube setup installs a user-local Python 3.12 through uv and the npm dependencies.

scripts/flycube setup       # environment + viewer deps and build
scripts/flycube prepare     # data (below)
scripts/flycube demo        # http://127.0.0.1:8766/
scripts/flycube test        # FlyCube tests (55 of them)
scripts/flycube stop

Opening the demo without prepared data prints exactly what to run instead of crashing; the same is true for training and evaluation.

Reproduce the connectome

scripts/flycube prepare runs four reproducible steps:

  1. download MaleCNS v1.0 annotations, neurotransmitters and edges (about 1.1 GB) and verify SHA-256 against data-provenance/malecns_v1/source.lock.json;
  2. import 166,700 retained neurons and 25,582,938 edges with exact IDs and soma voxel coordinates (data-provenance/malecns_v1/import-report.json);
  3. build the full signed CSR graph (graph_full-report.json);
  4. extract the training subgraph with guaranteed input→readout paths (graph_train-report.json).

Data lives in connectome_data/ and runs/ (gitignored). Set FLYCUBE_DATA_DIR=/path/to/data to reuse a prepared copy.

Train

scripts/flycube train --run-dir runs/train/main_v2 --seed 0 --time-budget 3600
scripts/flycube train --run-dir runs/train/main_v2 --resume

Curriculum: exhaustive exact labels for all states up to depth 4, sampled Kociemba upper-bound labels for depths 5-6, time-based phases that keep earlier levels in the mixture, held-out validation, checkpoints with optimizer state, graph hash, seeds, elapsed time and learning curves. flycube/train/dagger.py adds policy-guided relabeling and fine-tuning. New runs record propagation: pre_to_post.

Evaluate

scripts/flycube evaluate --checkpoint checkpoints/legacy-reversed-connectivity/flycube-main-best.npz \
    --modes policy,search --n-uniform 100 --short-per-depth 20 --quick

flycube/eval/benchmark.py produces per-mode solved/attempted, moves, latency, timeouts, expansions, hardware and graph size, plus determinism checks and the silenced / rewired / untrained controls and uninformed baselines. The propagation mode is read from checkpoint metadata (--propagation auto). flycube/eval/summarize.py turns reports into a markdown table.

Results

Legacy 60-minute model (all numbers from the reversed-connectivity checkpoints; full tables in docs/results.md):

suite policy brain + search
depth 1-2 (40 cases) 40/40 40/40
depth 3 (20 cases) 11/20 12/20
depth 4 (20 cases) 4/20 4/20
uniform random states (100) 0/100 0/100

Matched 15-minute controls: connectome 44 %/57 % (policy/search), rewired 10 %/14 %, silenced 0 %/14 %; random and uninformed-BFS baselines solve 6 %/14 % of the short suite (depth-1 only). Search overrode the policy on 4 committed steps out of 80. Determinism checks pass (same state → bit-identical policy), and the uniform test set has zero overlap with training pools.

Scientific limitations

Not fly vision (engineered sticker encoding), not deep cube solving (half-turn metric, exact only to depth 4; a 25-move scramble is not a distance-25 proof), no general claim that fly wiring helps, bounded search (no optimality claim), uniform random states unsolved, rate units rather than spikes, and a reduced training subgraph. The first session's connectivity-direction bug and its negative DAgger result are documented rather than hidden. Full list: docs/limitations.md.

Data, provenance, licenses

  • MaleCNS v1.0 (HHMI Janelia FlyEM and collaborators): download, CC BY 4.0 — see licenses/CC-BY-4.0.txt and data-provenance/. Not bundled.
  • DOOMFLY (nftechie, MIT): upstream revision 71ecf53d78eaffaf1a57ed7b0ccf5d458abc9f33; the connectome importer and transmitter-sign policy are adapted from it. Attribution and scope: THIRD_PARTY.md.
  • Kociemba (GPLv2) is an optional, lazily imported, training/evaluation-only teacher, never called by inference; pycuber (MIT) is the MIT cross-check.
  • three.js / Vite / TypeScript are npm dependencies under their own licenses; see THIRD_PARTY_NOTICES.md.
  • Large data and generated runs stay out of git; the only committed binaries are the small legacy checkpoints, the episode logs in examples/, and the demo media in docs/media/.

About

Measured MaleCNS connectome in the decision loop of a Rubik's cube solver: honest research code, viewer, checkpoints and results

Topics

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages