Skip to content

About

Interactive explainer: a fixed-size Hebbian synaptic matrix as short-term memory, mapped to Pathway's BDH and BDH-CQ.

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

One matrix, many memories

DataForge 2026 · Pathway Track · Topic: Synaptic Plasticity as Short-Term Memory (combined with Associative Memory and Fast Weights — one claim, one learning journey).

An interactive explainer in which a learner writes key→value associations into a 64×64 Hebbian synaptic matrix, reads them back, and watches recall fail through interference and decay rather than slot exhaustion — then sees exactly where this matrix is in Dragon Hatchling's state equation and in BDH-CQ's contextual memory.


The one-sentence claim

A fixed-size synaptic matrix written by a Hebbian outer-product rule stores in-context associations without allocating a slot per token, and recalls them exactly only while the stored keys stay nearly non-overlapping; as the number of stored pairs or the overlap between keys grows, recall degrades through interference, not by running out of slots.

How the artifact can falsify it. If recall stayed at 100% for all 40 pairs at high density (k = 14) with the standard seed, or if the failure point depended on matrix size rather than on key count and overlap, the claim would be wrong. The sweep chart recomputes the full recall-vs-m curve live for the current k and decay; the 60-second test asks the learner to predict the breaking point before finding it.

Intended learner and prerequisites

  • Audience: advanced undergraduates, ML engineers and data scientists who know what a Transformer's KV cache is.
  • Prerequisites: matrix–vector products; the idea of an outer product. No neuroscience, no BDH background.
  • Time: ~12 minutes guided, open-ended sandbox after.

Learning objectives

After using the artifact the learner can:

  1. State the Hebbian write rule σ ← σ + y xᵀ and the linear read σ·x, and say which of the two is where interference comes from.
  2. Predict the direction of change in recall when (a) stored pairs increase, (b) pattern density increases, (c) decay increases, (d) an older vs newer pair is probed.
  3. Derive the expected crosstalk (m−1)·k³/n² for random k-sparse patterns and compare it to a measurement.
  4. Locate the same mechanism in BDH's eq. 6 (σ update) and eq. 8 (ρ = Eσ), name what is state and what is parameter, and say what BDH-GPU adds that the toy omits.
  5. Explain BDH-CQ's S_t = U_θ(S_{t−1}, D_t) and why its linear-attention special case is the toy at u = 0, with correct evidence labels.
  6. Name three misconceptions (fast weights = KV cache; Hebbian writes = training; BDH = Mamba-style SSM) and one open problem (fast→slow consolidation).

Guided narrative (what the learner sees, in order)

  1. Five memories, one matrix — preset already running; recall exact; matrix size unchanged.
  2. Add memories until it breaks — m = 28; orange crosstalk bars cross the k-th-largest cutoff.
  3. Overlap is the real variable — m = 12, k = 14; fewer pairs, worse recall.
  4. Decay: the second way to forget — u = 0.2, probe oldest pair; (0.8)¹⁹ ≈ 1.4% of written strength.
  5. Sandbox — all controls open, each mapped to a BDH variable.

Then: Why it behaves this way (algebra), Where this lives in BDH (eq. 6/8/9 mapped line by line), And in BDH-CQ, Test the claim in sixty seconds, Three things this does not mean, Explain it back.

Architecture of the artifact

Single index.html, vanilla JavaScript, two <canvas> charts plus DOM strips. No dependencies, no fonts fetched, no storage, no network requests.

Component Role Live / precomputed / synthetic / animated
genPatterns() Draws 40 key patterns and 40 value patterns, each k of 64 neurons active (0/1). Seeded PRNG (mulberry32) so the guided walk is reproducible; "New random patterns" reseeds. Synthetic input, generated live
writeAll(m,u) σ ← (1−u)·σ, then σ += y_i x_iᵀ, for i = 1..m in order. Live computation
read(q) a = σ·x_q for the probe key. Live
recallOf() Top-k of a vs. true support of y_q; ties at the cutoff are reported as ambiguous, not exact. Live
σ heatmap 64×64 cells, opacity ∝ synapse strength, probe columns outlined, true-value rows ticked. Live rendering of real state
Activation bars a_j per neuron, teal = true value neurons, orange = all others, dashed = k-th largest. Live
Truth-beside-estimate strips y_q vs top-k recall. Live
Analytic readouts Signal (1−u)^age·k and mean crosstalk Σ wᵢ k³/n² computed from the formula in Why it behaves this way, shown beside the measured values. Live, formula is an expectation over random patterns
Sweep chart Average recall over all stored pairs for m = 1..40 at the current k, u. Recomputed on every control change (40 writes × ≤40 reads on 4,096 cells; well under 50 ms). Live
BDH module Equations transcribed from the BDH paper (eq. 6, 8, 9) and BDH-CQ report (§3.2); mapping table; evidence tags. Static text, primary-sourced
60-second test Learner prediction vs. first m at which the sweep drops below 100% for k = 8, u = 0 with the current patterns. Live

Nothing on the page is an animation of a scripted outcome. There are no precomputed results.

What is deliberately simplified relative to BDH

  • Random 0/1 k-sparse patterns; BDH's x, y are learned, non-negative, and sparse (~5% reported).
  • No encoder/decoder (E, D_x, D_y), no ReLU-lowrank step, no LayerNorm, no per-neuron rotation (RoPE); U is scalar decay (ALiBi-like).
  • Single layer, G_s = all-ones (the BDH-GPU correspondence, eq. 9), n = 64 vs n = 32,768 at 25M parameters.
  • Read is σ·x then top-k; BDH reads ((G_y σ x)⁺ ⊙ x).

These are named in the artifact under Limitations of the toy, stated plainly. The toy is an independent educational reimplementation of the stated equations. It is not an official BDH or BDH-CQ model and contains no Pathway code.

Run locally

git clone https://github.com/aryanbajaj868/bdh-synaptic-memory.git
cd bdh-synaptic-memory
python3 -m http.server 8000     # or just open index.html in a browser
# visit http://localhost:8000

No install step. Works offline. Tested logic in Node 18+:

node tests/core_test.js

Expected output includes measured mean crosstalk 2.345 expected 2.375 (k = 8, m = 20, seed 11) and a monotone decrease of average recall with m and with k.

Deploy

Hosted on GitHub Pages from the main branch root: https://aryanbajaj868.github.io/bdh-synaptic-memory/ . Any push to main redeploys within a couple of minutes.

Reproducing the results quoted in the text

  • "(0.8)¹⁹ ≈ 1.4%": 0.8**19 = 0.0144.
  • "two random keys share about k²/n neurons": expected |x_i ∩ x_q| for two independent k-subsets of n is k·(k/n) = k²/n. At k = 14, n = 64: 3.06.
  • "measured 2.35 vs expected 2.38": run node tests/core_test.js.
  • All BDH/BDH-CQ figures: page/section references in SOURCES_AND_LICENSES.md.

Primary papers (2022–2026) that use, extend, test, or rely on the concept

Cited beside the technical claims in index.html and CONCEPT_SUMMARY.md:

  1. Kosowski, Uznański, Chorowski, Stamirowska, Bartoszkiewicz. The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain. arXiv:2509.26507, 2025. — defines σ as Hebbian synaptic attention state (eq. 6, 8).
  2. Engdahl, Kosowski, Chorowski, Stamirowska, Uznański, Jiang, Phadke, Kinas, Zhong. BDH-CQ: In-Context Learning with Recurrent Latent Reasoning. arXiv:2608.09888, 2026. — contextual memory S_t = U_θ(S_{t−1}, D_t); linear-attention special case (§3.2).
  3. Behrouz, Zhong, Mirrokni. Titans: Learning to Memorize at Test Time. arXiv:2501.00663, 2024. — fast-weight neural memory updated at inference, with learned forgetting.
  4. Yang, Kautz, Hatamizadeh. Gated Delta Networks: Improving Mamba2 with Delta Rule. arXiv:2412.06464, ICLR 2025. — fixed-size associative state with a delta-rule write that removes the interference this toy exhibits.
  5. Sun, Li, Dalal, Xu, Vikram, Zhang, Dubois, Chen, Wang, Koyejo, Hashimoto, Guestrin. Learning to (Learn at Test Time): RNNs with Expressive Hidden States. arXiv:2407.04620, 2024. — hidden state as a model updated by gradient steps at test time.

Background (pre-2022): Hinton & Plaut 1987; Ba et al. 2016; Katharopoulos et al. 2020; Schlag, Irie & Schmidhuber 2021.

Credits and licences

Code and text: MIT (see LICENSE). No third-party code, fonts, images, or data are bundled; system fonts only. Full record in SOURCES_AND_LICENSES.md. AI assistance: see AI_DISCLOSURE.md.

Team and mentorship

  • Team: hackerz
  • Member: Aryan Bajaj, IIT (BHU) Varanasi
  • Mentorship: none

About

Interactive explainer: a fixed-size Hebbian synaptic matrix as short-term memory, mapped to Pathway's BDH and BDH-CQ.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages