DataForge 2026 · Pathway Track · Topic: Synaptic Plasticity as Short-Term Memory (combined with Associative Memory and Fast Weights — one claim, one learning journey).
An interactive explainer in which a learner writes key→value associations into a 64×64 Hebbian synaptic matrix, reads them back, and watches recall fail through interference and decay rather than slot exhaustion — then sees exactly where this matrix is in Dragon Hatchling's state equation and in BDH-CQ's contextual memory.
- Live artifact: https://aryanbajaj868.github.io/bdh-synaptic-memory/ (opens without sign-in; source is
index.html, single file, no build, no network, no precomputed data) - Source repository: https://github.com/aryanbajaj868/bdh-synaptic-memory
- One-page concept summary:
CONCEPT_SUMMARY.pdf(source:CONCEPT_SUMMARY.md) - Sources and licences:
SOURCES_AND_LICENSES.md - AI-assistance disclosure:
AI_DISCLOSURE.md - Live-defense preparation:
DEFENSE_NOTES.md
A fixed-size synaptic matrix written by a Hebbian outer-product rule stores in-context associations without allocating a slot per token, and recalls them exactly only while the stored keys stay nearly non-overlapping; as the number of stored pairs or the overlap between keys grows, recall degrades through interference, not by running out of slots.
How the artifact can falsify it. If recall stayed at 100% for all 40 pairs at high density (k = 14) with the standard seed, or if the failure point depended on matrix size rather than on key count and overlap, the claim would be wrong. The sweep chart recomputes the full recall-vs-m curve live for the current k and decay; the 60-second test asks the learner to predict the breaking point before finding it.
- Audience: advanced undergraduates, ML engineers and data scientists who know what a Transformer's KV cache is.
- Prerequisites: matrix–vector products; the idea of an outer product. No neuroscience, no BDH background.
- Time: ~12 minutes guided, open-ended sandbox after.
After using the artifact the learner can:
- State the Hebbian write rule σ ← σ + y xᵀ and the linear read σ·x, and say which of the two is where interference comes from.
- Predict the direction of change in recall when (a) stored pairs increase, (b) pattern density increases, (c) decay increases, (d) an older vs newer pair is probed.
- Derive the expected crosstalk (m−1)·k³/n² for random k-sparse patterns and compare it to a measurement.
- Locate the same mechanism in BDH's eq. 6 (σ update) and eq. 8 (ρ = Eσ), name what is state and what is parameter, and say what BDH-GPU adds that the toy omits.
- Explain BDH-CQ's S_t = U_θ(S_{t−1}, D_t) and why its linear-attention special case is the toy at u = 0, with correct evidence labels.
- Name three misconceptions (fast weights = KV cache; Hebbian writes = training; BDH = Mamba-style SSM) and one open problem (fast→slow consolidation).
- Five memories, one matrix — preset already running; recall exact; matrix size unchanged.
- Add memories until it breaks — m = 28; orange crosstalk bars cross the k-th-largest cutoff.
- Overlap is the real variable — m = 12, k = 14; fewer pairs, worse recall.
- Decay: the second way to forget — u = 0.2, probe oldest pair; (0.8)¹⁹ ≈ 1.4% of written strength.
- Sandbox — all controls open, each mapped to a BDH variable.
Then: Why it behaves this way (algebra), Where this lives in BDH (eq. 6/8/9 mapped line by line), And in BDH-CQ, Test the claim in sixty seconds, Three things this does not mean, Explain it back.
Single index.html, vanilla JavaScript, two <canvas> charts plus DOM strips. No dependencies, no fonts fetched, no storage, no network requests.
| Component | Role | Live / precomputed / synthetic / animated |
|---|---|---|
genPatterns() |
Draws 40 key patterns and 40 value patterns, each k of 64 neurons active (0/1). Seeded PRNG (mulberry32) so the guided walk is reproducible; "New random patterns" reseeds. | Synthetic input, generated live |
writeAll(m,u) |
σ ← (1−u)·σ, then σ += y_i x_iᵀ, for i = 1..m in order. | Live computation |
read(q) |
a = σ·x_q for the probe key. | Live |
recallOf() |
Top-k of a vs. true support of y_q; ties at the cutoff are reported as ambiguous, not exact. | Live |
| σ heatmap | 64×64 cells, opacity ∝ synapse strength, probe columns outlined, true-value rows ticked. | Live rendering of real state |
| Activation bars | a_j per neuron, teal = true value neurons, orange = all others, dashed = k-th largest. | Live |
| Truth-beside-estimate strips | y_q vs top-k recall. | Live |
| Analytic readouts | Signal (1−u)^age·k and mean crosstalk Σ wᵢ k³/n² computed from the formula in Why it behaves this way, shown beside the measured values. | Live, formula is an expectation over random patterns |
| Sweep chart | Average recall over all stored pairs for m = 1..40 at the current k, u. Recomputed on every control change (40 writes × ≤40 reads on 4,096 cells; well under 50 ms). | Live |
| BDH module | Equations transcribed from the BDH paper (eq. 6, 8, 9) and BDH-CQ report (§3.2); mapping table; evidence tags. | Static text, primary-sourced |
| 60-second test | Learner prediction vs. first m at which the sweep drops below 100% for k = 8, u = 0 with the current patterns. | Live |
Nothing on the page is an animation of a scripted outcome. There are no precomputed results.
- Random 0/1 k-sparse patterns; BDH's x, y are learned, non-negative, and sparse (~5% reported).
- No encoder/decoder (E, D_x, D_y), no ReLU-lowrank step, no LayerNorm, no per-neuron rotation (RoPE); U is scalar decay (ALiBi-like).
- Single layer, G_s = all-ones (the BDH-GPU correspondence, eq. 9), n = 64 vs n = 32,768 at 25M parameters.
- Read is σ·x then top-k; BDH reads ((G_y σ x)⁺ ⊙ x).
These are named in the artifact under Limitations of the toy, stated plainly. The toy is an independent educational reimplementation of the stated equations. It is not an official BDH or BDH-CQ model and contains no Pathway code.
git clone https://github.com/aryanbajaj868/bdh-synaptic-memory.git
cd bdh-synaptic-memory
python3 -m http.server 8000 # or just open index.html in a browser
# visit http://localhost:8000No install step. Works offline. Tested logic in Node 18+:
node tests/core_test.jsExpected output includes measured mean crosstalk 2.345 expected 2.375 (k = 8, m = 20, seed 11) and a monotone decrease of average recall with m and with k.
Hosted on GitHub Pages from the main branch root: https://aryanbajaj868.github.io/bdh-synaptic-memory/ . Any push to main redeploys within a couple of minutes.
- "(0.8)¹⁹ ≈ 1.4%":
0.8**19= 0.0144. - "two random keys share about k²/n neurons": expected |x_i ∩ x_q| for two independent k-subsets of n is k·(k/n) = k²/n. At k = 14, n = 64: 3.06.
- "measured 2.35 vs expected 2.38": run
node tests/core_test.js. - All BDH/BDH-CQ figures: page/section references in
SOURCES_AND_LICENSES.md.
Cited beside the technical claims in index.html and CONCEPT_SUMMARY.md:
- Kosowski, Uznański, Chorowski, Stamirowska, Bartoszkiewicz. The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain. arXiv:2509.26507, 2025. — defines σ as Hebbian synaptic attention state (eq. 6, 8).
- Engdahl, Kosowski, Chorowski, Stamirowska, Uznański, Jiang, Phadke, Kinas, Zhong. BDH-CQ: In-Context Learning with Recurrent Latent Reasoning. arXiv:2608.09888, 2026. — contextual memory S_t = U_θ(S_{t−1}, D_t); linear-attention special case (§3.2).
- Behrouz, Zhong, Mirrokni. Titans: Learning to Memorize at Test Time. arXiv:2501.00663, 2024. — fast-weight neural memory updated at inference, with learned forgetting.
- Yang, Kautz, Hatamizadeh. Gated Delta Networks: Improving Mamba2 with Delta Rule. arXiv:2412.06464, ICLR 2025. — fixed-size associative state with a delta-rule write that removes the interference this toy exhibits.
- Sun, Li, Dalal, Xu, Vikram, Zhang, Dubois, Chen, Wang, Koyejo, Hashimoto, Guestrin. Learning to (Learn at Test Time): RNNs with Expressive Hidden States. arXiv:2407.04620, 2024. — hidden state as a model updated by gradient steps at test time.
Background (pre-2022): Hinton & Plaut 1987; Ba et al. 2016; Katharopoulos et al. 2020; Schlag, Irie & Schmidhuber 2021.
Code and text: MIT (see LICENSE). No third-party code, fonts, images, or data are bundled; system fonts only. Full record in SOURCES_AND_LICENSES.md. AI assistance: see AI_DISCLOSURE.md.
- Team: hackerz
- Member: Aryan Bajaj, IIT (BHU) Varanasi
- Mentorship: none