DataForge 2026 · Pathway Track· Concept: Synaptic Plasticity as Short-Term Memory

The synapse is the cache

Dragon Hatchling's attention has no softmax. Remove it and the sum over past tokens can be reassociated — so “compare the current token against every previous one” and “carry one fixed-size matrix of synapse strengths” stop being alternatives. They become the same arithmetic.

The panel on the right is running both, right now, on the kernel released in pathwaycom/bdh. Scroll and you can break it.

Live · both forms, same inputrunning
—
relative difference between
● pairwise form   ● synaptic form
Computing…
tokens T
—
neurons n
—
values d
—
The claim
One falsifiable sentence

Because BDH's attention applies no softmax, it is a Hebbian synaptic memory: one fixed-size matrix σ, written once per token by an outer product and read once per token by a dot product, computing exactly what all-pairs attention computes — but with a bounded capacity, so it forgets through interference unless its activations stay sparse.

Two halves, two ways to prove me wrong. Both are controls on this page.

  • “Exactly” — switch the softmax on in Lab 1. If the equality survives a nonlinearity, my explanation of why it holds is wrong. It does not survive: the relative difference jumps from ~10−16 to ~101.
  • “Forgets through interference” — in Lab 2, push the number of stored associations up while holding n fixed. If retrieval stayed perfect, σ would be a lossless cache and the claim would be false. It degrades, and how fast depends on sparsity.

Who this is for. You have implemented or read scaled dot-product attention and you know what an outer product is. You do not need to have read the BDH paper. Everything below runs in your browser in float64; nothing is pre-rendered video.

Where the claim comes from

Four lines of released code

This is Attention.forward from bdh.py, the reference implementation Pathway published with the Dragon Hatchling paper (MIT licence). Reproduced verbatim:

QR = self.rope(r_phases, Q)
KR = QR                                   # note: K is Q
scores = (QR @ KR.mT).tril(diagonal=-1)   # strictly causal, s < t
return scores @ V

What is absent is the load-bearing part. There is no softmax, no /√d, no row normalisation. scores goes straight into a matmul with V. So for one query position t:

out_t = Σs<t (qr_t · kr_s) · v_s        # as written: every pair
      = qr_t · ( Σs<t kr_s ⊗ v_s )      # the sum moves inside
      = qr_t · σ_t

That middle step is the entire argument. A dot product is linear in its second argument, so the weighted sum of value vectors can be replaced by a single dot product against an accumulated matrix. Softmax would put a nonlinearity between the score and the sum and the step would be invalid.

Which leaves a two-line recurrence, one pass, no history:

read  out_t   = qr_t · σ_t
write σ_(t+1) = σ_t + kr_t ⊗ v_t     # rank-1 Hebbian update

Read before write, because tril(diagonal=-1) excludes the diagonal — a token may not attend to itself. σ has shape n×d whatever the sequence length. That is the claim's mechanism in full.

Why RoPE does not spoil it. The rotation is applied to queries and keys at their own absolute positions before the product, so R_t q and R_s k can be accumulated independently — the relative rotation appears in the dot product for free. A decay term would need more care; a position-dependent rotation does not.
Lab 1 · Equivalence

Run both. Then break one.

Both implementations below are written independently in src/bdh-kernel.js — quadratic() builds the full T×T score matrix; recurrent() never allocates it and carries σ instead. Same inputs, same float64, no shared code path.

Controls shrunken toy · not an official BDH model
How many tokens the layer reads.
Key/query width. In BDH this is very large.
Values stay in the small model dimension.
Deterministic. Same seed, same numbers.
relative diff
—
max abs diff
—
output scale
—
ReLU nonzero
—
compute time
—
—
● pairwise form — output rows, one per token
● synaptic form — same matrix, computed from σ
Signed difference, autoscaled to its own maximum. With softmax off this is pure rounding noise — the colour scale is amplifying ~10−16.
Read the cost, not just the equality. σ is n×d numbers no matter how long the sequence is; the score matrix is T×T and grows. At small T the score matrix is smaller — the crossover on the current settings is at — tokens. The recurrent form is an asymptotic win, not a free one, and saying otherwise would be overclaiming.
The state, made visible

Watch the wiring change

Each token adds one rank-1 outer product to σ. Nothing else ever happens to it. Drag the scrubber to step through the sequence and watch the memory accumulate — row i is neuron i's outgoing synapses, column j a coordinate of the value it writes.

σ at token t — live Hebbian write
Read happens before write, so σ at t excludes token t itself.
Rows = neurons (keys). Columns = value coordinates. Magenta positive, teal negative, brightness proportional to magnitude. Rows that stay blank are neurons the ReLU never switched on.
writes so far
0
‖σ‖ Frobenius
0
σ entries
—
rows ever written
—
Lab 2 · Interference

Where it breaks, and what saves it

A fixed-size σ cannot store unlimited associations. Write K key–value pairs into one σ by the same Hebbian rule, then read each key back. The retrieved vector is the true value plus cross-talk from every other pair:

k_b · σ = (k_b · k_b) v_b + Σa≠b (k_b · k_a) v_a
           ↥ what you want   ↥ interference

Here is the part that makes this specifically about BDH. Its keys are ReLU outputs, so they are non-negative. Two random non-negative vectors have a strictly positive dot product — they cannot be near-orthogonal. Dense non-negative keys are close to the worst case for superposition. Sparsity is the repair: it is what buys the capacity back.

Retrieval fidelity vs. load truth beside estimate
How many associations share one σ.
Fraction of neurons active per key. BDH reports ~5% in trained models.
The width the associations superpose into.
mean fidelity
—
worst pair
—
mean |k_a·k_b|
—
σ entries
—
—
Fidelity as K grows, at the current density. Dot marks your K. Recomputed live, not a stored curve.
One retrieved value (magenta) against the value that was actually stored (teal). The gap is the lesson.
What this is and is not. This lab uses random synthetic keys to isolate the capacity effect, so the numbers describe superposition in a fixed-size matrix, not the accuracy of any trained BDH model. A trained model chooses its keys; random keys are the pessimistic case. The mechanism is BDH's; the numbers are this toy's.
The BDH module

Which system, and what is actually changing

Dragon Hatchling is described at two levels, and conflating them is the most common mistake made about it.

 BDHBDH-GPU
What it isA graph of n neurons with local interaction rules — the biologically-motivated model.The tensorised special case that is actually trained on GPUs.
The stateσ(i,j) on every edge: an n×n synaptic matrix.Never materialises σ. Keeps ρ = Eσ, shape n×d, reached by linear attention.
Synapse graphG_s is a trainable parameter.Effectively complete and not trainable — a genuine restriction.
Where the memory isOn connections.In a low-rank compression of the same object.

The σ on this page is the BDH-GPU object — n×d, built by linear attention. That is what the released code computes, and it is why the heatmap above is a tall thin rectangle rather than a square.

Why this is not “an SSM” in the Mamba sense

The paper does call BDH a state-space architecture, but in the broad sense it defines — parameters plus an evolving state — under which a Transformer with a KV cache also qualifies. Mechanically it is a linear-attention associative memory, and two specifics rule out the Mamba reading: the state is a matrix built from outer products rather than a per-channel vector recurrence, and the positional operator is a fixed rotation, not an input-dependent selective gate. There is no selectivity mechanism, and no HiPPO-style initialisation. Calling it “a Mamba variant” gets the mechanism wrong.

Sparsity is load-bearing

The paper reports roughly 5% non-zero activations in trained BDH-GPU, emergent rather than imposed — L1 regularisation was explicitly disabled. Lab 2 shows what that buys: at the same width and load, moving from dense to ~5% density lifts retrieval fidelity several-fold. Sparsity is not a compression convenience here. It is the thing that makes a non-negative associative memory work at all.

The toy running everywhere else on this page is untrained, so its ReLU passes about half its coordinates. Rather than cite the 5% and move on, we trained our own small BDH to see whether sparsity emerges — because “unless its activations stay sparse” is part of the claim, and a claim should rest on a measurement where one is affordable.

Precomputed · our own tiny BDH, trained offline precomputed · not an official BDH model
parameters
—
neurons n
—
val loss
—
nonzero before
—
nonzero after
—
Mean fraction of non-zero ReLU activations across all four layers, measured on held-out data during training. L1 regularisation disabled, matching the paper's setup. Dashed line is the ~5% the paper reports for trained BDH-GPU — shown as a reference we did not reach, not as a target we hit.
What this does and does not show. Sparsity moved in the reported direction under training — from — to — non-zero, with no sparsity penalty in the loss. It did not reach ~5%, and we are not claiming to have reproduced the paper's figure. Our model is roughly 0.23M parameters against their 10M–1B, and our corpus is a small synthetic grammar the model nearly saturates (validation loss — nats/byte), so there is little for a token to be “busy” about — and the paper notes sparsity tracks how much work a token requires. The honest reading: the direction reproduces at toy scale; the magnitude is a property of scale and data we cannot test here. Regenerate with python train/train_tiny_bdh.py.

Where BDH-CQ fits — and where it does not

BDH-CQ is a later system in the same family, evaluated on ARC-AGI-1. Its report describes a contextual state S updated per demonstration with fixed weights, and names linear attention as “the conceptually simplest standalone realization” of that update, in the special case S_t = S_(t−1) + U(D_t) — additive accumulation, structurally the same move as the σ recurrence on this page. Beyond that structural echo I make no claim: U, E, F and G are not published, so nothing here reproduces BDH-CQ. The honest statement is that this page explains the class of update BDH-CQ's report points at, not BDH-CQ itself.

Evidence discipline

What is proven, measured, or merely announced

Every BDH number in circulation sits at a different evidential level. Mixing them is how an explainer becomes marketing.

ClaimLevelWhere it actually comes from
Pairwise form = synaptic form, no softmaxAlgebraic identityFollows from linearity. Re-derived and measured on this page (~10−16); reproducible via node src/verify-equivalence.js.
Fixed σ interferes as load growsMeasured hereLab 2, synthetic keys, deterministic seed. A property of superposition, not of any trained model.
~5% activation sparsityDeveloper-reportedDragon Hatchling paper, §6.4. Emergent, L1 disabled. Not independently reproduced.
Matches GPT-2-class scaling, 10M–1BDeveloper-reportedSame paper, §4.2, against a nanoGPT/Transformer-XL-style baseline on byte-level data.
Monosemantic synapsesNarrow evidenceSame paper, §6.3 — a small number of named synapses. Suggestive, not a systematic interpretability result.
Sudoku Extreme 97.4%Company blog onlyNot in the arXiv paper — the string “Sudoku” does not appear in it. Reported on Pathway's research blog, from an internal implementation that is not the public repo.
BDH-CQ: 29.5% pass@2 on ARC-AGI-1 at $0.0007/taskDeveloper-reportedarXiv:2608.09888. Public evaluation set, 400 tasks — not semi-private, not ARC-AGI-2. Model internals proprietary.
Scaling “1B to 600B parameters”Asserted, no dataTwo sentences in the BDH-CQ report with no table, curve, or dataset.
Amazon SageMaker HyperPodPartnership announcementPress-release material. Training infrastructure — not a benchmark, not a deployment, not an evaluation.

The last four rows matter most. A benchmark is not a deployment, a partnership is not an independent evaluation, and a result reported by its developer is not an external reproduction. As of writing I could find no independent reproduction of BDH's headline results.

No hidden limits

What this page is

ComponentStatus
Both attention forms, Lab 1Live — float64, computed in your browser on every control change.
σ heatmap and scrubberLive — real snapshots of the accumulating state.
Interference curve, Lab 2Live — every point recomputed, averaged over repeats.
WeightsRandom, untrained, from a seeded PRNG. This is a mechanism demo, not a trained model.
Everything elseProse and tables. Nothing on this page is a pre-rendered animation.

Stated caps. T ≤ 128, n ≤ 256 (512 in Lab 2), d ≤ 32, chosen so the slowest setting stays under ~0.3 s on a laptop and remains usable on a phone. The kernel is the released attention block only — the surrounding MLP, the gating x⊙y, LayerNorms and the multi-layer stack are not reproduced, because the claim does not need them. Values here are the raw embeddings, matching V=x in the released code.

The honest caveat about scale. Real BDH-GPU runs n in the tens of thousands. At n=128 you are seeing the mechanism, not the regime. Sparse superposition gets better with width, so this toy understates how well σ holds up.

Explain it back

Three questions

Answer before opening. If you can do these, you have the concept.

Why can't you do this trick to a standard Transformer's attention?

Softmax normalises scores across all source positions, so the weight on v_s depends on every other score in the row. The sum cannot move inside the dot product, and there is no fixed-size object to accumulate — which is exactly why the KV cache has to keep every past key and value.

You doubled T and the fidelity in Lab 2 didn't move. Why not?

Different labs, different variables. Lab 1's T is sequence length; Lab 2's K is how many associations share one σ. Interference is driven by how much is superposed relative to width n and how orthogonal the keys are — not by wall-clock sequence length on its own.

If σ is fixed-size, why isn't it always cheaper than a KV cache?

Because n×d is a constant, not a small one. Below the crossover reported in Lab 1 the T×T score matrix is genuinely smaller. The win is asymptotic: σ stops growing while the cache never does. Claiming an unconditional memory saving would be false, and the page tells you the crossover so you can check.

Sources

Primary, and where to go next

Every technical claim on this page sits beside the source it comes from. If a sentence has no source and is not something you can reproduce with a control, treat it as my inference and check it.