Dragon Hatchling's attention has no softmax. Remove it and the sum over past tokens can be reassociated — so “compare the current token against every previous one” and “carry one fixed-size matrix of synapse strengths” stop being alternatives. They become the same arithmetic.
The panel on the right is running both, right now, on the kernel released in pathwaycom/bdh. Scroll and you can break it.
Because BDH's attention applies no softmax, it is a Hebbian synaptic memory: one fixed-size matrix σ, written once per token by an outer product and read once per token by a dot product, computing exactly what all-pairs attention computes — but with a bounded capacity, so it forgets through interference unless its activations stay sparse.
Two halves, two ways to prove me wrong. Both are controls on this page.
Who this is for. You have implemented or read scaled dot-product attention and you know what an outer product is. You do not need to have read the BDH paper. Everything below runs in your browser in float64; nothing is pre-rendered video.
This is Attention.forward from bdh.py, the reference implementation Pathway published with the Dragon Hatchling paper (MIT licence). Reproduced verbatim:
QR = self.rope(r_phases, Q) KR = QR # note: K is Q scores = (QR @ KR.mT).tril(diagonal=-1) # strictly causal, s < t return scores @ V
What is absent is the load-bearing part. There is no softmax, no /√d, no row normalisation. scores goes straight into a matmul with V. So for one query position t:
out_t = Σs<t (qr_t · kr_s) · v_s # as written: every pair
= qr_t · ( Σs<t kr_s ⊗ v_s ) # the sum moves inside
= qr_t · σ_t
That middle step is the entire argument. A dot product is linear in its second argument, so the weighted sum of value vectors can be replaced by a single dot product against an accumulated matrix. Softmax would put a nonlinearity between the score and the sum and the step would be invalid.
Which leaves a two-line recurrence, one pass, no history:
read out_t = qr_t · σ_t write σ_(t+1) = σ_t + kr_t ⊗ v_t # rank-1 Hebbian update
Read before write, because tril(diagonal=-1) excludes the diagonal — a token may not attend to itself. σ has shape n×d whatever the sequence length. That is the claim's mechanism in full.
Both implementations below are written independently in src/bdh-kernel.js — quadratic() builds the full T×T score matrix; recurrent() never allocates it and carries σ instead. Same inputs, same float64, no shared code path.
Each token adds one rank-1 outer product to σ. Nothing else ever happens to it. Drag the scrubber to step through the sequence and watch the memory accumulate — row i is neuron i's outgoing synapses, column j a coordinate of the value it writes.
A fixed-size σ cannot store unlimited associations. Write K key–value pairs into one σ by the same Hebbian rule, then read each key back. The retrieved vector is the true value plus cross-talk from every other pair:
k_b · σ = (k_b · k_b) v_b + Σa≠b (k_b · k_a) v_a
↥ what you want ↥ interference
Here is the part that makes this specifically about BDH. Its keys are ReLU outputs, so they are non-negative. Two random non-negative vectors have a strictly positive dot product — they cannot be near-orthogonal. Dense non-negative keys are close to the worst case for superposition. Sparsity is the repair: it is what buys the capacity back.
Dragon Hatchling is described at two levels, and conflating them is the most common mistake made about it.
| BDH | BDH-GPU | |
|---|---|---|
| What it is | A graph of n neurons with local interaction rules — the biologically-motivated model. | The tensorised special case that is actually trained on GPUs. |
| The state | σ(i,j) on every edge: an n×n synaptic matrix. | Never materialises σ. Keeps ρ = Eσ, shape n×d, reached by linear attention. |
| Synapse graph | G_s is a trainable parameter. | Effectively complete and not trainable — a genuine restriction. |
| Where the memory is | On connections. | In a low-rank compression of the same object. |
The σ on this page is the BDH-GPU object — n×d, built by linear attention. That is what the released code computes, and it is why the heatmap above is a tall thin rectangle rather than a square.
The paper does call BDH a state-space architecture, but in the broad sense it defines — parameters plus an evolving state — under which a Transformer with a KV cache also qualifies. Mechanically it is a linear-attention associative memory, and two specifics rule out the Mamba reading: the state is a matrix built from outer products rather than a per-channel vector recurrence, and the positional operator is a fixed rotation, not an input-dependent selective gate. There is no selectivity mechanism, and no HiPPO-style initialisation. Calling it “a Mamba variant” gets the mechanism wrong.
The paper reports roughly 5% non-zero activations in trained BDH-GPU, emergent rather than imposed — L1 regularisation was explicitly disabled. Lab 2 shows what that buys: at the same width and load, moving from dense to ~5% density lifts retrieval fidelity several-fold. Sparsity is not a compression convenience here. It is the thing that makes a non-negative associative memory work at all.
The toy running everywhere else on this page is untrained, so its ReLU passes about half its coordinates. Rather than cite the 5% and move on, we trained our own small BDH to see whether sparsity emerges — because “unless its activations stay sparse” is part of the claim, and a claim should rest on a measurement where one is affordable.
BDH-CQ is a later system in the same family, evaluated on ARC-AGI-1. Its report describes a contextual state S updated per demonstration with fixed weights, and names linear attention as “the conceptually simplest standalone realization” of that update, in the special case S_t = S_(t−1) + U(D_t) — additive accumulation, structurally the same move as the σ recurrence on this page. Beyond that structural echo I make no claim: U, E, F and G are not published, so nothing here reproduces BDH-CQ. The honest statement is that this page explains the class of update BDH-CQ's report points at, not BDH-CQ itself.
Every BDH number in circulation sits at a different evidential level. Mixing them is how an explainer becomes marketing.
| Claim | Level | Where it actually comes from |
|---|---|---|
| Pairwise form = synaptic form, no softmax | Algebraic identity | Follows from linearity. Re-derived and measured on this page (~10−16); reproducible via node src/verify-equivalence.js. |
| Fixed σ interferes as load grows | Measured here | Lab 2, synthetic keys, deterministic seed. A property of superposition, not of any trained model. |
| ~5% activation sparsity | Developer-reported | Dragon Hatchling paper, §6.4. Emergent, L1 disabled. Not independently reproduced. |
| Matches GPT-2-class scaling, 10M–1B | Developer-reported | Same paper, §4.2, against a nanoGPT/Transformer-XL-style baseline on byte-level data. |
| Monosemantic synapses | Narrow evidence | Same paper, §6.3 — a small number of named synapses. Suggestive, not a systematic interpretability result. |
| Sudoku Extreme 97.4% | Company blog only | Not in the arXiv paper — the string “Sudoku” does not appear in it. Reported on Pathway's research blog, from an internal implementation that is not the public repo. |
| BDH-CQ: 29.5% pass@2 on ARC-AGI-1 at $0.0007/task | Developer-reported | arXiv:2608.09888. Public evaluation set, 400 tasks — not semi-private, not ARC-AGI-2. Model internals proprietary. |
| Scaling “1B to 600B parameters” | Asserted, no data | Two sentences in the BDH-CQ report with no table, curve, or dataset. |
| Amazon SageMaker HyperPod | Partnership announcement | Press-release material. Training infrastructure — not a benchmark, not a deployment, not an evaluation. |
The last four rows matter most. A benchmark is not a deployment, a partnership is not an independent evaluation, and a result reported by its developer is not an external reproduction. As of writing I could find no independent reproduction of BDH's headline results.
| Component | Status |
|---|---|
| Both attention forms, Lab 1 | Live — float64, computed in your browser on every control change. |
| σ heatmap and scrubber | Live — real snapshots of the accumulating state. |
| Interference curve, Lab 2 | Live — every point recomputed, averaged over repeats. |
| Weights | Random, untrained, from a seeded PRNG. This is a mechanism demo, not a trained model. |
| Everything else | Prose and tables. Nothing on this page is a pre-rendered animation. |
Stated caps. T ≤ 128, n ≤ 256 (512 in Lab 2), d ≤ 32, chosen so the slowest setting stays under ~0.3 s on a laptop and remains usable on a phone. The kernel is the released attention block only — the surrounding MLP, the gating x⊙y, LayerNorms and the multi-layer stack are not reproduced, because the claim does not need them. Values here are the raw embeddings, matching V=x in the released code.
The honest caveat about scale. Real BDH-GPU runs n in the tens of thousands. At n=128 you are seeing the mechanism, not the regime. Sparse superposition gets better with width, so this toy understates how well σ holds up.
Answer before opening. If you can do these, you have the concept.
Softmax normalises scores across all source positions, so the weight on v_s depends on every other score in the row. The sum cannot move inside the dot product, and there is no fixed-size object to accumulate — which is exactly why the KV cache has to keep every past key and value.
Different labs, different variables. Lab 1's T is sequence length; Lab 2's K is how many associations share one σ. Interference is driven by how much is superposed relative to width n and how orthogonal the keys are — not by wall-clock sequence length on its own.
Because n×d is a constant, not a small one. Below the crossover reported in Lab 1 the T×T score matrix is genuinely smaller. The win is asymptotic: σ stops growing while the cache never does. Claiming an unconditional memory saving would be false, and the page tells you the crossover so you can check.
Every technical claim on this page sits beside the source it comes from. If a sentence has no source and is not something you can reproduce with a control, treat it as my inference and check it.