A sequence model that is honestly worse
A language model built on a completely different mechanism — closer to chemistry spreading across a surface than to the attention maths everything else uses. It performs worse than a transformer, and its own front page says so. Not slightly worse: four orders of magnitude on perplexity (PPL), the standard measure, against Generative Pre-trained Transformer 2 (GPT-2) as the baseline, and printed on this page rather than in a footnote. It is published as a research log because a documented failure is worth something and a deleted one is not.
Training perplexity of 450,000 against GPT-2’s 20 is not a near miss. It is four orders of magnitude, and no amount of framing changes that.
The value is the E0–E26 experiment ladder with its failures annotated: which architectural ideas collapsed, how, and what the diagnostics said. A negative result with a documented cause is a contribution. A negative result quietly deleted is a waste of the work.
What this is, in one minute
Almost every AI that handles language is built on the same idea: let every word look at every other word. It works. Nobody knows how much of that is because it is the best idea available and how much is because it is the one everybody tried.
Cell AI takes that out and puts something else in: two chemicals spreading and reacting across a surface, the way stripes form on an animal, with the connections changing while it reads rather than afterwards. It was then trained at full size and measured. It came out four orders of magnitude worse than GPT-2. That is the headline of this page, not a line in an appendix.
Researchers who want to know what happens down the road nobody took. Scientists who already model living systems, because the mechanism underneath is the one that makes patterns in chemistry and markings on animals. Teachers, because published research is almost all successes, and a careful account of a failure is rarer and more useful.
The dynamics underneath the architecture
This is a Gray–Scott reaction-diffusion system — the same class of dynamics the CellularPDE core is built on. Two chemicals diffuse at different rates and react; feed and kill rates decide whether you get spots, stripes, dividing cells or nothing at all. The architecture’s bet is that this kind of self-organising field can carry computation. Watch it and judge the plausibility for yourself.
CellularPDE — reaction-diffusion core
Click or drag on the field to inject chemical B and perturb the pattern.
Partition state
N = 4 partitions of the field, as in the CellularPDE core. Height is each partition’s mean activation; leakage λ = 0.01 couples them.
Gray–Scott with Da=1.0, Db=0.5, nine-point Laplacian, explicit Euler. This is the dynamical substrate, not the trained model — the trained model is 125.8M parameters and does not fit in a browser tab.
CellularPDE is built from. The architecture itself is PyTorch.
What replaces attention
CellularPDE
N = 4 partitions of D = 256-dimensional state, leakage λ = 0.01, sigmoid reaction terms — a discrete reaction-diffusion system standing where self-attention would be.
MetaplasticityLayer
Hebbian learning in the Bienenstock–Cooper–Munro (BCM) style, with α = 0.1, β = 0.01, applied during the forward pass rather than after it. This is the boldest idea in the architecture and the source of its documented gradient-starvation problem.
SpectralPDE & SparseHebbian
SpectralPDE cuts complexity from O(D²) to O(D log D). SparseHebbian compresses updates from 65,536 weights to 1,024. MultiScalePartitions run fast and slow streams; AnnealedRouter uses Gumbel-Softmax scheduling.
The full stack
cl100k_base embedding → CellularPDE → MetaplasticityLayer → MemoryFormation → ResonanceSystem → CrystalLattice (K = 3) → optional Kuramoto coupling → MultiModalModel routing → logits. The resonance stage couples phases with a Fast Fourier Transform (FFT).
E0 to E26, failures annotated
The architecture-search programme is the actual artefact here. Twenty-seven experiments, each with what was tried and what went wrong, including the ones that went backwards.
| Finding | Outcome |
|---|---|
| PerFreqResonance | Documented issues, did not deliver |
| Routing | Collapsed to class degeneracy in some configurations |
| E26 vs E25 | Regression — the later experiment was worse |
| ResonanceSystem | Minimal gradient contribution |
| Generation (v1) | Lacks proper autoregressive sampling |
| v3 Cell-Fungal Harmonic | Speculative, unintegrated |
The gap, stated plainly
An experiment in building AI a completely different way
Almost every AI today is built on the same underlying idea. Cell AI is a full-scale attempt at a different one, borrowed from how patterns form in nature — published with its failures intact.
Researchers looking past the current approach
Progress needs someone to try the road not taken and write down what happened. This is that attempt, at full size, with the dead ends recorded rather than quietly dropped.
Scientists modelling living systems
The mechanism underneath is the one that produces stripes on animals and patterns in chemistry. If that is the world you work in, the model speaks your language.
Teachers who want an honest failure
Published research is overwhelmingly about things that worked. This is a careful account of something that did not, which is far more useful for learning than another success story.
Where these ideas went instead
Cell AI is the research line that did not converge. Its siblings did better: Cypha became a working first-principles architecture, and the long-context memory work became Unified Hash-Predictive Memory (UHPM).