An AI architecture built from first principles, not a fork
Almost all modern Artificial Intelligence (AI) is a variation on the same few designs. Cypha starts somewhere else: one C++ type, with a learning rule worked out from mathematics rather than borrowed, that updates on every example it sees instead of being trained once and frozen. It is a research project, and the failures are published alongside the results.
- AIXI / Solomonoff — Minimum Description Length (MDL) priors decide what the model is allowed to believe cheaply.
- Information geometry — updates follow the natural gradient, so learning respects the curvature of the parameter space.
- Active inference / free energy — gives the structural decomposition into prior, differential and context.
- Information bottleneck — sets the encoder objective: keep what predicts, discard the rest.
What this is, in one minute
Nearly every AI system in use is a variation on a handful of designs. They are trained once on a large machine and then shipped frozen, so keeping up with something that moves — a new kind of fraud, a new fault — means going back and training again. They are also big, which puts them in a data centre rather than on the equipment, and makes explaining any single decision hard.
Cypha is one C++ type that classifies, predicts numbers and generates sequences. Its learning rule is derived from four bodies of mathematics rather than copied from an existing architecture, and it updates on every example it sees — no batch, no epoch, no retraining run. It is small enough to run on the hardware that collects the data.
Today, researchers and engineers willing to try a different design. The problems it is aimed at are the ones where the data keeps moving — fraud, equipment faults, changing behaviour — where the model has to sit on the device rather than in a data centre, and where somebody has to explain the decision afterwards.
Random features, measured rather than described
Cypha’s case for random Fourier features (RFF) is that they let a linear head
separate what a linear head cannot. That is a measurable claim, so this measures it
— with cypha::rff_features compiled to WebAssembly, not a retelling
in JavaScript.
cypha::rff_features — WebAssembly
not loadedThe exact Radial Basis Function (RBF) kernel over n sampled points gives a matrix K. Random features reconstruct an approximation K̂. What is plotted is ‖K − K̂‖F as the number of features grows: more features, closer approximation.
chi_d so an orthogonal row matches the
Gaussian row it replaces. Rows were
√d too short, so the features approximated the wrong kernel and the error
plateaued instead of converging. One line. It is now the best of the three.
Seven layers, each doing one job
Every component below exists because one of the four programmes demands it — not because it appeared in a paper that quarter.
Pluggable front end
Raw input to feature vector. Ships VectorEncoder,
RFFEncoder (random Fourier features) and ConcatEncoder.
Swap it without touching anything downstream.
EncoderProjection
Features into latent space via Fisher–Rao contrastive updates, with Frobenius-norm capping so a single outlier cannot blow the geometry apart.
WorldPrior
A shared diagonal Gaussian fitted online by Welford and Exponential Moving Average (EMA) updates. This is the “infinite context” that never forgets — and whose movement is the drift signal.
ClassDifferential
Per-class natural-parameter offsets, attracted toward observations and pulled back by MDL decay. A class is a displacement from the world, not a separate model.
DIFMemory
Computes log-likelihood ratios under generalised-hyperbolic posteriors, with Bessel-ratio lookup tables so the heavy tails do not cost a transcendental per sample.
TieredContextBuffer
Short, mid and long tiers weighted by NIGField confidence, so recent evidence can dominate without erasing what the long tier established.
Train a classifier by clicking
A faithful 2-D miniature of the Cypha pipeline: a world prior fitted online, per-class differentials attracted toward what you place, and classification by log-likelihood ratio. It learns from every single click — there is no batch, no epoch, no restart.
cypha::Cypha — online classifier
Click the canvas to add a sample of the selected class. Drag to paint a cluster.
Class to place
Load a dataset
Last inference
Class differentials Δₖ
Results, with the comparison that matters
Figures from the repository’s own benchmark runs after the diagnostic fix. Correctness is proved by a CTest matrix matching committed fixture goldens — not by a leaderboard position.
Abbreviations in the table: Stochastic Gradient Descent (SGD), Long Short-Term Memory (LSTM), Backpropagation Through Time (BPTT).
| Dataset | Task | Cypha | Online SGD | Note |
|---|---|---|---|---|
| Linearly separable | 2-class | 0.783 | 0.644 | Same online budget |
| Iris | 3-class | 0.900 | — | Classical baseline set |
| Wine | 3-class | 0.969 | — | |
| Digits | 10-class | 0.922 | — | |
| Breast cancer | 2-class | 0.957 | — | |
| WikiText-2 | Sequence, 300k tokens | 2.664 BPC | — | Hybrid GRIA+LSTM L2+Wave2 BPTT |
| XOR | 2-class, latent RFF | ~0.763 | — | Linear LLR alone caps near chance |
Temperature-scaled, field-conditioned, latent-boundary interpolation, adversarial (entropy-maximising), OOD sampling, MDL-constrained, ancestral, and KDE sampling from the replay buffer.
Anomaly scores from gate values, active-query scores as entropy × boundary proximity, and drift detection read directly off world-prior movement.
A 10,000-capacity priority buffer weighted by recency and surprise, replayed at a 0.30 ratio — so the rare, informative sample is not drowned by the common one.
Profiled, not guessed
These come out of a profiled medium-grid optimisation and ship as the reference configuration.
What Cypha is not
- Not a wrapper. Bespoke from first principles — which also means it does not inherit anyone else’s tuning.
- Linear LLR has a ceiling. XOR sits near chance without the latent RFF encoder, which lifts it to roughly 76.3%.
- Validation is parity-based. CTest against committed fixture goldens. No leaderboard claims are made.
- CUDA is inference-only. Training stays on the Central Processing Unit (CPU), because on this architecture the CPU is faster.
- Theory lives elsewhere. Harmonic-spectrum and NMP work is a separate compression-algorithms paper; Cypha is the implementation layer.
Cypha, distilled from a real chess engine
26,568 positions labelled with a conventional alpha-beta engine’s own search evaluations, fitted with the same WorldPrior whitening and natural-gradient updates used everywhere else in the architecture. It reproduces the teacher’s evaluation at R² 0.866 on held-out positions, and scores 5W–19L–6D against that teacher at equal search depth.
A small AI that keeps learning
Most AI is trained once in a data centre and shipped frozen. Cypha is small enough to run on ordinary hardware and carries on learning from each new example it sees.
Devices too small for big AI
Sensors, cameras, controllers and other hardware with no room for a data centre behind it. Cypha is small enough to run on the device itself, so nothing has to leave the building.
Anything that has to keep up with change
Fraud patterns, equipment faults, shifting customer behaviour. A model trained last year is already out of date. This one updates as it goes, without being pulled offline and retrained.
Work that has to be explained afterwards
Banks, insurers and health services often have to justify a decision. Cypha is small enough to inspect, and it publishes what it is bad at rather than hiding it.
Three commands
$ cmake -S native -B /tmp/cypha_build -DCMAKE_BUILD_TYPE=Release -G Ninja $ cmake --build /tmp/cypha_build --parallel $ ctest --test-dir /tmp/cypha_build -R native_ --output-on-failure # REST service $ cypha_rest --listen 127.0.0.1:8099 --cypha fixtures/reference.cypha # Qt shell (build with -DCYPHA_BUILD_QT=ON) $ cypha_qt_shell