Cypha · v2.4.0

An AI architecture built from first principles, not a fork

Almost all modern Artificial Intelligence (AI) is a variation on the same few designs. Cypha starts somewhere else: one C++ type, with a learning rule worked out from mathematics rather than borrowed, that updates on every example it sees instead of being trained once and frozen. It is a research project, and the failures are published alongside the results.

LanguageC++
Releasev2.4.0
WikiText-2, Bits Per Character (BPC)2.664 BPC
Last pushrecently
Four programmes, one learning rule
  • AIXI / Solomonoff — Minimum Description Length (MDL) priors decide what the model is allowed to believe cheaply.
  • Information geometry — updates follow the natural gradient, so learning respects the curvature of the parameter space.
  • Active inference / free energy — gives the structural decomposition into prior, differential and context.
  • Information bottleneck — sets the encoder objective: keep what predicts, discard the rest.
In plain English

What this is, in one minute

The problem

Nearly every AI system in use is a variation on a handful of designs. They are trained once on a large machine and then shipped frozen, so keeping up with something that moves — a new kind of fraud, a new fault — means going back and training again. They are also big, which puts them in a data centre rather than on the equipment, and makes explaining any single decision hard.

The solution

Cypha is one C++ type that classifies, predicts numbers and generates sequences. Its learning rule is derived from four bodies of mathematics rather than copied from an existing architecture, and it updates on every example it sees — no batch, no epoch, no retraining run. It is small enough to run on the hardware that collects the data.

Who it is for

Today, researchers and engineers willing to try a different design. The problems it is aimed at are the ones where the data keeps moving — fraud, equipment faults, changing behaviour — where the model has to sit on the device rather than in a data centre, and where somebody has to explain the decision afterwards.

Judge it as research. The figures further down come from the repository’s own benchmark runs, checked against committed test fixtures rather than a leaderboard position. A section headed What Cypha is not sets out the limits, including the standard test where it sits near chance until one particular component is switched on. The rest of this page is the architecture and the measurements.
The library itself

Random features, measured rather than described

Cypha’s case for random Fourier features (RFF) is that they let a linear head separate what a linear head cannot. That is a measurable claim, so this measures it — with cypha::rff_features compiled to WebAssembly, not a retelling in JavaScript.

cypha::rff_features — WebAssembly

not loaded

The exact Radial Basis Function (RBF) kernel over n sampled points gives a matrix K. Random features reconstruct an approximation K̂. What is plotted is ‖K − K̂‖F as the number of features grows: more features, closer approximation.

Library
—
Compiling this found a bug. Orthogonal Random Features (ORF) — the dense variant — were normalising every row to unit length, when Yu et al. draw the row norm from chi_d so an orthogonal row matches the Gaussian row it replaces. Rows were √d too short, so the features approximated the wrong kernel and the error plateaued instead of converging. One line. It is now the best of the three.
Single-threaded on purpose: threaded WebAssembly needs SharedArrayBuffer, which needs the Cross-Origin-Opener-Policy (COOP) and Cross-Origin-Embedder-Policy (COEP) headers GitHub Pages cannot send.
Architecture

Seven layers, each doing one job

Every component below exists because one of the four programmes demands it — not because it appeared in a paper that quarter.

The Cypha pipeline in order: encoder, projection, world prior, class differentials, memory and tiered context, producing a class with confidence, anomaly score and an out-of-distribution flag
scroll to see the whole diagram →
One type, four jobs. The demo further down this page is this exact pipeline in two dimensions — the world prior is the dashed ellipse, the differentials are the bars.
Encoder

Pluggable front end

Raw input to feature vector. Ships VectorEncoder, RFFEncoder (random Fourier features) and ConcatEncoder. Swap it without touching anything downstream.

Projection

EncoderProjection

Features into latent space via Fisher–Rao contrastive updates, with Frobenius-norm capping so a single outlier cannot blow the geometry apart.

θ₀

WorldPrior

A shared diagonal Gaussian fitted online by Welford and Exponential Moving Average (EMA) updates. This is the “infinite context” that never forgets — and whose movement is the drift signal.

Δₖ

ClassDifferential

Per-class natural-parameter offsets, attracted toward observations and pulled back by MDL decay. A class is a displacement from the world, not a separate model.

Memory

DIFMemory

Computes log-likelihood ratios under generalised-hyperbolic posteriors, with Bessel-ratio lookup tables so the heavy tails do not cost a transcendental per sample.

Context

TieredContextBuffer

Short, mid and long tiers weighted by NIGField confidence, so recent evidence can dominate without erasing what the long tier established.

Classification, end to end: encode → project → score each class’s log-likelihood ratio against the world prior → return the argmax, together with a confidence, an anomaly score and an out-of-distribution flag. The demo below is that exact pipeline, in miniature.
Interactive

Train a classifier by clicking

A faithful 2-D miniature of the Cypha pipeline: a world prior fitted online, per-class differentials attracted toward what you place, and classification by log-likelihood ratio. It learns from every single click — there is no batch, no epoch, no restart.

cypha::Cypha — online classifier

Click the canvas to add a sample of the selected class. Drag to paint a cluster.

Class to place

Load a dataset

Samples
0
Accuracy
—
Max LLR
—
Prior drift
0.00

Last inference

Idle
Hover the canvas to classify a point without training on it.

Class differentials Δₖ

A 2-D teaching implementation written for this page — same structure as the real pipeline, none of its scale. The C++ implementation is in native/.
Try Exclusive OR (XOR) with the RFF encoder off. Accuracy collapses toward chance, and that is not a bug in this demo — it is the documented limitation of a linear log-likelihood ratio, stated plainly in the Cypha README. Switch the latent RFF encoder on and watch it recover. Publishing the failure mode next to the fix is the whole point.
Measured

Results, with the comparison that matters

Figures from the repository’s own benchmark runs after the diagnostic fix. Correctness is proved by a CTest matrix matching committed fixture goldens — not by a leaderboard position.

Abbreviations in the table: Stochastic Gradient Descent (SGD), Long Short-Term Memory (LSTM), Backpropagation Through Time (BPTT).

DatasetTaskCyphaOnline SGDNote
Linearly separable2-class0.7830.644Same online budget
Iris3-class0.900—Classical baseline set
Wine3-class0.969— 
Digits10-class0.922— 
Breast cancer2-class0.957— 
WikiText-2Sequence, 300k tokens2.664 BPC—Hybrid GRIA+LSTM L2+Wave2 BPTT
XOR2-class, latent RFF~0.763—Linear LLR alone caps near chance
Generation

Temperature-scaled, field-conditioned, latent-boundary interpolation, adversarial (entropy-maximising), OOD sampling, MDL-constrained, ancestral, and KDE sampling from the replay buffer.

Anomaly & active learning

Anomaly scores from gate values, active-query scores as entropy × boundary proximity, and drift detection read directly off world-prior movement.

Replay

A 10,000-capacity priority buffer weighted by recency and surprise, replayed at a 0.30 ratio — so the rare, informative sample is not drowned by the common one.

Reference defaults

Profiled, not guessed

These come out of a profiled medium-grid optimisation and ship as the reference configuration.

Feature dim128
RFF budget256
Replay ratio0.30
Context window32
World prior LR0.008
Class diff LR0.05
Encoder LR0.002
MDL lambda0.001
Honest framing

What Cypha is not

  • Not a wrapper. Bespoke from first principles — which also means it does not inherit anyone else’s tuning.
  • Linear LLR has a ceiling. XOR sits near chance without the latent RFF encoder, which lifts it to roughly 76.3%.
  • Validation is parity-based. CTest against committed fixture goldens. No leaderboard claims are made.
  • CUDA is inference-only. Training stays on the Central Processing Unit (CPU), because on this architecture the CPU is faster.
  • Theory lives elsewhere. Harmonic-spectrum and NMP work is a separate compression-algorithms paper; Cypha is the implementation layer.
Product demo

Cypha, distilled from a real chess engine

26,568 positions labelled with a conventional alpha-beta engine’s own search evaluations, fitted with the same WorldPrior whitening and natural-gradient updates used everywhere else in the architecture. It reproduces the teacher’s evaluation at R² 0.866 on held-out positions, and scores 5W–19L–6D against that teacher at equal search depth.

Who it is for

A small AI that keeps learning

Most AI is trained once in a data centre and shipped frozen. Cypha is small enough to run on ordinary hardware and carries on learning from each new example it sees.

01

Devices too small for big AI

Sensors, cameras, controllers and other hardware with no room for a data centre behind it. Cypha is small enough to run on the device itself, so nothing has to leave the building.

02

Anything that has to keep up with change

Fraud patterns, equipment faults, shifting customer behaviour. A model trained last year is already out of date. This one updates as it goes, without being pulled offline and retrained.

03

Work that has to be explained afterwards

Banks, insurers and health services often have to justify a decision. Cypha is small enough to inspect, and it publishes what it is bad at rather than hiding it.

Recognise your situation here? This is open for beta testing now, and the people it is built for are the ones whose feedback actually changes it. Become a beta tester →
Build it

Three commands

$ cmake -S native -B /tmp/cypha_build -DCMAKE_BUILD_TYPE=Release -G Ninja
$ cmake --build /tmp/cypha_build --parallel
$ ctest --test-dir /tmp/cypha_build -R native_ --output-on-failure

# REST service
$ cypha_rest --listen 127.0.0.1:8099 --cypha fixtures/reference.cypha
# Qt shell (build with -DCYPHA_BUILD_QT=ON)
$ cypha_qt_shell
Licensing terms