Cypha / Chess

Product demo

Play chess against Cypha

Cypha never saw a chess game. It learned by watching an ordinary chess program think — 26,568 positions, each one labelled with that program’s own verdict on who was winning — until Cypha could produce the verdict itself.

What you play below is that, running in your browser inside a shallow search. It is beatable, and it still loses to the program it learned from. Every number on this page came out of the training run that produced the weights the page loads.

Measured, not claimed

Root Mean Square Error (RMSE), below, is the usual measure of how far off a set of guesses was; the smaller it is, the closer the model sits to its teacher.

Held-out R²
—
RMSE
—
Positions
—
vs teacher
—

R² is against held-out positions the model never saw during fitting. The match line is Cypha at search depth 2 against the teacher engine at the same depth, with randomised openings and colours alternated. All four numbers describe the model as distilled — they are fixed, and they are not measurements of your copy.

Your copy is still learning. Cypha is an online learner, so the head does not stop fitting when the download finishes. Every move it plays against you is one more training pair, and the panel below the board shows how far your copy has moved away from the one that shipped.
And it does not yet make it play better. Measured, not assumed. Over 12 self-play games the correction reduces the head’s error against its own search from 0.76 to 0.67 pawns — and where the frozen head drifts 4.6% worse across a session, the learning one gets 5.2% better. So the fitting is real. But a copy trained on 1,119 positions then played against the shipped copy, 24 games at depth 2 with colours alternated, finishes 6W–7L–11D: a dead heat.

That is the expected result and it is worth stating rather than hiding. The signal is the model’s own depth-2 search, which is barely deeper than the static head’s own view — learning to agree with yourself is not the same as getting stronger. A thousand positions against the 26,568 it was distilled on is also nothing. The honest claim is that the online learner works and is doing what it says; the claim that it will beat the shipped model has not been earned.
In plain English

What this is, in one minute

The problem

Every new Artificial Intelligence (AI) design arrives with an explanation of why it ought to work. Reading the explanation tells you nothing about whether it does. Chess settles that, because the answer is not a matter of opinion: either it plays sensible legal chess or it does not, and you can sit down and find out.

The solution

An ordinary chess program judges a position by searching ahead through the moves. Cypha was shown 26,568 of those judgements and fitted a single formula to reproduce them — one line of maths over 388 numbers describing the board. Nothing was added to the architecture for chess. Only the numbers fed into it changed.

Who it is for

Anyone weighing up Cypha who would rather check than be told. People who work with machine learning and want to watch a model carry on learning while it is being used. And anyone who simply fancies a game — it all runs in your browser, and nothing you play is sent anywhere.

Copying a chess engine’s judgement is not learning chess. Nothing on this page claims it is. The rest of the page is how it was done, what each number means, and the four things this design cannot do.

cypha::Cypha — chess head learning

Click a piece, then click where it goes. Legal destinations are marked.

Position

Loading model…
Fetching the distilled weights.
Cypha eval
—
Nodes
—

What Cypha considered

Moves

Learning from you

Rating at this depth
—
Model in use
depth 2
Positions learned
0
Games finished
0
Drift from shipped
0.00%
Last correction
—

Drift is the size of the change to the weight vector, relative to the distilled one. Last correction is how wrong the static head was about the position it just moved from, in pawns, measured against its own deeper search.

Rating is on an internal ladder, not the Fédération Internationale des Échecs (FIDE) scale — there is no honest way to put a FIDE number on something that has never played a rated human. The reference engine at depth 1 is anchored to 1000 and everything is fitted from 210 games by Bradley-Terry maximum likelihood, so it does not depend on the order they were played. The measured ladder: reference 1000 / 1171 / 1245 at depths 1–3, Cypha 798 / 1010 / 1161. Cypha runs about one search depth behind the engine it was distilled from, and the error bars are roughly ±50, which is what 210 games buys you.

Each strength setting keeps its own model. A head playing at depth 1 is corrected toward a depth-1 search; at depth 3 the target is a different, better one. Training both into one weight vector would just have them argue. Switch strength and the counters change with it.

The evaluation is a linear function of 388 features. It has no search depth of its own, no opening book and no endgame knowledge. Beating it is expected; the interesting part is that a single whitened linear head reproduces an engine’s search evaluation at all — and then keeps correcting itself while you play.
Engine, features, distilled weights and player all run locally. Nothing is sent anywhere, and no model is downloaded beyond a few kilobytes of coefficients. What your copy learns is kept in your browser, belongs to you, and is never uploaded — so nobody else’s games can move your model, and yours cannot move theirs.
Method

How an engine becomes a Cypha model

This is knowledge distillation, done with the same components as every other Cypha application. Nothing chess-specific was added to the architecture — only the features.

The distillation chain: a perft-verified reference engine labels 26,568 positions, which become 388 features, fitted into a Cypha head of 9.4 kilobytes achieving held-out R squared of 0.866
scroll to see the whole diagram →
Distillation, not discovery. Every number in the last box was measured by the training run that produced the weights this page loads.
Step 1

The teacher

A 0x88 engine with full legal move generation, material and piece-square evaluation, Most Valuable Victim – Least Valuable Attacker (MVV-LVA) ordering and quiescence search. Verified against five standard perft positions including kiwipete, so the rules are provably right before anything is learned from it.

Step 2

The data

Positions drawn from self-play with randomised openings and deliberate noise, so the set spans real games rather than one narrow line. Each is labelled with the teacher’s own search score, clipped to ±12 pawns to keep mate scores from dominating.

Step 3

The features

388 dimensions: a 384-wide piece-square occupancy difference (6 types × 64 squares) plus bishop pair, doubled pawns, mobility and game phase. Always from the side to move’s perspective, so one weight vector serves both colours.

Step 4

The fit

WorldPrior θ₀ fitted online by Welford, whitening as the natural-gradient metric, normalised Least Mean Squares (LMS) weight updates, Minimum Description Length (MDL) decay pulling weights back toward the prior, and a surprise-weighted replay buffer at Cypha’s reference 0.30 ratio.

Why this is a fair demonstration

The architecture did the work, not the domain

Every component in step 4 is a real Cypha component doing its real job. The WorldPrior is the same online Gaussian that gives the 2-D classifier on the Cypha page its decision surface. The whitening is the information-geometry claim made concrete: updates follow the natural gradient rather than the raw one. MDL decay is the Solomonoff prior expressed as a shrinkage term.

What changed for chess is the feature map and the head — regression instead of classification. That is the argument the Cypha README makes about being one type that does several jobs, tested on a domain where the answer is unambiguous: either it plays legal, reasonable chess or it does not.

What this does not show

The honest limits

  • It is linear. A linear function of piece-square features cannot represent king safety, pawn structure interactions, or anything requiring a product of two features. It has the same ceiling the Cypha README describes for Exclusive OR (XOR).
  • The teacher is modest. The reference engine is a competent classical searcher, not Stockfish. Distilling it well means matching a modest evaluator well.
  • No search depth of its own. Strength here comes mostly from the alpha-beta wrapper, exactly as it does for the teacher.
  • Distillation is not discovery. Cypha reproduced an evaluation function it was shown. It did not learn chess from scratch, and nothing on this page claims it did.
Reproduce it: node tools/chess/perft.js verifies the engine against the standard test positions, and node tools/chess/train.js regenerates the model from scratch with a fixed seed. The numbers at the top of this page are read directly out of the file that run produces.