Research shelf / AI & machine learning / VDJ

AI & machine learning

The immune system’s combinatorics, borrowed for embedded pattern recognition

A vertebrate immune system generates recognition capability for antigens it has never seen, from a finite segment library, with no training corpus. That is an appealing property for anything that has to work on scarce data in a small memory budget, and it is a combinatorial mechanism before it is a biological one.

Result-bearing AGPL-3.0+ / commercial
Evidence level

Experiments were run and the numbers are reported here.

FolderVDJ Inspired Algorithm
FieldAI & machine learning
StatusResult-bearing on performance. Reference implementation, instrumented profile.
What it is

V(D)J recombination abstracted into five modules for one-shot learning and combinatorial generation, profiled to the millisecond and the kilobyte for embedded deployment.

Four properties of the RAG1/RAG2 recombination machinery are abstracted into software: combinatorial assembly from a finite segment library, geometric progression weighting across recombination depth, pattern-driven state transitions, and single-example generalisation. Nothing here is a neural network; it is a combinatorial system with a biological derivation.

The framework is five modules — OneShotLearner, PatternRecognizer, CombinatorialGenerator, MetaPatternProcessor, SpaceExplorer — plus seven supporting subsystems, all communicating through one typed Pattern dataclass. The mathematical foundations of each module are derived, and the whole is profiled by instrumented execution rather than estimated.

The geometric 1/2ᵏ weighting produces a clean engineering consequence: marginal information gain from deeper combinations falls below 1.6% past depth r = 6, which is a mathematical justification for a cap rather than a tuned hyperparameter. The target deployment is defence, embedded and real-time work where data scarcity, interpretability and resource limits rule out large-scale statistical learning.

The r = 6 cap is the nicest result here. Most systems cap search depth because someone tried a few values. This one caps it because the geometric weighting makes the marginal information gain past depth six smaller than 1.6% — the hyperparameter falls out of the construction. That is what a good abstraction buys you.
Claims ledger

Every number, and what stands behind it

A claim is only worth the evidence attached to it. Each row below carries its basis: measured on the author’s own hardware, derived from the construction, measured on synthetic data, projected from literature, or simply cited.

Breakdown of this page’s claims by what stands behind each one
scroll to see the whole chart →
Every claim, weighted by its evidence. The table below is the same data row by row.
ClaimFigureBasisContext
Full pipeline latency13.0 ms (σ = 4.4 ms)Measuredn = 16, r = 5; NumPy 2.4.2, Python 3.12, CPU only
Peak memory footprint997 KBMeasuredSame configuration
Marginal gain past depth r = 6<1.6%DerivedConsequence of the geometric 1/2ᵏ weighting
Topological fingerprinting + spatial exploration<2.5 ms at n = 64MeasuredEffectively free at all tested input sizes
Module count5 primary + 7 supportingDerivedArchitecture as specified
Learning regimesingle exampleDerivedInherited from the V(D)J abstraction

Measured — author-run experiment on the stated setup. Synthetic — measured, but on synthetic rather than real data. Derived — follows from the stated construction or proof. Projected — paper-stated projection, not an author-run benchmark. Cited — taken from external literature.

Methods

How it works

  • Segment-library recombination. Finite V, D and J segment sets recombined combinatorially, as the biological mechanism does.
  • Geometric depth weighting. 1/2ᵏ decay across recombination depth, which both scores candidates and bounds useful depth.
  • Typed Pattern dataclass. A single typed interchange format between all twelve subsystems — the reason the modules compose.
  • Topological fingerprinting. Scale-invariant shape features, measured to be near-free at the tested sizes.
Stated limitations

What it does not do

Taken from the folder’s own README. Nothing here has been softened.

  • All performance figures are single-machine, CPU-only, NumPy. No embedded-target measurements despite embedded being the stated use case.
  • Combinatorial systems scale by construction: n = 64 is the largest size profiled and the growth beyond it is not characterised empirically.
  • One-shot generalisation is demonstrated by the mechanism, not benchmarked against few-shot statistical baselines.
  • The biological derivation motivates the design; it is not evidence that the design is good.
  • No accuracy results on a standard pattern-recognition dataset.
Use it

Free under AGPL-3.0+ for almost everyone

Personal use, charities, education and organisations under AUD 50,000 a year pay nothing. A tiered commercial licence covers everyone else.