Research shelf / AI & machine learning / UHPM

AI & machine learning

UHPM — memory retrieval and inference under one loss function

If memory retrieval and inference are both Bayesian operations, they should share one loss function rather than being welded together at runtime. UHPM is what happens when you take that seriously: an LSH-based memory hierarchy and hierarchical predictive coding derived from a single variational free-energy functional.

Result-bearing AGPL-3.0+ / commercial
Evidence level

Experiments were run and the numbers are reported here.

FolderLong Reasoning and Thinking NN
FieldAI & machine learning
StatusPaper, 60-page derivation, Python implementations, and a runnable demonstration on synthetic data.
What it is

Locality-sensitive-hash memory and hierarchical predictive coding unified under a single free-energy functional, reporting a 289× query-latency speedup over full attention at 100K tokens.

UHPM (Unified Hash-Predictive Memory) merges LSH-based memory with hierarchical predictive coding under one variational framework. The argument is structural: two systems that are both doing Bayesian work should not be optimising two different objectives and reconciling the results at inference time.

The memory is a three-level hierarchy over 100 / 1,000 / 10,000-token segments with 64-bit LSH, compressed to a centroid dimension of d′ = 64. Inference runs roughly ten Bayesian refinement iterations with a stopping criterion of ‖Δs‖ < 10⁻³.

The efficiency numbers are large, and so are the caveats attached to them — the benchmarks are synthetic topic-cluster corpora with fixed, non-overlapping segments, and retrieval fidelity is 80–90% of exact attention rather than 100%.

Read the 289× carefully. It is a real measurement on a synthetic corpus with fixed, non-overlapping segments and static random hashes — close to the best case this design can face. The honest headline is not "289× faster than attention"; it is "289× on a benchmark constructed to suit it, with 80–90% fidelity, and slower than k-NN under 7,500 tokens".
Claims ledger

Every number, and what stands behind it

A claim is only worth the evidence attached to it. Each row below carries its basis: measured on the author’s own hardware, derived from the construction, measured on synthetic data, projected from literature, or simply cited.

Breakdown of this page’s claims by what stands behind each one
scroll to see the whole chart →
Every claim, weighted by its evidence. The table below is the same data row by row.
ClaimFigureBasisContext
Query-latency speedup at 100K tokens289× (8.1 ms vs 2,340 ms)Syntheticvs full attention, synthetic corpus
Memory reduction744× (2.2 MB vs 1,638 MB)Syntheticvs full attention, synthetic corpus
Retrieval fidelity80–90% of exact attentionSyntheticExplicitly not 100%
Memory hierarchy100 / 1,000 / 10,000-token segments, 64-bit LSHDerivedArchitecture parameter
Compressed centroid dimensiond′ = 64DerivedArchitecture parameter
Refinement iterations~10, stopping at ‖Δs‖ < 10⁻³DerivedInference procedure
Crossover pointSlower than simple k-NN below ~7,500 tokensSyntheticHonest negative result

Measured — author-run experiment on the stated setup. Synthetic — measured, but on synthetic rather than real data. Derived — follows from the stated construction or proof. Projected — paper-stated projection, not an author-run benchmark. Cited — taken from external literature.

Methods

How it works

  • Single free-energy functional. Memory and inference derived from one objective rather than composed at runtime.
  • Three-level LSH hierarchy. 64-bit hashes over nested segment sizes.
  • Iterative Bayesian refinement. Roughly ten iterations with an explicit convergence criterion.
  • 60-page derivation. The mathematical framework is written out, not asserted.
Stated limitations

What it does not do

Taken from the folder’s own README. Nothing here has been softened.

  • Experiments use synthetic topic-cluster corpora with fixed, non-overlapping segments — not natural text.
  • Hashes are random and static rather than learned. Learned hashing is flagged as future work.
  • Retrieval fidelity is 80–90% of exact attention, not equivalence.
  • Slower than simpler k-NN approaches below roughly 7,500 tokens.
  • Preprint dated March 2026. Not peer-reviewed.
Use it

Free under AGPL-3.0+ for almost everyone

Personal use, charities, education and organisations under AUD 50,000 a year pay nothing. A tiered commercial licence covers everyone else.