Research shelf / AI & machine learning / UHPM
Written Updated
UHPM — memory retrieval and inference under one loss function
If memory retrieval and inference are both Bayesian operations, they should share one loss function rather than being welded together at runtime. UHPM is what happens when you take that seriously: an LSH-based memory hierarchy and hierarchical predictive coding derived from a single variational free-energy functional.
Experiments were run and the numbers are reported here.
Locality-sensitive-hash memory and hierarchical predictive coding unified under a single free-energy functional, reporting a 289× query-latency speedup over full attention at 100K tokens.
UHPM (Unified Hash-Predictive Memory) merges LSH-based memory with hierarchical predictive coding under one variational framework. The argument is structural: two systems that are both doing Bayesian work should not be optimising two different objectives and reconciling the results at inference time.
The memory is a three-level hierarchy over 100 / 1,000 / 10,000-token segments with 64-bit LSH, compressed to a centroid dimension of d′ = 64. Inference runs roughly ten Bayesian refinement iterations with a stopping criterion of ‖Δs‖ < 10⁻³.
The efficiency numbers are large, and so are the caveats attached to them — the benchmarks are synthetic topic-cluster corpora with fixed, non-overlapping segments, and retrieval fidelity is 80–90% of exact attention rather than 100%.
Every number, and what stands behind it
A claim is only worth the evidence attached to it. Each row below carries its basis: measured on the author’s own hardware, derived from the construction, measured on synthetic data, projected from literature, or simply cited.
| Claim | Figure | Basis | Context |
|---|---|---|---|
| Query-latency speedup at 100K tokens | 289× (8.1 ms vs 2,340 ms) | Synthetic | vs full attention, synthetic corpus |
| Memory reduction | 744× (2.2 MB vs 1,638 MB) | Synthetic | vs full attention, synthetic corpus |
| Retrieval fidelity | 80–90% of exact attention | Synthetic | Explicitly not 100% |
| Memory hierarchy | 100 / 1,000 / 10,000-token segments, 64-bit LSH | Derived | Architecture parameter |
| Compressed centroid dimension | d′ = 64 | Derived | Architecture parameter |
| Refinement iterations | ~10, stopping at ‖Δs‖ < 10⁻³ | Derived | Inference procedure |
| Crossover point | Slower than simple k-NN below ~7,500 tokens | Synthetic | Honest negative result |
Measured — author-run experiment on the stated setup. Synthetic — measured, but on synthetic rather than real data. Derived — follows from the stated construction or proof. Projected — paper-stated projection, not an author-run benchmark. Cited — taken from external literature.
How it works
- Single free-energy functional. Memory and inference derived from one objective rather than composed at runtime.
- Three-level LSH hierarchy. 64-bit hashes over nested segment sizes.
- Iterative Bayesian refinement. Roughly ten iterations with an explicit convergence criterion.
- 60-page derivation. The mathematical framework is written out, not asserted.
What it does not do
Taken from the folder’s own README. Nothing here has been softened.
- Experiments use synthetic topic-cluster corpora with fixed, non-overlapping segments — not natural text.
- Hashes are random and static rather than learned. Learned hashing is flagged as future work.
- Retrieval fidelity is 80–90% of exact attention, not equivalence.
- Slower than simpler k-NN approaches below roughly 7,500 tokens.
- Preprint dated March 2026. Not peer-reviewed.
Free under AGPL-3.0+ for almost everyone
Personal use, charities, education and organisations under AUD 50,000 a year pay nothing. A tiered commercial licence covers everyone else.