SENTINEL · Phases 1–2 complete

Investigative leads that can be traced back

Investigators are rarely short of data. What is hard is finding the pattern inside thousands of records that look unremarkable one at a time. SENTINEL is a desktop tool that reads public crime data, runs the statistics over it and ranks what is worth a closer look — and every result traces back to the exact records, model and settings it came from. If an analyst cannot show why a lead ranked where it did, it should not be a lead.

LanguageC++23
User Interface (UI)Qt 6, 9-page dashboard
Tests495 automated
StorageSQLite
Last pushrecently
Stated up front

This is analytical software that ranks locations and associations from public crime data. It is a tool for directing an analyst’s attention, not a system that decides anything about a person.

The calibration and fairness pages exist because a model that is well ranked and badly calibrated is more dangerous than no model. Phase 3–5 planning is explicitly fairness-first, and published benchmarks are a stated deliverable rather than a marketing claim.

In plain English

What this is, in one minute

The problem

A burglary, a theft from a car, an insurance claim. Each one is unremarkable on its own, and what connects them only shows up when somebody lines hundreds of them up and looks. That is slow work by hand, and the analytics contracts that do it at scale are priced out of reach of most forces.

The solution

SENTINEL reads the records in and runs established statistical models over them: where crime is dense, which incidents look like one series rather than many, which area a series was most likely run from. It ranks what is worth a closer look. Every stage writes down what it did, so a lead arrives carrying its source records, the model and the settings behind it.

Who it is for

Smaller and regional forces, where the large analytics contracts are out of reach — this runs on one ordinary desktop computer and is built for a single analyst rather than a department. Fraud and repeat-offender investigators, where the pattern sits across incidents nobody filed together. Researchers, who can retrace any result to the records that produced it.

It points an analyst at something worth checking. It does not decide anything about a person. The rest of this page is the engineering: two of the real models running in your browser, the pipeline that feeds them, what works today, and what does not.
The models themselves

Two of the real models, compiled to WebAssembly

models/KDEHotspot and models/HawkesProcess — the same C++ an analyst runs on the desktop, compiled with Qt for WebAssembly and fed points you place yourself. Click the map to add incidents.

KDEHotspot + HawkesProcess — WebAssembly

not loaded

2.4 megabytes (MB) because Qt6Core comes with it. That is the honest cost of running the real model rather than a lookalike, and nothing downloads until you ask.

Incidents
0
Branching ratio
—
Branching ratio is what the Hawkes model learned: the expected number of follow-on incidents each incident triggers. Near zero means the points carry no self-excitation — scatter them evenly and it should say so.
The models only. Ingest, the database, provenance and the dashboard are the application, and the application is not here — SENTINEL’s own build needs Qt6::Test, which needs Qt6::Concurrent, which a single-threaded WebAssembly Qt does not have.
Pipeline

Ingest → Natural Language Processing (NLP) → features → models → inference

Each stage writes what it did into the provenance chain, so a lead at the end carries the record of how it got there.

The SENTINEL pipeline from ingest through rule-based NLP, statistical models and inference to ranked leads, with a provenance rail connected to every stage
scroll to see the whole diagram →
The rail underneath is the product. Every stage writes what it did into the provenance chain, so a lead arrives with the record of how it got there.
Ingest

UK Police Application Programming Interface (API) and historical feeds, Comma-Separated Values (CSV) import with auto-detected UK / US-city / generic layouts, weather via Open-Meteo, and data-quality scoring with quarantine thresholds.

NLP

Classical, rule-based modus-operandi extraction and crime classification. No language model in the loop — deliberately, because the output has to be explainable.

Models

Poisson counts, Hawkes self-exciting processes, series detection by Density-Based Spatial Clustering of Applications with Noise (DBSCAN), KDE hotspots, Gaussian process regression, Bayesian and ensemble layers.

Inference

Rossmo geographic profiling, MO analysis, evidence scoring, anomaly detection, co-offending PageRank and community detection, and the hint engine.

Interactive

Two models over the same incidents

Place incidents on the map. KDE asks where crime is dense. Rossmo’s formula asks something different and more interesting: given that offenders avoid a buffer zone around home, where does the offender most likely live? Both run live, over the same points you place.

sentinel — spatial inference

Click to add an incident. The surfaces recompute on every point.

Incidents
0
Peak density
—

Top-ranked lead

No leads
Place at least three incidents.

Provenance chain

Rossmo is contested. Geographic profiling performs well on series by a single offender with a stable anchor point, and poorly when those assumptions break. Treating its peak as evidence about a person would be a misuse of it.
Live implementations of Gaussian KDE and Rossmo’s criminal geographic targeting formula, written for this page. The shipping tool adds Hawkes processes, DBSCAN series detection and Gaussian process regression over real ingested data.
Complete

What works today

  • End-to-end pipeline from ingest to ranked leads
  • 495 automated tests
  • Real-data validation on UK and US datasets
  • Windows and Linux packaging
  • Local read-only Representational State Transfer (REST) API
  • Nine-page dashboard: events, map, calibration, cases, co-offending graph, leads, audit log, settings, debug
Limitations

What does not

  • Rule-based NLP only. Modus-operandi extraction is classical, with the recall limits that implies.
  • Single-machine desktop. No multi-user deployment, no shared case state across analysts.
  • Limited live connectors. Few real-time data sources are wired up.
  • Basic map. Custom-rendered, not a Geographic Information System (GIS). Optional GIS integration is a Phase 3–5 item.

Phases 3–5, planned

Multi-jurisdiction support, optional GIS integration, published benchmarks, fairness-first deployment, and a possible managed service. The backlog is in REMAINING.md rather than in a roadmap deck.

Who it is for

Finds patterns in crime data, and shows its working

Analysts drown in incident records. This looks for the patterns across them and ranks what is worth a closer look, while keeping a trail back to the evidence behind every suggestion.

01

Smaller and regional police forces

The large analytics contracts are out of reach on most budgets. This runs on one ordinary desktop computer and is meant for a single analyst rather than a department.

02

Fraud and repeat-offender investigators

Insurance fraud and organised retail theft look like unrelated incidents until someone lines them up. This does the lining up, and points at where the next one is likely to be.

03

Researchers studying crime and policing

Every result can be traced back to the exact records and settings that produced it, so another researcher can check the work rather than take it on trust.

Recognise your situation here? This is open for beta testing now, and the people it is built for are the ones whose feedback actually changes it. Become a beta tester →
Provenance

Every lead carries its own receipt

Source record, model, parameters, and the transformation at every stage. An audit log viewer is one of the nine dashboard pages, not a feature request. For software that ranks where police attention goes, that is the minimum bar — and it is the reason the research prototype behind it, ARIA-INTEL, is labelled as not ready for the same use.