Predictive Policing in the United Kingdom: NDAS, HART, the Gangs Violence Matrix, and AI Crime Mapping
Research cutoff: August 2, 2026
Executive findings, scope, and definitions
“Predictive policing” in the United Kingdom is not a single national system. It is a loose family of technologies and analytical practices that differ fundamentally in what they predict, whom they affect, and how closely they are connected to coercive police decisions. The systems examined in this report fall into four functional categories:
- Place-based forecasting, which identifies streets, cells, or neighborhoods where recorded crime is expected to concentrate.
- Person-based prediction, which estimates an individual’s future offending, victimization, or vulnerability.
- Watchlisting or prioritization, which ranks named people on the basis of recorded conduct, intelligence, associations, or biometric similarity.
- Conventional intelligence analysis, which searches, links, summarizes, or maps existing information without necessarily estimating a future event.
These distinctions matter legally and technically. A machine-learning forecast that an arrestee will commit serious violence within two years is different from an intelligence spreadsheet ranking people by previous offenses; both are different again from a map identifying streets with repeated robberies. Yet all three may redirect patrols, surveillance, safeguarding, diversion, stop-and-search activity, or investigative resources, and therefore may produce material consequences even where no computer makes a formally binding decision.
The principal findings are as follows.
Durham Constabulary’s Harm Assessment Risk Tool, or HART, was the clearest documented example of operational person-based prediction in British policing. Its random-forest model classified an arrested person as low, moderate, or high risk according to predicted offending over the following two years. The model was trained on approximately 104,000 custody events from 2008–2012, used 34 predictors, and combined 509 decision trees. Most predictors concerned criminal history, but age, gender, and two postcode-derived variables were also used. Moderate-risk classifications were relevant to eligibility for the Checkpoint deferred-prosecution program, although custody officers retained discretion and the tool was officially described as advisory rather than determinative. Public evidence confirms use from 2016 through 2021; no sufficiently reliable public evidence was found that HART remained operational on August 2, 2026. citeturn13search2turn13search4turn13search6
The National Data Analytics Solution, or NDAS, is best understood as a multi-use data and analytics capability, not one stable “pre-crime algorithm.” It was sponsored by the Home Office, hosted by West Midlands Police, developed with Accenture, and designed to ingest police and potentially partner data for separate use cases. Some early work did contemplate person-level predictions about serious violence, including whether people already known to police might commit a first gun- or knife-related serious offense. That model was not deployed: the West Midlands Police ethics committee rejected the Most Serious Violence model in 2020 after its reported accuracy remained below 50 percent. By contrast, the publicly documented modern-slavery use case used natural-language processing, behavioral analysis, and network analysis to surface potentially overlooked cases and patterns. Official material therefore supports a mixed classification: NDAS included person-prediction research, but its viable workstreams were also forms of investigative discovery, network analysis, prevalence estimation, and improved analytical access to existing records. citeturn18view2turn18view3turn18view1
The Metropolitan Police Gangs Violence Matrix was principally a watchlist and association-based intelligence system, not a machine-learning predictive model. It ranked named “gang nominals” using offense records and police intelligence, including information about weapons, violence, victimization, and gang association. Its practical importance came less from computational sophistication than from its circulation through policing and partner institutions. Inclusion could influence targeting, stop and search, enforcement, safeguarding, prosecution narratives, and third-party treatment. In 2016, 78 percent of listed people were Black, 80 percent were aged 12–24, and 15 percent were minors. In October 2017, 3,806 people were listed, but fewer than 5 percent were categorized red and 64 percent green; Amnesty reported that 35 percent had no recorded serious offense and 75 percent had themselves been victims of violence. citeturn23view0turn14search3
The Matrix’s governance failed in ways that demonstrate why “human intelligence” is not necessarily safer than machine learning. The Information Commissioner found serious data-protection breaches, including deficient governance, retention, accuracy, differentiation, and information-sharing practices. A 2022 judicial-review settlement followed the Metropolitan Police’s acknowledgment that the Matrix required “wholesale change,” that its operation infringed Article 8, and that Black people were disproportionately represented. The Matrix ceased operation in February 2024. Its successor, the Violence Harm Assessment, or VHA, is more rules-based, has explicit inclusion thresholds, quarterly refreshes, central governance, a published standard operating procedure, and restrictions on sharing the list itself. Nevertheless, it still incorporates intelligence concerning gang association and remains highly disproportionate: as of March 7, 2026, 60.55 percent of the 1,574 listed individuals were recorded by officers as Black, 62.96 percent were under 25, and 94.16 percent were male. citeturn14search6turn14search1turn16view0turn16view3turn17view4
The national AI-supported crime-map initiative announced in August 2025 was still a research-and-development program, not an operational national prediction engine, at the cutoff date. The initial £4 million Concentrations of Crime Data Challenge sought prototypes by April 2026 and envisaged an operational system across England and Wales by 2030. In April 2026, UK Research and Innovation described a second phase within a wider £50 million Safer Streets research program. Public descriptions refer to street-level mapping, police and partner data, incident locations, crime records, and behavioral patterns of known offenders, with intended applications to knife crime, violence against women and girls, theft, retail crime, and anti-social behavior. No public model card, complete data schema, national DPIA, validation report, supplier award, false-positive analysis, or operational deployment record was identified through August 2, 2026. citeturn22view0turn22view1
The initiative differs from earlier place-based systems principally in scale, data integration, and central coordination—not yet in demonstrated predictive capability. Kent’s PredPol deployment, Metropolitan Police Risk Terrain Modelling, and the national Grip/Hot Spot Response programs all direct attention toward places. The proposed national map would add cross-agency data and standardized national infrastructure, potentially supported by the new Police.AI center. That expansion also increases the risk that a nominally place-based map will become a conduit for person-based profiling where offender histories or “behavioral patterns” are displayed alongside neighborhood forecasts. The government’s 2026 policing reform plan promises £115 million for Police.AI, a national testing-and-scaling function, responsible-deployment controls, and a public registry, but those commitments were prospective rather than an enacted comprehensive regulatory regime at the cutoff. citeturn22view2turn22view14turn22view15
This report includes a force-level system where it satisfies the following standard: it was used or seriously piloted by a UK police force or police-led partnership after 2010; it applied statistical, machine-learning, biometric, network, or automated analytical methods to estimate risk, identify concentrations, prioritize named people or places, or generate a match; and its output was capable of affecting deployment, investigation, safeguarding, diversion, enforcement, stop and search, or watchlist treatment. Purely descriptive dashboards, back-office productivity tools, and academic prototypes with no operational pathway are excluded, although failed prototypes are included where they materially shaped national policy or reveal how governance worked.
Throughout this report, “operational” means that an output was available for use in live policing or a live diversion process. It does not mean that the output automatically determined an outcome. “Pilot” means a bounded live evaluation or controlled deployment. “Prototype” means research or model development without evidence that outputs were acted upon. No system examined here has been shown to make a legally consequential police decision entirely without human involvement. That does not make the systems harmless: an advisory score may determine which cases receive scrutiny, what information reaches the decision-maker, and which options appear institutionally legitimate.
Chronological map and comparative status
Chronological map of principal programs
| Period | Program and development | Evidentially supported classification |
|---|---|---|
| 2011–2012 | Following the 2011 disturbances and renewed political focus on “gangs,” the Metropolitan Police brought the Gangs Violence Matrix into operation at the beginning of 2012. It ranked named individuals by violence-related harm and association. citeturn23view0 | Operational London-wide watchlist and intelligence-prioritization system. |
| 2012 | Durham Constabulary began developing HART with University of Cambridge researchers. citeturn13search16turn13search5 | Person-based prediction research. |
| 2013–2018 | Kent Police used PredPol to produce small-area forecasts from historical recorded-crime patterns. The force ended the arrangement in 2018 after concluding that it was difficult to demonstrate a meaningful crime-reduction benefit. | Operational place-based forecasting, later discontinued. |
| 2015–2016 | Durham established the Checkpoint deferred-prosecution experiment and began using HART to identify a moderate-risk cohort. The operational process involved custody, model scoring, eligibility review, and initially randomized assignment for evaluation. citeturn13search4turn13search14 | Live person-risk decision support within an experimental diversion program. |
| 2016–2017 | HART entered routine live use in Durham custody. Public Matrix figures showed overwhelming representation of young Black men and substantial numbers of low-harm “green” nominals. citeturn13search6turn23view0 | HART operational; Matrix operational. |
| 2018 | Amnesty published Trapped in the Matrix. The ICO investigated and issued an enforcement notice concerning the Matrix. Kent ended PredPol. citeturn2search1turn23view0 | Major regulatory and civil-society scrutiny. |
| 2018–2019 | The Home Office funded and developed NDAS, with West Midlands Police as lead and Accenture as a private-sector processor and partner. The West Midlands ethics committee was established in early 2019 and began reviewing NDAS workstreams. citeturn18view1turn18view3 | National analytics platform under development; separate proof-of-concept use cases. |
| 2019–2020 | NDAS explored Most Serious Violence, modern slavery, firearms, sexual offending, organized exploitation, and organized-crime communications. Metropolitan Police Risk Terrain Modelling linked environmental features to geographically anchored crime data. citeturn19view1turn22view14 | NDAS prototypes and controlled pilots; RTM operational analytical support. |
| 2020 | The West Midlands ethics committee unanimously advised against proceeding with NDAS’s Most Serious Violence model because reported accuracy was below 50 percent. citeturn18view3 | Person-prediction prototype stopped before deployment. |
| March 2021 | An NDAS DPIA stated that a modern-slavery controlled pilot had gone live and that additional use cases were intended for operationalization. Home Office evidence described natural-language, behavioral, and network analysis to identify cases that might otherwise go undetected. citeturn19view1turn18view2 | Controlled operational pilot, principally investigative discovery and network analysis. |
| 2021–2022 | Essex Police and the University of Essex developed and tested the Knife Crime and Violence Model/Fearless Futures, estimating which people already known to police were most likely to commit knife-enabled violence within 12 months. HART remained in use through at least 2021. citeturn22view12turn22view13turn13search6 | Essex person-prediction pilot; HART operational through documented period. |
| 2021–2023 | The Home Office funded Grip and bespoke hot-spot policing in 20 forces. The first-year evaluation estimated a 7 percent reduction in violence and robbery in hotspots on patrolled days; the following year’s meta-analysis found no statistically significant reduction. citeturn22view15 | National place-based deployment program; analytics varied by force and was not necessarily machine learning. |
| November 2022 | The Matrix judicial-review case settled. The Met accepted that the system required wholesale change, agreed to remove green nominals, and agreed to provide more information to people who asked whether they were listed and where their data had been shared. citeturn14search6turn14search14 | Operational watchlist under legally compelled reform. |
| 2023 | A New Met for London committed the force to moving away from the Matrix toward evidence-based prioritization. Essex’s pilot evaluation found no statistically reliable evidence of crime reduction in the limited pilot period. citeturn14search19turn22view13 | Transition and limited evaluation. |
| February 2024 | The Met discontinued the Gangs Violence Matrix and launched the VHA across London. citeturn14search1turn14search7 | Operational rules-based person prioritization/watchlisting. |
| 2024–2025 | The Home Office replaced Grip grants with Hot Spot Response, expanding funded hotspot policing to all 43 territorial forces in England and Wales. citeturn22view15 | Nationwide place-based resource deployment, with heterogeneous local analytics. |
| August 2025 | Government announced the £4 million Concentrations of Crime Data Challenge to build an AI-supported national crime map, with prototypes expected in 2026 and an aspirational 2030 operational date. citeturn22view0 | National R&D initiative, not yet operational. |
| April 2026 | UKRI announced phase two within a £50 million Safer Streets research-and-innovation package. The same month, the Divisional Court upheld the Metropolitan Police’s revised 2024 live-facial-recognition policy, while emphasizing that the ruling concerned the safeguards and constraints in that policy. citeturn22view1turn23view3 | National crime-map research; LFR operational under revised policy. |
| March–August 2026 | The VHA contained 1,574 individuals as of March 7. Government proposed Police.AI, £115 million over three years, a testing-and-scaling platform, a new regulatory framework, and a public police-AI registry. citeturn16view3turn22view2 | VHA operational; Police.AI and registry prospective. |
Status and evidence table
| System | Purpose, reach, and period | Inputs and output | Decision affected and human review | Validation and governance | Status on August 2, 2026 |
|---|---|---|---|---|---|
| HART / Checkpoint | Predict future offending of people brought into Durham custody; County Durham and Darlington; operational evidence from 2016–2021. | Thirty-four variables, mainly custody and offending history, plus age, gender, and postcode/Mosaic variables; low, moderate, or high two-year risk category. | Moderate category opened possible Checkpoint diversion; custody officer conducted eligibility and retained discretion. Successful completion could avoid prosecution; failure could restore prosecution. | Random-forest development and validation; later live comparison with officers. ALGO-CARE emerged partly from HART scrutiny. | No reliable evidence of continued operation after the documented period; probably retired or superseded, but no clear public decommissioning notice was identified. |
| NDAS Most Serious Violence | Intended to identify people at risk of first serious gun- or knife-related violence; participating forces, national-development ambition; 2018–2020 research. | Police-controlled records and intelligence; intended risk insight or ranked cohort. | Would have informed preventive or investigative prioritization, with officers making final decisions. | West Midlands ethics review; accuracy below 50 percent. | Prototype stopped before deployment. citeturn18view3 |
| NDAS modern slavery | Surface hidden modern-slavery cases, patterns, networks, vulnerability, and potentially prevalence; led by West Midlands Police with West Yorkshire pilot and wider participating agencies. | Structured and free-text police data; natural-language processing, behavioral and network analysis; reports, flags, linked entities, or analytical insights. | Analysts and investigators review outputs; officers remain final decision-makers. | DPIA, joint-controller arrangements, ethics review, manual bias sampling, Accenture processing contract. | Publicly documented controlled pilot/“Acceleration Stage”; current nationwide operational extent unresolved. citeturn18view1turn18view2turn19view3 |
| Gangs Violence Matrix | Rank suspected gang-associated people for violence harm across London; 2012–February 2024. | Crime records, weapons and violence information, intelligence, victimization, gang association; red/amber/green and harm ranking. | Influenced tasking, stop and search, enforcement, safeguarding, prosecution context, and partner responses. Human intelligence officers managed entries and actions. | ICO enforcement, MOPAC reviews, equality assessments, 2022 litigation settlement. | Decommissioned. Matrix datasets were scheduled for deletion after closure, creating a dispute between privacy-based deletion and preservation for legal challenge. citeturn14search10turn14news34 |
| Violence Harm Assessment | Prioritize people involved or likely to be involved in serious street violence across London; launched February 2024. | Crime reports, graded intelligence, Cambridge Harm Index and ONS sentencing weights, CRIMINT including gang-association intelligence, victim, wanted, prison, county-lines and offender-management datasets. | Local detective superintendents and specialist units decide interventions; listing does not mandate action. | Published SOP, DPIA, equality assessment, quarterly refresh, central scoring, user authorization, three-year archives, case-by-case sharing. | Operational; 1,574 people listed March 7, 2026. citeturn16view0turn16view3 |
| National AI crime map | Identify street-level concentrations and support interventions across England and Wales; announced 2025, phase two 2026, target 2030. | Proposed police, council, social-service, incident-location, criminal-record, environmental, and known-offender behavioral data; interactive map or concentration forecasts. | Intended to influence patrols and local interventions; exact review workflow unpublished. | UKRI challenge governance; no public national DPIA, model card, supplier architecture, or operational validation found. | Research and development, not established operational use. |
| Kent PredPol | Forecast small geographic cells with elevated near-term recorded-crime risk; Kent, approximately 2013–2018. | Prior incident type, location, and time; hotspot boxes. | Patrol deployment; officers and supervisors selected tactics. | Force evaluation did not establish persuasive crime-reduction benefit. | Discontinued. |
| Met Risk Terrain Modelling | Identify environmental conditions statistically associated with selected crime types across London. | Geocoded crime and environmental datasets; weighted risk terrain. | Further analyst assessment, patrol and tactical prioritization. | Internal analytical governance; public FOI description, but no comprehensive independent validation located. | Available as analytical support; extent and frequency of current use unresolved. citeturn22view14 |
| Grip / Hot Spot Response | Direct visible patrols and problem-solving to serious-violence and ASB hotspots; initially 20 forces, later all 43 territorial forces. | Local incident and intelligence data, hotspot definitions, patrol records; place lists and patrol plans. | Resource allocation to neighborhoods; human supervisors control deployment. | Home Office national evaluations; heterogeneous local implementation. | Operational funding model nationwide, but not one standardized AI product. citeturn22view15 |
| Essex Knife Crime and Violence Model | Identify people known to police who may commit knife or serious violence within 12 months; Essex. | Age, gender, drug marker, criminal associates/co-accused, prior suspect and victim records, weapons, violence and drug indicators, and geographic district; probability score and top cohort. | Multi-agency review and offers of support or intervention; humans applied further priority assessment. | University partnership, ALGO-CARE review, DPO and ethics input; limited pilot evaluation found no statistically reliable crime-reduction result. | Public current status unresolved; evidence supports pilot use, not established permanent force-wide deployment. citeturn22view12turn22view13 |
| Live facial recognition | Compare faces in public-camera feeds against police watchlists; primarily Metropolitan and South Wales Police, with expansion to other forces announced. | Live biometric templates and watchlist images; similarity alert. | Engagement officers review alerts, approach people, verify identity, and determine further action. | Bridges invalidated earlier South Wales safeguards; Thompson and Carlo upheld the Met’s revised 2024 policy. | Operational under force policies; biometric watchlist matching, not future-crime prediction. citeturn22view7turn23view3 |
The evidence is uneven. HART has unusually detailed academic documentation, but its precise retirement date is unclear. The Gangs Matrix and VHA have comparatively rich official material because litigation and public pressure forced disclosure. NDAS has high-level privacy and ethics documentation but limited public information about production architecture, operational users, procurement terms, source code, model performance, or whether specific post-2021 workstreams survived. The national map has political announcements but little technical detail. This unevenness is itself a material governance problem: the systems with the greatest potential scale are not necessarily the systems with the strongest public evidence.
Detailed HART model audit and the Checkpoint process
Purpose and decision pathway
HART was designed to answer a specific predictive question: following a custody event, what category of offending was the person expected to commit during the next 24 months? The three modeled outcomes were:
- Low: predicted to commit no new offense during the following two years.
- Moderate: predicted to reoffend, but only through non-serious offending.
- High: predicted to commit a serious offense, including homicide, attempted homicide, grievous bodily harm or similarly serious violence, robbery, sexual offending, or firearms offending.
The model’s operational importance came from Checkpoint. Checkpoint was a deferred-prosecution or diversion program directed primarily at people classified as moderate risk. A qualifying person entered a structured process with a “navigator,” underwent a needs assessment, and agreed to a four-month contract addressing relevant pathways such as substance use, housing, finances, relationships, or behavioral problems. Successful completion could result in no prosecution for the triggering offense; failure or serious non-compliance could result in the original case returning to prosecution. citeturn13search2turn13search4turn13search14
HART did not itself dismiss a case, impose punishment, or place a person into Checkpoint. In the documented process, the model was embedded in the COMET offender-management application. A forecast was run after arrest and interview; officers then checked legal and program eligibility, considered offense and evidential factors, and in the experimental phase randomized eligible participants between Checkpoint and the conventional process. Custody officers could depart from the model’s result, although the institutionally important question is whether they were practically encouraged to treat the risk category as the starting point. citeturn13search4turn13search20
The consequences were therefore asymmetric. A moderate classification could create access to a beneficial alternative to prosecution. A high classification could exclude a person from that route because the perceived public-safety risk was judged too great. A low classification could also exclude the person because the intensive intervention was not considered necessary or likely to produce a detectable effect. The classification did not merely describe risk: it helped define who was an appropriate subject for leniency, treatment, and experimental resources.
Training data and model structure
The model was trained on approximately 104,000 Durham custody events from January 2008 through December 2012 and was separately tested on approximately 15,000 custody events from 2013. The unit was a custody event rather than necessarily a unique person, so repeat arrestees could appear more than once. Outcomes were established through subsequent police records over a two-year period. citeturn13search2turn13search5
HART used a random forest composed of 509 classification and regression trees. Each tree was built from a randomized subset of cases and predictors; the forest’s category was determined by the plurality of tree “votes.” Twenty-nine of the 34 fields concerned prior criminal behavior. The remaining fields included age, gender, and residential geography. citeturn13search2turn13search7
This architecture offered several practical advantages over a single regression equation. It could capture nonlinear relationships, interactions, and rare outcomes without requiring officers to specify a simple additive formula. It also allowed Durham to assign different costs to different mistakes. But the same structure complicated explanation: a custody officer could be told the result and perhaps the most influential factors, but could not readily reconstruct why 509 trees collectively assigned a specific person to one category.
Deliberate error-cost choices
Durham explicitly treated prediction errors as unequally harmful. A person forecast as low risk who then committed a serious offense was treated as a “dangerous” false negative. A person forecast as high risk who did not commit serious offending was treated as a “cautious” false positive. The published account describes the dangerous error as being assigned approximately eight times the cost of the cautious error during model construction. citeturn13search0turn13search2
That choice was not an objective property of the data. It was a policy judgment embedded in the model. It reduced the probability that someone who would later be recorded for serious offending would be placed in the low category, but it necessarily increased the number of people overclassified as high. From a public-safety perspective, Durham could describe this as prudent. From the affected person’s perspective, the cost was the possible loss of diversion and the institutional stigma of a high-risk forecast.
The key legal and ethical point is not that asymmetric costs are inherently unlawful. Criminal-justice decisions routinely weigh risks asymmetrically. The problem is that the cost choice determines who bears error and should therefore be publicly justified. A police force should disclose the false-positive burden, the foregone-benefit consequence, subgroup effects, and alternatives considered—not merely overall accuracy.
Postcode and commercial Mosaic data
HART did not use race as an explicit input, but it used two forms of residential postcode data. One field was associated with Experian’s Mosaic geodemographic product, which classifies small areas or households according to commercial and socioeconomic characteristics. Disclosed procurement information indicated that Durham paid Experian for Mosaic information and used a reduced local set of common socio-geodemographic categories. citeturn13search1turn1search4
Removing race does not remove the possibility of racial or class-based effects. Postcode can encode deprivation, housing tenure, unemployment, household composition, migration history, local police intensity, and the cumulative effects of residential segregation. Commercial classifications may add inferences about consumer behavior or lifestyle that the person never provided to police and may not know are being used.
The strongest available criticism is therefore a proxy and legitimacy argument, not a proven finding that HART produced a particular racial error ratio. No public independent audit was identified that reported calibration, false-positive rates, and false-negative rates across all protected racial groups. That absence is significant. A claim that HART was empirically “race neutral” would be unsupported; an equally categorical claim that a particular racial disparity was statistically established by a published HART audit would also overstate the evidence.
Accuracy and validation
A later operational comparison followed people taken into custody between September 2016 and October 2017 for two years. HART identified 89.8 percent of those later brought back into custody for a serious offense as high risk, compared with 81.2 percent for custody officers. Among those forecast low, the algorithm’s prediction was correct 78.8 percent of the time, compared with 66.5 percent for officers. Across all categories, however, HART’s overall accuracy was 53.8 percent, only slightly above officers’ 52.2 percent. citeturn13search9
These figures should not be collapsed into one claim that the model was either accurate or inaccurate. The high-risk sensitivity and low-risk predictive performance were relevant to the intended triage task, while the modest overall accuracy reflected the difficulty of distinguishing three future-outcome categories. The model and officers also agreed only a little more than half the time in one comparison, demonstrating that HART was not merely automating existing professional judgment. citeturn1search14
The outcome measure also creates a serious construct-validity issue. HART predicted future offending through future appearances in police and custody data. A person who offended but was not detected could be labeled a model success in the “no offending” category. A person subject to intensive policing could be more likely to generate a recorded failure. The model therefore predicted recorded criminal-justice contact, not an omniscient measure of actual offending.
Label leakage and feedback risk
HART’s strongest predictors were prior custody and offending variables. Those variables are legitimately associated with future recorded offending, but they also reflect past police deployment, discretionary arrest, charging, recording quality, and neighborhood enforcement. If repeated police attention raises the probability of future detection, the label partly contains the effects of the institution using the model.
A particularly important form of leakage arises where an input is temporally or institutionally close to the outcome. For example, the number and recency of prior custody events may reveal not only underlying conduct but current police interest. If the model is retrained on cases whose treatment was itself influenced by earlier scores, prediction and intervention can become circular.
Checkpoint somewhat mitigated the harshest version of this loop because a moderate score could lead to support rather than intensified enforcement. Yet diversion creates its own evaluation complication: a successful intervention may prevent the very outcome the model predicted. A model can appear “wrong” because the intervention worked, while a model can appear “right” because high-risk classification led to more surveillance and detection. Proper validation must therefore separate predictive performance from treatment effects.
Human discretion and automation bias
Durham repeatedly characterized HART as decision support. Officers were expected to apply professional judgment and could override a result. This meant the tool did not constitute a solely automated legal determination in the ordinary sense. citeturn13search20turn18view4
Nevertheless, formal override power is not equivalent to meaningful human control. An officer may defer to an apparently scientific score because disagreeing creates perceived personal risk, requires additional reasons, or conflicts with organizational expectations. Conversely, broad override discretion can reintroduce inconsistent human bias and undermine any standardization benefit. A defensible process therefore requires recorded reasons both for following and departing from a score, analysis of override patterns, and auditing for whether discretion is used differently across groups.
Governance, transparency, and current status
HART was unusually visible because researchers published its structure and used the project to develop the ALGO-CARE framework: advisory use, lawful authority, granularity, ownership and control, challengeability, accuracy, responsibility, and explainability. The National Police Chiefs’ Council later encouraged ALGO-CARE as good practice, but it was not a statutory authorization or binding code. citeturn18view4
Freedom-of-information material indicates that more than 12,000 people were assessed between 2016 and 2021, in more than 22,000 assessments. Fair Trials reported 3,292 high, 12,262 moderate, and 7,111 low classifications. Those numbers derive from an advocacy organization’s publication of FOI results rather than a contemporaneous force dashboard, but they are the best publicly documented scale figures located. citeturn13search6
No clear Durham publication was identified stating the exact date on which HART ceased live operation, whether the model was retrained, whether Mosaic remained in use throughout, or whether a successor tool replaced it. The defensible status finding is therefore: operational use is documented through 2021; continued use as of August 2026 is unproven. It should not be described as a currently operating national or even current Durham system without new evidence.
NDAS architecture, workstreams, and governance
What NDAS was designed to be
Official descriptions define NDAS as a scalable and flexible UK law-enforcement analytics capability sponsored and mandated by the Home Office, led by West Midlands Police, and developed with participating forces, the National Crime Agency, and Accenture. Its purpose was to combine data from participating agencies, apply advanced analytics to designated high-priority problems, and provide reports or insights supporting evidence-based interventions. citeturn18view1turn18view2
The architecture was use-case based. NDAS was not one model with one output. Its central capability was intended to ingest structured and tabular data from crime, intelligence, and custody systems, including names, dates of birth, addresses, telephone numbers, email addresses, criminal records, and free text. A 2019 DPIA said the system would not process photographs, fingerprints, genetic data, or other multimedia for the work then assessed. It also anticipated data matching across datasets and acknowledged that much of the processing would be invisible to the data subject. citeturn19view1turn19view3
The West Yorkshire privacy notice went further by describing the ability to incorporate data from social care, local authorities, education, emergency services, private organizations, and open sources. It identified chief officers of police forces in England and Wales, the Home Secretary, and other participating law-enforcement heads as joint controllers under a joint-controller agreement. Accenture was described as a processor operating under a data-processing contract. citeturn18view1
This architecture permits several analytically distinct operations:
- entity resolution, in which references to the same person, address, telephone, or organization are linked;
- natural-language processing, in which narrative reports are searched or classified;
- network analysis, in which relationships among people, addresses, communications, or events are mapped;
- risk-factor analysis, in which variables statistically associated with an outcome are identified;
- person-level classification, in which a named individual receives a risk estimate;
- prevalence or demand estimation, in which aggregate hidden harm is estimated; and
- operational reporting, in which analysts or investigators receive a list, graph, alert, or case summary.
Calling all of these “predictive policing” obscures the relevant legal question. Searching narrative reports for indicators of labor exploitation is not equivalent to estimating that a named person will commit murder. Both require lawful, accurate, and proportionate processing, but the evidential and rights thresholds should differ.
Violent-crime workstream
Early public and leaked descriptions of NDAS emphasized gun and knife crime. The Most Serious Violence work aimed to examine whether people already known to police could be identified as at risk of committing a first qualifying serious-violence offense within a future period. That is person-based prediction rather than place forecasting or ordinary intelligence retrieval.
The model did not become an operational national risk score. In July 2020, the West Midlands ethics committee concluded that its below-50-percent accuracy was unacceptable and unanimously recommended that it not proceed. The House of Lords later highlighted the episode as an example of an oversight body preventing deployment of an insufficiently reliable system. citeturn18view3
This distinction resolves a persistent public dispute. Statements that “NDAS predicts which people will commit knife crime” accurately describe an intended prototype, but inaccurately imply that the prototype was successfully deployed throughout British policing. Conversely, official claims that NDAS merely improves analysis omit the fact that person-level serious-violence prediction was genuinely pursued. The evidence supports both propositions in sequence: NDAS attempted such prediction, and the known model was stopped before operational use.
The rejection also demonstrates the limits of generic accuracy. Serious violence is a low-base-rate outcome. Even a model that captures many eventual cases may produce a large pool of false positives. A below-50-percent three-category accuracy figure does not alone show every metric was poor, but the ethics committee evidently concluded that the package was not adequate for the proposed purpose.
Modern-slavery and organized-exploitation work
The modern-slavery use case is better documented as an operational pilot. Home Office evidence stated that it used natural-language processing, behavioral analysis, and network analysis to identify cases that might otherwise remain undetected. During ethics review, NDAS reportedly conducted manual sampling to examine whether algorithmically surfaced events were being flagged because of protected characteristics and demonstrated that protected characteristics would not be used as primary or secondary natural-language-processing rules before operational evaluation. citeturn18view2
The West Yorkshire notice described the use case as being in an “Acceleration Stage,” with data shared for research, development, validation, and operational purposes. The intended goals included preventing, detecting, and prosecuting modern-slavery crimes and safeguarding victims. citeturn18view1
This workstream could produce person-level consequences without being a future-offending prediction. An analytical flag might cause investigators to reopen a report, connect victims to an address or employer, prioritize a network, or initiate safeguarding. The relevant output is closer to a lead or investigative hypothesis than a forecast that the subject will offend in a specified period.
NDAS-related work also informed an estimate, published with the Centre for Social Justice, that the United Kingdom might contain at least 100,000 potential modern-slavery victims. That was an aggregate prevalence estimate and should not be conflated with a list of 100,000 identified people or a validated individual-risk model.
The DPIA stated that a controlled modern-slavery pilot went live in March 2021 and that NDAS intended further use cases involving modern slavery, sexual offenses, firearms, organized exploitation, and organized-crime communications. It also said the project was fully funded through March 2022. citeturn19view1
Does NDAS predict offending, victimization, locations, or networks?
The answer depends on the workstream:
| Question | Finding |
|---|---|
| Did NDAS research individual-offending prediction? | Yes. The Most Serious Violence model sought a person-level forecast, but the known model was rejected before deployment. |
| Did NDAS examine victimization or vulnerability? | Yes. Modern-slavery work involved identifying possible victims, hidden vulnerability, and exploitation indicators; some early violent-crime descriptions also referred to risk of becoming a victim. |
| Did NDAS forecast crime locations? | No strong public evidence was identified that the core NDAS workstreams became an operational national place-forecasting product. Place analysis may be technically possible within the platform, but the later national crime-map initiative was announced separately. |
| Did NDAS analyze organized networks? | Yes. Network and behavioral analysis were expressly associated with the modern-slavery work, and planned work included organized exploitation and organized-crime communications. |
| Did it improve analytical access to existing data? | Yes. Centralized ingestion, text analysis, matching, and reporting were foundational functions. |
| Was it a single automated decision-maker? | No. The DPIA stated that officers and staff would make final decisions and that existing procedures would not be replaced. citeturn19view3 |
Human review and consequences
The NDAS DPIA expressly stated that force officers and staff would always make the final decision based on their own assessment and that there would be no automated decision-making. It marked profiling or automated decisions about access to a service, opportunity, or benefit as outside the assessed operation. citeturn19view3
That statement should be interpreted narrowly. NDAS could still profile people in the broader technical sense, generate ranked leads, and influence which records receive attention. Human review reduces but does not eliminate error. An analyst presented with a network link or modern-slavery alert may treat it as corroboration even where the underlying text is ambiguous, outdated, or based on an unverified intelligence report.
Consequences can include investigation, safeguarding contact, surveillance, intelligence retention, cross-force dissemination, or a decision to allocate specialist resources. For neighborhoods, aggregated insights can redirect enforcement or prevention funding. For people incorrectly linked to exploitation, the result may be repeated police contact or inaccurate association with an organized network.
Data protection and retention
NDAS relies on the premise that participating agencies already hold data for lawful policing purposes and that applying data science is an additional means of processing it. The DPIA described the legal basis as necessity for a law-enforcement purpose and asserted strict necessity for relevant sensitive processing. citeturn19view2
That reasoning is incomplete unless applied at the level of each use case. Lawful original collection does not automatically make every subsequent cross-dataset inference necessary and proportionate. A force may lawfully hold a witness report but still require a distinct justification before using it to train a model of future violence. Purpose limitation, data minimization, accuracy, category differentiation, and sensitive-processing rules must be reassessed whenever datasets are combined or a new target variable is defined.
The West Yorkshire notice states that current datasets may be overwritten weekly and that older versions can be archived for 12 months. The public documents do not reveal a single retention period governing all source data, derived features, model-training copies, audit logs, or analytical outputs. That is an unresolved architecture issue: deleting a current analytical table does not necessarily delete the underlying police record, model artifact, extracted feature, or copied lead.
Governance and procurement
West Midlands Police’s independent ethics committee reviewed NDAS projects throughout their lifecycle. The committee had independent volunteer members, published papers and minutes, and could advise the Chief Constable and Police and Crime Commissioner. It was not a statutory regulator and its recommendations were formally advisory, although the Most Serious Violence rejection was followed. citeturn18view3
ALGO-CARE was incorporated into NDAS project initiation. The modern-slavery project also underwent a DPIA, bias sampling, and operational evaluation planning. These are stronger controls than existed for many local intelligence spreadsheets.
However, important gaps remain. The reviewed public material does not disclose:
- the complete Accenture procurement route, total contract value attributable to NDAS, subcontractors, or intellectual-property allocation;
- production source code or sufficiently detailed model cards;
- complete participating-force and agency schedules at each stage;
- operational user numbers and query logs;
- workstream-specific false-positive and false-negative results;
- how individuals can identify and challenge a derived NDAS inference;
- whether discontinued models and training data were deleted;
- whether the “Acceleration Stage” modern-slavery notice remains substantively current or is merely an undated legacy page; or
- whether NDAS functions have migrated into the Police Digital Service, Police.AI, a force data platform, or another successor capability.
Consequently, NDAS’s current status is unresolved. Official evidence establishes a platform, prototypes, and at least one controlled operational pilot. It does not establish that a national person-risk engine was operating on August 2, 2026.
The Gangs Violence Matrix lifecycle, legal chronology, and VHA replacement
Creation and design
The Gangs Violence Matrix began operating in early 2012 in the political aftermath of the 2011 disturbances. Its declared purpose was to identify and risk-assess gang members involved in violence, identify people at risk of victimization, safeguard exploited people, and prioritize preventive or enforcement activity. Individuals on the database were called “gang nominals.” citeturn23view0turn14search3
The Matrix used a traffic-light structure—red, amber, and green—together with a harm score. Inputs included recorded violence and weapons offenses, intelligence about access to weapons, involvement in or exposure to gang violence, and association information. Unlike HART, it was not publicly documented as a trained statistical model that estimated a specified future probability. Its “prediction” arose from human-coded intelligence and scoring rules.
Association was central. Police could treat friendship, family connection, co-location, social-media activity, music, or contact with someone already regarded as gang-associated as relevant intelligence. This created a risk of guilt by association, especially where young people lived, attended school, socialized, or made music in heavily policed neighborhoods.
Population and disproportionality
In October 2017, the Matrix contained 3,806 people. Fewer than 5 percent were red and 64 percent were green. A July 2016 demographic analysis found that 87 percent were from Black, Asian, or minority ethnic backgrounds and 78 percent were Black. Eighty percent were aged 12–24, 15 percent were minors, and 99 percent were male. Amnesty reported that 35 percent had never committed a serious offense and 75 percent had themselves been victims of violence. citeturn23view0
These figures do not by themselves prove that every individual was wrongly included. Serious violence and victimization are not evenly distributed across London, and a lawful safeguarding system may focus on a nonrepresentative group. But the disparity demanded a documented explanation connecting inclusion criteria to a legitimate purpose. The large green population, zero or low harm scores, and reliance on association weakened the claim that the database was narrowly confined to the most dangerous people.
The Matrix also blurred victim and suspect categories. A young person could appear because police believed he posed violence risk, faced violence risk, or associated with someone involved in violence. Treating those categories as interchangeable can convert vulnerability into suspicion. Data-protection law requires competent authorities, where applicable, to distinguish among suspected offenders, convicted people, victims, witnesses, and other persons.
Operational consequences and third-party sharing
Matrix status could affect local police tasking, stop and search, enforcement plans, intelligence monitoring, gang injunctions, housing or tenancy interventions, safeguarding meetings, probation treatment, and information provided to prosecutors. Amnesty documented sharing with local authorities, probation providers, and other partners. citeturn2search2turn23view0
The critical point is that a database need not automatically issue a sanction to produce severe effects. Once a “gang” label appears in police and partner records, later decision-makers may treat it as an established fact rather than a contested intelligence assessment. A school, housing provider, youth service, immigration authority, prosecutor, or probation officer may not know the evidential quality behind the label.
Third-party sharing also amplified errors. A person could correct or outgrow the original basis for inclusion while copies, meeting notes, or local databases retained the association. The more widely a score is shared, the less realistic it becomes to guarantee correction or deletion.
Notice and challenge
For much of the Matrix’s life, people were generally not proactively told that they were included. Some discovered their status through criminal proceedings, partner action, a subject-access request, or litigation. Without notice, an individual could not effectively contest mistaken identity, obsolete intelligence, false association, or an inaccurate gang label.
Police may lawfully withhold intelligence where disclosure would prejudice an investigation, expose sources, or endanger others. But total secrecy is not the only model. A system can provide delayed notice, a neutral summary, independent review, or confirmation after operational sensitivity has passed. The absence of a workable challenge mechanism allowed low-quality information to acquire long-term authority.
The 2022 settlement required the Met to provide more meaningful responses to people who asked whether they were on the Matrix and to explain what had been shared, subject to lawful exemptions. The settlement was important, but it did not produce a general statutory right to contemporaneous notice for police watchlists. citeturn14search6
ICO enforcement and public-law chronology
The ICO concluded in 2018 that the Matrix had a valid policing purpose but was operated in serious breach of data-protection requirements, with potential to cause damage and distress. The identified problems included deficient governance, inconsistent inclusion and review, inadequate differentiation of data-subject roles, excessive or inaccurate retention, insufficient transparency, and unsafe or inadequately controlled sharing. The ICO issued an enforcement notice and a specific checklist for forces using gangs databases. citeturn14search0turn14search8
MOPAC subsequently reviewed the Matrix and monitored reforms. The Met reduced the population, centralized management, revised review practices, and eventually removed green nominals. By 2022, however, litigation still alleged that the system breached Article 8, produced indirect racial discrimination, and failed the Public Sector Equality Duty. The case settled after the Met accepted that the system’s regulation and operation had been unlawful under Article 8, acknowledged disproportionate representation of Black people, and committed to wholesale reform. citeturn14search6turn14search14
Because the case settled, there is no final judicial judgment determining every pleaded Equality Act and Article 14 issue. The Met’s concessions and agreed reforms are legally and politically significant, but they should not be described as a court judgment establishing every alleged act of discrimination.
The Met ceased using the Matrix in February 2024. A later dispute concerned deletion. Police planned to destroy the decommissioned database after a retention period, while campaigners argued that preservation was necessary to identify affected people, investigate wrongful convictions, and support claims. citeturn14search10turn14news34
This dispute illustrates competing data-protection values. Indefinite retention perpetuates stigma and creates breach risk. Immediate deletion can destroy evidence needed to prove past unlawful treatment. The preferable safeguard is controlled preservation by an independent archive or litigation custodian, with operational access disabled, rather than either unrestricted police retention or irreversible destruction.
The Violence Harm Assessment
The VHA was launched across London in February 2024. The Met describes it as an internal intelligence tool used to identify and risk-assess people involved, or likely to be involved, in violence. It focuses on harmful violent individuals rather than formally on “gangs,” though its inputs can still include gang-association intelligence. citeturn14search7turn16view0
The VHA’s published eligibility rules are more explicit than the Matrix’s. Within a rolling three-year period, relevant offenses include homicide, attempted murder, gun discharges, knife injury, grievous bodily harm, kidnapping, affray or violent disorder, gun and knife possession, robbery, threats to kill, acid or ammonia possession, rape, and theft from the person. A person generally must appear in at least four separate reports, two of which are crime reports, or in three crime reports whose combined harm score is at least 2,500. At least one relevant violent crime must identify the person as a suspect within the preceding 12 months, subject to an adjustment where the person has been in custody. Domestic-abuse offenses are generally discounted because the tool focuses on street violence. citeturn16view0
Crime scores decay by one third after each 12-month period. Intelligence from the preceding six months is graded according to confidence. Scores use a hybrid of the Cambridge Harm Index and Office for National Statistics sentencing weights, partly because a single Cambridge score would not distinguish weapon-enabled robbery from a lower-harm non-weapon robbery. citeturn16view0
Additional datasets include CRIMINT, including gang-association intelligence; crime reports; MERLIN child and safeguarding information; wanted-offender data; county-lines intelligence; victim data for stabbings and shootings; prison notifications; and offender-management records. The system is refreshed quarterly, stored in read-only form on the Met’s BOX platform, and older versions are archived for three years for operational, research, and equality-assessment purposes. citeturn16view0
The VHA is therefore better characterized as a rules-based harm-ranking and watchlisting system than a machine-learning forecast. It combines previous events, sentencing-based weights, intelligence confidence, recency decay, and inclusion thresholds. It does not publicly claim to estimate a calibrated probability that a person will offend within a specified future period.
Human review, sharing, and individual consequences
The VHA produces a single score allowing prioritization by threat, harm, and risk. Being listed does not mandate or guarantee action. Local Detective Superintendents heading criminal-investigation functions have oversight of how listed people are managed, while specialist and territorial units may use the ranking for tasking, disruption, safeguarding, or threat-to-life work. Local units can request removal where professional judgment indicates that a person should no longer remain. citeturn16view0
The Met does not proactively inform people that they are listed, reasoning that notice could alter behavior, frustrate investigations, or elevate gang status. Requests for access are considered case by case under statutory exemptions. The fact of being listed is not to be used as evidence in court, although prosecutors should be told where VHA use formed part of the case so that underlying material can be addressed through disclosure. citeturn16view0
The VHA list itself must not be shared with partners. Underlying intelligence can be shared case by case for a policing or safeguarding purpose, under legislation or a data-sharing agreement. For children, the SOP requires safeguarding consideration and says that underlying information—not the fact of VHA membership—should be shared. citeturn17view2turn16view0
This is a material improvement over routine circulation of a gang list. But it does not eliminate propagation risk. A partner told that a child is linked to specified violent incidents or associations may receive substantially the same stigmatizing message without the label “VHA.”
Disproportionality and continuity with the Matrix
As of March 7, 2026, the VHA contained 1,574 people, including 437 in prison. Since December 22, 2025, 295 new people had met the criteria and 192 had ceased to meet them. Self-defined ethnicity records showed 40.53 percent Black, 9.47 percent mixed, 6.04 percent Asian, 17.09 percent White, and 22.24 percent not stated. Officer-observed records classified 60.55 percent as Black and 23.83 percent as White North or South European. Twenty-three percent were under 18, another 39.64 percent were aged 18–24, and 94.16 percent were male. citeturn16view3turn17view3turn17view4
The VHA is not simply the old Matrix under a new name. It has no formal gang-membership criterion, uses documented report thresholds, decaying offense scores, graded intelligence, centralized scoring, a published SOP, quarterly updates, restricted sharing, and express review arrangements.
There is nevertheless substantial continuity. It ranks named people; incorporates police intelligence and gang-association material; guides disruption and safeguarding; withholds proactive notice; and disproportionately contains young Black men. Whether the new thresholds are substantively fair cannot be resolved by publishing demographic counts alone. The Met needs to report comparison populations, inclusion rates among similarly situated suspects, subgroup removal rates, intelligence contribution to total scores, downstream interventions, and error audits.
National AI crime mapping and selected force-level systems
The 2025–2026 national initiative
The government announced the Concentrations of Crime Data Challenge on August 15, 2025. The initial £4 million program proposed an interactive, AI-supported map capable of identifying crime concentrations at street level. Government descriptions referred to combining information held by police, local authorities, and social services, including criminal records, incident locations, and behavioral patterns of known offenders. Early prototypes were expected by April 2026, with a stated ambition for an operational system across England and Wales by 2030. citeturn22view0
On April 7, 2026, UKRI announced a broader £50 million Safer Streets Challenges program and described phase two of the crime-data challenge. The intended themes included knife crime, violence against women and girls, theft, retail crime, and anti-social behavior. Funding was to flow through research councils and Innovate UK in partnership with the Home Office, Ministry of Justice, and other departments. citeturn22view1
The public announcements use “AI” expansively. They do not establish whether the eventual map will employ supervised machine learning, spatiotemporal point processes, risk-terrain regression, graph analytics, large language models for text extraction, conventional hotspot density estimation, or a combination. Nor do they specify the geographic unit, forecast horizon, update frequency, minimum incident count, protected-characteristic treatment, uncertainty display, or operational threshold.
As of the cutoff, the initiative was therefore a funded national research program with intended operational deployment, not an operational national prediction service. References to police “catching criminals before they strike” are political framing, not evidence that a validated system was making individual preemptive decisions.
Material differences from earlier programs
The initiative could differ materially from earlier British hotspot programs in four respects.
Scale: PredPol and RTM were force-specific. Grip was nationally funded but locally implemented. The proposed map is intended as common infrastructure across England and Wales.
Data diversity: Traditional hotspot mapping relies heavily on offense time and location. The national proposal contemplates police, council, social-service, and person-related behavioral data.
Granularity and speed: The intended system is described as street-level and potentially real-time or frequently refreshed, enabling more dynamic tasking.
Institutional integration: The 2026 policing reform plan proposes Police.AI as a national center for identifying, testing, and scaling technology, backed by £115 million over three years and a public-facing registry. citeturn22view2
These differences can improve consistency and evaluation, but they also magnify errors. A flawed local map affects one force. A standardized national model can reproduce the same proxy or label error across 43 forces. Cross-agency data may improve context but also extends police inference into education, housing, family, health, and social-care domains.
Kent PredPol
Kent Police’s PredPol deployment was the clearest British example of commercial place-based predictive policing. The system used historical recorded-crime events—principally type, time, and location—to identify small geographic boxes with elevated near-term risk. It drew on self-exciting point-process logic, treating crime as involving both persistent background risk and short-term repeat or near-repeat effects.
The output affected patrol allocation rather than individual legal status. Officers still decided whom to observe, stop, or engage within the box. Yet place prediction can become person targeting in practice: once extra officers are sent to a small area, residents experience more observation and more recorded low-level incidents.
Kent ended its use in 2018 after concluding that it could not demonstrate sufficient crime-reduction value. The episode demonstrates the difference between predictive accuracy and policy effectiveness. A system may identify where recorded crime will recur but fail to produce additional reduction beyond ordinary analyst knowledge or established hotspot policing.
Metropolitan Police Risk Terrain Modelling
Risk Terrain Modelling identifies environmental conditions statistically correlated with a crime type. The Met’s official FOI response describes connecting multiple data sources to geographic places and calculating which factors correlate with the outcome of interest. Analysts then use the output to prioritize further analysis and possible deployment or tactics. citeturn22view14
RTM differs from pure repeat-crime forecasting because it asks why places appear risky. Potential variables can include transport hubs, licensed premises, schools, retail sites, ATMs, offender anchor points, or other environmental features. This can support problem-oriented interventions—lighting, licensing, guardianship, transport design—rather than patrol alone.
The risks are substantial. Some environmental features act as proxies for population composition or poverty. A known offender’s home address can turn private residence into an enduring area-risk factor. Correlation does not establish that the feature causes crime, and acting on the wrong mechanism may intensify surveillance without addressing underlying harm.
The public evidence located does not provide a complete current data dictionary, model-validation study, frequency of use, or neighborhood-level equality analysis. RTM should therefore be classified as operational analytical capability with unresolved scale, not as a fully evidenced London-wide automated deployment system.
Grip and Hot Spot Response
Grip began as Home Office funding for 18 forces with high serious-violence volumes, with two additional forces receiving bespoke funding. It combined visible patrol in identified hotspots with problem-oriented policing. The first evaluation estimated a 7 percent reduction in violence and robbery in hotspots on days when patrols occurred during the year ending March 2022. The subsequent evaluation found no statistically significant pooled effect for the year ending March 2023. citeturn22view15
For the year ending March 2025, Hot Spot Response replaced the narrower Grip grant and expanded funding to all 43 territorial forces, adding anti-social behavior and perceptions of safety. Approaches remained heterogeneous: the Home Office described effectively evaluating numerous different local implementations rather than one model. citeturn22view15
Grip is frequently grouped with predictive policing, but that label requires qualification. Hotspot identification may use algorithms, geographic information systems, analyst judgment, historical counts, or kernel-density methods. The program’s defining feature is place-focused patrol, not necessarily machine learning.
Its main individual consequence is exposure. People living, working, socializing, or traveling through selected streets encounter more officers, stops, searches, intelligence checks, and recorded interactions. Neighborhood-level evaluation must therefore include not only crime changes but stop rates, complaints, demographic impacts, displacement, and community trust.
Essex’s Knife Crime and Violence Model
The Essex Knife Crime and Violence Model, also called Fearless Futures, sought to identify people already known to Essex Police who were most likely to commit knife-enabled or serious violence during the next 12 months. The force’s privacy notice described a probability score used to identify a very small cohort for multi-agency support aimed at helping individuals leave criminality. citeturn22view12turn22view13
Inputs included age, gender, a drug warning marker, number of criminal associates or co-accused people, previous suspect status for violence with injury, weapons and drug offenses, previous victimization for violence with injury, and the district associated with recent offending. All people in the relevant police system with an Essex address could be processed to generate the cohort, after which human reviewers applied operational criteria.
The model illustrates how victimization can become a risk factor for future perpetration. That relationship may be statistically real in cycles of retaliatory violence, but it creates a danger that a victim seeking police protection is later treated as an offender risk. The force should distinguish supportive safeguarding from coercive disruption and prevent the score from contaminating unrelated charging or bail decisions.
The pilot involved University of Essex researchers, data-protection and legal review, ALGO-CARE, and an independent ethics structure. A limited evaluation reported no statistically reliable reduction during the observed pilot period. citeturn22view13
An official Essex FOI response reportedly resisted describing the program as “predictive,” while the privacy notice and research report expressly described prediction of future knife violence. This is a terminological rather than substantive dispute. A model that estimates who will commit a specified event in the next 12 months is person-based prediction regardless of institutional branding.
No sufficiently clear evidence was identified that the original model remained in permanent force-wide use in August 2026. It should be classified as a serious operational pilot whose long-term status is unresolved.
Live facial recognition
Live facial recognition does not forecast crime. It compares biometric representations of passers-by against a police watchlist and generates a similarity alert. It meets this report’s inclusion standard because it is algorithmic, watchlist based, and capable of causing immediate police engagement.
In Bridges v Chief Constable of South Wales Police, the Court of Appeal held in 2020 that South Wales Police’s earlier deployment framework was not sufficiently constrained as to who could be placed on a watchlist and where the technology could be used. The related DPIA was deficient, and the force had not complied with the Public Sector Equality Duty. citeturn22view7
The Metropolitan Police subsequently adopted a more detailed policy. In R (Thompson and Carlo) v Commissioner of Police of the Metropolis, handed down on April 21, 2026, the Divisional Court dismissed a challenge to the revised 2024 policy. The court concluded that the policy had sufficient clarity and foreseeability and adequate safeguards against arbitrary use for the purposes of Articles 8, 10, and 11. The claimants did not advance a standalone PSED ground, and the court emphasized that a discriminatory or arbitrary policy could still be unlawful. citeturn23view3turn21search2
The ruling was not a declaration that all police facial recognition is lawful. Its significance is policy-specific. Watchlist criteria, deployment location, matching threshold, demographic testing, notice, data deletion, engagement-officer review, and post-alert verification remain legally relevant.
LFR sits closest to the FBI Threat Screening Center model among current British technologies because both combine a person list with encounter-time matching. The principal difference is that British LFR watchlists are force- and deployment-specific rather than one consolidated national terrorism-and-gang screening database.
Legal framework, technical risks, comparative assessment, and reform
Human Rights Act and common-law police powers
Police rely principally on common-law duties and powers to prevent and detect crime, preserve public order, protect life, and apprehend offenders, supplemented by statutory powers for specific activities. The Supreme Court in Catt accepted that police may collect and retain information for legitimate policing purposes, while emphasizing that systematic collection and retention of personal information engage Article 8 and must be proportionate. citeturn22view8turn23view2
Under the Human Rights Act 1998, Article 8 protects private and family life, home, and correspondence. Police databases, risk scores, biometric scans, association records, and cross-agency sharing can engage Article 8 even where information originates in public activity. The interference must have a sufficiently accessible and foreseeable legal basis, pursue a legitimate aim, and be necessary and proportionate.
Articles 10 and 11 become important where intelligence systems use political expression, music, online speech, protest attendance, or association. Article 14 prohibits discrimination in the enjoyment of Convention rights. A gang database disproportionately affecting Black young men, for example, may engage Articles 8 and 14 even where race is not an explicit field.
Bridges demonstrates that broad common-law policing powers do not by themselves resolve every legality question. A sufficiently detailed policy must constrain discretion. Thompson and Carlo shows the other side: a court may uphold advanced surveillance where the revised rules specify watchlist and deployment criteria and create safeguards. citeturn22view7turn23view3
For predictive systems, proportionality should examine the seriousness and likelihood of the harm addressed, model reliability, breadth of the affected population, consequences of false classification, availability of less intrusive means, retention, notice, and review. A system used only to guide aggregate research may require fewer safeguards than one used to exclude someone from diversion or trigger persistent surveillance.
Data Protection Act, UK GDPR, and the Data (Use and Access) Act
Part 3 of the Data Protection Act 2018 governs processing by a competent authority for law-enforcement purposes: prevention, investigation, detection, or prosecution of criminal offenses, execution of criminal penalties, and safeguarding against threats to public security. It is separate from the UK GDPR regime. Police processing for employment, procurement, public communications, or other non-law-enforcement purposes may instead fall under UK GDPR. citeturn22view4
Core Part 3 requirements include:
- lawfulness and fairness;
- specified, explicit, and legitimate purposes;
- adequacy, relevance, and non-excessiveness;
- accuracy and updating;
- retention no longer than necessary;
- appropriate security;
- differentiation, where applicable, among suspects, convicted people, victims, witnesses, and other persons;
- differentiation between facts and personal assessments;
- strict necessity and an applicable Schedule 8 condition for sensitive processing;
- a DPIA where processing is likely to present high risk; and
- records, access controls, and logging.
These duties apply to inputs and outputs. A derived risk category, similarity alert, network link, or inferred gang association is personal data. Police cannot avoid responsibility by saying that the vendor generated the inference or that an analyst, rather than a computer, finally read it.
The Data (Use and Access) Act 2025 amended both UK GDPR and law-enforcement processing. Government guidance describes a more permissive framework for solely automated decisions with legal or similarly significant effects, accompanied by safeguards including information, representations, challenge, and human intervention. The law-enforcement regime includes limited postponement or exemption possibilities where immediate safeguards would obstruct an inquiry or affect national security, with later meaningful human reconsideration. citeturn21search0turn21search9
The amendments do not convert nominal human sign-off into meaningful review. A decision is not genuinely human merely because an officer clicks “approve.” The reviewer must understand the relevant information, have authority to depart, and actively reconsider the substance.
Most systems examined here are not solely automated. But Part 3 still applies fully to advisory analytics. Data minimization, accuracy, fairness, purpose limitation, and strict necessity are not limited to automated final decisions.
The ICO can investigate, audit, require information, issue enforcement notices, order changes or cessation, and impose penalties where statutory conditions are met. The Matrix enforcement action demonstrates that police intelligence systems are within active regulatory jurisdiction. citeturn14search0turn14search8
Equality Act
The Equality Act 2010 prohibits direct and indirect discrimination in relevant public functions and imposes the Public Sector Equality Duty in section 149. Police must have due regard to the need to eliminate discrimination, advance equality of opportunity, and foster good relations.
Indirect discrimination can arise where an apparently neutral criterion—postcode, prior police contact, association count, neighborhood, or recorded victimization—puts a racial or other protected group at a particular disadvantage and cannot be justified as a proportionate means of achieving a legitimate aim.
The PSED is a continuing procedural duty. It must inform design, procurement, pilot selection, deployment, retraining, geographic expansion, and response to evidence of disparity. In Bridges, the Court of Appeal found that South Wales Police had not done enough to investigate possible demographic bias before and during deployment. citeturn22view7
An equality impact assessment should not merely list demographic percentages. It should identify the causal pathway from data to output to action; test subgroup error rates; analyze proxies; compare similarly situated populations; examine intersectional effects; and document mitigation. For VHA, publishing that 60.55 percent of listed people were officer-observed as Black is only the beginning of the analysis.
Judicial review and procedural fairness
A police-analytics policy can be challenged through judicial review for illegality, irrationality, failure to take relevant considerations into account, procedural unfairness, improper purpose, breach of legitimate expectation, or disproportionality where Convention rights are engaged.
Potential grounds include:
- absence of sufficiently foreseeable rules;
- deploying a model outside its validated context;
- relying on materially inaccurate data;
- failing to assess equality impact;
- treating an advisory score as binding;
- inadequate reasons or review;
- unlawful data sharing;
- retaining data beyond policy periods;
- using a tool for a new purpose without reassessment; and
- irrational reliance on a model with known poor performance.
Individual challenge is difficult where the person does not know that a system was used. Disclosure obligations, subject-access rights, complaints to the ICO, criminal-proceedings disclosure, and targeted FOI requests partially mitigate secrecy but do not create a comprehensive right to explanation.
Procurement
Procurements commenced on or after February 24, 2025 are governed by the Procurement Act 2023 and Procurement Regulations 2024. Earlier procurements generally remain governed by the previous regime and transitional rules. citeturn22view6
Police procurement of an algorithm should specify:
- access to training and validation documentation;
- audit rights for the force, ICO, courts, and authorized independent researchers;
- ownership and portability of input and derived data;
- access to source code or an escrow mechanism where necessary;
- subgroup testing and minimum performance thresholds;
- notice of model or vendor changes;
- incident reporting;
- limits on secondary use;
- retraining, drift, and decommissioning duties;
- deletion and return of data;
- assistance with subject-access and litigation disclosure;
- prohibition on undisclosed subcontracting; and
- termination where legality, accuracy, or equality requirements are not met.
Vendor intellectual property is not a lawful basis for withholding information necessary to understand a public decision. A force that cannot audit a supplier’s model cannot demonstrate that its own processing is necessary, fair, accurate, or proportionate.
Procurement Policy Note 017 addresses transparency concerning supplier use of AI in covered procurement, but its mandatory scope is principally central-government organizations rather than every territorial police force. It is therefore not a substitute for a police-specific statutory standard. citeturn22view6
Algorithmic transparency requirements
The Algorithmic Transparency Recording Standard is mandatory for UK government departments and arm’s-length bodies delivering public or frontline services or directly interacting with the public, where a tool significantly influences a decision with public effect or directly interacts with the public. It is recommended across the wider public sector, and police forces may publish records voluntarily, but territorial forces are not generally within the standard’s mandatory central-government scope. citeturn22view5
This produces an accountability gap. A central department using an algorithm for a significant benefit decision may have to publish a structured ATRS record, while a police force using a person-risk model may not.
The government’s 2026 policing reform paper promises a public-facing registry of police AI and a new regulatory framework with strong oversight and accountability. It also proposes Police.AI as a national testing and scaling center. Those commitments could close part of the gap, but by August 2, 2026 they were policy promises rather than a fully implemented statutory registry. citeturn22view2
The NPCC AI strategy, policing AI playbook, ALGO-CARE, local ethics committees, DPIAs, equality assessments, and professional guidance are valuable but fragmented. None supplies a single legally enforceable preauthorization process for high-risk police algorithms.
Technical and social risks
Label leakage
Police models frequently use outcomes such as arrest, custody return, recorded offense, intelligence report, or charge. These labels measure police detection and recording as well as underlying behavior. If an input includes current investigative attention, bail status, or a variable generated close to the outcome window, the model may appear predictive while merely rediscovering the institution’s existing decision.
HART’s return-to-custody outcome and prior-custody inputs create this concern. VHA’s use of multiple crime reports and intelligence is less a forecast than a formalized summary of already accumulated police attention. A future national crime map trained on enforcement-generated incident data could reproduce where police have historically looked.
Postcode and socioeconomic proxies
Postcode, district, housing, school, transport, household, and commercial lifestyle classifications can reveal socioeconomic and racialized structure even without race. HART’s Mosaic variable is the clearest example. Essex’s district variable and place-based maps create related risks.
Proxy analysis should ask not only whether a variable predicts an outcome but whether it adds enough legitimate value to justify the social meaning it imports. Forces should publish ablation tests showing performance with and without postcode and commercial data.
Racialized gang intelligence and false association
Gang intelligence can encode cultural judgments, neighborhood stereotypes, music, friendship, and family relationships. Once recorded, association becomes both an input and a reason for further surveillance. The Matrix’s demographic profile and the continued inclusion of gang-association intelligence in VHA show that changing a system’s name does not automatically remove the underlying epistemic risk. citeturn23view0turn16view0
An association should never by itself establish risk. Systems should record the source, date, confidence, specific relevance, and whether the relationship indicates coercion, family, victimization, or offending. A child exploited by adults should not acquire the same inference as an organizer.
Automation bias
Officers may over-trust a score because it appears objective or under-trust it where it conflicts with intuition. Both patterns can be discriminatory. Meaningful review requires training on base rates, false positives, uncertainty, and known limitations, plus recorded reasons and audit of overrides.
Retention
Risk classifications can outlive the conduct that generated them. HART forecasts covered two years, but public information about model-output deletion is incomplete. VHA archives remain for three years after quarterly replacement. Matrix deletion created the opposite problem of destroying evidence needed for redress. citeturn16view0turn14news34
Retention policy should distinguish operational data, audit copies, litigation preservation, research datasets, and source records. Independent custodianship is preferable where past illegality must be investigated.
Function creep
A model built for diversion can migrate into bail, charging, or offender management. A modern-slavery discovery system can be repurposed to immigration enforcement. A violence watchlist can influence housing. A crime map can acquire person overlays.
Every materially new purpose should trigger a new legal basis analysis, DPIA, equality assessment, validation, public record, and approval.
Feedback loops
Place-based deployment produces more police observations in selected areas. Those observations become new data, strengthening the next forecast. Person lists similarly produce more intelligence reports about listed people. The system can thereby confirm itself.
Evaluation should use victimization surveys, hospital data, calls for service, independent community information, and randomized or phased designs where ethically appropriate—not recorded police incidents alone.
Vendor opacity
Commercial secrecy can prevent disclosure of weights, source data, preprocessing, error rates, and system changes. This obstructs challenge and can make the police dependent on one supplier. NDAS’s public documents identify Accenture as processor, but do not disclose enough to reconstruct procurement and intellectual-property arrangements. citeturn18view1
Inconsistent local implementation
Grip’s national evaluation found materially different approaches across forces. Local variation can be useful, but it creates unequal treatment: the same data pattern may trigger intensive patrol in one area, safeguarding in another, and no intervention elsewhere. citeturn22view15
National systems should standardize minimum safeguards and reporting, not necessarily every tactic.
Comparison with the FBI Threat Screening Center
The FBI’s Threat Screening Center, renamed from the Terrorist Screening Center in March 2025, maintains and shares a consolidated federal watchlist. The center’s stated expansion included cartel and gang members connected to newly designated foreign terrorist organizations. The list includes identifiers such as names, dates of birth, and fingerprints. Government agencies nominate people under intelligence criteria, and the FBI generally does not confirm an individual’s watchlist status. citeturn22view9turn22view10
The TSC concept has three defining features:
- a centralized national person list;
- identity resolution across agencies and screening systems; and
- encounter-time alerts or consequences when a person interacts with border, aviation, law-enforcement, or other screening infrastructure.
No examined UK predictive-policing system is a full equivalent.
- The national crime map is place based and still under development.
- NDAS is a data and analytics platform, not a single national encounter watchlist.
- HART was a local custody forecast.
- The Matrix and VHA most closely resemble watchlisting, but are Metropolitan Police systems rather than consolidated national screening infrastructure.
- Live facial recognition resembles the encounter function, but watchlists are assembled for particular force purposes and deployments.
- National Police systems such as the Police National Computer and Police National Database provide shared records and intelligence, but the systems reviewed here do not establish a TSC-style universal threat nomination and automatic encounter-alert regime.
The closest British analogue would arise if VHA-like person prioritization, NDAS-style multi-force analytics, national biometric matching, and Police.AI infrastructure were combined. That possibility supports regulating architectures and data flows, not only named products.
Proposed statutory safeguards
A defensible statutory framework should regulate “high-impact police analytics,” defined by function and consequence rather than by whether the supplier calls a product AI. It should cover any computational or structured scoring system that materially influences policing of a person or small geographic area, including watchlists and biometric matching.
The statute should require the following.
Predeployment authorization. A force should submit a public impact dossier to an independent police-technology regulator before live use. The dossier should include purpose, legal basis, inclusion criteria, data dictionary, model type, forecast horizon, expected decision, alternatives, validation, equality testing, retention, procurement, and human-review design. Urgent temporary use should require time-limited authorization and later review.
Clear categories. Systems should be registered as place forecasting, person prediction, watchlisting, biometric identification, investigative discovery, or administrative support. A change of category should require new approval.
Minimum scientific evidence. Person-based models should demonstrate calibration, discrimination, subgroup error rates, external validation, temporal validation, base-rate analysis, and operational benefit. Accuracy below an approved threshold should prohibit deployment. Claims should be tested by an independent evaluator with access to code and data.
Outcome validity. The force should explain whether the label represents offending, arrest, custody return, intelligence recording, victimization, or another event. It should report how police behavior affects the label and test for feedback.
Proxy controls. Postcode, commercial segmentation, association, social-media, protected-characteristic proxies, and third-party social data should require specific necessity findings. Commercial lifestyle data should be presumptively prohibited in person-risk scoring unless Parliament expressly authorizes it and independent evidence shows indispensable value.
Association safeguards. No adverse prioritization should be based solely on friendship, family, shared location, music, social-media contact, or association with a listed person. Association intelligence should expire unless renewed with recorded evidence and independent supervisory review.
Human review. The law should define meaningful human involvement: access to the reasons and source quality, authority to depart, adequate time, training, and a recorded decision. Override and concurrence patterns should be audited by subgroup.
Notice and challenge. Where immediate notice would not prejudice a legitimate operation, the subject should be informed that a high-impact system materially influenced a decision. Where notice must be delayed, an independent reviewer should represent the individual’s interest, and notice should follow when sensitivity ends. People should be able to correct source data, contest associations, request reconsideration, and obtain a comprehensible explanation.
Strict downstream-use limits. A diversion score should not be used for charging or bail without separate authorization. A safeguarding flag should not automatically become immigration or enforcement intelligence. The fact of watchlist inclusion should not be shared where the underlying verified information is sufficient.
Public registers. Police.AI’s proposed registry should be statutory, cover all forces and suppliers, and include retired systems and failed pilots. Entries should identify operational period, force, purpose, model, supplier, contract route, data categories, validation, equality findings, incidents, and decommissioning.
Procurement transparency. Contracts should be published with narrowly justified redactions. Supplier IP claims should not prevent regulatory, judicial, defense, or independent audit. Forces should retain sufficient rights to explain and contest outputs.
Independent audits and sunset clauses. Approval should expire after a fixed period, geographic expansion, material model change, or purpose change. Annual audits should assess performance, disparities, drift, intervention outcomes, complaints, and community impact.
Retention and evidence preservation. Operational scores should expire with the forecast horizon unless renewed. Where a system may have caused unlawful outcomes, an independent custodian should preserve an immutable audit copy for claims while preventing operational reuse.
Remedies. Individuals should have a statutory right to obtain reconsideration, correction, deletion where appropriate, disclosure in criminal proceedings, compensation for material damage or distress, and suspension of a system where systemic illegality is shown.
Community impact review. Place-based systems should report patrol exposure, stops, searches, arrests, use of force, complaints, displacement, and trust—not only crime counts.
Final functional ranking
The following ranking assesses resemblance, not lawfulness or social harm. A system can resemble conventional intelligence and still be deeply intrusive.
| Rank within category | System | Closest functional type | Reason |
|---|---|---|---|
| Place-based forecasting | Kent PredPol | Strongest resemblance | Explicitly forecast near-term crime risk in small geographic cells from historical events. |
| National AI-supported crime map | Strong intended resemblance, not yet operational | Designed to identify street-level concentrations nationally, but technical design and live use remained unproven. | |
| Met Risk Terrain Modelling | Place-risk modeling | Estimates environmental conditions correlated with crime rather than simply repeating past incident density. | |
| Grip / Hot Spot Response | Place-based deployment, weaker “AI” resemblance | Uses hotspot analysis to direct patrols, but is a funding and operational strategy rather than one predictive model. | |
| Person-based prediction | HART | Strongest documented operational resemblance | Produced low, moderate, or high forecasts of an arrestee’s offending over two years and affected diversion eligibility. |
| Essex Knife Crime and Violence Model | Strong person-prediction resemblance | Estimated 12-month probability of knife or serious violence among people known to police; pilot status and long-term use unresolved. | |
| NDAS Most Serious Violence | Strong prototype resemblance | Intended person-level serious-violence prediction, but rejected before deployment. | |
| NDAS modern-slavery analytics | Mixed vulnerability detection and investigative discovery | May identify persons or networks at risk, but public evidence emphasizes hidden-case discovery rather than calibrated future-offending prediction. | |
| Watchlisting and encounter screening | Gangs Violence Matrix | Strongest UK watchlist resemblance | Named people, ranked harm, association intelligence, broad operational and partner consequences, limited notice and challenge. |
| Violence Harm Assessment | Rules-based watchlisting and harm prioritization | Named London-wide cohort, standardized thresholds and scoring, no calibrated probability, operational targeting and safeguarding. | |
| Live facial recognition | Biometric encounter watchlisting | Matches passers-by against a curated watchlist and can produce immediate engagement; does not predict future crime. | |
| Conventional intelligence analysis | NDAS modern-slavery and network work | Strongest advanced-analytics resemblance | Text search, behavioral patterning, entity linking, and network analysis improve discovery within existing police data. |
| VHA | Structured intelligence analysis with ranking | Formalizes crime and intelligence records using harm weights and recency rules. | |
| Risk Terrain Modelling | Geographic intelligence analysis | Links environmental features and crime data to inform analyst and deployment decisions. | |
| National crime map, if limited to descriptive concentration mapping | Conventional analytical support | Classification will depend on whether final outputs are forecasts, causal-risk estimates, or descriptive maps. |
The overall conclusion is that the United Kingdom has not operated one comprehensive national “pre-crime” system. It has instead developed a fragmented ecosystem in which a local custody forecast, a national analytics platform, association-based watchlists, biometric screening, and hotspot maps coexist under overlapping but incomplete controls.
HART most clearly crossed the line into operational prediction of an individual’s future offending. The Gangs Matrix most clearly demonstrated the dangers of watchlisting, association, secrecy, racial disproportionality, and uncontrolled downstream sharing. NDAS most clearly demonstrated the ambiguity of a platform whose use cases range from legitimate investigative discovery to attempted person-level prediction. The 2025–2026 national crime-map initiative most clearly presents the next governance challenge: national scale is being planned before the public can see the model, data architecture, validation design, procurement, or rights framework.
The most important policy distinction is therefore not between “AI” and “non-AI.” It is between systems that merely help retrieve information and systems that allocate suspicion, attention, opportunity, or coercion. Any technology that does the latter should be governed as a high-impact public decision system, even where a human officer formally makes the final call.