Hot Spots, Feedback Loops, and Vendor Claims: A Technical Audit of Place-Based Predictive Policing
Executive findings and scope
This report audits place-based predictive policing as a sociotechnical intervention, not merely as a forecasting algorithm. The relevant causal chain is:
\[ \text{historical data} \rightarrow \text{risk estimate} \rightarrow \text{patrol assignment} \rightarrow \text{officer behavior} \rightarrow \text{crime, reporting, and enforcement outcomes} \rightarrow \text{future data}. \]
A technically accurate forecast can fail operationally because officers do not visit the forecast area, because the recommended intervention is ineffective, or because the outcome measure does not distinguish deterrence from reporting and recording changes. Conversely, conventional hot-spot patrol can reduce crime even when an algorithm contributes little or nothing beyond recent-event counts, kernel-density maps, or an experienced analyst. The National Institute of Justice’s Shreveport evaluation explicitly concluded that it remained unclear whether predictive maps were more useful than traditional maps of prior crime locations; the study found no statistically significant incremental reduction in property crime. citeturn21view0turn21view1
The necessary technical distinctions
The following categories should not be collapsed into one label merely because they use maps, databases, artificial intelligence, or centralized police technology.
| Function | What it does | Appropriate analytical question | Is it place-based predictive policing? |
|---|---|---|---|
| Descriptive crime mapping | Plots already recorded incidents, calls, arrests, or other events. | Where were recorded events located? | No. It describes observations without estimating future risk. |
| Retrospective hot-spot analysis | Identifies places with unusually high historical event concentrations, often through counts, kernel-density estimation, clustering, or long-term rankings. | Where has crime recently or persistently concentrated? | Not necessarily. It becomes forecasting only when the historical surface is explicitly treated as a prediction for a future period. |
| Short-term place forecasting | Estimates elevated future event risk for specified places and forecast windows. | Where and when are designated offenses expected next? | Yes. PredPol/Geolitica, HunchLab, self-exciting point-process maps, and prospective near-repeat maps fit here when used operationally. |
| Patrol optimization | Selects routes, missions, staffing levels, or patrol durations under resource constraints. | How should available officers be allocated? | Not by itself. It may consume predictive scores, historical hot spots, or managerial priorities. |
| Environmental risk-terrain modeling | Estimates how land uses, facilities, transportation nodes, vacancy, or other environmental features combine to form settings associated with crime. | Which environmental configurations are associated with elevated risk, and what conditions might be changed? | Sometimes. A static diagnostic RTM is not necessarily a short-term prediction; a prospective RTM used to allocate future patrol is functionally predictive. |
| Gunshot detection | Detects and locates a suspected acoustic event after it has apparently occurred. | Did a sound consistent with gunfire just occur, and where? | No. ShotSpotter is an event-detection and dispatch system, not a forecast, even though its corporate owner also sells predictive patrol products. Chicago’s inspector general described it as a network of acoustic sensors that triggers responses to suspected gunshots. citeturn17search11 |
| Real-time crime center | Integrates feeds such as computer-aided dispatch, cameras, license-plate readers, records, gunshot alerts, and analyst communications. | What is happening now, and what information can support response or investigation? | Not inherently. An RTCC becomes part of predictive policing only when it generates or operationalizes forward-looking risk estimates. |
| Person-based prediction | Scores individuals, groups, victims, suspects, or alleged networks for future offending, victimization, or police attention. | Who is believed to pose or face future risk? | Predictive policing in a broader sense, but analytically distinct from place-based systems and associated with different due-process and error problems. |
A map can therefore be nonpredictive, and a product marketed as “resource management,” “precision policing,” “mission planning,” or “risk reduction” can remain predictive if it ranks future place-time risk and directs patrol accordingly. Functional behavior, not branding, is the controlling criterion.
Principal conclusions
The strongest evidence supports focused policing at genuinely high-crime microplaces, not the proposition that proprietary predictive software substantially improves on transparent hot-spot methods. A Campbell systematic review covering 65 studies and 78 tests found a small but statistically significant crime-control effect from hot-spots policing; displacement was not inevitable, and nearby areas sometimes experienced diffusion of benefits. Those results establish the value of concentration and focused intervention, but they do not establish the incremental value of PredPol, HunchLab, RTM software, or machine learning. citeturn3search2turn3search6turn22search11
The best-known favorable PredPol field experiment found that an epidemic-type aftershock sequence model forecast between 1.4 and 2.2 times as much crime as analyst-produced maps and estimated a 7.4 percent reduction in crime as a function of patrol time in Los Angeles treatment conditions. The experiment used day-level random assignment and visually identical treatment and control maps. It is an important result, but several authors were PredPol co-founders or shareholders, the analysis intertwined model selection and patrol delivery, and the article itself called for better GPS measurement and research on the feedback between patrol, crime, and algorithms. citeturn18view2turn19view0turn19view1turn27view0
An independently funded Philadelphia randomized trial of HunchLab found a 31 percent reduction in expected property-crime counts for a dedicated marked-car intervention, with an estimated 40 percent reduction during the following eight-hour period. Mere awareness of forecast areas and a dedicated unmarked-car strategy did not show comparable benefits; no intervention produced a persuasive violent-crime reduction. The result is better interpreted as evidence about visible patrol dosage in selected microplaces than as proof that HunchLab’s proprietary forecasting added value over a simple baseline, because the experiment did not randomize HunchLab forecasts against recent-event or conventional analyst hot spots. citeturn18view3turn19view2turn27view2turn27view3
Independent evidence from Plainfield, New Jersey, is far less favorable to PredPol/Geolitica. Of 5,722 robbery or aggravated-assault forecasts without recorded patrol dosage, 32 forecast records corresponded to 20 of 139 reported incidents. That implies alert-level precision of 0.56 percent and incident-level recall of 14.4 percent. Only 8 of 116 burglaries occurred in the predicted place-time windows, for 6.9 percent recall; the published alert success rate was approximately 0.1 percent. Officers entered forecast boxes during only 129 of 23,760 forecast instances, and department officials said the system was rarely, if ever, used to direct patrol. Thus the case is simultaneously an unfavorable forecast evaluation and an implementation failure—not a valid test of whether faithful PredPol patrol reduces crime. citeturn26view0turn26view1turn27view5
The Shreveport Predictive Intelligence-Led Operational Targeting experiment likewise found no statistically significant incremental crime reduction. Treatment districts reportedly spent 6–10 percent less than controls, largely because of lower overtime, but heterogeneous implementation, few districts, low event counts, and limited statistical power prevented a strong conclusion. RAND recommended a direct comparison in which treatment and control receive equivalent interventions and differ only in whether locations come from an algorithm or traditional crime mapping. citeturn21view0turn21view1turn21view2
Risk-terrain modeling has credible evidence that environmental features can identify places with elevated future risk. A systematic review found 25 studies and reported that roughly half of future cases were captured in the highest-risk 10 percent of cells across the pooled literature. That is evidence of spatial discrimination, not automatically of calibration or causal crime reduction. Promotional claims associated with Atlantic City and Rutgers frequently cite reductions around 20–30 percent, but the prominent Atlantic City comparisons were before-and-after figures accompanied by business checks, lighting changes, vacant-building interventions, and other measures. Without a randomized or credible matched counterfactual, those changes cannot be attributed uniquely to RTM. citeturn22search1turn22search4turn22search20turn22search28
Vendor consolidation has not eliminated the underlying functions. Azavea sold HunchLab to ShotSpotter in 2018; it became ShotSpotter Missions and later ResourceRouter. PredPol renamed itself Geolitica in 2021 and ceased operating independently at the end of 2023, after SoundThinking hired its engineering team and acquired technology, intellectual property, and customer-related assets. SoundThinking’s 2025 annual filing described HunchLab as renamed ResourceRouter and said the product supports strategically planned patrols. Geolitica’s corporate exit therefore did not mark the disappearance of predictive patrol allocation; it marked an acquisition and migration into a broader policing platform. citeturn17search1turn17search3turn17search6turn17search20turn18view4turn26view0
The central accountability problem is not only algorithmic opacity. It is the frequent absence of preserved forecast outputs, patrol-dosage records, denominator definitions, baseline models, preregistered evaluation plans, contract deliverables, and auditable links between forecast, dispatch, officer presence, enforcement activity, and subsequent outcomes. NYPD records, for example, indicated that its in-house predictive tool did not retain or recreate historical outputs, making retrospective validation substantially more difficult. citeturn24view1
Technical taxonomy and vendor genealogy
Traditional hot spots, near repeats, and self-excitation
Crime is spatially concentrated, but concentration has at least three analytically different components.
The first is persistent heterogeneity: some street segments, facilities, and small areas have chronically higher recorded crime than others because of land use, opportunity structures, traffic, guardianship, population exposure, and enduring social conditions. A historical count, moving average, or kernel-density estimate can exploit this persistence without making assumptions about contagion.
The second is repeat victimization: a previously victimized address or target experiences temporarily increased risk. The third is near-repeat concentration: after an event, nearby places experience elevated risk for a limited period. Prospective hot-spotting operationalizes these patterns by creating temporary risk buffers or surfaces around recent incidents. Research supports near-repeat structure for some property offenses, particularly burglary and vehicle crime, but predictive usefulness varies by offense, geography, event volume, spatial threshold, and forecast horizon. Short-lived near-repeat signals can also be operationally difficult because a forecast may expire before an agency can deliver a meaningful preventive response. citeturn22search2turn22search22turn22search26turn22search29
A self-exciting point process formalizes this idea. In simplified form, the conditional event intensity at location \(s\) and time \(t\) is:
\[ \lambda(s,t) = \mu(s) + \sum_{i:t_i<t} g(s-s_i,t-t_i), \]
where \(\mu(s)\) represents persistent background risk and \(g(\cdot)\) represents the temporary increase associated with prior events. PredPol’s intellectual ancestry lies in earthquake aftershock models: an incident is treated analogously to a main shock that may be followed by nearby events, although the criminological mechanisms can include repeat targeting, offender foraging, shared opportunity, reporting artifacts, and police detection rather than literal contagion. The foundational crime-modeling article was published by Mohler and colleagues in 2011. citeturn0search14turn18view2
The distinction between background and triggering components is useful, but it does not solve the label problem. The process models recorded events, not latent criminal behavior directly. If patrol increases drug arrests in a forecast area, the resulting event pattern may be self-exciting in the database even when underlying drug use is geographically similar elsewhere. If residents in one area report burglary more often, the fitted background rate can reflect reporting access and insurance requirements as well as victimization.
PredPol and Geolitica
PredPol emerged from collaboration among researchers associated with UCLA and police agencies in Los Angeles and Santa Cruz. Its classic implementation used a deliberately narrow input schema—incident type, location, and time, plus identifiers and modification timestamps for data management—and produced ranked small boxes indicating elevated near-term risk. Records disclosed during NYPD’s vendor evaluation described crime type, geographic coordinates, occurrence time, incident identifier, and optional record-update time as the core fields. citeturn24view2
The standard public description emphasized roughly 500-by-500-foot prediction boxes, refreshed for daily or shift-level patrol. The product’s appeal was operational simplicity: officers received a small number of red boxes and were encouraged to enter them when not handling calls. Some implementations recorded “dosage,” meaning time spent inside a forecast box, through vehicle location, application interaction, or dispatcher records. That dosage was central to the vendor’s causal story: a forecasted event that did not occur after patrol could be counted as possible deterrence rather than a false positive.
That logic makes dosage indispensable to evaluation. It is impossible to interpret the absence of a subsequent incident without knowing whether treatment was delivered. Yet The Markup identified prediction data for 38 jurisdictions and obtained meaningful long-duration dosage records from only Plainfield; most agencies said the records were unavailable or nonpublic. citeturn27view5
PredPol renamed itself Geolitica in 2021. Geolitica’s independent operations ended in 2023, and customers were offered migration to SoundThinking’s platform. Contemporary accounts distinguished among the engineering employees, patents or other intellectual property, customer relationships, and the original source code; SoundThinking stated that it had not acquired the old crime-prediction source code, while its securities filings described acquisition of Geolitica’s primary technology and intellectual property. The most defensible conclusion is that the corporate entity and original product ended, while selected personnel, intellectual property, and customer functions were absorbed into a successor patrol-allocation ecosystem. citeturn17search3turn17search6turn17news23turn26view0
HunchLab, Missions, and ResourceRouter
HunchLab began with a Philadelphia “Crime Spike Detector,” developed with the police department and the U.S. Attorney’s Office. Azavea received National Science Foundation small-business funding and expanded the prototype between approximately 2008 and 2011. Early functions included localized spike detection, seasonal crime-load forecasting, and near-repeat methods. HunchLab 2.0 later combined multiple feature families and a machine-learning forecasting model. citeturn18view5turn17search20
Its documented inputs were substantially broader than PredPol’s. They could include historical crime at several look-back windows; recent three-, seven-, and fourteen-day patterns; zoning and land-use variables; schools, hospitals, transit, roads, water, and public-safety facilities; income, rent, housing, vacancy, population, and vehicle measures; temporal cycles; events; and weather. In the Philadelphia experiment, HunchLab used gradient-boosted decision trees, cross-validation, and smoothing to convert model scores into expected counts for 500-by-500-foot cells. The highest-scoring cells became “mission grids.” The specific variable importance logs were no longer available to the evaluators after the ownership transfer, illustrating how acquisition can sever the evidentiary chain needed for later audits. citeturn27view2
HunchLab also went beyond forecasting into mission generation. It ranked expected crime per unit of patrol effort, considered available officers and vehicles, and distributed printable or mobile missions. Its promotional interface contemplated tactics such as neighborhood patrol, vehicle stops, reviewing known offenders, or interviewing residents. The inclusion of tactic recommendations is important: two agencies could use identical forecast scores but generate very different community impacts depending on whether officers provide visible presence, engage businesses, conduct stops, seek arrests, or coordinate environmental remediation. citeturn18view5turn19view4
ShotSpotter acquired HunchLab in 2018 and marketed it as ShotSpotter Missions. SoundThinking’s later filings identify ResourceRouter as the renamed descendant. The genealogy is therefore:
\[ \text{Crime Spike Detector} \rightarrow \text{HunchLab} \rightarrow \text{ShotSpotter Missions} \rightarrow \text{ResourceRouter}. \]
The names changed, and product components evolved, but the continuing function—estimating risk and recommending patrol missions—places the forecasting and allocation modules within the predictive-policing audit boundary. citeturn17search1turn17search10turn18view4
Risk-terrain modeling
Risk-terrain modeling begins from the environmental-criminology proposition that crime opportunities are produced partly by configurations of places: bars, transit nodes, vacant properties, street networks, commercial facilities, poorly managed parcels, and other features may attract targets, facilitate offending, or weaken guardianship. RTM creates standardized spatial layers representing proximity, density, or exposure and combines selected factors into a risk surface. Its original shooting applications were explicitly presented as forecasts of where future shootings would be distributed. citeturn22search12turn22search16turn22search31
RTM differs from event-driven systems in three respects. First, its features are often more stable than yesterday’s incident pattern. Second, its intended intervention can be environmental—lighting, code enforcement, property management, business engagement, or removal of an opportunity structure—rather than repeated patrol. Third, the model can have an explanatory or diagnostic purpose even when it does not generate short-term probability forecasts.
RTM’s strengths are transparency of spatial factors and potential linkage to nonenforcement remedies. Its weaknesses include researcher degrees of freedom in selecting features, distances, operationalizations, and model-selection criteria; risk of encoding segregated land-use patterns; and frequent reliance on observational validation. Capturing many future incidents in the top risk decile establishes ranking ability, but not necessarily well-calibrated probabilities or a causal effect from acting on the map. citeturn22search4turn22search28
Comparable police-built systems
Police-built systems deserve the same functional scrutiny as commercial products. NYPD integrated predictive algorithms into its Domain Awareness System after earlier hand-built and computer-generated hot-spot mapping. From 2015 to 2016, it evaluated PredPol, HunchLab, and Keystats in a no-cost trial requiring five small forecast boxes by precinct, platoon, and crime category. PredPol withdrew, and NYPD ultimately used an internally developed algorithm. The department’s refusal and delayed production of records—and the reported inability to recreate old outputs—made independent evaluation difficult. citeturn23search2turn24view1turn24view2
NYPD’s Patternizr, by contrast, should not be treated as a place forecast. It is an investigative pattern-matching tool for connecting robbery, burglary, and larceny reports. Its use of machine learning and maps does not convert it into a system predicting future place-time risk.
LAPD’s Operation LASER also illustrates category overlap. LASER combined chronic-location analysis with person-focused “chronic offender” mechanisms. Its place component belongs in a place-based audit, but its individual scoring must be separately assessed under person-based due-process and bias standards. LAPD’s inspector general reviewed LASER, PredPol, and related data-driven practices and found inadequate documentation and insufficient evidence to isolate their crime effects. LAPD ended LASER in 2019. citeturn1search13turn1search17turn1search25turn23search13
Shreveport’s PILOT program is another useful comparator because it was built around agency analytics rather than a major commercial license. It combined historical crime locations, calls, field interviews, and prior hot spots to identify small property-crime risk areas, then directed locally selected prevention activities. Its evaluation demonstrates that in-house status does not ensure effectiveness, and proprietary status is not the defining issue. citeturn21view1turn21view2
Deployment audit and evidence matrix
The tables below distinguish documented facts from missing information. “Not publicly located” is not equivalent to zero cost, no training window, or no validation; it means the available record did not support a reliable entry. Operational statuses are stated as of August 2, 2026, where records permit.
City-by-city technical and operational table
| Jurisdiction and program | Inputs | Spatial and temporal unit | Target offenses | Training window and refresh | Output and patrol instructions |
|---|---|---|---|---|---|
| Santa Cruz, California — PredPol | Police-recorded incident type, location, and time. | Small place boxes, commonly described as about 500 by 500 feet; shift- or day-oriented forecasts. | Initially burglary and vehicle-related property offenses; configurations changed during the pilot period. | Historical incident window not fully disclosed in public deployment records; forecasts regenerated for operational shifts. | Maps of high-risk boxes. Officers were encouraged to visit boxes during available patrol time. Historical outcome claims were based largely on before-and-after city trends rather than randomized controls. citeturn5search21turn23search4 |
| Los Angeles — PredPol field experiment | Recorded crime events available to both ETAS model and analysts; incident location and time were central. | Dynamic 150-by-150-meter hot spots in the published experiment; daily assignment, with entire division-days randomized to algorithmic or analyst maps. | Burglary, vehicle theft, and theft from vehicles in the Kent comparison; configured crime categories in LAPD divisions. | Sliding and historical event data; daily forecasting. The article compared ETAS with analyst maps and simple three- and seven-day count maps. | Treatment and control maps looked alike. Officers focused free patrol time in displayed boxes; call logs measured patrol dosage. citeturn18view2turn19view1 |
| Plainfield, New Jersey — PredPol/Geolitica | Incident reports containing offense type, coordinates/address, and time. | Approximately 500-by-500-foot boxes; four locally defined shifts of just under 12 hours; 80 forecast records per day. | Motor-vehicle theft, robbery/aggravated assault, residential/nonresidential burglary, and gun crime, grouped into four forecast categories. | Vendor training window not publicly documented. Daily/shift output from at least February–December 2018 in the principal audit period. | Reports listed date, squad, crime category, and patrol boxes. Dosage was recorded when officers reported location, but officials said the product rarely or never directed patrol. citeturn26view1 |
| Philadelphia — HunchLab randomized trial | Crime records; multiple historical windows; recent-event indicators; census and housing variables; zoning, roads, transit, schools, hospitals, public facilities, and other environmental features. | 500-by-500-foot cells; three mission grids per district per day; one eight-hour treatment shift; district-week outcome aggregation. | Separate three-month phases for property crime and violent crime. | Several years of crime data; features included 28-, 56-, 84-, 112-, 168-, and 364-day historical levels and 3-, 7-, and 14-day near-repeat windows; daily mission selection. | Roll-call awareness, dedicated marked car, dedicated unmarked car, or business as usual depending on randomized district condition. citeturn18view3turn27view2 |
| Shreveport — PILOT | Historical crime locations, calls for service, field-interview information, previously identified hot spots, and agency intelligence. | Small forecast areas within six districts; three treatment and three control districts during the 2012 experiment. | Property crime. | In-house statistical model; precise public training horizon and refresh specification are incompletely documented in summary records. | Risk maps were paired with district-selected interventions intended to gather intelligence, solve prior offenses, create presence, and deter future crime. Implementation varied materially across districts. citeturn21view1turn21view2 |
| NYPD — vendor trial and in-house forecast system | Vendor trial: crime type, location, time; HunchLab requested at least five years of complaints plus environmental data; PredPol requested minimal incident fields. In-house inputs were not fully disclosed. | Trial requirement: five 300-by-300-foot boxes per precinct, platoon, and offense category. | Robbery, assault, burglary, theft from person, vehicle-related larceny, and shootings in the 2016 trial. | Vendor pilot ran approximately April–June 2016. Historical window for the later in-house system was not fully disclosed. | Comparative forecast outputs informed resource-allocation development. Historical predictions were reportedly not retained in a form that allowed comprehensive later production. citeturn24view1turn24view2 |
| Kent Police, United Kingdom — ETAS/PredPol-related trial | Recorded burglary, vehicle theft, and theft-from-vehicle incidents. | 150-by-150-meter boxes; forecasts evaluated at short future horizons. | Burglary, car theft, theft from vehicle. | Dynamic event history; three- and seven-day count maps were included as transparent comparators. | Silent forecast comparison against analyst hot spots; not all components constituted a patrol-outcome experiment. ETAS substantially outperformed short-window count maps in the reported tests. citeturn18view2turn19view1 |
| Newark and Atlantic City — RTM-informed policing | Crime events plus environmental features such as facilities, vacant properties, businesses, transport, and other local risk factors. | Raster or microplace risk surfaces; generally medium- to long-term risk rather than shift-by-shift aftershock boxes. | Shootings, robbery, and violent crime in prominent applications. | Models may use multiyear events and relatively stable environmental layers; updates are slower and project-specific. | Risk maps supported patrol changes, business checks, property interventions, lighting, and cross-agency responses. Public descriptions often combine model use with several simultaneous interventions. citeturn22search0turn22search1turn22search20 |
| Chicago — HunchLab/ShotSpotter Missions/ResourceRouter | Product architecture could combine reported incidents, environmental and temporal factors, staffing, and other agency data. Deployment-specific feature and model documentation was not publicly located in sufficient detail. | Missions used gridded risk and patrol assignments; vendor descriptions referred to cells around 250 meters for Missions-era products. | Agency-configured place-based crime priorities. | Deployment-specific window and update frequency not publicly documented sufficiently for an independent audit. | Forecast-driven patrol missions and dosage tracking were marketed. This must be distinguished from Chicago’s separate ShotSpotter acoustic dispatch program, which ended in September 2024. citeturn17search1turn17search10turn17search13turn17news22 |
Procurement, validation, code, and status table
| Jurisdiction and program | Vendor claims | Documented cost | Validation method | Public code or model? | Operational status |
|---|---|---|---|---|---|
| Santa Cruz PredPol | Early accounts associated the pilot with burglary reductions and more efficient patrol, but causal denominators and controls were inadequate. | Reliable complete contract cost not located; early partnerships included research and pilot support. | Primarily historical comparison and operational reporting, not a randomized test in Santa Cruz. | No complete production source code. Academic point-process methods were published, but the commercial implementation was proprietary. | Police use entered a moratorium in 2017. Santa Cruz enacted a predictive-policing ban in June 2020; the system is inactive there. citeturn1search12turn1search16turn23search7turn23search10 |
| LAPD PredPol | Higher forecast accuracy than analysts and additional crime deterrence under fixed patrol resources. | PredPol-specific lifetime cost was not cleanly isolated in the inspector-general record available for this audit; LAPD’s broader data-driven applications involved substantial staff, platform, and integration costs. | Randomized division-days in the published field experiment; later OIG review found agency records insufficient to isolate long-term program effects. | Academic algorithmic formulation published; operational platform closed. Historical forecast and dosage data were not routinely public. | LAPD ended or moved away from PredPol-era programs; LASER ended in 2019. Geolitica itself ended independent operation in 2023. citeturn18view2turn23search13turn27view0 |
| Plainfield Geolitica | Forecast boxes plus patrol dosage would deter crime and improve deployment efficiency. | $20,500 for the first annual term; $15,500 for a one-year extension. | Independent record linkage of forecasts, crimes, and dosage. No causal crime-reduction test because treatment was essentially not implemented. | Vendor code closed. The Markup released data-processing scripts, notebooks, derived data, dosage data, and a reproducible build process. | Contract ended; department reported little practical use. Geolitica ceased independent operations in 2023. citeturn26view0turn26view2 |
| Philadelphia HunchLab | Unified use of crime, near-repeat, temporal, environmental, and demographic signals; optimized missions and patrol tactics. | Earlier Philadelphia procurement reporting identified a one-year predictive-software contract around $80,000, but allocation across pilot, customization, support, and the specific experimental period requires caution. citeturn17search12 | NIJ-funded randomized district trial of patrol responses, not a head-to-head randomization of HunchLab against a simple forecast baseline. | Model and production code closed; input families and modeling architecture described. Authors had no financial investment in Azavea or ShotSpotter. | HunchLab was sold in 2018 and rebranded; the original product is no longer independently sold. Successor functionality exists in ResourceRouter. citeturn17search1turn27view2turn27view3 |
| Shreveport PILOT | Better identification of property-crime risk and more efficient targeting. | Treatment cost was estimated to be approximately $13,000–$20,000 below status quo, or 6–10 percent lower, chiefly because of reduced overtime; this is a comparative program-cost estimate, not a software price. | NIJ-sponsored randomized district assignment with process, impact, and cost evaluation. Low power and treatment heterogeneity. | In-house model; complete production code was not publicly located. | Historical 2012 experiment; not identified as an active current product. citeturn21view0turn21view2 |
| NYPD in-house system | Resource-allocation improvement relative to hand-produced hot spots and vendor alternatives. | Vendor trial was required to be free; cost of internal development and maintenance was not transparently isolated. | Internal vendor comparison; no sufficiently public independent causal evaluation. | No public production code. Historical outputs were not preserved or produced comprehensively. | The historical in-house predictive function was deployed through the Domain Awareness environment; current module-level status could not be established with sufficient public specificity. citeturn23search2turn24view1 |
| Kent ETAS trial | More crime captured in small forecast areas than analyst or simple recent-count maps. | Not publicly located as a separable procurement cost in the trial publication. | Out-of-sample silent tests; transparent count-map and analyst comparisons. | Mathematical method published; commercial implementation closed. | Historical evaluation. Geolitica no longer operates independently. citeturn18view2turn27view0 |
| RTM applications | Identifies environmental risk factors, forecasts microplaces, supports targeted risk reduction, and has been promoted as contributing to reductions around 20–30 percent in selected cities. | Project costs vary; complete city-specific procurement totals were not consistently public. Some implementations are research partnerships rather than conventional software subscriptions. | Forecast-validation studies and six-city or matched-area designs exist, but prominent operational claims often rest on before-and-after comparisons with bundled interventions. | Core methodology is extensively published; commercial RTMDx and implementation tools are not fully open source. | RTM remains an active method and commercial/research practice; individual city programs vary. citeturn22search1turn22search21turn22search24turn22search28 |
| Chicago Missions/ResourceRouter | Forecast-driven missions optimize scarce patrol resources; corporate materials frame the product as crime deterrence through strategically planned patrol. | Missions-specific Chicago price was not publicly isolated from broader SoundThinking relationships in the records reviewed. Chicago’s separate acoustic gunshot contract cost tens of millions of dollars and must not be attributed to predictive patrol software. | No independent Chicago causal evaluation of HunchLab/Missions located. | Closed system. | HunchLab/Missions survives as ResourceRouter under SoundThinking. Chicago ended acoustic ShotSpotter use in 2024, but that action does not by itself establish the status of every SoundThinking software module. citeturn17search11turn17news22turn18view4 |
Independent and vendor-connected evidence matrix
| Study or claim | Design | Independence and conflicts | Main result | Audit interpretation |
|---|---|---|---|---|
| Mohler et al., Los Angeles/Kent | Out-of-sample forecast tests and randomized division-days comparing ETAS with analyst maps. | Several authors co-founded or held stock in PredPol; public research grants also supported the work. citeturn27view0 | ETAS captured 1.4–2.2 times as much crime as analyst maps; LAPD treatment estimated at 7.4 percent crime reduction as a function of patrol time. | Stronger than testimonial evidence, but conflict disclosure, dosage measurement, analytical complexity, and lack of an independently reproduced field trial reduce certainty. |
| Philadelphia HunchLab experiment | Twenty districts randomized among awareness, marked car, unmarked car, and control; property and violent-crime phases. | NIJ-funded; authors reported no financial interest or HunchLab incentive. citeturn27view3 | Marked-car condition reduced expected property-crime counts by 31 percent; no comparable general effect for other conditions. | Good patrol-strategy evidence. Does not identify HunchLab’s incremental value over a recent-count or analyst baseline. |
| Shreveport RAND evaluation | Three treatment and three control districts; process, impact, and cost analysis. | NIJ-sponsored, independent RAND evaluation. | No significant incremental property-crime reduction; lower estimated treatment cost; low power and inconsistent delivery. | A credible null result with wide uncertainty, and a model for evaluating implementation fidelity. |
| Plainfield Markup audit | Independent linkage of 23,760 forecasts, crime reports, and dosage; treated forecasts excluded for accuracy analysis. | Independent journalism; reproducible repository released. | Less than 0.5 percent alert success overall; very low dosage; approximately 10 percent of forecast-eligible incidents captured across categories. | Strong evidence of poor operational value in this jurisdiction, but not a causal test of faithful forecast-led patrol. |
| RTM forecasting review | Systematic review and proportional meta-analysis of 25 studies. | Academic synthesis; much underlying RTM literature comes from developers or close collaborators. | Approximately half of future events in top 10 percent of risk cells. | Suggests useful ranking performance, but pooled proportions obscure offense, area, horizon, and denominator differences. |
| Atlantic City RTM claims | Before-and-after operational reporting, with environmental and patrol interventions. | Developer-associated and promotional descriptions are prominent. | Violent crime and robbery declined after implementation. | Cannot separate RTM from regression to the mean, citywide trend, patrol, business checks, lighting, demolition, and recording changes. |
| Hot-spots policing meta-analysis | Randomized and quasi-experimental studies of small-area policing. | Independent evidence synthesis across many interventions. | Small significant reduction; little evidence of inevitable displacement; some diffusion. | Strongest general basis for focused place intervention, but not evidence that proprietary forecasting is necessary. |
| Lum and Isaac; Ensign et al.; Akpinar et al. | Simulations and empirical-data demonstrations of data bias and feedback mechanisms. | Independent academic work. | Arrest-data targeting can reproduce racialized enforcement geography; discovered incidents generate runaway loops; differential victim reporting can shift forecast hot spots. | Establishes plausible and mathematically demonstrable failure mechanisms, though local disparity magnitude must be measured empirically. citeturn2search0turn25search3turn26view5 |
Methodological audit and public-data reproduction
Forecast metrics and their denominators
Predictive-policing claims frequently use “accuracy” without identifying the unit of analysis. At least five quantities can be called a success rate:
\[ \text{Incident hit rate} = \frac{\text{unique future incidents inside forecast areas}} {\text{all future incidents}} \]
\[ \text{Alert precision} = \frac{\text{forecast records matched to an incident}} {\text{all forecast records}} \]
\[ \text{Cell precision} = \frac{\text{forecast cells containing at least one incident}} {\text{forecast cells issued}} \]
\[ \text{Coverage} = \frac{\text{forecasted geographic area}} {\text{eligible study area}} \]
\[ \text{Prediction Accuracy Index} = \frac{\text{incident hit rate}}{\text{coverage}}. \]
The National Institute of Justice used PAI in its Real-Time Crime Forecasting Challenge. Entries forecast between 0.25 and 0.75 square miles within a 147.71-square-mile Portland study area, and PAI rewarded the proportion of events captured relative to the proportion of area designated. citeturn26view3turn8view0
PAI is useful for comparing forecasts at similar scales, but it is not a probability-accuracy measure. A tiny forecast area can obtain a large PAI from very few events. PAI also changes if the denominator includes parks, industrial land, water, highways, or areas with no meaningful population exposure. Street length, address count, businesses, households, resident population, ambient population, or patrol-accessible area may be more appropriate denominators for particular questions. Published methodological critiques show that PAI rankings can depend materially on these choices. citeturn3search8turn3search19
Precision and recall answer different operational questions. Precision measures forecast burden: how many issued alerts correspond to the target outcome? Recall measures coverage of harm: how many target incidents were captured? A system can achieve high recall by designating much of a city, but that produces low precision and little prioritization. A system can achieve high PAI by designating a tiny area around one persistent location while missing most citywide incidents.
Calibration asks whether predicted probabilities correspond to observed frequencies. For cells assigned risk \(p\), approximately \(p\) of equivalent cell-periods should contain an event. Calibration can be evaluated with reliability plots, calibration intercept and slope, Brier score, and log loss. A product that releases only ranked red boxes without probabilities cannot be fully calibrated by outsiders. “Twice average risk” is not interchangeable with “a 20 percent probability.”
Forecast lead time is the gap between model issuance and the target period. A forecast made after a late report may inadvertently include information unavailable at the intended decision time. Data pipelines must use immutable “known as of” timestamps, not merely occurrence dates, or they risk leakage from revised reports, delayed entry, or retrospective geocoding.
Baseline comparison must include more than random allocation. At minimum, a proposed system should be compared with:
\[ \begin{aligned} B_1 &: \text{recent event count by cell},\\ B_2 &: \text{historical average by cell and season},\\ B_3 &: \text{kernel-density estimate},\\ B_4 &: \text{last-period persistence},\\ B_5 &: \text{analyst-selected hot spots},\\ B_6 &: \text{same patrol budget allocated to conventional hot spots}. \end{aligned} \]
The Los Angeles/Kent publication did include analysts and three- and seven-day count maps. Analysts were statistically indistinguishable from seven-day count maps in the reported Kent comparisons, while ETAS performed better. This is precisely the type of transparent benchmark needed, although it should be independently replicated on multiple cities and locked evaluation periods. citeturn18view2turn19view1
Reproduction of the Plainfield evaluation
The Markup’s public repository contains prediction data, dosage data, census-block-group demographics, preparation scripts, an analysis notebook, dependency files, and a make reproduce workflow. It joins Geolitica forecasts to Plainfield police reports and calculates success rates. This is unusually strong reproducibility practice for a police-technology evaluation. citeturn26view2
The arithmetic below reproduces core published metrics from the released denominators. It is a denominator-level reproduction rather than a claim to have re-geocoded every address independently.
Robbery and aggravated assault
Published counts:
\[ N_{\text{alerts}} = 5{,}722 \]
\[ N_{\text{matched alert records}} = 32 \]
\[ N_{\text{actual incidents}} = 139 \]
\[ N_{\text{unique incidents captured}} = 20. \]
Therefore:
\[ \text{Alert precision} = \frac{32}{5{,}722} = 0.005592 = 0.559\%. \]
\[ \text{Incident recall} = \frac{20}{139} = 0.1439 = 14.4\%. \]
The alert workload was:
\[ \frac{5{,}722}{20} = 286.1 \]
forecast records per unique captured robbery or aggravated assault. Because overlapping predictions sometimes matched the same event, alert precision and incident recall cannot be inferred from one another. citeturn27view5
Burglary
Published counts:
\[ N_{\text{alerts}} = 10{,}141, \quad N_{\text{actual burglaries}} = 116, \quad N_{\text{captured burglaries}} = 8. \]
Thus:
\[ \text{Incident recall} = \frac{8}{116} = 6.90\%. \]
There were approximately:
\[ \frac{10{,}141}{8} = 1{,}267.6 \]
burglary forecast records for each unique burglary captured. The publication reported alert success of approximately 0.1 percent; the exact number of matched burglary forecast records, as distinct from unique incidents, is required for a more precise alert-level calculation. citeturn27view5
Overall target-incident coverage and patrol delivery
Across the forecast-eligible target categories, the public methodology reported 34 of 336 relevant incidents captured:
\[ \text{Overall incident recall} = \frac{34}{336} = 10.12\%. \]
After treated prediction records were excluded, 23,631 forecast records remained:
\[ \frac{23{,}631}{336} = 70.3 \]
untreated forecast records per relevant incident occurring anywhere in the city during the period.
Actual patrol occurred during 129 of 23,760 forecast instances:
\[ \text{Observed forecast-box dosage rate} = \frac{129}{23{,}760} = 0.543\%. \]
Dosage records averaged about 20 minutes, although the agency stated that those overlaps were coincidental rather than forecast-directed. citeturn26view1turn27view5
These calculations support three conclusions.
First, Plainfield’s forecast output had very low alert precision and low incident recall under The Markup’s matching rules. Second, the department delivered almost no forecast-linked treatment, so the procurement produced negligible operational exposure. Third, the case cannot resolve whether a properly implemented PredPol patrol strategy would reduce crime, because the treatment mechanism was absent.
A full independent replication should additionally recalculate results under alternative spatial buffers, exact versus interval-censored occurrence times, address-geocoding uncertainty, offense-classification rules, unique cell-period rather than record-level precision, and simple recent-count baselines. The original methodology already used a generous 600-by-600-foot matching box around nominal 500-foot forecasts and documented ambiguous addresses, report-time uncertainty, and offense hierarchy limitations. citeturn15view0turn26view1
Forecast accuracy is not crime reduction
A forecast evaluation observes whether future recorded events fall inside designated places. A crime-reduction evaluation asks whether acting on the forecast changes outcomes relative to a credible counterfactual.
Let \(F_j\) denote the forecast method, \(D_j\) patrol dosage, \(T_j\) the tactics delivered, and \(Y_j\) subsequent crime. A simple causal structure is:
\[ F_j \rightarrow D_j \rightarrow T_j \rightarrow Y_j. \]
But the same forecast can produce different \(D_j\) and \(T_j\), while call demand, staffing, weather, neighborhood conditions, and existing patrol priorities affect both dosage and outcomes. Comparing high-dosage boxes with low-dosage boxes is confounded because officers may spend more time where risk is visibly greater.
The cleanest factorial evaluation would randomize both forecast source and intervention:
| Forecast assignment | Patrol assignment | Question identified |
|---|---|---|
| Transparent baseline | Business as usual | Baseline outcome |
| Proprietary model | Business as usual | Whether merely informing officers changes behavior |
| Transparent baseline | Standardized focused patrol | Effect of conventional hot-spot patrol |
| Proprietary model | Same standardized focused patrol | Incremental effect of proprietary location selection |
| Random or placebo boxes | Same patrol dosage | Whether any visible patrol at comparable places works |
| No displayed boxes | Same total staffing | Systemwide effects and contamination |
The contrast between proprietary-model patrol and transparent-baseline patrol estimates algorithmic value only if patrol time, tactics, staffing, forecast-area size, offense definition, and temporal coverage are held constant. Shreveport failed to maintain fully comparable interventions across districts, which is why RAND could not cleanly isolate map quality. citeturn21view0
Philadelphia showed a related principle: awareness alone was insufficient, while a dedicated marked car produced a property-crime effect. The active ingredient appears to have included visible, sustained patrol rather than the simple existence of a risk score. citeturn18view3turn19view2
Displacement, diffusion, and contamination
A treatment-area decline is not enough. Offending may move to adjacent places or later times:
\[ \text{Net effect} = \Delta Y_{\text{treatment}} + \Delta Y_{\text{buffer}} + \Delta Y_{\text{rest of jurisdiction}}. \]
Spatial displacement occurs if crime shifts into nearby untreated areas. Temporal displacement occurs if it shifts beyond the treated shift. Target or offense displacement occurs if offenders switch victims or offense categories. Diffusion of benefits occurs if deterrence extends beyond the treated area or time.
The Philadelphia experiment reported a reduction during the subsequent eight-hour period after marked-car property-crime patrol, consistent with temporal diffusion. Its evaluators also expanded the treatment unit to adjacent cells because officers could not realistically remain inside one 500-foot square and had to travel along surrounding streets. That adjustment illustrates why nominal forecast geometry is not the same as the actual treatment footprint. citeturn18view3turn27view2
Treatment contamination arises when control officers enter treatment areas, commanders independently target the same places, citywide operations affect both groups, or forecast maps reveal information used outside assigned conditions. GPS or automated vehicle-location data should measure all officer presence, not only self-declared “missions.” The Los Angeles study acknowledged that call logs were less precise than in-car GPS for this purpose. citeturn27view0
Patrol dosage and dose response
Dosage should be measured as more than a binary visit:
\[ D_{it} = \sum_o \left( \text{officer-minutes}_{oit} \times \text{visibility}_{oit} \times \text{tactic weight}_{oit} \right). \]
Required components include officer-minutes, number of officers, marked versus unmarked presence, stationary versus moving patrol, pedestrian versus vehicle presence, stops, searches, business contacts, arrests, calls handled, and interruptions. “Entered box” is not equivalent to meaningful treatment.
Analysis should estimate a dose-response curve and test diminishing returns. Repeated fifteen-minute patrol episodes may differ from one uninterrupted hour. High dosage can deter visible street offenses while increasing recorded drug, weapons, traffic, and disorder enforcement. Outcomes must therefore distinguish victim-reported offenses from police-initiated detections.
Data generation, feedback loops, and racial geography
Recorded crime is a selected measurement process
Underlying victimization is not directly observed in police records. A simplified measurement model is:
\[ R_{ist} = U_{ist} \times P(\text{reported}\mid U,X_{ist}) \times P(\text{recorded}\mid \text{report},Z_{ist}), \]
where \(U\) is underlying victimization, \(X\) includes victim trust, offense severity, insurance requirements, language access, immigration concerns, and police accessibility, and \(Z\) includes agency recording rules, classification, workload, and discretion.
The National Crime Victimization Survey measures both reported and unreported nonfatal victimization, whereas police incident systems measure crimes known to law enforcement. In 2024, only about three in ten property victimizations were reported to police, including approximately 41 percent of burglaries. For 2020–2023, BJS estimated that about 38 percent of urban violent victimizations were reported, compared with 43 percent in suburban and 51 percent in rural areas; reporting also varied sharply by offense. citeturn25search1turn25search5turn25search16turn25search17turn25search20
Reported burglary is generally closer to a victim-driven measure than a narcotics arrest, but it remains selected. Residents may report because of insurance, property value, confidence that police will respond, surveillance-camera availability, or institutional requirements. Conversely, fear of retaliation, prior negative police contact, informal resolution, or belief that police cannot help can suppress reporting.
Calls for service are another distinct process. They can measure victimization, disorder, concern, nuisance complaints, duplicate reports, proactive officer activity, or differences in willingness and ability to call. A neighborhood with more calls may have more harm, more trust, more surveillance by residents, more conflict over public space, or some combination.
Arrests and stops are even more endogenous. They require police presence, observation, legal discretion, investigative priorities, and enforcement choices. For offenses such as drug possession, public drinking, loitering, traffic violations, and weapons possession, increased patrol directly increases opportunities to discover records that are then treated as evidence of future risk.
Feedback-loop causal diagram
Structural conditions and segregation
│
├──────────────► Underlying victimization ◄──────────────┐
│ │ │
│ ▼ │
│ Victim reporting behavior │
│ │ │
│ ▼ │
│ Police-recorded incidents │
│ │ │
│ ▼ │
└──────────────► Forecast / risk model │
│ │
▼ │
Patrol allocation and tactics │
│ │ │ │
│ │ │ │
▼ ▼ ▼ │
Deterrence or Police-discovered Stops, │
displacement offenses/arrests searches,
│ │ force
│ └──────► records ◄─────┘
│
▼
Future underlying victimization
Patrol experience also affects trust and future reporting, creating
a second loop through the measurement process.
Ensign and colleagues mathematically demonstrated that repeatedly updating a model with discovered crime can create runaway allocation to the same neighborhoods even when true crime rates do not justify it. Resident-reported incidents attenuate the loop but do not necessarily eliminate it. citeturn26view5turn25search14
Lum and Isaac’s Oakland analysis supplied recorded drug events to a PredPol-like process and compared the resulting geography with survey-based estimates of drug use. The simulated patrol allocation concentrated in lower-income, minority neighborhoods and would have targeted Black residents at roughly twice the rate of white residents, despite survey evidence that drug use was much more broadly distributed. The lesson is not that every burglary forecast has the same disparity, but that police-generated labels can transform enforcement geography into apparent risk geography. citeturn2search0turn2search4turn16search7
Differential reporting creates a different failure mode. Akpinar, De-Arteaga, and Chouldechova showed that even systems trained on victim reports rather than arrests can shift hot spots away from high-victimization, low-reporting areas and toward medium- or high-victimization areas with higher reporting. The resulting allocation can simultaneously overpolice one area and underserve another. citeturn25search3turn25search11
Dirty data and historical police misconduct
Richardson, Schultz, and Crawford documented predictive-policing deployments in jurisdictions where source data had been produced during periods of unlawful, biased, corrupt, or otherwise unreliable police practice. Their “dirty data” framework is broader than ordinary statistical noise. It asks whether records embody unconstitutional stops, fabricated evidence, discriminatory enforcement, manipulated classifications, or institutional incentives that make them unsuitable for secondary algorithmic use. citeturn2search2turn2search6turn25search6
A predeployment audit must therefore identify:
\[ \text{record source} \rightarrow \text{legal authority} \rightarrow \text{collection practice} \rightarrow \text{known misconduct} \rightarrow \text{cleaning or exclusion decision}. \]
Deleting race from the model is not enough. Location is correlated with residential segregation, zoning, wealth, transit access, surveillance density, calls, and prior enforcement. Environmental features can operate as proxies for racialized urban policy without any explicit race field. Conversely, excluding all environmental variables does not eliminate bias if incident labels already reflect unequal enforcement or reporting.
Service patterns and repeat police presence
Police presence affects both the numerator and the meaning of crime rates. Consider two neighborhoods with equal underlying rates of low-level drug possession. If one receives twice the officer-hours, it offers twice the detection opportunity. Feeding resulting arrests into the model can justify still more patrol.
For victim-reported violent or property offenses, repeated presence may increase reporting by making police more accessible. A rise in recorded incidents after deployment could therefore mean increased crime, increased reporting, increased detection, or reduced trust accompanied by changes in which events are reported. A fall could mean deterrence, displacement, underreporting, victim withdrawal, or recording changes.
For this reason, evaluation should report separate outcome families:
| Outcome family | Examples | Primary interpretation problem |
|---|---|---|
| Victim-initiated | Reported robbery, burglary, assault calls | Underreporting, trust, occurrence-time uncertainty |
| Police-initiated | Drug arrests, weapons possession, traffic and disorder enforcement | Exposure and discretion are treatment-induced |
| Harder-to-manipulate severe outcomes | Homicide, shooting injury, emergency medical treatment | Rare events, low power, incomplete linkage |
| Community survey outcomes | Victimization, perceived safety, trust, avoidance | Sampling cost and response bias |
| Operational outcomes | Response time, officer-hours, calls cleared | May improve efficiency without reducing harm |
| Police-contact outcomes | Stops, searches, arrests, force, complaints | Benefits and burdens must be assessed jointly |
Racial geography in actual deployments
The Markup’s broader Geolitica investigation found that forecast areas across 38 jurisdictions disproportionately covered neighborhoods with higher shares of low-income, Black, and Latino residents relative to the jurisdiction as a whole. That comparison does not alone establish unjustified discrimination because crime, population exposure, land use, and reporting also vary geographically. It does establish a distributive burden requiring explanation and causal measurement—especially when forecast-linked patrol may entail stops, searches, surveillance, and repeated presence. citeturn27view5
An adequate racial-geography audit should calculate, for each demographic group \(g\):
\[ \text{Forecast exposure}_g = \frac{\sum_{c,t} I(c \text{ forecast}) \times \text{population}_{gc}} {\sum_c \text{population}_{gc}}, \]
\[ \text{Patrol exposure}_g = \frac{\sum_{c,t} \text{officer-minutes}_{ct} \times \text{population}_{gc}} {\sum_c \text{population}_{gc}}, \]
and contact, stop, search, arrest, and force rates per resident, per ambient population, per reported victimization, and per officer-hour. No single denominator is sufficient. Resident population is weak for commercial districts; ambient population is difficult to estimate; reported crime inherits reporting bias; officer-hour denominators can normalize an already unequal allocation. A robust audit presents several denominators and explains their assumptions.
The relevant benefit distribution must also be measured. A heavily patrolled Black neighborhood may receive reduced victimization and faster response while bearing more stops and surveillance. Accountability requires estimating both rather than treating either enforcement intensity or crime reduction as the sole welfare measure.
Procurement, transparency, and minimum field-testing protocol
Procurement failures
Predictive-policing contracts have often been treated as ordinary software subscriptions even though the product affects government deployment of coercive authority. Conventional procurement questions—price, uptime, support, integration, and cybersecurity—are insufficient.
The Plainfield contract illustrates a basic value-for-money failure. The department paid $20,500 for the first year and $15,500 for an extension, generated tens of thousands of forecasts, but officials said the system was rarely or never used to direct patrol. A technically sophisticated product has no operational value if training, command policy, workflow integration, staffing, or officer acceptance prevents treatment delivery. citeturn26view0turn26view1
NYPD’s records illustrate a transparency failure. Vendor pilots were structured as free “gifts,” subject to nondisclosure agreements, and the department initially denied public-record requests by invoking nonroutine techniques, security, and proprietary interests. After litigation, records emerged piecemeal, but the in-house system reportedly did not store historical outputs. A free pilot can create switching costs, shape internal development, or transfer methods without appearing as a conventional procurement expenditure. citeturn24view1turn24view2
Acquisitions create another accountability gap. HunchLab’s transfer from Azavea to ShotSpotter meant that logs identifying variable importance in the Philadelphia models were unavailable to later evaluators. Geolitica’s closure and asset transfer likewise complicated the distinction among source code, patents, engineering knowledge, customer contracts, and successor functionality. Contracts that do not require escrow, documentation transfer, output retention, and post-termination audit access make public validation dependent on vendor continuity. citeturn26view0turn27view2
The most common deficiencies are:
- no publication of forecast denominators or historical outputs;
- no contractually specified baseline;
- “accuracy” warranties without a defined metric, area, offense, or horizon;
- vendor evaluation of its own product;
- nondisclosure provisions that obstruct public oversight;
- no requirement to preserve model versions and feature logs;
- no linkage between forecasts and patrol dosage;
- no audit of source-data legality and reliability;
- no disparate-impact or community-contact analysis;
- no termination trigger for nonuse or failure to beat baseline;
- no assignment of responsibility when a vendor is acquired or shuts down.
Contract requirements
A procurement should state that all forecasts, scores, model versions, input snapshots, configuration changes, officer-facing outputs, and dosage records are government records subject to retention and audit, with narrow redaction for genuinely sensitive operational details. Public evaluation can aggregate or delay release to avoid facilitating evasion; secrecy about active locations does not justify permanent secrecy about historical performance.
The contract should define a forecast event and specify whether multiple overlapping alerts count separately. It should state the eligible area, excluded land, temporal window, offense taxonomy, occurrence-time rules, geocoding procedure, late-report policy, and treatment of revised classifications. It should require forecast probabilities or comparable scores—not only colored boxes—so calibration can be tested.
Payment should be staged. Initial compensation may cover integration and a silent test. Continued payment should depend on data quality, operational use, independent baseline performance, and evaluation completion, not on unverified crime trends. Crime reduction itself should not be a simple performance bonus, because agencies and vendors could influence classifications, reporting, or forecast coverage.
Intellectual-property protection should not prevent:
- independent execution on a locked test set;
- inspection of feature definitions and transformations;
- subgroup and neighborhood performance testing;
- access to model cards and version histories;
- evaluation of calibration and false positives;
- public disclosure of aggregate metrics;
- preservation of evidence after contract termination.
Minimum field-testing protocol
The following protocol is the minimum defensible sequence before operational deployment.
| Phase | Required procedure | Pass criterion |
|---|---|---|
| Problem definition | Specify the harm to be reduced, affected population, offense definition, decision being supported, and why existing analysis is inadequate. | A forecastable, operationally actionable problem with a lawful and proportionate response. |
| Data provenance | Inventory every field, collection mechanism, legal basis, missingness pattern, geocoding error, reporting pathway, known misconduct period, and modification timestamp. | No unresolved source whose inclusion could materially contaminate results; police-initiated and victim-initiated events are separated. |
| Baseline construction | Implement recent counts, historical averages, seasonal persistence, kernel density, and analyst maps under the same area and forecast budget. | Proprietary model must improve preregistered metrics over the strongest transparent baseline, not merely random allocation. |
| Silent prospective test | Freeze model and baseline before outcomes occur; issue no patrol instructions; retain all outputs and input snapshots. | Statistically and operationally meaningful improvement in recall, precision, PAI, calibration, and lead-time utility across multiple periods. |
| Robustness test | Vary grid origin, cell size, time window, geocoding assumptions, offense definitions, delayed reports, and coverage area. | Performance is not an artifact of a single boundary, denominator, or leakage-prone timestamp. |
| Equity audit | Measure forecast and expected patrol exposure by race, income, age, housing tenure, and neighborhood, using several denominators. | Any disparity has a documented relation to legitimate harm reduction and survives less burdensome alternatives analysis. |
| Community review | Publish delayed/aggregated methods, intended tactics, rights implications, complaint process, and sunset terms; solicit affected-community input. | Governing body makes an informed, public authorization rather than leaving adoption solely to police or procurement staff. |
| Randomized field trial | Randomize comparable place-times to proprietary forecast, transparent baseline, placebo or business-as-usual conditions; standardize patrol dosage and tactics. | Incremental crime or harm reduction with uncertainty intervals, not merely within-group before-and-after decline. |
| Dosage and contamination monitoring | Use automatic location data and activity records to measure all officer-minutes, calls, stops, searches, arrests, force, and spillovers. | Sufficient separation between conditions and documented treatment fidelity. |
| Displacement and diffusion analysis | Examine adjacent rings, later shifts, alternative offenses, and jurisdictionwide totals. | Net benefit remains after accounting for movement in place, time, target, and offense. |
| Community-impact analysis | Measure victimization, perceived safety, trust, calls, stops, searches, arrests, complaints, force, and avoidance behavior. | Benefits exceed enforcement and legitimacy costs under a publicly stated welfare framework. |
| Replication and sunset | Independent team reproduces results; contract expires automatically absent renewed evidence. | Replication confirms value and no material drift, data contamination, or disparate harm. |
Statistical design requirements
Power analysis must precede the trial. Microplace crime is sparse and overdispersed; a study with very few districts or events can miss policy-relevant effects. The unit of randomization should minimize spillover while providing enough units for inference. If commanders cannot prevent officers from crossing nearby grid boundaries, randomizing larger clusters or time blocks may be preferable.
The primary analysis should be intention-to-treat:
\[ Y_{it} = \alpha_i + \gamma_t + \beta Z_{it} + \epsilon_{it}, \]
where \(Z_{it}\) is randomized assignment. A secondary treatment-on-treated estimate can use assignment as an instrument for actual dosage if exclusion and monotonicity assumptions are plausible.
Count outcomes should generally use Poisson or negative-binomial models with exposure offsets and cluster-robust or randomization-based uncertainty. Rare violent outcomes may require longer periods, pooled prespecified categories, or Bayesian partial pooling, but post hoc category expansion should not rescue a failed primary endpoint.
Forecast and causal endpoints must be separately preregistered. A model can win the forecast endpoint but fail the crime endpoint. A patrol intervention can win the crime endpoint even when the proprietary model does not beat the baseline. These are substantively different findings.
Public reporting template
Every published evaluation should include:
\[ \begin{array}{ll} \textbf{Data:} & \text{source, report type, exclusions, missingness, geocoding error};\\ \textbf{Forecast:} & \text{unit, horizon, lead time, coverage, probability or rank};\\ \textbf{Baseline:} & \text{method, tuning rules, analyst information};\\ \textbf{Accuracy:} & \text{precision, recall, PAI, calibration, uncertainty};\\ \textbf{Operations:} & \text{officer-minutes, tactics, compliance, contamination};\\ \textbf{Causality:} & \text{randomization or identification strategy};\\ \textbf{Spillovers:} & \text{displacement and diffusion};\\ \textbf{Equity:} & \text{forecast, patrol, contact, and benefit distribution};\\ \textbf{Cost:} & \text{license, integration, analyst, officer, evaluation, exit};\\ \textbf{Reproducibility:} & \text{code, data schema, outputs, model versions};\\ \textbf{Governance:} & \text{authorization, complaints, audit, sunset, termination}. \end{array} \]
Public release need not disclose active patrol boxes, current staffing vulnerabilities, or real-time response rules. Historical outputs can be delayed, spatially generalized where necessary, and evaluated inside secure environments. Tactical sensitivity is compatible with meaningful algorithmic accountability.
Final assessment
What adds value beyond conventional hot-spot policing?
Conventional hot-spot identification plus focused, procedurally constrained patrol has the strongest evidence base. The aggregate literature indicates modest crime reductions and does not support the assumption that crime simply moves next door. Visible patrol, problem-oriented responses, and environmental changes can work when delivered with sufficient dosage. citeturn3search2turn3search6turn22search11
Near-repeat forecasting adds plausible short-term information for selected high-volume offenses, especially burglary and vehicle crime, where recent events can temporarily elevate nearby risk. Its incremental value is likely greatest where event times are timely and reliable, near-repeat chains are common, patrol can respond quickly, and the baseline is not already a high-quality dynamic hot-spot map. Its value is likely small where reporting is delayed, event volume is sparse, or chronic locations dominate.
Self-exciting point-process models have demonstrated forecast gains in important trials, particularly the Los Angeles/Kent work. The reported gains over analysts and short-window count maps should not be dismissed. However, the evidence base is too narrow and too connected to PredPol’s developers to establish a general, jurisdiction-independent advantage. Independent multicity replication against tuned transparent baselines remains the missing test. citeturn18view2turn27view0
HunchLab showed that richer features can be operationalized, but the strongest field evidence points more clearly to patrol strategy than to proprietary model superiority. In Philadelphia, a dedicated marked car reduced property crime, while lower-dosage or less-visible conditions did not. Since HunchLab was not randomized against simple recent-count hot spots, the study cannot determine whether gradient boosting, demographic and environmental variables, or mission optimization selected materially better places. citeturn18view3turn27view2
Risk-terrain modeling offers the most distinctive conceptual contribution when it changes environments rather than merely intensifying patrol. It can identify modifiable features and support lighting, property management, business practices, code enforcement, or place redesign. Its forecasting literature suggests meaningful spatial ranking ability. Its causal crime-reduction claims remain weaker when interventions are bundled and evaluated through before-and-after trends. RTM should be judged most favorably as a transparent diagnostic framework linked to noncoercive remediation, and more skeptically when it functions mainly as another rationale for concentrated stops and enforcement. citeturn22search1turn22search28turn22search31
PredPol/Geolitica’s public evidence is mixed and ultimately insufficient to justify broad vendor claims. A developer-connected experiment reported meaningful performance and patrol effects. An independent Plainfield audit found extremely low alert precision, low incident recall, and almost no operational use. Santa Cruz stopped using the system and later banned predictive policing. LAPD’s inspector general could not isolate program effects from available records. Geolitica’s corporate shutdown and asset transfer further complicated accountability. citeturn23search7turn23search13turn26view0turn27view0
Shreveport provides the clearest caution against attributing value to a map without controlling operations. Its randomized design was undermined by low power and heterogeneous implementation, and it found no significant incremental reduction. Its central recommendation remains the right standard: compare predictive and traditional maps while holding interventions and effort constant. citeturn21view0turn21view1
Police-built tools are not presumptively more accountable than vendor products. They may avoid subscription costs and trade-secret barriers, but undocumented code, absent version control, inaccessible outputs, and internal-only validation can make them equally opaque. NYPD’s inability to recreate historical outputs is an example. citeturn24view1
Overall evidentiary rating
| Method | Forecast evidence | Causal crime-reduction evidence | Incremental value over simple hot spots | Accountability assessment |
|---|---|---|---|---|
| Recent counts, persistence, KDE, analyst hot spots | Strong as baseline; performance varies by offense and scale | Substantial broader hot-spots literature when paired with intervention | Baseline by definition | High transparency and low implementation cost |
| Near-repeat rules | Moderate for selected property offenses | Limited evidence that the forecast component itself causes additional reduction | Plausible but context-dependent | Can be implemented transparently |
| Self-exciting point processes | Promising, including favorable LA/Kent tests | One prominent favorable patrol trial, with developer conflicts | Demonstrated in limited settings, not broadly replicated | Mathematical form public; production systems historically closed |
| PredPol/Geolitica | Mixed: favorable developer-linked trials, poor Plainfield results | Insufficient independent causal replication | Unproven as a general proposition | Weak historical transparency; company exited and assets transferred |
| HunchLab/Missions/ResourceRouter | Technically sophisticated; limited independent forecast benchmarking | Philadelphia supports marked patrol in forecast places | Algorithmic increment not identified | Closed model; ownership transfer caused loss of some logs |
| Risk-terrain modeling | Moderate-to-strong ranking evidence across studies | Limited clean causal attribution; bundled interventions common | Potentially valuable for environmental diagnosis | More interpretable factors, but implementation and software may remain proprietary |
| In-house police forecasting | Variable and often undocumented | Sparse independent evidence | Unknown without mandatory baselines | No inherent transparency advantage |
| Gunshot detection | Not a forecast method | Separate detection/response evidence question | Not applicable | Should not be represented as predictive policing |
| RTCC integration | Not inherently predictive | Depends on specific modules and interventions | Not applicable without a forecast component | Requires module-by-module audit |
Final judgment
The technically defensible conclusion is narrower than either vendor promotion or categorical rejection.
Place-based crime concentration is real enough to support focused prevention. Transparent recent-event counts, historical persistence, kernel-density maps, and analyst knowledge already capture much of that signal. Near-repeat and self-exciting models can improve short-term ranking for some offenses, and environmental models can identify stable risk settings. Yet the public evidence does not show that proprietary predictive-policing products reliably and generally produce crime reductions beyond well-implemented conventional hot-spot policing.
Where positive effects have appeared, they are often inseparable from visible patrol dosage, problem-oriented intervention, or environmental remediation. Where products have failed, the causes include poor precision, weak recall, nonuse, missing dosage, incompatible workflows, low event volume, and inadequate evaluation. The relevant procurement question is therefore not “Does the algorithm predict crime?” It is:
\[ \text{Does this system, compared with a transparent alternative, produce enough additional public benefit to justify its financial cost, enforcement burden, data risks, and loss of accountability?} \]
On the evidence available through August 2, 2026:
- A jurisdiction can justify conventional microplace hot-spot policing when it uses lawful, proportionate, evidence-based tactics and monitors displacement and community harm.
- Near-repeat or self-exciting forecasts may merit a controlled trial for timely, high-volume property offenses, but only after they beat tuned recent-count baselines in a prospective silent test.
- RTM may add the most public value when it redirects policy toward modifiable environmental conditions rather than repeated enforcement.
- HunchLab-style mission optimization may improve operational discipline, but its forecasting and routing components must be evaluated separately from the effect of dedicated patrol.
- PredPol/Geolitica did not establish a sufficiently independent, reproducible, multicity evidence base before its corporate exit.
- Rebranding predictive functions as “precision policing,” “resource routing,” or “mission management” does not remove them from audit.
- Acoustic gunshot detection, crime mapping, real-time crime centers, and investigative pattern matching should not be called predictive policing unless they actually estimate future risk and influence prospective allocation.
- No system should be deployed without preserved outputs, complete dosage measurement, independent baseline comparison, randomized or credible quasi-experimental evaluation, racial-geography analysis, public governance, and an enforceable sunset clause.
The minimum responsible default is therefore transparent hot-spot analysis first, proprietary prediction only after demonstrated incremental value, and coercive patrol only after separate evidence that the chosen intervention reduces harm more than it creates.