CleanZone Field Brief
Cancer Clusters and the Superfund Next Door
The US Environmental Protection Agency maintains two lists that every homebuyer should know: the National Priorities List (NPL) of Superfund sites, and the Toxics Release Inventory (TRI) of reported industrial emissions. Neither is a death sentence. Neither is irrelevant.
At a glance
- The US EPA tracks two different registries: the National Priorities List (NPL) for legacy Superfund contamination, and the Toxics Release Inventory (TRI) for ongoing self-reported emissions. The EU equivalent is the E-PRTR.
- Pollutant concentration falls off with distance from a stack or discharge point — but rarely to zero, and rarely as a neat circle once wind direction is added.
- Most "cancer cluster" claims that start from a map fail once tested against age-adjusted registry data — that's the Texas-sharpshooter fallacy plus ordinary small-number noise, not evidence of a shared cause.
- A genuine exposure concern needs three things together — a plausible pathway, a measurable dose, and a biologically consistent cancer type — not proximity alone.
- Groundwater plumes can keep migrating for decades after a plant is demolished, so "no factory here anymore" isn't the same as "no contamination here anymore".
As of mid-2026, the EPA NPL contains roughly 1,300 sites across the United States, of which approximately 430 have no formal remedy in place and roughly 55 are proposed for addition. The list is not static — sites are deleted once cleanup goals are achieved and long-term monitoring is institutionalised. The TRI, meanwhile, captures annual self-reported releases from roughly 21,000 industrial facilities nationwide, covering nearly 800 chemicals and chemical categories. In the European Union, the equivalent framework is the European Pollutant Release and Transfer Register (E-PRTR), covering roughly 30,000 facilities. None of these registries, on their own, tells you whether one specific household nearby carries elevated health risk — that takes pathway and dose, which is the harder half of this brief.
NPL vs. TRI: two different conversations
The NPL is about legacy contamination
Superfund sites are places where past disposal or accidental release has created a hazard requiring long-term remediation. They include abandoned mines, former smelters, defunct chemical plants, and military bases with solvent or fuel plumes. The critical question for a nearby resident is not whether a site exists, but where the contamination has travelled and whether the migration pathway intersects with human exposure points — groundwater wells, surface-water intakes, or vapour intrusion into basements.
The EPA publishes Record of Decision (ROD) documents for each NPL site that describe the contaminant of concern, the medium (soil, groundwater, sediment, air), and the remedy selected. A site with a completed soil cap and a groundwater pump-and-treat system in operation for ten years is a different proposition from a site whose groundwater plume has been detected off-property but not yet delineated. Crucially, a plume keeps moving on its own schedule, long after the source that created it is gone.
Neither registry ranks risk for you. A TRI facility reporting a large tonnage of a low-toxicity compound can matter less than a smaller facility releasing a compound on the EPA IRIS or IARC Group 1 carcinogen lists. Read the chemical, not just the tonnage.
The TRI is about ongoing emissions
Toxics Release Inventory data are self-reported by facilities that meet EPA thresholds — typically ten or more full-time employees and manufacture or process listed chemicals above weight thresholds that vary by substance. The reports include on-site releases to air, water, and land, plus off-site transfers for disposal. The data are not real-time; they are annual summaries released the following autumn. They also do not capture illegal releases, accidents during transport, or small facilities below the employment threshold.
For a household, the most useful TRI fields are stack and fugitive air emissions within a given radius, because those are the releases most likely to reach a residential property without warning. Water discharges matter if your drinking water source is downstream. Land releases matter if you are eating vegetables grown in soil adjacent to the facility boundary.
Cancer clusters: what the public data can and cannot say
A cancer cluster is defined by the US National Cancer Institute as a greater-than-expected number of cancer cases among a group of people in a defined geographic area over a specific time period. Identifying a genuine cluster requires three things that are hard to satisfy together: a clear elevated rate relative to age-adjusted expected rates, a known exposure pathway, and a plausible biological mechanism linking the exposure to the cancer type observed. Most concerned-neighbour reports that reach a state health department satisfy none of the three, and that isn't a dismissal — it's how the maths of rare events behaves.
Why most "cluster" maps don't survive scrutiny
The statistical trap has a name: the Texas-sharpshooter fallacy, after the apocryphal marksman who fires at the side of a barn and then paints a target around the tightest group of holes. Applied to disease mapping, it works the same way — a handful of nearby diagnoses get noticed, someone draws a boundary around those households after the fact, and the boundary is sized precisely to make the count look alarming. Any scatter of random points, health data included, contains tighter and looser patches somewhere. Finding one proves nothing until the boundary and the comparison population are fixed before the data are looked at.
The second trap is small-number noise. Cancer incidence is a low-probability event per person per year, so in a neighbourhood of a few hundred people the expected count in any given year might be one or two cases. A single additional diagnosis can double the local rate and look catastrophic on a map, while remaining well inside the range random variation produces on its own. This is exactly why the CDC's National Program of Cancer Registries (NPCR) and its European counterparts publish incidence at county or regional level, age-adjusted, over multi-year windows — smaller geographies and shorter windows are too noisy to interpret reliably.
Multiple comparisons compound both problems. Test enough neighbourhoods, cancer types, or time windows and one combination will look statistically "elevated" by chance alone. This is why the US National Cancer Institute's cluster-investigation protocol requires a single cancer type, a population and period fixed in advance, and comparison against an external expected rate — not a self-selected local one.
Distance, dose, and why proximity is still a legitimate concern
None of this means proximity to an emitting facility is irrelevant — it means the burden of proof runs through exposure, not through a map coincidence. Airborne pollutant concentration falls off with distance from a point source in a broadly predictable way, governed by plume physics: the US EPA's AERMOD dispersion model formalises how stack height, wind speed and direction frequency, atmospheric stability, and terrain combine to set ground-level concentration at any given distance. The general shape is steep near the source and flattens further out — but it rarely reaches zero at typical residential setbacks, and because prevailing wind carries the plume preferentially in one direction, the real footprint is closer to an elongated fan than a neat circle.
A legitimate exposure concern needs three things together, not one: a plausible pathway (air, groundwater, soil, or vapour intrusion that actually reaches the property), a measurable or estimable dose (TRI/E-PRTR release volume and the compound's toxicity profile), and a cancer type with a biologically consistent latency and mechanism (checked against IARC Monographs or the EPA's Integrated Risk Information System, IRIS). Distance alone is a screening heuristic, not a verdict.
An apparent cluster
- Starts from a coloured map or a handful of neighbours' diagnoses
- Boundary drawn after the cases were already known
- Small local population compared to itself, not to an age-adjusted expected rate
- Different cancer types pooled together for a bigger, scarier number
A confirmed cluster
- Starts from a registry-wide, age-adjusted incidence rate for one cancer type
- Boundary and time window fixed before the data were examined
- Rate is statistically elevated against the external expected rate
- A plausible exposure pathway and biologically consistent latency exist
A red circle on a map is not epidemiology — it's a hypothesis, and the state cancer registry is where you go to test it, not confirm it.
Booking-value risk adjustment: the practical framing
From a property-valuation perspective, proximity to a NPL or TRI site introduces a stigma discount that often outlasts the actual environmental risk. The hedonic-pricing literature on homes near hazardous sites is real but modest and contested: meta-analyses across dozens of studies put the typical value effect in the low single digits per mile of distance, and some of the largest studies (Greenstone & Gallagher, 2008) found no statistically significant price effect from NPL designation at all from the NPL. The discount tracks media coverage more closely than it tracks the underlying remedy status.
This means that even if the health risk turns out to be negligible once pathway and dose are checked, the financial risk of buying near a listed site is real. The CleanZone approach is to flag the proximity and let the buyer price the uncertainty, rather than to assert a health outcome that public data cannot support at 25 km² resolution.
The threshold table
We score a cell on four industrial-health metrics:
| Measure | Reference point | Reading it |
|---|---|---|
| Industrial proximity | Straight-line distance in km to nearest NPL site, TRI-reporting facility, or E-PRTR facility | < 1 km to active NPL → high stigma and possible exposure risk; 1–3 km → moderate, pathway-dependent; > 5 km → generally below most hedonic-discount thresholds |
| Toxic release | TRI or E-PRTR annual air + water + land release (kg/year) within 5 km radius | > 10,000 kg/year air releases within 5 km → elevated exposure potential for respiratory and sensitisation endpoints; cross-check against dominant chemical toxicity profiles |
| Superfund site | EPA NPL status: proposed / listed / deleted; remedy status from ROD | Listed with no remedy selected → highest uncertainty; remedy in construction → discount persists but risk is managed; deleted with long-term monitoring → residual stigma only |
| Cancer incidence | CDC NPCR state registry age-adjusted incidence rate per 100,000 at county/region level | County rate > 1.2× expected national rate for a specific, plausibly linked cancer, sustained over a multi-year window → warrants deeper investigation; a single-year spike in one neighbourhood is usually small-number noise, not a cluster |
How to use the data in a transaction
Most disclosure regimes do not require a seller to volunteer proximity to a NPL site unless there is confirmed contamination on the property itself. A buyer who runs the EPA EnviroMapper search for the address and a 5 km buffer is doing more due diligence than the average transaction. The next step is to read the ROD abstract for each nearby NPL site and the TRI form R for each nearby facility. Both documents are public and free.
If the nearest NPL site is 2.5 km away and its groundwater plume has been delineated to 1.8 km, the risk to your specific property is likely low unless your water supply is drawn from the affected aquifer. If the nearest TRI facility is 800 metres away and reports persistent benzene fugitive emissions, the exposure risk is not zero, and the hedonic recovery is uncertain. Both facts belong in the decision, not as headlines, but as weights.
- Run EPA EnviroMapper or EEA E-PRTR for the target address with a 5 km buffer
- Read the ROD abstract for every NPL site within that radius
- Pull TRI Form R or E-PRTR release data for all facilities within 5 km for the last 3 years
- Check state cancer registry age-adjusted rates for the county by cancer type, over a multi-year window
- Cross-check dominant released chemicals against known carcinogenicity (IARC Monographs, EPA IRIS)
- Ask whether the state health department has an open formal cluster investigation — not just a neighbourhood rumour
- Factor stigma discount into valuation even when health risk appears managed
The metrics above — haz_industrial, toxic release, superfund site, and cancer incidence — are part of the CleanZone cell card. We report public listing status and registry rates. We do not calculate individual lifetime cancer risk. That requires site-specific exposure modelling that is beyond the resolution of open data.