WaterQ

Data Sources

WaterQ is built on publicly available federal data. Here is a detailed overview of our data sources and how we use them.

EPA Safe Drinking Water Information System (SDWIS)

Primary Data Source

SDWIS is the EPA's database of public water system information and is our primary data source. It holds records for more than 150,000 public water systems serving roughly 300 million Americans. WaterQ currently carries 7,929 of those systems — the ones serving the most populous 3,000 US cities — obtained from the bulk SDWA files EPA publishes through ECHO.

Data We Use

  • Water System Inventory: System names, IDs (PWSID), locations, service populations, water source types, and system classifications
  • Contaminant Test Results: Measured concentrations, test dates, Maximum Contaminant Levels (MCLs), and violation flags
  • Violation Records: Violation types, severity classifications, dates, resolution status, and enforcement actions
Format: REST API / CSV Update: Quarterly Coverage: All US states & territories

Open-Meteo

Supplementary Source

City pages show current precipitation for the area, fetched live from the Open-Meteo forecast API. Heavy rainfall increases surface runoff, which is one of the routes contaminants reach a water supply.

This is contextual weather data only. It is not a water quality measurement, and it does not feed into any system's score or grade.

What this data does not cover

Being straight about the gaps matters more than looking comprehensive. These are the limits of what WaterQ can tell you:

  • Private wells are not covered at all. SDWIS regulates public water systems. An estimated 13 million US households draw from private wells, which no federal agency tests or regulates — testing is entirely the owner's responsibility.
  • Sampling data covers lead and copper only. The Lead and Copper Rule sample file reports 90th-percentile results for those two metals. Other contaminants appear through violation records rather than measured concentrations.
  • We only publish a score when records back it. Cities without system-level records in our dataset show no score and say so, rather than displaying a number with nothing behind it.
  • System-level, not tap-level. Every figure describes water as delivered to the property line. Lead and copper are usually picked up from a building's own plumbing after that point, so only a test at your tap can tell you what you are drinking.
  • Displayed samples are recent ones. We show results from the last ten years and drop records dated in the future, which the source data does contain.

Update Schedule

Data Type Frequency Notes
Water System Inventory Quarterly Jan, Apr, Jul, Oct
Contaminant Test Results Quarterly Aligned with EPA reporting cycles
Violations Quarterly Including resolution status updates
Scores & Grades After each data update Recalculated when new data is available

Data Processing

Raw data from our sources goes through several processing steps before being presented on WaterQ. You can also browse the resulting contaminant reference pages to see how these inputs become user-facing health context.

  1. 1
    Ingestion

    Raw data is fetched from federal APIs and loaded into our processing pipeline

  2. 2
    Validation

    Records are checked for completeness, format consistency, and data quality

  3. 3
    Normalization

    Data from multiple sources is mapped to a unified schema with consistent units and identifiers

  4. 4
    Scoring

    Water quality scores and grades are calculated using our scoring methodology

  5. 5
    Aggregation

    System-level data is aggregated to city, county, and state levels using population-weighted averages

Data Accuracy

While we take every effort to ensure accuracy, WaterQ relies on data reported by water systems to federal agencies. Reporting delays, data entry errors, and testing gaps may affect the completeness of the information presented. If you believe any data is incorrect, please contact us so we can investigate.