NWGS Grid™

Methods & Data Sources

Transparency about our data sources, processing methods, and analytical approaches.

System Architecture

Data Sources

NWSS
GISAID
PDB
CDC
Census

Ingestion

ETL Pipelines
Validation
Normalization

Analysis

Variant Calling
Burden Calc
Modeling

Outputs

Dashboards
API
Reports
Wastewater Data & Normalization

NWGSGrid ingests wastewater viral concentration data from state and local partners and national programs, including the CDC National Wastewater Surveillance System (NWSS). Viral concentrations are normalized and converted into wastewater viral activity levels (WVAL) that can be compared across time and space, similar to the categories used in existing federal dashboards.

Normalization accounts for variations in flow rate, population served, and laboratory methods to enable meaningful comparisons between sites and over time.

Genomics & GISAID

Genomic analyses leverage sequences and metadata shared via GISAID and other repositories. Variant calling and lineage assignment follow established pipelines (Pangolin, Nextclade) to classify sequences into WHO-designated lineages.

In line with GISAID's access agreements and sharing policy, NWGSGrid presents derived results only—such as variant prevalence, spike mutation frequencies, and EPI_SET identifiers—without redistributing underlying raw sequences. This approach ensures compliance with data-sharing principles while providing actionable genomic intelligence.

Protein Structures

Protein structure visualizations are built on experimentally determined and computed structures from the Protein Data Bank (PDB). Interactive exploration uses modern web-based viewers such as Mol* for visualization of spike and other viral proteins.

When mutations are highlighted on structures, their positions are mapped to reference structure coordinates. Note that some structures may not include all protein regions; flexible regions or those not resolved in experiments may not be visible.

Clinical & Demographic Data

Clinical burden and demographic context are derived from public health surveillance systems and population datasets, including:

• State health department reports and hospitalization data • CDC respiratory illness surveillance systems • U.S. Census Bureau and American Community Survey (ACS) population estimates • HHS hospitalization and healthcare capacity data

Data are aggregated to appropriate geographic and temporal levels to protect privacy while enabling meaningful analysis of burden and equity across demographic groups.

Data Access & Privacy

NWGSGrid is committed to responsible data use and privacy protection:

  • GISAID Compliance: We present derived metrics only (variant prevalence, mutation frequencies) without redistributing raw sequence data, in accordance with GISAID's terms of use.
  • Privacy Protection: Clinical and demographic data are aggregated to appropriate geographic and temporal levels to prevent individual identification.
  • Data Attribution: All data sources are properly attributed. Users should cite original data providers when using NWGSGrid outputs in publications.
  • API Access: Programmatic access to NWGSGrid data is available through our API. See the Reports & Downloads page for documentation.

How to Cite NWGSGrid

National Wastewater Genomics Smart Grid (NWGSGrid). [Year]. Available at: https://nwgsgrid.org. Data sources: CDC NWSS, GISAID, RCSB PDB, state health departments.
base44
Edit with Base44