A curated whole-genome reference panel for shotgun-metagenomic aquatic invasive species surveillance in the Laurentian Great Lakes & Lake Winnipeg basin

A pilot proposal to bridge the gap between published reference genomes and operational AIS detection.
Principal investigator: [PI NAME], [INSTITUTION]
Co-investigators: [CO-I NAMES]
Submission target: [FUNDER / PROGRAM]
Duration: [X] months
Total requested: [$X]
Date: 2026-05-29

Summary

Aquatic invasive species (AIS) cause billions of dollars in damage annually to North American freshwater ecosystems and the economies that depend on them. Existing DNA-based surveillance relies almost exclusively on PCR-amplified marker assays (qPCR or amplicon metabarcoding), which constrain detection to a small predefined target list and depend on primer compatibility. Shotgun metagenomic sequencing — sequencing all DNA in a water sample without amplification — would let a single library serve any reference-genome–bearing taxon, but is not currently used operationally for AIS detection. The blocker is not the sequencing chemistry; the blocker is that no curated panel of whole-genome reference assemblies exists for the species AIS managers actually need to detect.

We propose to construct, validate, and publicly release this missing panel. We will (1) systematically pair the NOAA GLANSIS watchlist and species inventory against the NCBI Datasets API to enumerate which AIS have published whole-genome assemblies; (2) integrate those assemblies into an existing real-time nanopore pipeline (danaSeq, already deployed at microscape.app) as a versioned, reproducible reference set; (3) demonstrate operational detection by reprocessing [N] existing nanopore metagenomic libraries from [STUDY REGION]; and (4) publish the panel, the inventory methodology, and a recommendations document for the AIS taxa whose reference-genome gaps most urgently need closing.

1. Background

1.1 Current state of DNA-based AIS surveillance

DNA-based monitoring of aquatic invasive species has matured rapidly over the past decade. Resource managers now routinely deploy:

What is conspicuously missing from this stack is a curated, versioned, genome-level reference resource matched to the species watchlists.

1.2 The case for shotgun metagenomics

Shotgun sequencing of environmental DNA samples — sequencing total DNA without PCR — has three advantages over amplicon-based approaches for AIS detection:

  1. No primer bias. Detection sensitivity is determined by sequencing depth and reference availability, not by primer fit. A single library can be re-queried for new species without re-sequencing as the reference panel grows.
  2. Per-position alignment evidence. Each read carries CIGAR, mapq, and identity information. Hits that cluster on one repetitive element are visually distinguishable from hits spread uniformly across a target genome — a discrimination amplicon counts cannot provide.
  3. Single library, many taxa. Fish, mussels, crustaceans, plants, pathogens, and the background microbiome are all sequenced together. Cost per detected taxon falls as the reference panel grows.

The recently published call by McCartney et al. (Frontiers in Environmental Science, 2023) — "Time to invest in the worst" — argues for closing the reference-genome gap specifically to enable this kind of shotgun-based monitoring. The authors report that approximately 55% of the IUCN "100 worst invasive species" lack a reference genome. The gap is the reason a panel-based shotgun pipeline does not exist; closing it is a tractable next step.

1.3 What is missing today

  1. A reproducible mapping between the species AIS managers care about (GLANSIS watchlists, state lists) and the genomic resources available for each (chromosome-scale assembly, scaffold-level draft, marker-only, nothing).
  2. A curated, versioned reference panel deployable in a standard bioinformatics pipeline.
  3. An operational pipeline that consumes that panel, runs at sample turnaround speeds compatible with field decision-making, and surfaces per-sample evidence (identity histograms, genome-position distributions) usable by non-bioinformaticians.
  4. A documented audit of which AIS most urgently need a reference assembly — feeding back into sequencing prioritisation by groups like the Earth BioGenome Project and Darwin Tree of Life.

2. Approach

Aim 1. Inventory the GLANSIS / regional watchlist species against NCBI to classify each as chromosome-scale WGS, fragmented draft, marker-pool, or no-data.

The GLANSIS database exposes its species list via the USGS IPT Darwin Core archive and via a web-based list generator. We will pull the canonical species list, normalise taxonomy (resolving synonyms via NCBI Taxonomy), and query the NCBI Datasets API per taxon. Each species is then assigned a tier:

TierDefinitionAction for the panel
AChromosome-scale RefSeq or GenBank reference assemblyInclude directly.
BScaffold- or contig-level draft assemblyInclude with sensitivity caveats in meta.json.
CMarker / mitochondrion-only GenBank records (COI, 18S, ITS, mtDNA)Build a multifasta marker pool; flag as "marker reference" in the SPA.
DNo usable public sequenceFlag in the gaps report (see Aim 4). Not in the panel.

The inventory will be published as an open TSV alongside the panel itself, re-runnable at any time as new assemblies are published.

Aim 2. Build, version, and publicly release the panel as a directory of references conforming to a documented schema.

Each reference subdirectory will contain a minimap2 ONT-preset index, a contig offset map (for downstream genome-position visualisation), the source FASTA, and a JSON meta.json file with taxonomic, accession, and provenance information. The directory becomes a single artifact that any downstream pipeline can ingest. We will version the panel by GLANSIS snapshot date and by panel-build date, allowing reproducible re-runs against historical states.

Aim 3. Demonstrate operational detection at scale using the existing danaSeq nanopore pipeline.

The nanopore_live pipeline already implements a generic --mapping_refs module (github.com/rec3141/danaSeq, commit c5c74e7, May 2026), which consumes any compatible reference directory and writes per-barcode alignment results into a DuckDB mapping table. A dashboard (currently deployed at microscape.app/complete/) surfaces per-sample hit counts, identity histograms, and genome-position distributions with a live identity-threshold slider.

We will reprocess [N] existing nanopore metagenomic libraries from [STUDY REGION] against the full panel and characterise:

Aim 4. Publish a "reference-genome gaps" recommendations document for AIS sequencing prioritisation.

The Tier-D species from Aim 1 — the AIS taxa for which no usable reference exists — are not a panel problem but a policy problem. We will compile a recommendations document ranking these species by management priority (GLANSIS impact scores, regional jurisdiction), invasion-front proximity, and phylogenetic isolation (where related-taxon references are unlikely to substitute). The document is intended for distribution to the Earth BioGenome Project, Darwin Tree of Life, and equivalent regional efforts (Canadian BioGenome Project).

3. Preliminary results

A pilot implementation of Aims 2 and 3 has been completed under unfunded development. We have:

Initial findings confirm the literature: sister-taxon cross-mapping is the dominant false-positive source (zebra and quagga share substantial signal below 95% identity), repetitive-element artifacts are visually identifiable on genome-position histograms, and a single identity threshold across all species under-performs species-specific calibration. These observations are the empirical basis for Aim 3's cutoff-calibration work.

4. Deliverables & timeline

MonthDeliverable
1–2Inventory pipeline (GLANSIS → NCBI Datasets → tiered classification). Open TSV + reproducible script.
3–6Tier-A/B panel build for GLANSIS species with assemblies. Public release on Zenodo and GitHub.
5–9Reprocessing of [N] existing libraries; per-species detection benchmarking; cutoff calibration.
9–11Tier-D gaps report. Manuscript drafting.
12Manuscript submission; final panel release.

5. Broader impacts

6. Budget summary

[Budget narrative — personnel (X months bioinformatician, Y months data manager, Z student months), sequencing reagents for validation runs, cloud compute / storage, publication open-access fees, travel for one stakeholder workshop.]

References

  1. McCartney M. et al. (2023). Time to invest in the worst: a call for full genome sequencing of the 100 worst invasive species. Front Environ Sci. link
  2. Rossi G. et al. (2022). Genomic data is missing for many highly invasive species, restricting our preparedness for escalating incursion rates. Mol Ecol Resour. link
  3. Buchner D. et al. (2022). Shotgun metagenomics of soil invertebrate communities reflects taxonomy, biomass, and reference genome properties. Nat Comms. link
  4. Urban L. et al. (2021). Freshwater monitoring by nanopore sequencing. eLife. link
  5. Spens J. et al. (2020). Speeding up the detection of invasive aquatic species using environmental DNA and nanopore sequencing. bioRxiv. link
  6. NOAA GLERL. Great Lakes Aquatic Nonindigenous Species Information System (GLANSIS). glerl.noaa.gov/glansis
  7. IUCN/ISSG. 100 of the World's Worst Invasive Alien Species. iucngisd.org/gisd/100_worst.php