Aquatic invasive species (AIS) cause billions of dollars in damage annually to North American freshwater ecosystems and the economies that depend on them. Existing DNA-based surveillance relies almost exclusively on PCR-amplified marker assays (qPCR or amplicon metabarcoding), which constrain detection to a small predefined target list and depend on primer compatibility. Shotgun metagenomic sequencing — sequencing all DNA in a water sample without amplification — would let a single library serve any reference-genome–bearing taxon, but is not currently used operationally for AIS detection. The blocker is not the sequencing chemistry; the blocker is that no curated panel of whole-genome reference assemblies exists for the species AIS managers actually need to detect.
We propose to construct, validate, and publicly release this missing panel. We will (1) systematically pair the NOAA GLANSIS watchlist and species inventory against the NCBI Datasets API to enumerate which AIS have published whole-genome assemblies; (2) integrate those assemblies into an existing real-time nanopore pipeline (danaSeq, already deployed at microscape.app) as a versioned, reproducible reference set; (3) demonstrate operational detection by reprocessing [N] existing nanopore metagenomic libraries from [STUDY REGION]; and (4) publish the panel, the inventory methodology, and a recommendations document for the AIS taxa whose reference-genome gaps most urgently need closing.
DNA-based monitoring of aquatic invasive species has matured rapidly over the past decade. Resource managers now routinely deploy:
What is conspicuously missing from this stack is a curated, versioned, genome-level reference resource matched to the species watchlists.
Shotgun sequencing of environmental DNA samples — sequencing total DNA without PCR — has three advantages over amplicon-based approaches for AIS detection:
The recently published call by McCartney et al. (Frontiers in Environmental Science, 2023) — "Time to invest in the worst" — argues for closing the reference-genome gap specifically to enable this kind of shotgun-based monitoring. The authors report that approximately 55% of the IUCN "100 worst invasive species" lack a reference genome. The gap is the reason a panel-based shotgun pipeline does not exist; closing it is a tractable next step.
The GLANSIS database exposes its species list via the USGS IPT Darwin Core archive and via a web-based list generator. We will pull the canonical species list, normalise taxonomy (resolving synonyms via NCBI Taxonomy), and query the NCBI Datasets API per taxon. Each species is then assigned a tier:
| Tier | Definition | Action for the panel |
|---|---|---|
| A | Chromosome-scale RefSeq or GenBank reference assembly | Include directly. |
| B | Scaffold- or contig-level draft assembly | Include with sensitivity caveats in meta.json. |
| C | Marker / mitochondrion-only GenBank records (COI, 18S, ITS, mtDNA) | Build a multifasta marker pool; flag as "marker reference" in the SPA. |
| D | No usable public sequence | Flag in the gaps report (see Aim 4). Not in the panel. |
The inventory will be published as an open TSV alongside the panel itself, re-runnable at any time as new assemblies are published.
Each reference subdirectory will contain a minimap2 ONT-preset index, a contig offset map (for downstream genome-position visualisation), the source FASTA, and a JSON meta.json file with taxonomic, accession, and provenance information. The directory becomes a single artifact that any downstream pipeline can ingest. We will version the panel by GLANSIS snapshot date and by panel-build date, allowing reproducible re-runs against historical states.
The nanopore_live pipeline already implements a generic --mapping_refs module (github.com/rec3141/danaSeq, commit c5c74e7, May 2026), which consumes any compatible reference directory and writes per-barcode alignment results into a DuckDB mapping table. A dashboard (currently deployed at microscape.app/complete/) surfaces per-sample hit counts, identity histograms, and genome-position distributions with a live identity-threshold slider.
We will reprocess [N] existing nanopore metagenomic libraries from [STUDY REGION] against the full panel and characterise:
The Tier-D species from Aim 1 — the AIS taxa for which no usable reference exists — are not a panel problem but a policy problem. We will compile a recommendations document ranking these species by management priority (GLANSIS impact scores, regional jurisdiction), invasion-front proximity, and phylogenetic isolation (where related-taxon references are unlikely to substitute). The document is intended for distribution to the Earth BioGenome Project, Darwin Tree of Life, and equivalent regional efforts (Canadian BioGenome Project).
A pilot implementation of Aims 2 and 3 has been completed under unfunded development. We have:
--mapping_refs module in nanopore_live (committed and pushed to github.com/rec3141/danaSeq).Initial findings confirm the literature: sister-taxon cross-mapping is the dominant false-positive source (zebra and quagga share substantial signal below 95% identity), repetitive-element artifacts are visually identifiable on genome-position histograms, and a single identity threshold across all species under-performs species-specific calibration. These observations are the empirical basis for Aim 3's cutoff-calibration work.
| Month | Deliverable |
|---|---|
| 1–2 | Inventory pipeline (GLANSIS → NCBI Datasets → tiered classification). Open TSV + reproducible script. |
| 3–6 | Tier-A/B panel build for GLANSIS species with assemblies. Public release on Zenodo and GitHub. |
| 5–9 | Reprocessing of [N] existing libraries; per-species detection benchmarking; cutoff calibration. |
| 9–11 | Tier-D gaps report. Manuscript drafting. |
| 12 | Manuscript submission; final panel release. |
[Budget narrative — personnel (X months bioinformatician, Y months data manager, Z student months), sequencing reagents for validation runs, cloud compute / storage, publication open-access fees, travel for one stakeholder workshop.]