Back to Research Notes and Tutorials
Tutorials

Making Nextstrain useful for routine outbreak communication

How to turn phylogenetic builds into focused communication products with disciplined metadata, clear questions, and careful interpretation.

Nextstrain Phylogenetics Dashboards

The tree is not the message

Phylogenetic dashboards are powerful because they make pathogen evolution visible. They can show clustering, lineage movement, sampling gaps, and time structure. But a tree is only useful when it answers a question. Without a question, it becomes a complex figure that invites overinterpretation.

For routine outbreak communication, I start by writing the questions before building the visualization: Which lineages are circulating? Are recent samples clustering with local or imported context? Is there evidence of sustained transmission? Which samples require epidemiological follow-up?

Metadata discipline

Nextstrain-style builds depend on metadata quality. Dates should be consistent, locations should use stable names, and restricted fields should be separated from public-display fields. Missing metadata is not just empty space; it affects interpretation of geographic and temporal patterns.

Subsampling also matters. A tree can look different depending on which contextual genomes are included. I prefer to document the contextual sampling rule so readers know whether the tree is intended for local comparison, national overview, or global placement.

  • Use consistent sample dates and location levels.
  • Keep private metadata out of public builds.
  • Document contextual sequence selection.
  • Write an interpretation note beside the tree.
  • Avoid implying transmission from tree proximity alone.

What a good dashboard should do

A good dashboard helps a reader decide what to ask next. It should connect phylogeny to sampling time, geography, lineage, and relevant mutations without overwhelming the user. It should also make uncertainty visible: low coverage, missing dates, sparse context, and uneven sampling can all shape the story.

The best public-health phylogenetics outputs are not the most visually crowded ones. They are the ones that help teams act carefully.

A dashboard should have a written question

Before building a Nextstrain or similar dashboard, I write the question in plain language. If the question is lineage surveillance, the metadata and coloring should support lineage interpretation. If the question is local clustering, the contextual dataset must be chosen carefully. If the question is importation, geographic labels and sampling density become central.

Without this discipline, a dashboard can look impressive while saying very little. The visualization should reduce ambiguity for the reader, not add new visual complexity.

Sampling context belongs beside the tree

A phylogeny displays sampled genomes, not every infection that occurred. Uneven sequencing across locations, dates, and patient groups can make a cluster look more isolated or more dominant than it is. Context-sequence selection can also change the apparent nearest neighbours.

I would publish a small sampling summary with the dashboard: counts by period and location, the rule used to select contextual sequences, the amount of missing metadata, and any important surveillance gaps. The interpretation note should state what the tree supports and which conclusions remain limited by sampling.

  • Show sample counts across time and geography.
  • Choose contextual sequences with a reproducible rule.
  • Keep private metadata separate from public display metadata.
  • Avoid implying direct transmission from tree proximity alone.
  • Keep branch support, QC exclusions, and missing-data caveats visible.

References and Further Reading

  1. Nextstrain: real-time tracking of pathogen evolution
  2. Nextstrain automates real-time phylodynamic analysis
  3. TreeTime: maximum-likelihood phylodynamic analysis
  4. IQ-TREE 2 phylogenetic inference
  5. Phylogenetic tree building in the genomic age