Whole-capsid poliovirus NGS: QC, consensus, antigenic sites, and phylogeny
A workflow-oriented guide to reviewing whole-capsid poliovirus NGS outputs from read QC through antigenic-site summaries and phylogenetic placement.
Whole-capsid analysis is a chain of evidence
A whole-capsid poliovirus NGS workflow connects several different kinds of evidence: read quality, primer or adapter handling, reference mapping, depth masking, consensus sequence quality, amino acid changes, and tree placement. Weakness in one part of the chain should be visible in the report.
I prefer to review the workflow in the same order the evidence is generated. That makes it easier to explain why a sample passed, why a consensus is partial, or why a tree placement should be treated cautiously.
Define the comparison before mapping
Whole-capsid analysis should begin with a declared reference accession, expected capsid coordinates, primer scheme when applicable, and a consistent amino-acid numbering convention for VP1, VP2, and VP3. Those choices determine how variants, masked sites, and antigenic-site changes will be reported.
Reference choice should be checked against the sample's serotype and surveillance question. When mapping is weak or the best reference is uncertain, an assembly or broader reference screen can be used as a diagnostic step rather than forcing all reads onto one sequence.
QC and consensus review
The workflow in my public repository uses fastp for read QC, BWA-MEM for mapping, optional primer clipping, duplicate marking, depth calculation, FreeBayes variant calling, and BCFtools consensus generation with a depth mask. The report should preserve the configuration used for each of those steps rather than showing only the final FASTA.
- Check paired-read counts, quality trimming, duplicate handling, and mapping rate.
- Review depth across the capsid rather than relying only on mean coverage.
- Mask low-depth sites and report ambiguous regions clearly.
- Compare consensus names, sample metadata, and tree labels before publication.
- Keep private reads, BAM files, and sample-identifying metadata outside public repositories.
Interpret antigenic-site tables conservatively
Antigenic-site tables translate nucleotide variation into amino-acid changes at defined capsid positions. They are useful for screening and comparison, but a substitution in a named site is not by itself proof of antigenic escape or altered neutralization.
I would report the reference residue, sample residue, protein and coordinate, underlying coverage, and whether the call intersects a masked or ambiguous region. Interpretation should be tied to published poliovirus antigenic-site definitions and, when biologically important, laboratory evidence.
Phylogeny adds context, not direct transmission proof
Capsid sequences and contextual references should cover the same homologous region, pass comparable QC, and use transparent inclusion rules. I check alignment quality, sample labels, collection dates, model selection, and branch support before interpreting clusters.
Tree proximity can support questions about genetic relatedness and lineage context, but it does not establish a direct transmission event. Sampling gaps and uneven geographic coverage should travel with the interpretation.
A decision-ready output bundle
- Run and sample QC table with explicit pass, review, limited-use, or fail status.
- Depth profile and masking summary for the complete capsid region.
- Versioned consensus FASTA with traceable sample names and reference accession.
- Amino-acid and antigenic-site table with conservative interpretation.
- Alignment, tree, contextual-sequence manifest, and a short statement of limitations.
References and Further Reading
- Polio whole-capsid NGS analysis repository
- IQ-TREE 2 phylogenetic inference
- BCFtools consensus documentation
- High-throughput next-generation sequencing of polioviruses
- The role of genetic sequencing and analysis in polio eradication
- Evolution of poliovirus neutralizing antigenic sites
- Conserved antigenic structure of wild poliovirus type 1 strains in Pakistan