Back to Research Notes and Tutorials
Research Agenda Updated

Four research questions at the bench–code interface

A research agenda for earlier detection, transmission inference, sustainable genomic surveillance, and understanding infection patterns.

Earlier Detection Phylodynamics Implementation Science Pathogen Ecology Population Immunity

Why these questions belong together

Earlier detection is not only a laboratory problem. A sensitive assay or metagenomic signal still requires appropriate sampling, quality controls, reference databases, confirmation rules, and cautious reporting.

Transmission inference is not only a phylogenetic problem. A tree is shaped by which specimens were collected, which genomes passed quality thresholds, which contextual sequences were included, and how dates and locations were recorded.

Implementation is not only a software problem. A reproducible workflow must also fit the laboratory, computing environment, staff capacity, data-governance requirements, and decisions that the surveillance programme needs to make.

Understanding infection patterns also requires looking beyond the genomes. Population immunity, exposure and ecological conditions may contribute to what is observed, while changes in testing and sampling can alter the picture. These explanations need to be considered together.

I am interested in research that studies these connections rather than treating the bench, analysis, and public-health use as separate stages.

Question 1: How can genomic surveillance detect viral threats earlier?

The central challenge is to recognise a credible signal before the outbreak is obvious, without lowering the evidential standard so far that noise becomes a finding.

My work on the 2023 CV-A24v conjunctivitis outbreak showed why this is a complete-chain problem. When a suitable targeted test was not readily available, specimen processing, metagenomic sequencing, reference-database design, confirmatory analysis, and phylogenetic interpretation all contributed to the final conclusion.

The same principle applies beyond one virus. Clinical specimens, environmental samples, and wastewater each contain different mixtures of target material, host background, inhibitors, and biological complexity. Targeted sequencing may provide depth and efficiency when the agent is known. Metagenomic approaches may broaden discovery when it is not. Direct sequencing may shorten parts of the workflow but can change the balance between sensitivity, speed, and interpretability.

Testable directions

  1. Compare metagenomic, targeted-amplicon, and direct-sequencing strategies under defined input and background conditions.
  2. Develop a candidate-to-confirmation evidence framework combining controls, breadth, depth, mapping specificity, close-relative competition, replication, and orthogonal confirmation.
  3. Quantify how specimen type, viral input, storage history, and library strategy affect the practical detection boundary.
  4. Use public-safe synthetic datasets and benchmark mixtures to make performance and failure modes comparable across laboratories.
  5. Design reporting categories that distinguish a candidate signal, supported detection, confirmed finding, and unresolved result.

The aim would not be one universal threshold. It would be a transparent framework showing what a workflow can support under defined conditions.

Question 2: What can viral genomes reveal about transmission and evolution?

Genomes can help identify introductions, lineage replacement, local persistence, reassortment, and changing viral diversity. They cannot, on their own, prove a direct transmission event or fully represent an outbreak that was sampled unevenly.

My experience with SARS-CoV-2, dengue, CCHFV, RSV, influenza, and poliovirus has reinforced the need to connect sequence quality with evolutionary interpretation. Low coverage can weaken a lineage assignment. Reference choice can change apparent relatedness. A segmented virus requires separate histories for each segment. A public phylogeny needs sampling and metadata context beside it.

This is why I am interested not only in generating trees, but in measuring how robust an interpretation remains when sampling, metadata, contextual sequences, or quality thresholds change.

Testable directions

  1. Evaluate how geographic and temporal sampling imbalance affects inference about introductions and sustained transmission.
  2. Compare phylogenetic and phylodynamic conclusions across alternative context-selection and quality-control strategies.
  3. Integrate epidemiological metadata with genomic evidence while keeping uncertainty and missingness visible.
  4. Develop segment-aware approaches for identifying reassortment and comparing evolutionary histories across viral genome segments.
  5. Examine how longitudinal surveillance across multiple seasons changes conclusions about viral population structure and persistence.

The intended output would be more than a tree. It would be an interpretation package that states the question, sampling frame, quality boundaries, sensitivity analyses, and conclusions that the available evidence can genuinely support.

Question 3: How can advanced genomic methods become routine public-health practice?

A method has limited public-health value if it works only for the person who developed it, on one computer, for one paper.

Routine implementation requires controlled inputs, versioned configurations, acceptance criteria, logs, interpretable outputs, escalation rules, and training. It also requires a realistic understanding of compute, reagent, staffing, internet, and data-sharing constraints.

My public workflows, SOP and runbook work, Pakistan-focused SARS-CoV-2 Nextstrain Community Build, and practical training experience have made this implementation question increasingly important to me.

Testable directions

  1. Define a minimum reproducibility standard for public-health genomics workflows using public-safe test data, expected outputs, and automated checks.
  2. Measure whether structured runbooks and decision categories improve agreement between analysts reviewing the same result.
  3. Compare implementation models for laboratories with different compute, staffing, and data-governance constraints.
  4. Develop traceable sample-to-report records that preserve both automated results and human review decisions.
  5. Evaluate training approaches using practical outcomes such as successful independent runs, error recognition, documentation quality, and reproducibility.

The goal would be a surveillance method that can be inspected, taught, maintained, and improved locally—not a black box that creates long-term dependence.

Question 4: How do pathogen evolution, immunity and ecology shape infection patterns?

Surveillance can show that infection patterns have changed, but describing a change is different from explaining it. I am interested in why patterns differ between populations and over time, and how much those differences reflect pathogen evolution, population immunity, exposure or ecological conditions.

My experience with dengue and respiratory-virus surveillance provides a starting point. For example, a change in the dengue serotypes represented in referred samples raises questions about immunity and exposure, but also about who was tested and where samples came from. I would like to connect genomic and epidemiological evidence with demographic and immunological data, where available, to examine these competing explanations.

Testable directions

  1. Examine how observed infection patterns vary across seasons, locations and population groups using appropriately governed observational data.
  2. Investigate the contributions of population immunity, exposure and ecological conditions to those patterns without assuming that an association establishes causation.
  3. Assess how referral patterns, testing availability and changes in sampling influence apparent trends.
  4. Use statistical and mathematical modelling to compare plausible explanations, make uncertainty explicit and identify where the available evidence cannot distinguish them.
  5. Identify what additional epidemiological or immunological information would be most useful for interpreting the observations and informing public-health planning.

This is a direction in which I want to deepen my modelling training, building on my laboratory and bioinformatics experience. Pakistan offers practical insights that can inform questions relevant across settings, rather than defining a geographic boundary. Any use of local data should involve appropriate permissions, equitable collaboration and benefits for the communities and teams contributing to the research.

A possible doctoral programme

These four themes provide directions for a focused doctoral project, not a requirement to pursue four separate projects. The central question and balance of methods would be refined with a supervisor around the available evidence, expertise and data permissions.

Phase 1 — Establish the evidence framework

Build public-safe benchmark datasets and define laboratory, sequencing, and analytical quality dimensions. Compare candidate-to-confirmation rules and identify the failure modes most likely to produce false confidence.

Phase 2 — Test inference under realistic surveillance conditions

Apply the framework to selected clinical and environmental viral-genomics use cases. Evaluate how input quality, sampling design, reference choice, and context selection change detection and evolutionary interpretation.

For a project centred on Question 4, the emphasis would be on defining the population and observation process, then combining genomic and epidemiological evidence with modelling to compare explanations for changing infection patterns. The study would make sampling limitations and uncertainty explicit.

Phase 3 — Translate the method into routine practice

Package the strongest approach as a versioned workflow with runbooks, traceable outputs, training material, and an implementation evaluation across users or laboratory settings.

Potential outputs would include a benchmark dataset, a methods paper, a pathogen-specific application study, a reproducible software release, and an implementation framework.

What I would bring to this work

  • More than ten years across clinical and public-health genomics roles
  • Experience with molecular assays, library preparation, Illumina, Ion Torrent, and Oxford Nanopore sequencing
  • Bioinformatics experience spanning quality control, mapping, assembly, consensus and variant analysis, metagenomics, phylogenetics, phylodynamics, and public-database submission
  • First-author outbreak-genomics research, with additional contributions across multi-pathogen surveillance studies
  • Public, versioned workflows and a maintained Pakistan-focused Nextstrain Community Build
  • Experience writing SOPs, runbooks, review criteria, technical reports, and training material
  • Practical training and scientific collaboration across local and international settings

I would be particularly interested in doctoral environments connecting pathogen ecology and evolution, genomic epidemiology, population immunity, quantitative modelling and public-health implementation. I would bring practical laboratory and bioinformatics experience while seeking deeper training in statistical and mathematical approaches.

Closing perspective

The research direction I want to pursue is not limited to one virus or platform. It is centred on a recurring public-health challenge: turning uncertain genomic and epidemiological evidence into a clearer understanding of infection patterns and conclusions that are scientifically defensible, useful in practice, and reproducible by others.

That requires the bench, the code, the epidemiological context, and the implementation setting to remain part of the same research question.

If this agenda overlaps with your doctoral programme, laboratory, or research team, I would welcome a conversation about where my experience could contribute and develop further.

Contact me about research and doctoral opportunities →