Skip to lesson content

FASTQ to Phylogeny · Part 5 · Free beginner exercise

Which FastQC warnings actually matter?

Flags are clues. Context determines the action.

Read the pattern behind a warning, ask whether it fits the experiment, and decide what deserves attention before your next analysis.

Join the Part 5 discussion on LinkedIn ↗
Three small datasets5,000 synthetic 100-base reads in eachInterpret before processingCompare quality loss and adapter signalsContinue from Part 4Use your existing FastQC installation

Everything for Part 5

Download. Compare. Interpret.

Get the eight-slide lesson and all three practice datasets in one ZIP. You can start this exercise without downloading the Part 4 files separately.

Practice ZIP: about 502 KB · Slides: eight pages, about 26 KB

Inside the ZIP: fastqc_good.fastq.gz, fastqc_tail_drop.fastq.gz, fastqc_adapter_rich.fastq.gz, and README.txt.

These are synthetic teaching files, not biological sequencing data. Extract the outer ZIP and keep the three .fastq.gz files compressed.

Set up the practice folder and run the exercise →

Step 01 · Before you fix anything

Ask three questions first.

In Part 4, we generated our first FastQC reports. Now we will decide what their warnings mean. The goal is to recognize patterns that could affect the analysis, rather than make every module green.

  1. 1
    Flag

    A module shows WARN or FAIL.

  2. 2
    Plot

    Inspect the actual pattern.

  3. 3
    Context

    Ask whether it fits this library.

  4. 4
    Action

    Investigate, process, continue, or document.

Is it expected?
Amplicon, targeted, RNA-seq, small-RNA, and other libraries can legitimately look unusual.
Could it be technical?
Adapters, primers, poor-quality tails, contamination, or run problems may explain a warning.
Will it affect my next step?
Consider whether the pattern could bias mapping, assembly, variant calling, or the analysis you actually plan to perform.
Inspect the plot first.

Never make a processing decision from the PASS, WARN, or FAIL label alone.

Step 02 · Per base sequence quality

How serious is the quality drop?

Lower quality toward the end of a read can reflect lower-confidence calls in later cycles or a subset of reads with poor-quality tails. Look at how large the drop is, where it begins, and how many reads it affects.

The mean line alone does not describe the full spread. Use the box and whiskers to judge variation across the reads, then consider the read length and quality needed for the next analysis. Trimming may help when poor-quality tails would interfere with that analysis.

The plot shows a symptom.

Sequencing chemistry, signal quality, library quality, and platform behaviour can all contribute. The pattern alone does not identify the exact wet-lab cause.

See the official Per Base Sequence Quality guide for the plot elements and default thresholds.

Step 03 · Adapter content and overrepresentation

Identify the sequence before removing it.

When a biological insert is shorter than the sequencing read, sequencing can extend into the adapter. Confirm which adapter or unwanted primer is present before choosing an appropriate trimming approach.

An overrepresented sequence needs the same careful identification. It may be a technical sequence, but it can also reflect real biology, such as an abundant amplicon or transcript.

These modules ask different questions.

Adapter Content searches for configured adapter sequences. Overrepresented sequences looks for individual sequences that recur unusually often. A result in one module does not guarantee a matching result in the other.

Step 04 · Context first

Some warnings can be expected.

Sequence duplication
High duplication may be expected in an amplicon or targeted library. In a diverse whole-genome library, extreme duplication may deserve closer investigation. A duplication flag alone cannot distinguish technical copies from biological repetition.
GC distribution
GC content depends on the organism and library. An unexpected extra peak could reflect a mixed population, contamination, or a biased subset of sequences. Compare the observation with what the experiment should contain.
Per-base sequence content
Early-position bias can arise from primers, transposase activity, or library-preparation chemistry. It is not automatically evidence of a failed run.

FastQC is a general screening tool. Specialized libraries do not always resemble random whole-genome libraries, so the same flag can mean different things in different experiments.

Step 05 · Action guide

Match the pattern to the next step.

A decision starts with the plot and experimental context.
PatternWhat to consider next
Low-quality tailAssess its severity and extent. Trim or filter when the poor-quality sequence could affect the planned analysis.
Adapter signalConfirm adapter or primer identity, then trim the unwanted sequence with a suitable tool.
Overrepresented sequenceIdentify it before deciding whether it is technical or biological.
Duplication, GC, or base biasCompare it with the library design and expected biology. Investigate unexpected patterns.

On a small screen, scroll the table sideways.

Record the reason for your decision.

Note what you observed, what you expected, and why you trimmed, filtered, investigated, or chose to continue. Reproducibility includes the reasoning behind preprocessing.

Step 06 · Hands-on comparison

Compare three FastQC reports.

Use the FastQC installation from Part 4. All commands below belong in Ubuntu. Use a separate part5 folder to keep the earlier lesson files and reports intact.

Download the Part 5 practice ZIP and choose Extract All in Windows File Explorer. Then open your new practice folder:

Ubuntu terminal · Prepare the workspace
mkdir -p ~/bioinformatics_training/part5
cd ~/bioinformatics_training/part5
explorer.exe .

Copy the three .fastq.gz files from the extracted ZIP into the folder that opens. Keep the gzip files compressed and retain their filenames. Check that Ubuntu can see all three:

Ubuntu terminal · Check the inputs
cd ~/bioinformatics_training/part5
ls -lh fastqc_*.fastq.gz
fastqc --version

Run FastQC on all three files:

Ubuntu terminal · Compare the datasets
fastqc fastqc_good.fastq.gz fastqc_tail_drop.fastq.gz \
       fastqc_adapter_rich.fastq.gz

The single backslash at the end of the first line continues the command on the next line. Copy both lines together, with no spaces after the backslash.

Once analysis finishes, list the reports and open the folder:

Ubuntu terminal · Open the reports
ls -lh *_fastqc.html
explorer.exe .
Open these three HTML reports in your browser
fastqc_good_fastqc.html
fastqc_tail_drop_fastqc.html
fastqc_adapter_rich_fastqc.html

Each input also produces a ZIP archive containing its report data. Compare Per base sequence quality, Adapter Content, and Overrepresented sequences across the reports.

What should you see?

The website practice files were tested with FastQC 0.11.9. Each report shows 5,000 sequences, a 100-base read length, and Sanger / Illumina 1.9 (Phred+33) encoding.

These flags describe deliberate teaching patterns.
DatasetPer base sequence qualityAdapter Content
fastqc_good.fastq.gzPASS. High quality across the read, approximately Q40.PASS. No adapter signal detected in this test.
fastqc_tail_drop.fastq.gzFAIL. Quality declines from Q40 to Q10 at the read end.PASS. No adapter signal detected in this test.
fastqc_adapter_rich.fastq.gzPASS. High quality across the read, approximately Q40.FAIL. The adapter trace rises to 75%, with adapter sequence after position 60.

On a small screen, scroll the table sideways.

An empty overrepresentation table is expected here.

All three files pass Overrepresented sequences in this test. In the adapter-rich file, many reads share an adapter near the end but have different beginnings. For these 100-base reads, FastQC 0.11.9 checks the first 50 bases for overrepresentation, so it does not count that shared adapter suffix as a repeated read.

This is a useful comparison: Adapter Content can flag an adapter even when Overrepresented sequences is empty. If a different dataset lists an overrepresented sequence, identify it before choosing an action. The official Overrepresented Sequences guide explains the length rule.

Other modules can also flag the deliberate patterns in these synthetic files. A completed run with WARN or FAIL entries is a valid result; inspect the relevant plots rather than trying to make every section green.

Step 07 · 10-minute challenge

Interpret first. Decide second.

  1. A read has a quality drop only at the final bases. What would you inspect before trimming?
  2. FastQC reports adapter content. What should you confirm before removing sequence?
  3. An amplicon library has high duplication. Is that automatically a problem?
  4. A GC plot has an unexpected second peak. What possibilities would you investigate?
  5. Why is an all-green FastQC report not the goal of quality control?
Check your answers after trying the exercise
  1. Inspect how far quality drops, how many positions and reads are affected, the spread of scores, and whether the remaining read length and quality would suit the next analysis.
  2. Confirm the identity of the adapter or unwanted primer and whether it is expected from the library preparation. Choose the processing method after that check.
  3. No. Constrained sequence starts and repeated targets can make duplication expected in amplicon libraries. Interpret it against the design.
  4. Consider mixed populations, contamination, or a biased subset of reads. Compare with the expected sample and library; a second peak alone does not identify its cause.
  5. Default thresholds cannot describe every experiment. Quality control is about fitness for the intended analysis and justified decisions, not collecting green ticks.

If you get stuck

Share the command, error, or report pattern you are unsure about in the Part 5 LinkedIn discussion.

The command says “fastqc: command not found”.

Complete the Part 4 installation in Ubuntu, then check fastqc --version.

One of the input files is missing.

Run pwd and ls -lh. The command must run from the folder containing all three compressed FASTQ files. The Part 5 ZIP includes every input, so use those files together in the part5 folder.

Why did the “good” file in an earlier download show Q9?

A synthetic file containing only Q40 characters can make FastQC guess the older Phred+64 encoding. The website copies of the high-quality and adapter-rich files each include one Q30 character, which keeps the data high quality while making Phred+33 detection unambiguous. Their sequences and read counts are unchanged. The README records this adjustment.

Does the adapter-rich file need an overrepresented-sequence warning?

No. These modules measure different patterns. See the tested results above: this file has a strong adapter signal and an empty overrepresentation table.

Do I need to trim the practice files now?

This exercise is about reading reports and explaining a decision. Compare the patterns and record what you would investigate. Part 6 will consider when to trim, filter, or leave reads alone.

Further reading: Babraham Bioinformatics documentation for Adapter Content, Duplicate Sequences, GC Content, and Per Base Sequence Content.

Coming next · Part 6

Cleaning reads: when should we trim or filter?

And when should we leave them alone? Use the report and the experimental context to choose a sensible next step.