Skip to lesson content

FASTQ to Phylogeny · Part 4 · Free beginner exercise

One read is easy. What about millions?

Assessing sequencing data with FastQC

Move from individual quality scores to a whole dataset. Install FastQC in Ubuntu, compare two small FASTQ files, and learn what the report is helping you ask.

View the LinkedIn post and discussion ↗
Two small datasets5,000 synthetic reads in each fileSame sequences, different qualityCompare full-length quality with a tail dropLearn, then practiseEight slides and a 10-minute challenge

Everything for Part 4

Download. Run. Compare.

Get the eight-slide lesson and the two practice FASTQ datasets used in the exercise. Both downloads are free.

Practice ZIP: about 354 KB · Slides: eight pages, about 20 KB

Inside the ZIP: fastqc_good.fastq.gz, fastqc_tail_drop.fastq.gz, and README.txt.

These are synthetic teaching files, not biological sequencing data. Extract the outer ZIP, but keep the two .fastq.gz files compressed. FastQC can read them directly.

How to put the files in your Ubuntu workspace →

From one read to a whole dataset.

In Part 3, we decoded the quality characters for individual bases. FastQC summarizes patterns across a FASTQ file so we can decide what needs a closer look before downstream analysis.

  1. 1
    FASTQ file

    Reads and their quality scores.

  2. 2
    FastQC

    Checks patterns across the dataset.

  3. 3
    HTML report

    Plots, statistics, and flags to inspect.

FastQC inspects the data.

It does not trim, filter, or fix the reads.

Use the Ubuntu/WSL2 environment from Part 1. You will need internet access and your Ubuntu password for the one-time installation. All command blocks below belong in Ubuntu.

Keep three questions in mind

  • Does confidence stay high across the read?
  • Do the base and GC patterns look plausible for this library?
  • Are adapters, duplication, changing read lengths, or other unusual patterns visible?

Step 01 · One-time setup

Install FastQC in Ubuntu

Update the package list, install FastQC, and check the version. Run these lines in order:

Ubuntu terminal
sudo apt update
sudo apt install fastqc -y
fastqc --version

If a FastQC version is printed, the command-line installation is ready. The version number may differ between Ubuntu releases. If FastQC is already installed, you can simply check it with fastqc --version.

Step 02 · Practice datasets

Put the files in your training folder

High quality throughout

fastqc_good.fastq.gz

5,000 synthetic reads, each 100 bases long, with high quality across the full read.

Quality drops at the end

fastqc_tail_drop.fastq.gz

The same 5,000 reads, with lower quality toward the read end. Only the quality strings differ.

  1. Download the practice ZIP above. In Windows File Explorer, right-click it and choose Extract All.
  2. In Ubuntu, run the block below. It opens your training folder in Windows File Explorer.
Ubuntu terminal · Open your workspace
mkdir -p ~/bioinformatics_training
cd ~/bioinformatics_training
explorer.exe .

Copy both .fastq.gz files from the extracted download into the training folder you just opened. Keep their names unchanged. You do not need to extract the gzip files or copy the README over your existing workspace README.

Back in Ubuntu, confirm that both files are in place:

Ubuntu terminal · Check the input files
cd ~/bioinformatics_training
ls -lh fastqc_*.fastq.gz

You should see fastqc_good.fastq.gz and fastqc_tail_drop.fastq.gz. Each compressed file is about 177 KB.

Step 03 · Run FastQC

Generate the quality-control reports

From your training folder, run FastQC once for each file:

Ubuntu terminal · Analyse both datasets
fastqc fastqc_good.fastq.gz
fastqc fastqc_tail_drop.fastq.gz

Then list the files and open the folder in Windows:

Ubuntu terminal · Find and open the reports
ls -lh *fastqc*
explorer.exe .

Each input produces an HTML report for viewing and a ZIP archive containing the report data. Double-click each HTML report to open it in your browser.

Four new output files, alongside your two inputs
fastqc_good_fastqc.html
fastqc_good_fastqc.zip
fastqc_tail_drop_fastqc.html
fastqc_tail_drop_fastqc.zip

Step 04 · Read the report

Start with three sections

Open both reports side by side. You do not need to interpret every FastQC module at once.

In Basic Statistics, the practice files should be identified as Sanger / Illumina 1.9 encoding: the Phred+33 convention used in Part 3.

Basic Statistics
Check the number of reads, read length, and GC content. Both files contain 5,000 reads of 100 bases, so those values should agree.
Per base sequence quality
Compare quality from the beginning to the end of the read. Read position runs along the horizontal axis; the vertical axis shows the quality score.
Adapter Content
Inspect whether known adapter sequences become more common at particular positions. These two datasets have identical base sequences, so changing their quality scores alone does not create an adapter difference.
The Part 3 connection

The per-base plot summarizes Phred quality across many reads. The tail-drop file should show lower quality near the end; the high-quality file should remain high across its full length.

These deliberately simple teaching files isolate a quality difference. Real sequencing data can show more variation, so use them to learn how to read the report rather than as a template for every experiment.

Step 05 · Interpretation

Green, yellow, and red are clues.

FastQC uses thresholds to flag patterns for your attention. Read the module name and the plot together with its status.

PASS
The result is within that module’s default expectations.
WARN
The pattern is unusual enough to inspect.
FAIL
The pattern is strongly unusual by that module’s threshold.
A warning or failure does not automatically mean unusable data.

Ask whether the pattern fits the sample type, library design, sequencing platform, and expected biology. A flag points to something to investigate; it does not identify the exact cause.

Likewise, a PASS does not establish that the dataset is suitable for every downstream analysis. Interpret the report in the context of the question you are trying to answer.

Step 06 · 10-minute challenge

Your turn.

  1. Run FastQC on both FASTQ files.
  2. Open both HTML reports.
  3. Which dataset shows lower quality toward the end of the read?
  4. Which FastQC module makes that easiest to see?
  5. Does the report tell you the exact wet-lab or sequencing cause? Why not?
Check your answers after trying the exercise
  1. Run fastqc fastqc_good.fastq.gz and fastqc fastqc_tail_drop.fastq.gz from the folder containing the inputs.
  2. Open fastqc_good_fastqc.html and fastqc_tail_drop_fastqc.html.
  3. fastqc_tail_drop.fastq.gz shows the quality drop toward the read end.
  4. Per base sequence quality is the clearest module for comparing this pattern.
  5. No. FastQC summarizes the data it receives; it does not observe your experiment. Different causes can produce similar patterns, and these particular files are synthetic examples.

If you get stuck

Share the command, error, or step you are working on in the Part 4 discussion on LinkedIn.

Ubuntu says “fastqc: command not found”.

Run the installation commands in Ubuntu, then check fastqc --version. The command name is lowercase. If installation fails, keep the exact error message so you can identify the cause.

My password does not appear when I type it.

That is normal for sudo in Ubuntu. Type your Ubuntu password and press Enter; no letters or dots appear as you type.

The FASTQ files cannot be found.

Check that you extracted the outer practice ZIP and copied both .fastq.gz files into ~/bioinformatics_training. Run pwd to check your location and ls -lh to check the filenames. Files still in Windows Downloads are not automatically in your Ubuntu training folder.

Do I need to unzip the .fastq.gz files?

No. Keep those two files compressed. FastQC reads gzip-compressed FASTQ directly. The outer practice ZIP is just the download package; the ZIP files produced after running FastQC contain report data.

File Explorer did not open.

Make sure you are running explorer.exe . in Ubuntu under WSL2, with a space before the final dot. The dot means the current folder. See Microsoft’s guide to working across Windows and Linux file systems for another way to find your files.

Why is an HTML report flagged even though FastQC finished successfully?

A report flag describes a data pattern; it is not the same as a software error. Read the named module, compare its plot, and return to the experimental context before deciding what to do.

Further reading: Babraham Bioinformatics: FastQC, including the official documentation and example reports.

Coming next · Part 5

Which FastQC warnings matter?

What should you do next? Connect the report’s patterns to decisions about your data.