Skip to lesson content

Part 2 · Free beginner exercise

FASTA vs FASTQ.

Where do these files come from? Connect the wet lab to the sequence files on your computer, then create two tiny practice files and look inside them.

View the LinkedIn post and discussion ↗
Learn, then practiseA short overview and a 10-minute challengeNo new softwareUse Ubuntu from Part 1Slides includedThe complete eight-page lesson

Before you start

In Part 1, we set up WSL2, opened Ubuntu, and created a workspace. Now we can start looking at sequence data before using analysis tools.

There is no new software to install. The files in this exercise are deliberately small so we can focus on their structure.

What you should have at the end

Two files in your bioinformatics_training workspace: a FASTA file containing two sequences and a FASTQ file containing two reads. You will understand where these formats come from, recognise their structure, and count their records.

Start here · Connect the wet lab to the files

From the wet lab to FASTQ

Before learning file syntax, it helps to understand how sequence data reaches your computer.

  1. Sample

    Biological material containing DNA or RNA.

  2. Extraction

    Nucleic acid is extracted.

  3. Library prep

    Fragments are prepared for sequencing.

  4. Sequencing

    The instrument generates measurement signals.

  5. Base calling

    Signals are converted into sequence bases with quality estimates.

  6. FASTQ

    Reads together with their per-base quality scores.

The exact files depend on the platform

FASTQ is commonly produced after base calling and, when needed, demultiplexing. Demultiplexing separates reads by sample. FASTQ is the practical sequence file you will usually work with.

Illumina

BCL → FASTQ

BCL instrument files already contain base calls and quality scores. Conversion and demultiplexing produce FASTQ files.

Illumina: BCL Convert output files ↗

Nanopore

Raw signal → FASTQ

Raw signal is stored in POD5 or older FAST5 files. After base calling, reads and their quality scores can be exported as FASTQ.

Nanopore: base calling and FASTQ output ↗

FASTA vs FASTQ

FASTQ

Stores sequencing reads together with a per-base quality string. Our examples use four lines per read, beginning with an @ identifier.

@read_001
ATGCGTACGTAGC
+
IIIIIIIIIIIII

Reads + quality scores

FASTA

Stores sequence information without a per-base quality string. Each record has a header starting with >, followed by the sequence.

>virus_A
ATGCGTACGTAGC
>virus_B
ATGCGTACCTAGC

Sequence + identifier

Why does FASTQ need a quality line?

The sequencer is not equally confident about every base call. The quality string records that confidence so we can assess read quality before downstream analysis. There must be one quality character for every base.

Where can FASTA come from?

Reference databases
Genome, gene, or protein sequences downloaded for comparison or analysis.
Assembly
Reads assembled into longer contigs or genomes.
Consensus
A representative sequence generated after mapping reads.

FASTA is a general sequence format. It can store reads converted from FASTQ, as well as reference, assembled, and consensus sequences.

FASTA vs FASTQ at a glance
FeatureFASTAFASTQ
Starts with> identifier@ read identifier
StoresSequenceSequence + quality
Per-base qualityNoYes
Typical useReferences, genes, contigs, consensusSequencing reads
Record structureHeader + sequenceFour lines per read in this exercise

Step 01 · Ubuntu terminal

Create two tiny practice files

Open Ubuntu. Move into the workspace you created in Part 1:

Ubuntu · Open your workspace
cd ~/bioinformatics_training

Create example.fasta

Copy the entire block, including the final EOF, into Ubuntu and press Enter.

Ubuntu · Create the FASTA file
cat > example.fasta << 'EOF'
>virus_A
ATGCGTACGTAGC
>virus_B
ATGCGTACCTAGC
EOF

Create example.fastq

Now copy this complete block to create the FASTQ file:

Ubuntu · Create the FASTQ file
cat > example.fastq << 'EOF'
@read_001
ATGCGTACGTAGC
+
IIIIIIIIIIIII
@read_002
ATGCGTACCTAGC
+
BBBBBBBBBBBBB
EOF
What does EOF do?

Here, << 'EOF' lets us enter several lines of text into a file. The final EOF, on a line by itself, marks the end. Keep the commands and line breaks exactly as shown.

Check that both files exist:

Ubuntu · List the files
ls -lh

You should see example.fasta and example.fastq. Dates, file sizes, and your username may differ from someone else's output.

Step 02 · Read the records

Now look inside the files

The cat command prints a file's contents in the terminal. Run these two commands:

Ubuntu · Show both files
cat example.fasta
cat example.fastq

FASTA: header, then sequence

In example.fasta, >virus_A identifies the first sequence. The next line contains that sequence. The next > header begins the second record.

FASTQ: four lines in each example read

The first read in example.fastq
Line and exampleWhat it means
1 · Identifier@read_001The name of the read. The identifier line starts with @.
2 · SequenceATGCGTACGTAGCThe bases in the read.
3 · Separator+Separates the sequence from its quality information.
4 · Quality stringIIIIIIIIIIIIIOne quality character for every base. This first read uses capital I characters.

The sequence and quality string must have the same number of characters. Each practice read has 13 bases and 13 quality characters. The first uses capital I characters; the second uses capital B characters. We will explain what they mean in Part 3.

Step 03 · A little analysis

Count sequences and reads

1. Show only the first FASTQ read

Ubuntu · Show the first four lines
head -n 4 example.fastq

Our FASTQ records use four lines each, so this shows exactly one read:

Expected output
@read_001
ATGCGTACGTAGC
+
IIIIIIIIIIIII

2. Count sequences in the FASTA file

Ubuntu · Count FASTA headers
grep -c '^>' example.fasta

Expected result: 2. The pattern ^> matches lines that begin with >, and -c counts them.

3. Count reads in this four-line FASTQ

Ubuntu · Count the reads
awk 'END {print NR/4}' example.fastq

Expected result: 2. At the end of the file, this divides the total number of lines by four.

Two reads, with four lines each, give us eight lines in total.

Why NR/4?

For this toy FASTQ file, each read occupies four lines, so eight lines represent two reads. For real datasets, we will use dedicated bioinformatics tools instead of relying on NR/4.

10-minute challenge

Your turn.

Try these without scrolling back to the commands above:

  1. Open example.fastq and identify the read ID, sequence, separator, and quality string.
  2. Which file contains per-base quality information: FASTA or FASTQ?
  3. Count the sequences in example.fasta.
  4. Count the reads in example.fastq.
  5. Explain in one sentence where FASTQ commonly comes from and where FASTA may come from.
Check your answers
  1. head -n 4 example.fastq shows the first read: @read_001, its sequence, the + separator, and the quality string.
  2. FASTQ contains per-base quality information.
  3. The FASTA file contains 2 sequences.
  4. The FASTQ file contains 2 reads across 8 lines.
  5. FASTQ commonly comes from base calling and, when needed, demultiplexing; FASTA may come from reference databases, assembly, or consensus generation.
To remember

FASTQ usually contains sequencing reads plus quality scores. FASTA contains sequence information without per-base quality scores.

If you get stuck

Ubuntu says the workspace does not exist.

Complete the workspace step in Part 1, then return here. The folder used in both lessons is ~/bioinformatics_training.

I see a > prompt and the file is not finished.

Ubuntu is waiting for the end of the text block. Type EOF on its own line, without quotes or spaces, and press Enter. If you want to cancel and start again, press Ctrl+C and copy the complete block.

My FASTQ count is not 2.

Inspect the file with cat example.fastq. Compare it with the example above, including all eight lines. Extra blank lines or missing lines will affect the simple counting command. Re-running the creation block replaces that practice file with the example contents.

Can I use PowerShell for these commands?

Use the Ubuntu terminal you set up in Part 1. The file-creation syntax and commands in this exercise are written for that environment.

For more detail on the file formats, see the NCBI SRA File Format Guide.

Have a question? Share the command, error, or step you are working on in the Part 2 discussion on LinkedIn.

Keep a copy of the eight-slide lesson (PDF) for your next practice session.

Continue with Part 3

Can I trust this base?

Understand FASTQ quality scores, decode quality characters, and connect them to confidence in each base call.