Before you start
In Part 1, we set up WSL2, opened Ubuntu, and created a workspace. Now we can start looking at sequence data before using analysis tools.
There is no new software to install. The files in this exercise are deliberately small so we can focus on their structure.
Two files in your bioinformatics_training workspace: a FASTA file containing two sequences and a FASTQ file containing two reads. You will understand where these formats come from, recognise their structure, and count their records.
Start here · Connect the wet lab to the files
From the wet lab to FASTQ
Before learning file syntax, it helps to understand how sequence data reaches your computer.
- Sample
Biological material containing DNA or RNA.
- Extraction
Nucleic acid is extracted.
- Library prep
Fragments are prepared for sequencing.
- Sequencing
The instrument generates measurement signals.
- Base calling
Signals are converted into sequence bases with quality estimates.
- FASTQ
Reads together with their per-base quality scores.
FASTQ is commonly produced after base calling and, when needed, demultiplexing. Demultiplexing separates reads by sample. FASTQ is the practical sequence file you will usually work with.
Illumina
BCL → FASTQ
BCL instrument files already contain base calls and quality scores. Conversion and demultiplexing produce FASTQ files.
Illumina: BCL Convert output files ↗Nanopore
Raw signal → FASTQ
Raw signal is stored in POD5 or older FAST5 files. After base calling, reads and their quality scores can be exported as FASTQ.
Nanopore: base calling and FASTQ output ↗FASTA vs FASTQ
FASTQ
Stores sequencing reads together with a per-base quality string. Our examples use four lines per read, beginning with an @ identifier.
@read_001
ATGCGTACGTAGC
+
IIIIIIIIIIIIIReads + quality scores
FASTA
Stores sequence information without a per-base quality string. Each record has a header starting with >, followed by the sequence.
>virus_A
ATGCGTACGTAGC
>virus_B
ATGCGTACCTAGCSequence + identifier
Why does FASTQ need a quality line?
The sequencer is not equally confident about every base call. The quality string records that confidence so we can assess read quality before downstream analysis. There must be one quality character for every base.
Where can FASTA come from?
- Reference databases
- Genome, gene, or protein sequences downloaded for comparison or analysis.
- Assembly
- Reads assembled into longer contigs or genomes.
- Consensus
- A representative sequence generated after mapping reads.
FASTA is a general sequence format. It can store reads converted from FASTQ, as well as reference, assembled, and consensus sequences.
| Feature | FASTA | FASTQ |
|---|---|---|
| Starts with | > identifier | @ read identifier |
| Stores | Sequence | Sequence + quality |
| Per-base quality | No | Yes |
| Typical use | References, genes, contigs, consensus | Sequencing reads |
| Record structure | Header + sequence | Four lines per read in this exercise |
Step 01 · Ubuntu terminal
Create two tiny practice files
Open Ubuntu. Move into the workspace you created in Part 1:
cd ~/bioinformatics_trainingCreate example.fasta
Copy the entire block, including the final EOF, into Ubuntu and press Enter.
cat > example.fasta << 'EOF'
>virus_A
ATGCGTACGTAGC
>virus_B
ATGCGTACCTAGC
EOFCreate example.fastq
Now copy this complete block to create the FASTQ file:
cat > example.fastq << 'EOF'
@read_001
ATGCGTACGTAGC
+
IIIIIIIIIIIII
@read_002
ATGCGTACCTAGC
+
BBBBBBBBBBBBB
EOFHere, << 'EOF' lets us enter several lines of text into a file. The final EOF, on a line by itself, marks the end. Keep the commands and line breaks exactly as shown.
Check that both files exist:
ls -lhYou should see example.fasta and example.fastq. Dates, file sizes, and your username may differ from someone else's output.
Step 02 · Read the records
Now look inside the files
The cat command prints a file's contents in the terminal. Run these two commands:
cat example.fasta
cat example.fastqFASTA: header, then sequence
In example.fasta, >virus_A identifies the first sequence. The next line contains that sequence. The next > header begins the second record.
FASTQ: four lines in each example read
| Line and example | What it means |
|---|---|
1 · Identifier@read_001 | The name of the read. The identifier line starts with @. |
2 · SequenceATGCGTACGTAGC | The bases in the read. |
3 · Separator+ | Separates the sequence from its quality information. |
4 · Quality stringIIIIIIIIIIIII | One quality character for every base. This first read uses capital I characters. |
The sequence and quality string must have the same number of characters. Each practice read has 13 bases and 13 quality characters. The first uses capital I characters; the second uses capital B characters. We will explain what they mean in Part 3.
Step 03 · A little analysis
Count sequences and reads
1. Show only the first FASTQ read
head -n 4 example.fastqOur FASTQ records use four lines each, so this shows exactly one read:
@read_001
ATGCGTACGTAGC
+
IIIIIIIIIIIII2. Count sequences in the FASTA file
grep -c '^>' example.fastaExpected result: 2. The pattern ^> matches lines that begin with >, and -c counts them.
3. Count reads in this four-line FASTQ
awk 'END {print NR/4}' example.fastqExpected result: 2. At the end of the file, this divides the total number of lines by four.
Two reads, with four lines each, give us eight lines in total.
For this toy FASTQ file, each read occupies four lines, so eight lines represent two reads. For real datasets, we will use dedicated bioinformatics tools instead of relying on NR/4.
10-minute challenge
Your turn.
Try these without scrolling back to the commands above:
- Open
example.fastqand identify the read ID, sequence, separator, and quality string. - Which file contains per-base quality information: FASTA or FASTQ?
- Count the sequences in
example.fasta. - Count the reads in
example.fastq. - Explain in one sentence where FASTQ commonly comes from and where FASTA may come from.
Check your answers
head -n 4 example.fastqshows the first read:@read_001, its sequence, the+separator, and the quality string.- FASTQ contains per-base quality information.
- The FASTA file contains 2 sequences.
- The FASTQ file contains 2 reads across 8 lines.
- FASTQ commonly comes from base calling and, when needed, demultiplexing; FASTA may come from reference databases, assembly, or consensus generation.
FASTQ usually contains sequencing reads plus quality scores. FASTA contains sequence information without per-base quality scores.
If you get stuck
Ubuntu says the workspace does not exist.
Complete the workspace step in Part 1, then return here. The folder used in both lessons is ~/bioinformatics_training.
I see a > prompt and the file is not finished.
Ubuntu is waiting for the end of the text block. Type EOF on its own line, without quotes or spaces, and press Enter. If you want to cancel and start again, press Ctrl+C and copy the complete block.
My FASTQ count is not 2.
Inspect the file with cat example.fastq. Compare it with the example above, including all eight lines. Extra blank lines or missing lines will affect the simple counting command. Re-running the creation block replaces that practice file with the example contents.
Can I use PowerShell for these commands?
Use the Ubuntu terminal you set up in Part 1. The file-creation syntax and commands in this exercise are written for that environment.
For more detail on the file formats, see the NCBI SRA File Format Guide.
Have a question? Share the command, error, or step you are working on in the Part 2 discussion on LinkedIn.
Keep a copy of the eight-slide lesson (PDF) for your next practice session.
Continue with Part 3
Can I trust this base?
Understand FASTQ quality scores, decode quality characters, and connect them to confidence in each base call.