The fourth line has a job.
In Part 2, we created FASTA and FASTQ files and identified the four lines in a FASTQ read. Now we will focus on the quality string.
@read_demo
ATGCGTAC
+
IIII5?I+The second line contains the bases. The fourth line contains one quality character for every base. Both strings have eight characters.
A quality score describes the estimated confidence in a base call. It gives us a way to express how likely that call is to be wrong.
Step 01 · Wet lab to dry lab
From sequencing signal to confidence
A sequencer measures a physical signal. Software then decides which base is most likely and estimates how confident that call is.
- Signal measured
The instrument records fluorescence, an electrical signal, or another platform-specific measurement.
- Base called
Base-calling software interprets the signal as a base such as A, C, G, or T.
- Confidence estimated
The software assigns an estimated probability that the base call could be wrong.
- Stored in FASTQ
The sequence is stored with one quality character for every base.
A quality score tells us about confidence in the call. By itself, it does not explain why that confidence is high or low.
Step 02 · Read the probability
What does a Phred score mean?
FASTQ quality values are usually expressed as Phred scores, written as Q. A higher Q means a lower estimated probability of an incorrect base call.
Q = −10 log10(Perror)
Perror = 10−Q/10
You do not need to calculate logarithms for this exercise. Focus on what the values mean:
| Quality | Error probability | About 1 error in | Estimated accuracy |
|---|---|---|---|
| Q10 | 0.1 | 10 base calls | 90% |
| Q20 | 0.01 | 100 base calls | 99% |
| Q30 | 0.001 | 1,000 base calls | 99.9% |
| Q40 | 0.0001 | 10,000 base calls | 99.99% |
Every increase of 10 in Q corresponds to a tenfold decrease in the estimated error probability. Q30 has one tenth the estimated error probability of Q20.
Q30 means an estimated error probability of 0.1%, or about 1 in 1,000. It does not mean “30% quality.” This is an estimate for a base call, not a guarantee of exactly one error in a particular set of 1,000 bases.
The definition and values follow Illumina’s explanation of quality scores.
Step 03 · Phred+33 encoding
Why are there letters and symbols?
Modern FASTQ commonly uses Phred+33 encoding: the numerical quality score is stored as a printable character. These characters represent quality values; they are not additional nucleotides.
+Q10About 1 in 10 incorrect
90% estimated accuracy
5Q20About 1 in 100 incorrect
99% estimated accuracy
?Q30About 1 in 1,000 incorrect
99.9% estimated accuracy
IQ40About 1 in 10,000 incorrect
99.99% estimated accuracy
The I here is a capital letter I, and 5 is the digit five. Their meaning comes from their character codes.
What does the “+33” mean?
The character’s ASCII code equals the Phred score plus 33. To decode it, subtract 33: the ASCII code for capital I is 73, so 73 − 33 = 40, or Q40. You do not need to memorise ASCII codes; software handles this for real datasets.
This lesson uses Phred+33. Other encodings occur in older data; the NCBI SRA File Format Guide describes the main variants.
Step 04 · One read
Read one base at a time
Line up each base with the quality character at the same position. Our eight-base sequence, ATGCGTAC, pairs with IIII5?I+.
| Position | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|---|---|
| Sequence | A | T | G | C | G | T | A | C |
| Quality character | I | I | I | I | 5 | ? | I | + |
| Phred score | Q40 | Q40 | Q40 | Q40 | Q20 | Q30 | Q40 | Q10 |
On a narrow screen, scroll the table sideways to see all eight positions.
Which base has the lowest confidence?
Position 8: C. Its quality character is +, which represents Q10, the lowest score in this example. That means a 10% estimated probability that this call is wrong; it does not prove that this particular C is incorrect.
The + on the third FASTQ line separates sequence and quality. The final + on the fourth line is a quality character for the eighth base. Its role depends on which line it appears on.
Step 05 · Wet-lab perspective
Why should the wet lab care?
Quality scores connect the sequence files back to the experiment. Several factors can affect confidence, so a low score alone cannot identify the cause.
- Input material
- The condition of the nucleic acid can affect the library entering sequencing.
- Library preparation
- Suboptimal or uneven libraries can contribute to poorer or less consistent reads.
- Sequencing run
- Loading, chemistry, instrument conditions, and run performance can influence read quality.
- Read position and platform
- Quality patterns differ by sequencing technology and can change across a read.
- Base calling
- The software interprets the measured signal and estimates confidence in its calls.
Investigating the cause requires the sample, library, run, and platform context. A quality score shows where confidence is low, but does not establish which part of the experiment caused it.
Step 06 · Ubuntu terminal
Create a mixed-quality FASTQ read
No new software is needed. Open the Ubuntu terminal from Part 1 and use the same workspace as Part 2.
1. Move into your workspace
cd ~/bioinformatics_training2. Create quality_demo.fastq
Copy the entire block, including the final EOF, into Ubuntu and press Enter.
cat > quality_demo.fastq << 'EOF'
@read_demo
ATGCGTAC
+
IIII5?I+
EOFThe last EOF must be on its own line. Re-running this block replaces quality_demo.fastq with the same practice read.
3. Inspect the file
cat quality_demo.fastq@read_demo
ATGCGTAC
+
IIII5?I+Your task
Identify the lowest-quality base, the Q30 character, and why the sequence and quality strings must both contain eight characters.
Check the practice-file answers
- The lowest-quality base is C at position 8:
+means Q10. - The Q30 character is
?, paired with T at position 6. - Each of the eight bases needs its own quality character, so the two strings must have equal lengths.
Step 07 · 10-minute challenge
Your turn.
Try these without looking back at the explanations above:
- What does a FASTQ quality score describe?
- What does Q30 mean in plain language?
- Which has lower confidence: Q10 or Q40?
- In Phred+33, which character in our example represents Q20?
- Why can a low quality score not identify the exact wet-lab cause by itself?
Check your answers
- It describes the estimated confidence in a base call, expressed through the probability that the call is wrong.
- Q30 means an estimated error probability of about 1 in 1,000, or 0.1%.
- Q10 has lower confidence and a higher estimated error probability than Q40.
- The digit
5represents Q20 in Phred+33. - Different sample, library, sequencing, and software factors can affect quality. The score does not identify which factor caused a particular low-confidence call.
Each FASTQ quality character belongs to one base. Higher Phred scores mean lower estimated error probabilities.
If you get stuck
Have a question? Share the command, error, or step you are working on in the Part 3 discussion on LinkedIn.
Ubuntu says the workspace does not exist.
Complete the workspace step in Part 1, then return here. This exercise uses ~/bioinformatics_training.
I see a > prompt and the file is not finished.
Ubuntu is waiting for the end of the text block. Type EOF on its own line, without quotes or spaces, and press Enter. To cancel and start again, press Ctrl+C and copy the complete creation block.
My quality string does not have eight characters.
Compare it with IIII5?I+: four capital I characters, then 5, ?, one more capital I, and +. Keep this on one line with no spaces. Copy the complete creation block again if needed.
Does the + quality character mean the base is definitely wrong?
No. In Phred+33, + means Q10, an estimated error probability of 10%. It indicates lower confidence, not a confirmed error.
Should I use PowerShell or Ubuntu?
Use Ubuntu for the commands in this lesson. The file-creation syntax is the same as in Part 2.
Keep a copy of the eight-slide lesson (PDF) for your next practice session.
Continue with Part 4
One read is easy. What about millions?
Assessing sequencing data with FastQC: download two practice datasets, run the reports, and compare their quality patterns.