FASTQ to Phylogeny - Part 5 practice files

These are synthetic teaching datasets, not biological sequencing data.

fastqc_good.fastq.gz
  5,000 reads, 100 bp, consistently high quality.

fastqc_tail_drop.fastq.gz
  5,000 reads, 100 bp, lower quality toward the read end.

fastqc_adapter_rich.fastq.gz
  5,000 reads, 100 bp, most reads contain an Illumina-style adapter sequence after position 60.
  Use FastQC to inspect Adapter Content and Overrepresented sequences.

Author: Syed Adnan Haider
Website edition: 10 October 2026

All three files use Phred+33 quality encoding. In the high-quality and
adapter-rich files, the first base of the first read is Q30 rather than Q40.
This one-character adjustment in each synthetic file avoids FastQC
misreading uniformly high quality as older Phred+64 encoding. All read
identifiers, sequences, read lengths, and read counts are unchanged.
The tail-drop file is unchanged.

Expected comparison with FastQC 0.11.9:
- High-quality file: high per-base quality; no adapter signal.
- Tail-drop file: quality declines from Q40 to Q10 near the read end.
- Adapter-rich file: high per-base quality and a strong adapter signal.

The Adapter Content and Overrepresented sequences modules ask different
questions. The adapter-rich file shares an adapter after position 60,
but its first 50 bases are diverse. FastQC 0.11.9 uses the first 50 bases
of these 100-base reads for its overrepresentation check, so an empty
Overrepresented sequences table is expected despite the adapter signal.

Use a separate part5 folder so earlier practice files are preserved.
Lesson and downloads:
https://syedadnanhaider.com/learn/bioinformatics-on-windows/part-5-fastqc-warnings.html
