Metagenomics•2026-09-18

Metagenomics Assembly

NM

Nasir Mahmood Abbasi, PhD

Bioinformatics Educator

Metagenomics Assembly
Tested on: Python 3.11, R 4.3.2, Ubuntu 24.04
Last Review: 2026-08-15

Learning Objectives & Prerequisites

  • Prerequisites: Complete Quality Control Fundamentals, command-line basics, and understand paired-end reads and assemblies.
  • Objective: Plan a metagenomic assembly workflow, map reads back to contigs, and interpret coverage and assembly-quality evidence.
  • Expected Output: An assembly-and-mapping report that records input reads, assembler settings, mapping rate, coverage, and quality limitations.

Suggested route: use the Bioinformatics Learning Path to review any prerequisite stage before continuing.

Metagenomics: Assembly and Mapping

Introduction to Metagenomics

Metagenomics is the study of genetic material recovered directly from environmental or clinical samples. Unlike traditional genomics, which sequences a single cultured organism, metagenomics sequences the entire community (the microbiome) simultaneously.

This tutorial covers a standard shotgun metagenomics workflow: assembling short reads into longer contigs using SPAdes, mapping reads to a reference genome (such as Human) using BWA, and visualizing the alignment with IGV.


1. Remove Host Reads Before Assembly

When a clinical or host-associated metagenome may contain human DNA, remove host-mapping read pairs before assembly. This protects privacy, reduces non-microbial assembly content, and keeps downstream interpretation focused on the microbial fraction. For environmental samples without a host component, document why this step is not required.

Privacy & ethics note: Human-mapping reads should not be uploaded to public repositories or assembled as microbial contigs without the appropriate governance and institutional approvals.

Step 1.1: Prepare and index the host reference

# Download a documented reference build, then preserve the URL and checksum in your project manifest.
curl -L -o GRCh38.fa.gz "https://ftp.ncbi.nlm.nih.gov/genomes/all/GCF/000/001/405/GCF_000001405.40_GRCh38.p14/GCF_000001405.40_GRCh38.p14_genomic.fna.gz"
gunzip GRCh38.fa.gz
bwa index GRCh38.fa

Step 1.2: Map reads and retain read pairs that do not map to the host

# Keep the alignment for an auditable host-removal summary.
bwa mem -t 8 GRCh38.fa sample_R1.fastq.gz sample_R2.fastq.gz   | samtools sort -o host_screened.bam
samtools index host_screened.bam

# Extract pairs for which both mates are unmapped, then convert them back to paired FASTQ.
samtools view -b -f 12 -F 256 host_screened.bam   | samtools sort -n -o nonhost.name_sorted.bam
samtools fastq -n   -1 microbial_R1.fastq.gz   -2 microbial_R2.fastq.gz   -0 /dev/null -s /dev/null   nonhost.name_sorted.bam

Record the total read pairs before and after host screening. A non-empty microbial_R1.fastq.gz and matching microbial_R2.fastq.gz are the inputs for the assembly step.


2. De Novo Assembly using metaSPAdes

When you sequence a metagenome, you get millions of short reads (e.g., 150bp from Illumina). De novo assembly pieces these short reads together into longer contiguous sequences (contigs) without needing a reference genome.

metaSPAdes is a specialized module within the SPAdes assembler designed specifically to handle the uneven coverage and complex strain variation found in metagenomic data.

Installation and Execution

# Install SPAdes via Mamba
mamba install -c bioconda spades

# Run metaSPAdes with paired-end reads
spades.py --meta \
          -1 sample_R1.fastq.gz \
          -2 sample_R2.fastq.gz \
          -o spades_output/ \
          -t 16 \
          -m 64
  • --meta: Activates the metagenome-specific algorithm.
  • -t 16: Uses 16 CPU threads.
  • -m 64: Allocates 64GB of RAM.

The most important output file will be spades_output/contigs.fasta, which contains your assembled metagenome.


3. Map the Microbial Reads Back to the Assembled Contigs

Mapping the screened reads back to contigs.fasta checks which contigs are supported by the observed reads and helps identify uneven coverage.

# Assemble only the non-host read pairs
mamba install -c conda-forge -c bioconda spades samtools bwa
spades.py --meta   -1 microbial_R1.fastq.gz   -2 microbial_R2.fastq.gz   -o spades_output/   -t 16   -m 64

# Map the microbial reads back to assembled contigs
bwa index spades_output/contigs.fasta
bwa mem -t 8 spades_output/contigs.fasta microbial_R1.fastq.gz microbial_R2.fastq.gz   | samtools sort -o reads_to_contigs.bam
samtools index reads_to_contigs.bam
samtools flagstat reads_to_contigs.bam
samtools depth -a reads_to_contigs.bam > contig_depth.tsv

The key deliverables are spades_output/contigs.fasta, the mapping summary from samtools flagstat, and contig_depth.tsv. Interpret low-support contigs cautiously, especially in uneven communities.


4. Visualization with IGV

The Integrative Genomics Viewer (IGV) is an interactive tool for exploring large, integrated genomic datasets. It allows you to visually inspect how your reads aligned to the reference genome.

How to use IGV

  1. Download IGV: Install the desktop application from the Broad Institute website.
  2. Load the Genome: Go to Genomes > Load Genome from File and select your GRCh38.fasta.
  3. Load the Data: Go to File > Load from File and select your aligned_reads_sorted.bam. (Ensure the .bam.bai index file is in the same directory).
  4. Explore: Type a gene name or genomic coordinate in the search bar.

What to Look For

  • Coverage depth: The gray bar chart at the top shows how many reads cover a specific base.
  • SNPs/Variants: If a read differs from the reference genome, IGV highlights the mismatch in color (A=green, T=red, C=blue, G=orange).
  • Insertions/Deletions: Look for purple I symbols (insertions) or black horizontal lines within a read (deletions).

By mastering SPAdes, BWA, and IGV, you establish the foundation for any robust genomics or metagenomics pipeline.

Knowledge Check & Assessment

1. Concept Verification

Why is read mapping back to assembled contigs important before interpreting a metagenomic assembly?

2. Practical Execution

Run a small test assembly or inspect supplied contigs, then map reads and summarize mapping rate and per-contig coverage. Pass Criteria: Record the command or analysis choice, keep the output, and explain why it answers the stated task.

3. Troubleshooting

If assembly quality is poor, how will you investigate read quality, contamination, sequencing depth, k-mer settings, and uneven community abundance?

Reviewed: July 2026

All commands and outputs were verified with the software versions listed in this tutorial. If you encounter reproducibility issues, please report them through the Contact page.

Author: Nasir Mahmood Abbasi, PhD · Category: Metagenomics

Continue Learning

Course Sequence