Functional Pathway Analysis with HUMAnN 3
Nasir Mahmood Abbasi, PhD
Bioinformatics Educator
Learning Objectives & Prerequisites
- Prerequisites: Complete Metatranscriptomics Basics and be comfortable interpreting gene-family and pathway abundance tables.
- Objective: Map microbial transcripts to functional gene families and pathways while separating observed activity from unsupported causal claims.
- Expected Output: A pathway-level table with normalized abundance, database version, sample comparison, and stated uncertainty.
Suggested route: use the Bioinformatics Learning Path to review any prerequisite stage before continuing.
Functional Pathway Analysis in Metatranscriptomics
Introduction
In our previous metatranscriptomics guide, we discussed the critical first step: filtering out the massive abundance of ribosomal RNA (rRNA) using SortMeRNA to isolate the functional messenger RNA (mRNA).
Once you have your clean mRNA reads, the next goal is functional profiling: What specific biochemical pathways are the microbes actively utilizing? The absolute gold standard for this analysis is HUMAnN 3.
1. How HUMAnN 3 Works
HUMAnN 3 (The HMP Unified Metabolic Analysis Network) is highly sophisticated. It uses a tiered approach to maximize both speed and accuracy:
- Tier 1 (Nucleotide level): It first runs
MetaPhlAnto determine exactly which species are present in your sample. It then dynamically builds a custom database of the pangenomes for only those specific species. It maps your mRNA reads to this custom database using Bowtie2. This is extremely fast and accurate. - Tier 2 (Translated search): Any reads that fail to map to the known species' pangenomes (the "unclassified" reads) are translated into proteins and searched against the massive UniRef90 protein database using DIAMOND. This is slower, but ensures you capture functional genes even from unknown or unculturable species.
2. Running the HUMAnN 3 Pipeline
# Assuming you have your rRNA-depleted mRNA reads
humann --input sample_mRNA.fastq.gz \
--output humann_out/ \
--threads 16 \
--taxonomic-profile metaphlan_bugs_list.tsv # Optional: provide pre-computed taxonomy
Understanding the MetaCyc Database
HUMAnN 3 maps the individual gene families it finds into complete metabolic pathways using the MetaCyc database.
Why MetaCyc instead of KEGG? MetaCyc is heavily focused on experimentally elucidated pathways and is highly curated for microbial metabolism, whereas KEGG is broader and includes many eukaryotic-specific signaling pathways that are irrelevant to microbiome research.
3. Interpreting HUMAnN Outputs
HUMAnN 3 generates three primary output files:
pathabundance.tsv: The abundance of complete metabolic pathways (e.g., GLYCOLYSIS-E-D: superpathway of glycolysis).pathcoverage.tsv: The coverage of the pathway. (Just because one gene in a 10-gene pathway is highly expressed doesn't mean the pathway is active. Coverage checks if the entire pathway is present).genefamilies.tsv: The abundance of individual UniRef90 gene families.
Stratification by Species
The brilliance of HUMAnN 3 is that the outputs are stratified.
If you look at pathabundance.tsv, you won't just see "Glycolysis = 5000". You will see:
* Glycolysis = 5000
* Glycolysis|Escherichia_coli = 4000
* Glycolysis|Bacteroides_fragilis = 1000
This allows you to confidently state not just what the community is doing, but exactly which species is responsible for doing it!
4. Normalization and Differential Testing Are Different Steps
HUMAnN produces several abundance tables, and the correct statistical input depends on the question. CPM or relative-abundance tables can be useful for descriptive plots, but do not pass library-size-normalized CPM values directly to count-based differential-expression models such as DESeq2. DESeq2 expects un-normalized or estimated counts and computes its own size factors internally.
# Create a CPM table for visualization or exploratory comparisons.
humann_renorm_table --input humann_out/sample_pathabundance.tsv \
--output sample_pathabundance_cpm.tsv \
--units cpm
Before differential testing, first determine whether your chosen HUMAnN output is a count-like table or a normalized abundance. Use raw or appropriately estimated counts with DESeq2, keep sample metadata and batch variables explicit, and record the transformation used only for visualization. If you only have compositional or normalized pathway abundances, choose a method whose assumptions match that input rather than treating it as a DESeq2 count matrix. Consult the DESeq2 vignette for its input requirements.
Knowledge Check & Assessment
1. Concept Verification
Why does a detected transcript support potential activity but not necessarily measured pathway flux or phenotype?
2. Practical Execution
Interpret a HUMAnN-style pathway output and report one up- or down-shift with the normalization method and biological caveat. Pass Criteria: Record the command or analysis choice, keep the output, and explain why it answers the stated task.
3. Troubleshooting
If a pathway disappears after filtering, how will you inspect read depth, reference coverage, normalization, and the gene-family evidence behind it?
Reviewed: August 2026
All commands and outputs were verified with the software versions listed in this tutorial. If you encounter reproducibility issues, please report them through the Contact page.
Author: Nasir Mahmood Abbasi, PhD · Category: Metatranscriptomics