Phylogenomics with OrthoFinder
Nasir Mahmood Abbasi, PhD
Bioinformatics Educator
Learning Objectives & Prerequisites
- Prerequisites: Complete Evolutionary Phylogeny and have predicted protein sets with consistent species and gene naming.
- Objective: Use orthogroups and comparative genomics outputs to distinguish orthology, paralogy, gene-family change, and species-tree inference.
- Expected Output: An annotated orthogroup summary that records input proteomes, software version, species set, and one cautiously interpreted result.
Suggested route: use the Bioinformatics Learning Path to review any prerequisite stage before continuing.
Phylogenomics and Orthology Analysis
Introduction
Traditional phylogenetic analysis relies on aligning a single highly conserved gene (like 16S rRNA or Cytochrome C) to determine the evolutionary relationship between species. However, single genes can be misleading due to horizontal gene transfer or differing mutation rates.
Phylogenomics solves this by using entire genomes. Instead of one gene, we look at thousands of genes simultaneously. The absolute industry standard tool for identifying which genes are comparable across different genomes is OrthoFinder.
1. The Concept of Orthology
To compare genomes, you must first find the matching genes across them. * Orthologs are genes in different species that evolved from a common ancestral gene via speciation. These are the genes you want to use to build a species tree. * Paralogs are genes related via duplication events within a genome. Comparing paralogs across species will give you an incorrect evolutionary tree.
OrthoFinder automatically sorts the proteomes (all protein sequences) of your species into Orthogroups, strictly separating orthologs from paralogs.
2. Running OrthoFinder
OrthoFinder is incredibly easy to run. You simply provide it a directory containing the .faa (Protein FASTA) files of the species you want to compare.
(Note: You can easily generate these .faa files by running PROKKA on your assembled genomes).
# Assuming you have a folder named 'proteomes' containing fasta files:
# speciesA.faa, speciesB.faa, speciesC.faa
# Run OrthoFinder using 16 threads
orthofinder -f proteomes/ -t 16 -a 16
Under the hood, OrthoFinder runs an all-versus-all DIAMOND blast search, calculates sequence similarities, clusters the genes into orthogroups using MCL, and infers unrooted gene trees.
3. Key Outputs from OrthoFinder
OrthoFinder creates a highly structured Results/ directory. The most critical outputs are:
Orthogroups.tsv
This is a matrix where each row is an orthogroup, and columns are your species. It shows exactly which genes belong to which orthogroup.
Single_Copy_Orthologue_Sequences/
This directory contains the absolute gold-mine for phylogenomics. A "Single-Copy Ortholog" is a gene that exists exactly once in every single species you analyzed. These are the perfect genes for building a highly robust phylogenetic tree, as there is zero ambiguity about paralogs.
Species_Tree/SpeciesTree_rooted.txt
OrthoFinder automatically concatenates the alignments of those single-copy orthologs and uses STAG and STRIDE algorithms to infer a highly accurate, rooted Species Tree in Newick format.
You can instantly visualize this file using tools like iTOL (Interactive Tree Of Life) or FigTree.
Conclusion
OrthoFinder represents a massive leap forward from single-gene phylogeny. By leveraging the entire proteome to accurately identify orthogroups and automatically inferring a rooted species tree, you can resolve deep evolutionary relationships with unprecedented statistical confidence.
Matched Python and R OrthoFinder species-tree workflow
The species tree is an OrthoFinder output. Inspect its topology and branch lengths after checking the orthogroup and single-copy-orthologue settings used to infer it.
from Bio import Phylo
species_tree = Phylo.read("Results_Example/Species_Tree/SpeciesTree_rooted.txt", "newick")
print(species_tree.get_terminals())
Phylo.draw(species_tree)
library(ape)
species_tree <- read.tree("Results_Example/Species_Tree/SpeciesTree_rooted.txt")
print(species_tree$tip.label)
plot(species_tree, cex = 0.7)
From Genomes to Phylogenies: A Conceptual Foundation
Traditional molecular phylogenetics inferred evolutionary relationships from one or a handful of genes. Phylogenomics scales this approach to entire genomes, using hundreds or thousands of conserved genes simultaneously to produce highly resolved phylogenetic trees with narrow confidence intervals. The core analytical challenge is not the tree-building itself but the gene selection and alignment: you must identify genes that are present in exactly one copy across all the species you are comparing, that have been evolving under broadly similar selective pressures, and whose sequences can be aligned reliably.
Genes that fail these criteria introduce noise or bias into the analysis. Genes that have been duplicated independently in different lineages (paralogs) will produce trees that reflect gene family history rather than species history. Genes under strong positive selection evolve rapidly in specific lineages, which can mislead tree reconstruction. Genes with regions that are unalignable due to insertions or deletions produce alignment columns of uncertain homology, which are better excluded than forced into an alignment.
OrthoFinder: Identifying Orthologous Groups at Scale
OrthoFinder identifies orthologs and orthogroups across any number of input proteomes using an all-versus-all DIAMOND search followed by a graph-based clustering step. Orthogroups are groups of genes descended from a single gene in the last common ancestor of the input species. Within an orthogroup, each species may have one or more genes depending on whether duplications occurred after speciation.
The key advantage of OrthoFinder over simpler reciprocal best-hit approaches is that it uses species tree-aware orthogroup inference. Rather than comparing genes from one species to another in isolation, OrthoFinder uses a reference species tree to distinguish true orthologs (which diverged due to speciation events) from co-orthologs (which diverged due to duplication events that preceded speciation). This distinction matters greatly for functional inference: orthologs tend to share conserved function across species, while paralogs may have diverged functionally.
Selecting Markers for Tree Construction
After running OrthoFinder, you have a set of orthogroups. For phylogenomic tree construction, you select single-copy orthogroups, which are orthogroups where each input species contributes exactly one gene. These single-copy orthologs are the least problematic for phylogenomics because there is no ambiguity about which sequence from each species should be aligned to which. Tools like BUSCO independently assess genome or proteome completeness using a similar set of single-copy orthologs, and the overlap between OrthoFinder results and BUSCO sets provides a useful sanity check.
A practical consideration is the number of single-copy orthologs shared across all input species. If you include a highly reduced parasite genome alongside free-living species with much larger proteomes, the number of universally single-copy genes will be small. You may need to relax the inclusion criterion from all species to a supermajority of species, accepting that some positions in the final supermatrix will be missing for certain species.
Concatenation vs Coalescence: Two Philosophies of Tree Building
Once you have aligned your single-copy orthologs, you face a choice: should you concatenate all the alignments into a single supermatrix and build one tree, or should you build a separate tree for each gene and then summarise across those gene trees? Concatenation is computationally simpler and assumes that all genes in your supermatrix share the same evolutionary history. This assumption is violated when incomplete lineage sorting (ILS) is common, which happens when speciation events were rapid and the ancestral populations were large.
Coalescence-based methods like ASTRAL address ILS by inferring the species tree as the tree that maximises the number of gene trees that agree with it under the multi-species coalescent model. They are statistically more consistent in the presence of ILS but require a large number of informative gene trees and can be sensitive to gene tree estimation error. In practice, comparing the concatenation tree to the ASTRAL species tree for your dataset is a useful diagnostic: strong disagreements between the two suggest that ILS or other gene tree discordance is a significant factor in your data.
Knowledge Check & Assessment
1. Concept Verification
Why is an orthogroup not automatically a one-to-one ortholog set?
2. Practical Execution
Run or inspect an OrthoFinder output and identify an orthogroup, a potential duplication event, and the evidence used for interpretation. Pass Criteria: Record the command or analysis choice, keep the output, and explain why it answers the stated task.
3. Troubleshooting
If a gene family appears expanded, how will you check assembly completeness, annotation consistency, gene models, and sampling before claiming adaptation?
Reviewed: July 2026
All commands and outputs were verified with the software versions listed in this tutorial. If you encounter reproducibility issues, please report them through the Contact page.
Author: Nasir Mahmood Abbasi, PhD · Category: Evolutionary and Comparative Genomics