Evolutionary and Comparative Genomics•2026-09-14

Evolutionary Phylogeny

NM

Nasir Mahmood Abbasi, PhD

Bioinformatics Educator

Evolutionary Phylogeny
Tested on: Python 3.11, R 4.3.2, Ubuntu 24.04
Last Review: 2026-08-15

Learning Objectives & Prerequisites

  • Prerequisites: Complete Biological Data Formats and understand sequences, multiple sequence alignment, and homologous characters.
  • Objective: Build and interpret a phylogenetic tree while distinguishing tree topology, support, rooting, and evolutionary inference.
  • Expected Output: A reproducible tree-analysis record containing aligned input, model or method, root choice, support values, and figure caption.

Suggested route: use the Bioinformatics Learning Path to review any prerequisite stage before continuing.

Evolutionary Bioinformatics Analysis

Introduction to Evolutionary Bioinformatics

Evolutionary bioinformatics applies computational methods to understand the evolutionary relationships between biological sequences (DNA, RNA, or proteins). By analyzing these relationships, we can infer common ancestry, trace the history of gene families, and understand functional conservation.

Two of the most fundamental steps in evolutionary analysis are Multiple Sequence Alignment (MSA) and Phylogeny (Tree Construction).


1. Multiple Sequence Alignment using Kalign

Before building an evolutionary tree, homologous sequences must be aligned to identify conserved regions and mutations. While tools like Clustal Omega and MAFFT are popular, Kalign is highly efficient for aligning massive datasets rapidly.

Installing Kalign

Kalign is available via Conda/Mamba in the bioconda channel:

# Create a dedicated environment
mamba create -n phylogeny_env kalign
conda activate phylogeny_env

Running Kalign

To perform a multiple sequence alignment on a FASTA file containing unaligned sequences:

# Basic alignment
kalign -i unaligned_sequences.fasta -o aligned_sequences.fasta

# Specify output format (e.g., Clustal format instead of FASTA)
kalign -i unaligned_sequences.fasta -f clu -o aligned_sequences.aln

Kalign uses the Wu-Manber string-matching algorithm, making it exceptionally fast for large-scale phylogenomics projects where thousands of sequences are involved.


2. Phylogenetic Tree Construction

Once sequences are aligned, the next step is calculating the evolutionary distance and constructing a phylogenetic tree. There are several methods for tree construction: * Distance-based: Neighbor-Joining (NJ), UPGMA * Character-based: Maximum Parsimony (MP), Maximum Likelihood (ML) * Bayesian Inference: (e.g., MrBayes)

Using MEGA via Command Line (MEGA-CC)

Molecular Evolutionary Genetics Analysis (MEGA) is an industry standard for phylogeny. While famous for its graphical interface, MEGA also provides a command-line version called MEGA-CC (MEGA Computational Core), which is perfect for HPC environments or automated pipelines.

Generating a Configuration (.mao) File

MEGA-CC requires an analysis options file (.mao). You may create it in the MEGA GUI, but do not treat a GUI selection as the complete record of an analysis. Save the exact .mao file beside the alignment, record the MEGA-CC version, and commit both settings and command to the project repository.

For an auditable exercise, create a project folder containing:

phylogeny-project/
├── aligned_sequences.fasta
├── ML_Tree_Settings.mao
├── command.txt
└── results/

Write the exact megacc command into command.txt. If a learner cannot produce or inspect a .mao file, use the fully command-line IQ-TREE route below instead; it exposes model selection and bootstrap settings directly in the command.

Running MEGA-CC

Once you have your alignment file and your .mao configuration file, you can run the analysis on the command line:

# Run MEGA-CC for Maximum Likelihood tree construction
megacc -a ML_Tree_Settings.mao -d aligned_sequences.fasta -o output_tree

This will output several files, including the tree in Newick format (output_tree.nwk), which can be visualized using tools like FigTree or iTOL (Interactive Tree Of Life).


3. Alternative Command Line Phylogeny Tools

While MEGA is excellent, other highly optimized command-line tools are heavily used in modern bioinformatics:

IQ-TREE (Maximum Likelihood)

IQ-TREE is renowned for its speed and its "ModelFinder" feature, which automatically determines the best-fitting evolutionary model for your data before building the tree.

mamba install -c bioconda iqtree

# Run IQ-TREE with automatic model selection and 1000 ultrafast bootstraps
iqtree -s aligned_sequences.fasta -m MFP -B 1000

FastTree (Approximate Maximum Likelihood)

For massive alignments (e.g., tens of thousands of sequences), FastTree is often the only feasible option.

mamba install -c bioconda fasttree

# Build a tree using the General Time Reversible (GTR) model
FastTree -gtr -nt < aligned_sequences.fasta > output.tree

Summary Workflow

  1. Gather Sequences: Download homologous FASTA sequences.
  2. Align: Use kalign to generate a robust multiple sequence alignment.
  3. Build Tree: Use megacc or iqtree to compute evolutionary distances and infer the tree topology.
  4. Visualize: Upload the resulting .nwk file to iTOL for publication-ready visualization.

Matched Python and R tree-inspection workflow

Use code to inspect and plot an already inferred Newick tree; it does not substitute for model selection, alignment review, or branch-support assessment during inference.

from Bio import Phylo

species_tree = Phylo.read("species_tree.nwk", "newick")
print({"tips": species_tree.count_terminals(), "total_branch_length": species_tree.total_branch_length()})
Phylo.draw(species_tree)
library(ape)

species_tree <- read.tree("species_tree.nwk")
print(list(tips = Ntip(species_tree), total_branch_length = sum(species_tree$edge.length)))
plot(species_tree, cex = 0.7)

Knowledge Check & Assessment

1. Concept Verification

What does a branch pattern represent, and what cannot be inferred from branch placement without an explicit root and support assessment?

2. Practical Execution

Generate or inspect a tree from a provided alignment, identify sister groups and support values, and write a cautious interpretation. Pass Criteria: Record the command or analysis choice, keep the output, and explain why it answers the stated task.

3. Troubleshooting

If a tree changes after alignment trimming or rooting, what methodological choices should be reported before making evolutionary claims?

Reviewed: June 2026

All commands and outputs were verified with the software versions listed in this tutorial. If you encounter reproducibility issues, please report them through the Contact page.

Author: Nasir Mahmood Abbasi, PhD · Category: Evolutionary and Comparative Genomics

Continue Learning

Course Sequence