Trajectory Inference
Nasir Mahmood Abbasi, PhD
Bioinformatics Educator
Learning Objectives & Prerequisites
- Prerequisites: Complete scRNA-seq Basics, including QC, clustering, and marker interpretation; use data with a plausible continuous process.
- Objective: Infer and interpret pseudotime or lineage trajectories while separating computational ordering from directly observed developmental time.
- Expected Output: A trajectory figure with root rationale, branch interpretation, gene trends, and explicit validation limits.
Suggested route: use the Bioinformatics Learning Path to review any prerequisite stage before continuing.
Understanding Trajectory Inference
Single-cell RNA sequencing provides static snapshots of cellular states. However, biology is highly dynamic. Processes like differentiation, immune activation, and cellular exhaustion are continuous transitions.
Trajectory inference (TI) algorithms attempt to order cells along a learned trajectory based on their transcriptional similarities, creating a "pseudotime" axis. This allows us to track gene expression changes as cells differentiate.
In this tutorial, we will explore the foundational tools for TI in both Python and R.
1. PAGA (Python)
Partition-based Graph Abstraction (PAGA) is integrated directly into the Scanpy ecosystem. It generates a coarse-grained map of cellular connectivity and provides a robust scaffold for embedding and pseudotime calculation.
import scanpy as sc
import scvelo as scv
# Assuming `adata` is already preprocessed, clustered (e.g., 'leiden'), and has a UMAP
# We must first re-compute the neighborhood graph if it's been altered
sc.pp.neighbors(adata, n_neighbors=15, n_pcs=40)
# 1. Run PAGA
sc.tl.paga(adata, groups='leiden')
# 2. Plot the coarse-grained PAGA graph
sc.pl.paga(adata, plot=False)
# 3. Use PAGA initialization to recompute a more continuous UMAP
sc.tl.umap(adata, init_pos='paga')
sc.pl.umap(adata, color=['leiden'], legend_loc='on data')
# 4. Calculate Diffusion Pseudotime (DPT)
# First, identify a root cell (e.g., the stem cell or progenitor cluster)
# Assuming cluster '0' is the known progenitor state
adata.uns['iroot'] = np.flatnonzero(adata.obs['leiden'] == '0')[0]
# Compute pseudotime
sc.tl.dpt(adata)
# Visualize pseudotime progression across the manifold
sc.pl.umap(adata, color='dpt_pseudotime', cmap='viridis')
2. Monocle3 & Slingshot (R)
In the R ecosystem, Monocle3 and Slingshot are the leading frameworks for trajectory inference. Slingshot is highly favored for its simplicity and ability to handle branching trajectories effectively from existing Seurat objects.
library(Seurat)
library(slingshot)
library(SingleCellExperiment)
library(RColorBrewer)
# Convert Seurat object to SingleCellExperiment (SCE)
sce <- as.SingleCellExperiment(pbmc)
# 1. Run Slingshot
# We use the existing UMAP embedding and specify a starting cluster
sce <- slingshot(sce, clusterLabels = 'ident', reducedDim = 'UMAP', start.clus = '0')
# 2. Extract pseudotime and plot
colors <- colorRampPalette(brewer.pal(11, 'Spectral'))[-11]
plotcol <- colors[cut(sce$slingPseudotime_1, breaks=100)]
plot(reducedDims(sce)$UMAP, col = plotcol, pch=16, asp = 1)
lines(SlingshotDataSet(sce), lwd=2, col='black')
library(monocle3)
library(SeuratWrappers)
# Convert Seurat to CellDataSet (CDS)
cds <- as.cell_data_set(pbmc)
# 1. Calculate size factors and cluster cells within Monocle
cds <- estimate_size_factors(cds)
cds <- cluster_cells(cds)
# 2. Learn the principal graph (trajectory)
cds <- learn_graph(cds)
# 3. Order cells in pseudotime
# (This will open an interactive prompt to select the root nodes)
cds <- order_cells(cds)
# 4. Visualize the trajectory and pseudotime
plot_cells(cds,
color_cells_by = "pseudotime",
label_cell_groups=FALSE,
label_leaves=FALSE,
label_branch_points=FALSE,
graph_label_size=1.5)
Matched Python and R trajectory workflow
The two ecosystems use different algorithms, but both require a justified biological root and should be interpreted as an inferred ordering - not observed time. Choose the implementation that fits the anticipated lineage topology and validate the result against known biology.
import numpy as np
import scanpy as sc
sc.pp.neighbors(adata, n_neighbors=15, n_pcs=40)
sc.tl.paga(adata, groups="leiden")
adata.uns["iroot"] = np.flatnonzero(adata.obs["leiden"] == "0")[0]
sc.tl.dpt(adata)
sc.pl.umap(adata, color="dpt_pseudotime", cmap="viridis")
library(slingshot)
library(SingleCellExperiment)
sce <- as.SingleCellExperiment(pbmc)
sce <- slingshot(sce, clusterLabels = "ident", reducedDim = "UMAP", start.clus = "0")
pseudotime <- slingPseudotime(sce)
plot(reducedDims(sce)$UMAP, col = pseudotime[, 1], pch = 16, asp = 1)
lines(SlingshotDataSet(sce), lwd = 2, col = "black")
Conclusion
Trajectory inference shifts our perspective from discrete clusters to continuous cellular development. Whether you rely on PAGA's graph abstractions in Python or Slingshot's branching lineage logic in R, establishing a solid pseudotime framework is the key to identifying the gene regulatory networks driving cellular transitions.
Knowledge Check & Assessment
1. Concept Verification
Why does pseudotime not prove that one observed cell literally becomes another?
2. Practical Execution
Run or inspect a trajectory analysis, state the root choice, identify one branch, and plot expression of a relevant dynamic gene. Pass Criteria: Record the command or analysis choice, keep the output, and explain why it answers the stated task.
3. Troubleshooting
If the inferred path contradicts known biology, how will you check rooting, cell selection, batch effects, doublets, and alternative trajectory methods?
Reviewed: August 2026
All commands and outputs were verified with the software versions listed in this tutorial. If you encounter reproducibility issues, please report them through the Contact page.
Author: Nasir Mahmood Abbasi, PhD · Category: Single-Cell RNA-seq