Immune Repertoire Analysis
Nasir Mahmood Abbasi, PhD
Bioinformatics Educator
Learning Objectives & Prerequisites
- Prerequisites: Complete scRNA-seq Basics and understand TCR/BCR clonotypes, paired chains, and appropriate study-design metadata.
- Objective: Analyze immune-receptor repertoire features while distinguishing clone abundance, diversity, expansion, and antigen specificity claims.
- Expected Output: A repertoire summary with clonotype definition, chain handling, diversity metric, sample denominator, and cautious interpretation.
Suggested route: use the Bioinformatics Learning Path to review any prerequisite stage before continuing.
Immune Repertoire Analysis
Introduction
In immunology and oncology, tracking the expansion of specific T-cells or B-cells is crucial. Every T-cell has a unique T-Cell Receptor (TCR) sequence created by V(D)J recombination. When an immune cell recognizes a specific antigen (like a virus or a tumor), it rapidly clones itself.
By sequencing the TCRs at the single-cell level (scTCR-seq), we can measure clonal expansion. This is especially critical in T-cell lymphomas, such as Sézary Syndrome, where a single malignant T-cell clone expands uncontrollably.
This tutorial covers the basics of analyzing V(D)J single-cell data using the scRepertoire package in R.
1. Preparing V(D)J Data
Typically, 10x Genomics Cell Ranger provides a filtered_contig_annotations.csv file containing the TCR Alpha and Beta chain sequences for every cell barcode.
library(scRepertoire)
library(Seurat)
# Load the Cell Ranger V(D)J output
contig_data <- read.csv("filtered_contig_annotations.csv")
# Convert the contig list into a combined TCR format for scRepertoire
combined_tcr <- combineTCR(contig_data,
samples = "Patient1",
ID = "Timepoint1",
cells = "T-AB")
2. Visualizing Clonal Expansion
Once the TCR chains are combined, we can visualize the diversity of the immune repertoire. In a healthy patient, the repertoire is highly diverse. In a Sézary Syndrome patient, you will see massive clonal dominance.
# Plot the relative abundance of the top clones
quantContig(combined_tcr, cloneCall="gene+nt", scale = TRUE)
# Visualize clonal space homeostasis (the proportion of the repertoire occupied by rare vs hyperexpanded clones)
clonalHomeostasis(combined_tcr, cloneCall="gene+nt")
3. Integrating TCR Data with Seurat (scRNA-seq)
The true power of modern immune profiling is linking the identity of the clone (the TCR) with its phenotype (the RNA expression). We can overlay the clonal expansion data directly onto a Seurat UMAP.
# Assume 'seurat_obj' is your standard processed scRNA-seq object
# Merge the TCR data into the Seurat metadata using the cell barcodes
seurat_obj <- combineExpression(combined_tcr, seurat_obj, cloneCall="gene+nt")
# The Seurat object now has a 'cloneType' column!
# We can visualize where the highly expanded clones sit on the UMAP
DimPlot(seurat_obj, group.by = "cloneType") + scale_color_viridis_d()
Finding Markers of Malignant Clones
Once integrated, you can easily find the biological markers driving the disease by comparing the hyperexpanded clone against all other normal T-cells.
# Set the identity to the clonal grouping
Idents(seurat_obj) <- "cloneType"
# Find genes upregulated in the 'Hyperexpanded' clone compared to 'Rare' clones
malignant_markers <- FindMarkers(seurat_obj, ident.1 = "Hyperexpanded", ident.2 = "Rare")
head(malignant_markers)
This workflow forms the bioinformatics foundation for identifying targetable biomarkers in cutaneous T-cell lymphomas and other immune-driven diseases.
Matched Python and R clonotype workflow
Define clonotypes before calculating expansion. Record whether the definition uses nucleotide or amino-acid sequence, which receptor chains are required, and how dual chains are handled.
import scirpy as ir
ir.pp.index_chains(mdata)
ir.tl.chain_qc(mdata)
ir.pp.ir_dist(mdata)
ir.tl.define_clonotypes(mdata, receptor_arms="all", dual_ir="primary_only")
ir.tl.clonal_expansion(mdata)
library(scRepertoire)
combined_tcr <- combineTCR(contig_list, samples = sample_ids, cells = "T-AB")
combined_tcr <- combineExpression(
combined_tcr,
seurat_obj,
cloneCall = "aa",
group.by = "sample"
)
clonalProportion(combined_tcr, cloneCall = "aa", group.by = "sample")
The Biology Behind T Cell and B Cell Receptor Diversity
The adaptive immune system's ability to recognise virtually any antigen depends on the extraordinary diversity of T cell receptors and B cell receptors. This diversity is generated during lymphocyte development through V(D)J recombination, a somatic process in which gene segments are randomly joined with additional nucleotide insertions and deletions at the junctions. The result is that each lymphocyte carries a unique receptor sequence, and the total number of possible TCR or BCR sequences in a single individual exceeds 10 to the power of 15.
In practice, the circulating repertoire of any individual contains between 10 million and 100 million distinct clonotypes, each representing a unique receptor sequence present on one or more cells. When an antigen is encountered, cells with receptors that bind that antigen proliferate through clonal expansion, producing a large population of identical cells from the same precursor. Repertoire analysis measures the frequency of each clonotype in a sample, allowing you to detect clonal expansion, track antigen-specific responses, and compare repertoire diversity across conditions or time points.
Sequencing Strategies: Bulk vs Single-Cell
Bulk repertoire sequencing using targeted amplification of V(D)J gene segments provides deep coverage of clonotype frequencies at relatively low cost per sample. Tools like MiXCR, TRUST4, and VDJtools assemble clonotypes from bulk reads and estimate their frequencies. The limitation is that you cannot link a TCR sequence to a cell's full transcriptome or to a BCR sequence in the same cell.
Single-cell V(D)J sequencing, available through the 10x Genomics Chromium platform and others, sequences the full receptor chain sequences from individual cells alongside their whole-transcriptome gene expression profiles. This allows you to ask which transcriptional states are enriched in expanded clones, whether clonally related cells have diverged in their transcriptional programs, and whether paired heavy and light chain sequences or paired alpha and beta chain sequences share features across related clones. The cost is much higher per cell than bulk sequencing, and the clonotype depth is lower, but the paired information is unique.
Diversity Metrics: What They Measure and What They Do Not
Repertoire diversity is typically summarised using metrics borrowed from ecology, including Shannon entropy, Simpson's index, and the Chao1 estimator. Shannon entropy captures both richness (number of distinct clonotypes) and evenness (how uniformly they are distributed). A sample dominated by one or two expanded clones will have low Shannon entropy even if the total number of distinct clonotypes is large. This is the expected pattern in a strong antigen-specific response.
A critical technical consideration is sequencing depth. Diversity estimates from shallow sequencing are strongly influenced by sampling artefacts: rare clonotypes that are present in the repertoire may simply not be sequenced, leading to underestimated richness. Rarefaction, which involves subsampling all samples to the same number of reads before calculating diversity metrics, is mandatory for fair comparisons across samples with different sequencing depths. Never compare Shannon entropy values calculated from samples with 10-fold differences in read depth without rarefaction.
Clonotype Tracking Across Time Points and Tissues
One of the most powerful applications of repertoire analysis is tracking the same clonotype across multiple time points or tissue compartments. In cancer immunology, you can compare the TCR repertoire in tumour-infiltrating lymphocytes to the peripheral blood to determine whether tumour-specific clones are present in the blood at detectable frequencies. In infectious disease, you can track how the repertoire changes before, during, and after pathogen clearance to identify clones that expand and then contract after resolution.
Clonotype matching across samples requires careful consideration of the matching criteria. Two sequences are considered the same clonotype if their CDR3 amino acid sequences, V gene usage, and J gene usage all match. Matching on nucleotide sequence is more stringent and identifies cells from the same precursor with certainty. Matching on amino acid sequence is more permissive and groups functionally similar but independently arising clonotypes. Your choice should depend on whether you want to track identical cells or functionally convergent responses.
Knowledge Check & Assessment
1. Concept Verification
Why does clonal expansion not by itself identify antigen specificity or functional state?
2. Practical Execution
Calculate or inspect clone-size distributions for two samples and report the clonotype definition and normalization used. Pass Criteria: Record the command or analysis choice, keep the output, and explain why it answers the stated task.
3. Troubleshooting
If one sample appears oligoclonal, how will you check sequencing depth, cell recovery, doublets, chain pairing, and technical batch effects?
Reviewed: August 2026
All commands and outputs were verified with the software versions listed in this tutorial. If you encounter reproducibility issues, please report them through the Contact page.
Author: Nasir Mahmood Abbasi, PhD · Category: Advanced Single-Cell Analysis