Inferring Copy Number Variation (inferCNV)
Nasir Mahmood Abbasi, PhD
Bioinformatics Educator
Learning Objectives & Prerequisites
- Prerequisites: Complete scRNA-seq Basics and cell-type annotation; use matched reference cells and non-clinical training data where appropriate.
- Objective: Infer broad copy-number patterns from expression while understanding reference selection, smoothing, and the need for orthogonal validation.
- Expected Output: A CNV inference figure with reference-cell rationale, genomic patterns, sample context, and an explicit non-diagnostic interpretation.
Suggested route: use the Bioinformatics Learning Path to review any prerequisite stage before continuing.
Inferring Copy Number Variation with inferCNV
Introduction
In oncology and cancer single-cell RNA-seq, distinguishing malignant tumor cells from normal, healthy cells in the microenvironment is the most critical first step. Because cancer is fundamentally driven by genomic instability, malignant cells often have massive chromosomal amplifications or deletions.
inferCNV is a powerful R package that uses single-cell RNA expression as a proxy for DNA copy number variation (CNV). By comparing the expression of genes across the genome in a suspect cell against a known "normal" reference cell, it can identify these large-scale chromosomal alterations.
1. Preparing Data for inferCNV
You need three inputs for inferCNV: 1. A raw count matrix. 2. A metadata file mapping each cell to its annotation (e.g., "Tumor", "Normal_T_cell", "Normal_B_cell"). 3. A gene ordering file mapping each gene to its physical position on the chromosome.
library(infercnv)
# Assuming you have extracted counts from your Seurat object
raw_counts <- GetAssayData(seurat_obj, slot = "counts")
# Create the inferCNV object
infercnv_obj <- CreateInfercnvObject(
raw_counts_matrix = raw_counts,
annotations_file = "cell_annotations.txt",
delim = "\t",
gene_order_file = "gene_ordering_hg38.txt",
ref_group_names = c("Normal_T_cell", "Normal_B_cell") # Define the healthy reference!
)
2. Running the inferCNV Pipeline
inferCNV applies a series of smoothing steps, moving averages, and hidden Markov models (HMM) to denoise the RNA data and uncover the true DNA copy number signal.
# Run the core inferCNV algorithm
infercnv_obj <- infercnv::run(
infercnv_obj,
cutoff = 0.1, # 0.1 for 10x Genomics, 1 for Smart-seq2
out_dir = "infercnv_output/",
cluster_by_groups = TRUE,
denoise = TRUE,
HMM = TRUE,
num_threads = 8
)
Key Considerations
- Cutoff: Use
0.1for sparse data like 10x Genomics, and1.0for full-length methods like Smart-seq2. - Denoising: Always use
denoise = TRUEto remove background noise using the residual distributions. - HMM: The Hidden Markov Model predicts the actual integer copy number (e.g., 2 copies, 3 copies, or complete deletion).
3. Interpreting the Output
inferCNV automatically generates a heatmap in your output directory (infercnv_output/infercnv.webp).
- Rows are individual cells.
- Columns are genes, ordered strictly by their physical location from Chromosome 1 to Chromosome X/Y.
- Colors: Red indicates chromosomal amplification (e.g., Trisomy). Blue indicates chromosomal deletion.
Large, coordinated shifts across chromosome arms can be consistent with copy-number alteration and may help prioritize clusters for follow-up. They do not, by themselves, prove that cells are malignant: expression programs, technical effects, an unsuitable reference, or subclonal structure can all influence the pattern. A neutral-looking reference is useful, but it is not a guarantee of a diploid genome.
Treat expression-derived CNV inference as a screening and hypothesis-generation step. Before reporting a malignant population, validate the result with orthogonal DNA-based evidence when possible, such as matched bulk or single-cell DNA sequencing, shallow whole-genome sequencing, FISH, or clinically reviewed genomic data. Interpret the result together with pathology, sample context, QC, and canonical lineage markers.
Matched Python and R CNV-inference workflow
Define normal-reference cells before inference and interpret broad expression-derived CNV patterns cautiously. These methods do not replace DNA-based copy-number validation.
import infercnvpy as cnv
cnv.tl.infercnv(
adata,
reference_key="cell_type",
reference_cat=["B cell", "T cell"],
window_size=100,
key_added="cnv",
)
cnv_matrix = adata.obsm["X_cnv"]
library(infercnv)
infercnv_obj <- CreateInfercnvObject(
raw_counts_matrix = counts_matrix,
annotations_file = "cell_annotations.tsv",
delim = "\t",
gene_order_file = "gene_order.tsv",
ref_group_names = c("B cell", "T cell")
)
infercnv_obj <- infercnv::run(
infercnv_obj,
cutoff = 0.1,
out_dir = "infercnv_output",
cluster_by_groups = TRUE
)
Knowledge Check & Assessment
1. Concept Verification
Why is expression-derived CNV inference not a substitute for validated DNA-based copy-number measurement?
2. Practical Execution
Run or inspect an inferCNV output, identify the reference population, and report one broad pattern with its limitations. Pass Criteria: Record the command or analysis choice, keep the output, and explain why it answers the stated task.
3. Troubleshooting
If a broad CNV signal appears in many cell types, how will you investigate reference quality, normalization, cell-cycle effects, and technical artifacts?
Reviewed: September 2026
All commands and outputs were verified with the software versions listed in this tutorial. If you encounter reproducibility issues, please report them through the Contact page.
Author: Nasir Mahmood Abbasi, PhD · Category: Advanced Single-Cell Analysis