Bulk RNA-seq Deconvolution using scRNA-seq
Nasir Mahmood Abbasi, PhD
Bioinformatics Educator
Learning Objectives & Prerequisites
- Prerequisites: Complete scRNA-seq Basics and understand bulk RNA-seq normalization, reference quality, and cell-type definitions.
- Objective: Use a single-cell reference to estimate bulk mixture proportions and assess whether the reference covers the tissue and condition of interest.
- Expected Output: A deconvolution table with reference composition, input normalization, estimated fractions, uncertainty, and validation plan.
Suggested route: use the Bioinformatics Learning Path to review any prerequisite stage before continuing.
Cell Type Deconvolution: Bridging Bulk and Single-Cell
Introduction
Single-cell RNA-seq provides incredible resolution into cell states, but it is expensive and difficult to scale to hundreds or thousands of patients. Bulk RNA-seq is cheap and highly scalable, but the output is a "smoothie" - the gene expression is an average of all the thousands of cells in the tissue chunk.
Deconvolution is the mathematical process of taking a high-quality single-cell dataset (the reference "ingredients list") and using it to estimate the exact proportions of each cell type present in a bulk RNA-seq dataset (the "smoothie").
1. The Principle of Deconvolution
If a bulk RNA-seq sample shows high expression of CD8A and GZMB, is it because there are many CD8+ T cells, or because a few CD8+ T cells are expressing those genes at astronomically high levels?
Deconvolution algorithms like CIBERSORTx, MuSiC, or Cell2Location solve this by building a "signature matrix" from your single-cell reference, mapping exactly how much of each gene a typical cell of type X produces.
2. Using MuSiC in R
MuSiC (Multi-subject Single Cell deconvolution) is a fantastic tool because it accounts for biological variance across different subjects in your single-cell reference, rather than just taking a flat average.
Preparing the Data
library(MuSiC)
library(ExpressionSet)
# 1. Prepare your Bulk Data
# 'bulk_counts' is a matrix of genes (rows) by bulk samples (columns)
bulk.eset <- ExpressionSet(assayData = as.matrix(bulk_counts))
# 2. Prepare your Single-Cell Reference
# 'sc_counts' is the raw count matrix, 'sc_meta' has the cell types and patient IDs
sc.eset <- ExpressionSet(assayData = as.matrix(sc_counts),
phenoData = AnnotatedDataFrame(sc_meta))
Running the Deconvolution
# Estimate the cell type proportions in the bulk data
music_results <- music_prop(bulk.mtx = exprs(bulk.eset),
sc.eset = sc.eset,
clusters = 'cell_type', # The column in sc_meta with annotations
samples = 'patient_id') # The column indicating biological replicates
3. Visualizing the Results
The output of MuSiC is a matrix showing the estimated percentage of each cell type in every bulk sample. We can easily visualize this using standard ggplot2 stacked bar charts.
library(ggplot2)
library(reshape2)
# Extract the estimated proportions
estimated_props <- music_results$Est.prop.weighted
# Melt the data for ggplot
plot_data <- melt(estimated_props)
colnames(plot_data) <- c("Bulk_Sample", "Cell_Type", "Proportion")
# Create a stacked bar plot of cell type proportions across clinical samples
ggplot(plot_data, aes(x = Bulk_Sample, y = Proportion, fill = Cell_Type)) +
geom_bar(stat = "identity") +
theme_minimal() +
labs(title = "Deconvoluted Cell Type Proportions in Bulk Cohort",
y = "Relative Proportion",
x = "Patient Sample") +
theme(axis.text.x = element_text(angle = 45, hjust = 1))
Conclusion
Deconvolution is arguably the most powerful translational bioinformatics technique today. It allows you to take a 5-patient single-cell experiment, extract the signatures of a novel disease-driving cell state, and immediately scan for that state across historical 1,000-patient bulk RNA-seq clinical trials to test for survival outcomes.
What Deconvolution Is and Why It Matters
Bulk RNA sequencing measures the average gene expression across all cells in a sample. If your tumour sample contains 60 percent cancer cells, 20 percent T cells, and 20 percent fibroblasts, your bulk RNA-seq profile is a weighted average of the gene expression profiles of those three populations. A gene that is highly expressed in T cells but not in cancer cells will appear at moderate expression in the bulk profile, and without additional information you cannot distinguish this from a gene that is expressed at moderate levels in every cell.
Deconvolution attempts to reverse this mixing process, estimating the cellular composition of a bulk sample using a reference that specifies the gene expression signature of each cell type. The input is the bulk expression matrix and the reference signature matrix; the output is a proportion estimate for each cell type in each sample. This allows you to ask questions that bulk RNA-seq cannot answer on its own: does the fraction of cytotoxic T cells correlate with treatment response? Is the tumour macrophage proportion associated with prognosis?
Reference-Based vs Reference-Free Deconvolution
Reference-based methods like CIBERSORT, MuSiC, and DWLS require a signature matrix that specifies the expected expression of each marker gene in each cell type. This signature matrix can come from a published reference built from isolated cell populations, or from your own single-cell RNA-seq data from the same tissue type. The quality of the deconvolution result depends heavily on the quality and relevance of the signature matrix. A signature built from peripheral blood mononuclear cells will perform poorly when applied to tumour-infiltrating lymphocytes, which can have very different transcriptional states from their circulating counterparts.
Reference-free methods like NMF-based approaches do not require a pre-specified signature and instead learn the cell-type components directly from the data. They are useful when no appropriate reference exists, but the learned components are not labelled and must be interpreted post-hoc by examining which marker genes load strongly onto each component. This adds an interpretation step that can be subjective and requires domain knowledge.
Using Single-Cell Data as a Deconvolution Reference
The most principled approach in 2026 is to use single-cell RNA-seq data from the same tissue type as your deconvolution reference. Tools like MuSiC and SCDC can estimate cell-type proportions in bulk samples using a matched or closely related single-cell atlas as the reference. The single-cell data provides both the cell type labels and the per-cell-type expression profiles needed to build the signature matrix.
A critical practical consideration is that the single-cell reference should capture the same cellular states present in your bulk samples. If your bulk samples are from a disease condition and your single-cell reference was generated from healthy tissue, the disease-specific transcriptional states of infiltrating immune cells may not be well represented in the reference. In this situation, the deconvolution will estimate proportions that reflect the closest healthy cell type, potentially misclassifying activated or exhausted cells.
Validating Deconvolution Results
Deconvolution estimates should always be validated against an independent measurement where possible. Flow cytometry data from the same samples is the gold standard for immune cell proportions. If flow cytometry data is available for a subset of your samples, correlating the deconvolution estimates against the flow cytometry proportions for the same cell types tells you how accurate the deconvolution is in your specific data context. Correlation coefficients of 0.7 or higher between deconvolution estimates and flow cytometry across samples are generally considered acceptable.
When independent validation data is unavailable, assess internal consistency: do the estimated proportions sum to approximately one across cell types? Do biologically expected correlations hold, such as a negative correlation between tumour cell fraction and immune cell fraction? Do samples that are clinically annotated as T-cell-rich show high T cell deconvolution estimates? These sanity checks do not confirm accuracy but can detect gross failures of the deconvolution model.
Knowledge Check & Assessment
1. Concept Verification
Why can a precise-looking estimated fraction be biased when a cell type or activation state is missing from the reference?
2. Practical Execution
Run or inspect deconvolution for a small bulk dataset and compare estimated fractions with one independent expectation or marker pattern. Pass Criteria: Record the command or analysis choice, keep the output, and explain why it answers the stated task.
3. Troubleshooting
If fractions are negative, unstable, or biologically implausible, how will you inspect normalization, collinearity, reference coverage, and batch differences?
Reviewed: August 2026
All commands and outputs were verified with the software versions listed in this tutorial. If you encounter reproducibility issues, please report them through the Contact page.
Author: Nasir Mahmood Abbasi, PhD · Category: Advanced Single-Cell Analysis