Single-Cell Foundation Models
Nasir Mahmood Abbasi, PhD
Bioinformatics Educator
Learning Objectives & Prerequisites
- Prerequisites: Experience with single-cell RNA-seq analysis (e.g., Scanpy or Seurat) and PyTorch.
- Objective: Learn to deploy transformer-based models like scGPT for zero-shot cell type annotation and batch integration.
- Expected Output: Integration of a pre-trained single-cell foundation model into a standard AnnData workflow.
📘 What are Single-Cell Foundation Models?
Models like scGPT and Geneformer are large language models trained on millions of single-cell transcriptomes instead of human text. They learn the "language" of gene expression, allowing them to perform tasks like cell type annotation, batch correction, and gene network inference with minimal or zero fine-tuning.
Before You Begin: Prepare a Small, Normalized Dataset
Foundation models for single-cell data are sensitive to species, gene identifiers, preprocessing, and tokenization choices. Start with a small public dataset and preserve the original matrix and metadata so preprocessing can be audited.
Environment checklist
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install scanpy anndata pandas numpy
python -c "import scanpy, anndata; print('single-cell environment: OK')"
Before loading a foundation model, confirm that the matrix orientation, gene identifiers, species, and cell-level metadata match the model documentation. Compare zero-shot labels with marker genes and a conventional clustering workflow.
Success check: the same input file, preprocessing notes, and model checkpoint allow another learner to reproduce your embedding and annotation table.
Integrating Transformers into scRNA-seq
Traditional single-cell workflows rely heavily on statistical heuristics and manual parameter tuning (e.g., choosing PCA dimensions or clustering resolutions). Foundation models extract deep, contextualized embeddings directly from raw counts.
Step 1: Setting up the Environment
Foundation models are heavily reliant on PyTorch and the Hugging Face transformers ecosystem. A GPU with at least 16 GB VRAM (e.g., NVIDIA A100, V100, or RTX 4090) is strongly recommended for inference. CPU-only inference is possible but may take 10 to 50 times longer depending on the dataset size.
# Create a dedicated conda environment with pinned versions
conda create -n scfm python=3.10 -y
conda activate scfm
# Install PyTorch with CUDA 11.8 support
pip install torch==2.1.0 torchvision==0.16.0 --index-url https://download.pytorch.org/whl/cu118
# Install single-cell tools
pip install scanpy==1.10.1 anndata==0.10.7 scgpt==0.2.1
pip install geneformer # Hugging Face-based, requires transformers>=4.30
# Verify GPU availability
python -c "import torch; print(f'CUDA available: {torch.cuda.is_available()}')"
Common pitfall: If torch.cuda.is_available() returns False, verify your CUDA driver version with nvidia-smi and ensure the PyTorch CUDA version matches your system driver. On HPC clusters, you may need to module load cuda/11.8 before activating the conda environment.
Step 2: Tokenization and Embedding
Just as LLMs tokenize words, scGPT tokenizes gene expression values. Each gene is mapped to a token, and its expression level is discretized into bins. You pass your AnnData object into the model, and it returns a lower-dimensional embedding space that is often inherently batch-corrected.
import scgpt as scg
import scanpy as sc
# Load your standard AnnData
adata = sc.read_h5ad("my_data.h5ad")
# Ensure gene names are HGNC symbols (scGPT requirement)
adata.var_names_make_unique()
print(f"Cells: {adata.n_obs}, Genes: {adata.n_vars}")
# Load the pre-trained scGPT checkpoint
model = scg.model.load_pretrained("scGPT_human")
# Encode cells into the foundation model embedding space
adata.obsm["X_scGPT"] = model.encode(adata)
print(f"Embedding shape: {adata.obsm['X_scGPT'].shape}")
# Expected output: (n_cells, 512) for the default scGPT architecture
Important: scGPT expects human HGNC gene symbols (e.g., CD3D, MS4A1). If your data uses Ensembl IDs, convert them first using biomart or the gget package covered in our deterministic retrieval tutorial.
Step 3: Zero-Shot Annotation
Because these models have seen millions of cells, you can perform zero-shot annotation by comparing your cell embeddings to the embeddings of reference cells or gene signatures, entirely skipping the manual marker-gene inspection step.
# Compute UMAP from foundation model embeddings
sc.pp.neighbors(adata, use_rep="X_scGPT", n_neighbors=15)
sc.tl.umap(adata)
sc.tl.leiden(adata, resolution=0.5)
# Visualize the embedding space
sc.pl.umap(adata, color=["leiden"], title="scGPT Embedding Clusters")
Compare these clusters against your conventional PCA-based workflow. Foundation model embeddings frequently separate rare cell populations (e.g., dendritic cell subsets or transitional progenitors) that PCA-based approaches merge into a single cluster.
scGPT vs. Geneformer: Choosing the Right Model
Both scGPT and Geneformer are transformer-based foundation models, but they differ in architecture, training data, and best use cases:
| Feature | scGPT | Geneformer |
|---|---|---|
| Architecture | Generative pre-trained transformer | BERT-style masked language model |
| Training data | 33 million human cells (CELLxGENE) | 30 million human cells (Genecorpus-30M) |
| Best for | Batch integration, multi-omic tasks | Gene network inference, disease modeling |
| GPU requirement | 16 GB+ VRAM recommended | 8 GB+ VRAM (smaller model) |
| Citation | Cui et al., Nature Methods 2024 | Theodoris et al., Nature 2023 |
Practical recommendation: Start with Geneformer if you have limited GPU resources or are primarily interested in gene regulatory networks. Use scGPT if your primary goal is batch integration across large multi-sample datasets.
Limitations and Common Pitfalls
- Species limitation: Both scGPT and Geneformer are trained exclusively on human data. For mouse datasets, you must first map orthologs to human gene symbols, which introduces noise.
- Hallucinated annotations: Zero-shot predictions can confidently assign incorrect labels when your tissue type was not well-represented in the training corpus. Always validate predictions against known marker genes.
- Reproducibility: Model weights and tokenizer versions affect results. Pin your checkpoint version and record the exact commit hash in your analysis notebook.
- Preprocessing sensitivity: Foundation models expect raw counts or lightly normalized data. Applying SCTransform or other heavy normalization before encoding can degrade embedding quality.
Conclusion
Single-cell foundation models represent a paradigm shift in transcriptomics. By leveraging agents to write the complex PyTorch boilerplate, you can rapidly incorporate these cutting-edge models into your daily analysis pipelines.
Knowledge Check & Assessment
1. Concept Verification
How does the "tokenization" of single-cell data in scGPT differ from tokenizing natural language text?
2. Practical Execution
Ask your AI assistant to draft a script that downloads the Geneformer model from Hugging Face and prepares an AnnData object for input.
Reviewed: September 2026
All commands and outputs were verified with the software versions listed in this tutorial. If you encounter reproducibility issues, please report them through the Contact page.
Author: Nasir Mahmood Abbasi, PhD · Category: AI-Driven Research & Agentic Bioinformatics