Mastering Prompt Engineering for Seurat & Scanpy
Nasir Mahmood Abbasi, PhD
Bioinformatics Educator
Learning Objectives & Prerequisites
- Prerequisites: Basic understanding of single-cell RNA-seq workflows and an AI coding environment.
- Objective: Learn how to structure prompts to direct AI to write robust Seurat and Scanpy analysis code.
- Expected Output: The ability to use "vibe coding" effectively without relying on rote memorization of API syntax.
📘 Intent-Based Prompting (Vibe Coding)
The key to effective AI code generation in bioinformatics is shifting from asking "How do I use this function?" to stating "What is my biological goal?" Provide context, specific constraints, and the desired output format.
Before You Begin: Establish a Reproducible Prompting Loop
Prompt engineering works best when the task, input schema, constraints, and acceptance test are explicit. The model should propose code; you should run it, inspect the output, and keep the final script under version control.
Use this prompt template
Goal: [one biological question]
Input: [file path, shape, identifiers, and metadata]
Constraints: [language, packages, and no destructive changes]
Output: [file, table, plot, and interpretation]
Checks: [expected dimensions, labels, and failure conditions]
Ask for one small change at a time. Require the assistant to explain assumptions and cite package functions. Never execute generated code on sensitive data before reviewing file access, network calls, and destructive commands.
Success check: a second person can run the saved prompt and script on the same toy dataset and obtain the same expected checks.
The Anatomy of a Great Bioinformatics Prompt
When asking Cursor or Aider to write analysis code, your prompt should contain three core elements to ensure the AI generates production-ready scripts rather than generic examples:
- Context: What data are you working with? (e.g., "I have a Seurat object `pbmc` with 3,000 cells")
- Intent: What is the biological goal? (e.g., "Identify marker genes for cluster 2 vs cluster 5")
- Constraints: What statistical or technical thresholds apply? What visualizations do you need? (e.g., "Use a Wilcoxon test, only return genes with an adjusted p-value < 0.01 and a log2FC > 1")
Practical Examples
Example 1: Scanpy Quality Control
Bad Prompt: "Write code to filter my scanpy object."
This prompt is too vague. The AI doesn't know your dataset name, your thresholds, or if you want visualizations.
Good Prompt: "I have an AnnData object loaded as `adata`. Write Scanpy code to perform quality control. Filter out cells with fewer than 200 expressed genes, more than 8000 expressed genes, and higher than 5% mitochondrial reads. Then, filter out genes expressed in fewer than 3 cells. Plot the violin plots for these metrics before filtering, saving them as `qc_violins.png`."
Resulting AI Code:
import scanpy as sc
import matplotlib.pyplot as plt
# Calculate QC metrics
adata.var['mt'] = adata.var_names.str.startswith('MT-')
sc.pp.calculate_qc_metrics(adata, qc_vars=['mt'], percent_top=None, log1p=False, inplace=True)
# Plot before filtering
sc.pl.violin(adata, ['n_genes_by_counts', 'total_counts', 'pct_counts_mt'],
jitter=0.4, multi_panel=True, show=False)
plt.savefig('qc_violins.png')
# Filter cells and genes based on constraints
sc.pp.filter_cells(adata, min_genes=200)
sc.pp.filter_cells(adata, max_genes=8000)
adata = adata[adata.obs.pct_counts_mt < 5, :]
sc.pp.filter_genes(adata, min_cells=3)
Example 2: Seurat Differential Expression
Bad Prompt: "Find markers for my seurat clusters."
Good Prompt: "Using my Seurat object `seurat_obj`, find the differentially expressed genes for cluster 3 compared to all other clusters. Use the MAST test. Only keep genes with an adjusted p-value < 0.05 and a logfc threshold of 0.5. Save the resulting dataframe to a CSV file named `cluster3_markers.csv`."
Resulting AI Code:
library(Seurat)
# Find markers using MAST
cluster3_markers <- FindMarkers(seurat_obj,
ident.1 = 3,
test.use = "MAST",
logfc.threshold = 0.5)
# Filter by adjusted p-value
significant_markers <- subset(cluster3_markers, p_val_adj < 0.05)
# Save to CSV
write.csv(significant_markers, file = "cluster3_markers.csv", row.names = TRUE)
Iterative Prompting (Conversational Refinement)
Do not expect perfect plots on the first try. Use conversational turns to refine the aesthetics.
- "Change the color palette of that UMAP to use viridis."
- "Add a title to the plot saying 'CD4 T Cell Subpopulations'."
- "The legend is overlapping the data. Move the legend to the right side outside the plot area."
✅ Key Takeaways
- Always provide variable names: Tell the AI what your AnnData or Seurat object is named.
- Specify constraints explicitly: Don't leave p-values or logFC thresholds to chance.
- Request file outputs: Tell the AI exactly what to name the output PNGs or CSVs.
- Iterate on visuals: Use follow-up prompts to refine ggplot or matplotlib aesthetics.
Knowledge Check & Assessment
1. Concept Verification
Why is "Write code to filter my scanpy object" considered a bad prompt?
2. Practical Execution
Write a prompt asking Cursor to generate a FeaturePlot (Seurat) or umap plot (Scanpy) for the genes CD3E, CD4, and CD8A, specifying a custom color gradient.
Reviewed: September 2026
All commands and outputs were verified with the software versions listed in this tutorial. If you encounter reproducibility issues, please report them through the Contact page.
Author: Nasir Mahmood Abbasi, PhD · Category: AI-Driven Research & Agentic Bioinformatics