AI-Driven Research & Agentic Bioinformatics•2026-08-27

Mastering Prompt Engineering for Seurat & Scanpy

NM

Nasir Mahmood Abbasi, PhD

Bioinformatics Educator

Prompt Engineering for Data Analysis
Tested on: Cursor, Aider
Last Review: 2026-09-08

Learning Objectives & Prerequisites

  • Prerequisites: Basic understanding of single-cell RNA-seq workflows and an AI coding environment.
  • Objective: Learn how to structure prompts to direct AI to write robust Seurat and Scanpy analysis code.
  • Expected Output: The ability to use "vibe coding" effectively without relying on rote memorization of API syntax.

📘 Intent-Based Prompting (Vibe Coding)

The key to effective AI code generation in bioinformatics is shifting from asking "How do I use this function?" to stating "What is my biological goal?" Provide context, specific constraints, and the desired output format.

Before You Begin: Establish a Reproducible Prompting Loop

Prompt engineering works best when the task, input schema, constraints, and acceptance test are explicit. The model should propose code; you should run it, inspect the output, and keep the final script under version control.

Use this prompt template

Goal: [one biological question]
Input: [file path, shape, identifiers, and metadata]
Constraints: [language, packages, and no destructive changes]
Output: [file, table, plot, and interpretation]
Checks: [expected dimensions, labels, and failure conditions]

Ask for one small change at a time. Require the assistant to explain assumptions and cite package functions. Never execute generated code on sensitive data before reviewing file access, network calls, and destructive commands.

Success check: a second person can run the saved prompt and script on the same toy dataset and obtain the same expected checks.

The Anatomy of a Great Bioinformatics Prompt

When asking Cursor or Aider to write analysis code, your prompt should contain three core elements to ensure the AI generates production-ready scripts rather than generic examples:

  1. Context: What data are you working with? (e.g., "I have a Seurat object `pbmc` with 3,000 cells")
  2. Intent: What is the biological goal? (e.g., "Identify marker genes for cluster 2 vs cluster 5")
  3. Constraints: What statistical or technical thresholds apply? What visualizations do you need? (e.g., "Use a Wilcoxon test, only return genes with an adjusted p-value < 0.01 and a log2FC > 1")

Practical Examples

Example 1: Scanpy Quality Control

Bad Prompt: "Write code to filter my scanpy object."

This prompt is too vague. The AI doesn't know your dataset name, your thresholds, or if you want visualizations.

Good Prompt: "I have an AnnData object loaded as `adata`. Write Scanpy code to perform quality control. Filter out cells with fewer than 200 expressed genes, more than 8000 expressed genes, and higher than 5% mitochondrial reads. Then, filter out genes expressed in fewer than 3 cells. Plot the violin plots for these metrics before filtering, saving them as `qc_violins.png`."

Resulting AI Code:

import scanpy as sc
import matplotlib.pyplot as plt

# Calculate QC metrics
adata.var['mt'] = adata.var_names.str.startswith('MT-')
sc.pp.calculate_qc_metrics(adata, qc_vars=['mt'], percent_top=None, log1p=False, inplace=True)

# Plot before filtering
sc.pl.violin(adata, ['n_genes_by_counts', 'total_counts', 'pct_counts_mt'],
             jitter=0.4, multi_panel=True, show=False)
plt.savefig('qc_violins.png')

# Filter cells and genes based on constraints
sc.pp.filter_cells(adata, min_genes=200)
sc.pp.filter_cells(adata, max_genes=8000)
adata = adata[adata.obs.pct_counts_mt < 5, :]
sc.pp.filter_genes(adata, min_cells=3)

Example 2: Seurat Differential Expression

Bad Prompt: "Find markers for my seurat clusters."

Good Prompt: "Using my Seurat object `seurat_obj`, find the differentially expressed genes for cluster 3 compared to all other clusters. Use the MAST test. Only keep genes with an adjusted p-value < 0.05 and a logfc threshold of 0.5. Save the resulting dataframe to a CSV file named `cluster3_markers.csv`."

Resulting AI Code:

library(Seurat)

# Find markers using MAST
cluster3_markers <- FindMarkers(seurat_obj, 
                               ident.1 = 3, 
                               test.use = "MAST", 
                               logfc.threshold = 0.5)

# Filter by adjusted p-value
significant_markers <- subset(cluster3_markers, p_val_adj < 0.05)

# Save to CSV
write.csv(significant_markers, file = "cluster3_markers.csv", row.names = TRUE)

Iterative Prompting (Conversational Refinement)

Do not expect perfect plots on the first try. Use conversational turns to refine the aesthetics.

  • "Change the color palette of that UMAP to use viridis."
  • "Add a title to the plot saying 'CD4 T Cell Subpopulations'."
  • "The legend is overlapping the data. Move the legend to the right side outside the plot area."

✅ Key Takeaways

  • Always provide variable names: Tell the AI what your AnnData or Seurat object is named.
  • Specify constraints explicitly: Don't leave p-values or logFC thresholds to chance.
  • Request file outputs: Tell the AI exactly what to name the output PNGs or CSVs.
  • Iterate on visuals: Use follow-up prompts to refine ggplot or matplotlib aesthetics.

Knowledge Check & Assessment

1. Concept Verification

Why is "Write code to filter my scanpy object" considered a bad prompt?

2. Practical Execution

Write a prompt asking Cursor to generate a FeaturePlot (Seurat) or umap plot (Scanpy) for the genes CD3E, CD4, and CD8A, specifying a custom color gradient.

Reviewed: September 2026

All commands and outputs were verified with the software versions listed in this tutorial. If you encounter reproducibility issues, please report them through the Contact page.

Author: Nasir Mahmood Abbasi, PhD · Category: AI-Driven Research & Agentic Bioinformatics

Continue Learning

Course Sequence