Workflow Management and Containerization•2026-09-21

Snakemake & Nextflow

NM

Nasir Mahmood Abbasi, PhD

Bioinformatics Educator

Snakemake & Nextflow
Tested on: Python 3.11, R 4.3.2, Ubuntu 24.04
Last Review: 2026-08-15

Learning Objectives & Prerequisites

  • Prerequisites: Complete Advanced Shell Scripting and Conda/Mamba Environments; install one workflow manager in an isolated environment.
  • Objective: Translate manual pipeline steps into a dependency-aware Snakemake or Nextflow workflow with reproducible inputs and outputs.
  • Expected Output: A minimal workflow that runs on test data, records software requirements, and produces a declared result file.

Suggested route: use the Bioinformatics Learning Path to review any prerequisite stage before continuing.

Reproducible Bioinformatics Workflows

The Problem with Bash Scripts

When analyzing a single dataset, a simple Bash script (e.g., running FastQC, then BWA, then Samtools) works fine. However, as your projects grow to hundreds of samples, simple scripts fail: * If the pipeline crashes halfway, you have to manually figure out where to restart. * They don't automatically parallelize across High-Performance Computing (HPC) nodes. * They are difficult for other researchers to reproduce.

The solution is using a Workflow Manager. The two absolute industry standards in bioinformatics are Snakemake and Nextflow.


1. Snakemake: Python-Based Pipelines

Snakemake is built on top of Python. If you already know Python, the syntax will feel incredibly familiar. It uses a "make-like" logic: you define the final output files you want, and Snakemake works backward to find the rules required to create them.

Basic Structure of a Snakefile

You define pipelines in a file called Snakefile.

# A simple Snakemake rule to align reads using BWA
rule bwa_map:
    input:
        ref="genome.fa",
        reads="data/samples/{sample}.fastq.gz"
    output:
        "mapped_reads/{sample}.bam"
    threads: 8
    shell:
        "bwa mem -t {threads} {input.ref} {input.reads} | samtools view -Sb - > {output}"

Running Snakemake

# Run the pipeline locally using 16 cores
snakemake --cores 16

# Submit the pipeline to an HPC cluster (Slurm)
snakemake --profile slurm

Snakemake automatically tracks which samples have been processed. If sample 5 fails, you just re-run the exact same command, and Snakemake will only process sample 5!


2. Nextflow: Enterprise-Grade Scalability

Nextflow uses a Groovy-based Domain Specific Language (DSL2). While the learning curve is slightly steeper than Snakemake, Nextflow is the backbone of massive institutional pipelines (such as the nf-core project).

Nextflow excels at "dataflow" programming. You define processes, and data flows between them through channels.

Basic Structure of a Nextflow Script (main.nf)

// Define a process
process BWA_ALIGN {
    cpus 8

    input:
    tuple val(sample_id), path(reads)
    path reference

    output:
    path "${sample_id}.bam"

    script:
    """
    bwa mem -t ${task.cpus} $reference $reads | samtools view -Sb - > ${sample_id}.bam
    """
}

// Define the workflow pipeline
workflow {
    read_ch = Channel.fromFilePairs('data/samples/*_{1,2}.fastq.gz')
    ref_ch  = file('genome.fa')

    BWA_ALIGN(read_ch, ref_ch)
}

nf-core: The True Power of Nextflow

The biggest advantage of Nextflow is nf-core: a massive community repository of highly curated, peer-reviewed pipelines. Instead of writing your own RNA-seq pipeline, you can simply run the community standard:

# Run the gold-standard nf-core RNA-seq pipeline directly from GitHub
nextflow run nf-core/rnaseq -profile docker --input samplesheet.csv --outdir results/

Summary

  • Use Snakemake if you want to quickly wrap your existing Python/Bash scripts into a robust pipeline.
  • Use Nextflow if you are building enterprise-level pipelines, or if you want to utilize the incredible pre-built nf-core pipelines.

Why Shell Scripts Are Not Enough

Almost every bioinformatics project begins with a shell script. You write a few lines to trim reads, align to a reference, and count features. The script works. You run it for your 10 samples and move on to analysis. Six months later, a collaborator sends you 40 new samples and asks you to run the same pipeline. You re-run your script, and it fails partway through because one sample has a different file naming convention, one step was already completed and the output files exist but are incomplete, and you cannot easily tell which samples finished successfully.

Workflow managers like Snakemake and Nextflow solve these problems by separating the description of the pipeline from the execution logic. You define rules or processes that specify inputs, outputs, and the command to run. The workflow manager handles dependency resolution, partial re-runs, job scheduling on HPC clusters, and parallel execution automatically. You never manually manage which samples need to be re-processed; the tool determines this by checking which output files exist and which are missing or out of date.

Snakemake: A Natural Fit for Python Users

Snakemake is written in and extends Python. If you already write Python analysis scripts, the syntax for defining rules feels familiar. Each rule specifies an input function, an output file pattern, and a shell command or Python function. Snakemake resolves the order in which rules must run by working backwards from the target output files: it asks what files are needed to produce the final output, then what files are needed to produce those files, and so on, building a directed acyclic graph of jobs.

One of Snakemake's practical advantages is its integration with Conda. You can specify a Conda environment file for each rule, and Snakemake will create and activate that environment automatically when the rule runs. This means your STAR alignment step uses a pinned STAR version and your DESeq2 step uses a pinned Bioconductor version, with no manual environment management required. Combined with the --use-singularity flag, you can run each rule inside its own container for maximum isolation.

Nextflow: Portability Across Compute Environments

Nextflow uses a dataflow programming model where each process consumes from input channels and emits to output channels. This is conceptually different from Snakemake's file-based dependency graph. The practical consequence is that Nextflow pipelines are often more portable across compute environments: the same Nextflow script can run locally, on AWS Batch, on Google Cloud, on Azure, or on an HPC cluster with SLURM or PBS by changing a configuration profile. No changes to the pipeline logic are required.

The nf-core community has built a library of over 100 peer-reviewed, community-maintained Nextflow pipelines covering standard bioinformatics workflows including RNA-seq, ChIP-seq, methylation, metagenomics, and single-cell RNA-seq. Rather than building a pipeline from scratch, you can start with an nf-core pipeline and modify it for your specific use case. Each nf-core pipeline includes automated testing, documentation, and a consistent parameter interface, which significantly reduces the time from raw data to analysed results.

Which Tool Should You Choose?

The choice between Snakemake and Nextflow often comes down to team familiarity and compute environment. If your team works primarily in Python and runs jobs on a single HPC cluster, Snakemake is easier to learn and deploy. If your team needs pipelines that run across multiple cloud providers or you want to leverage the nf-core ecosystem, Nextflow is the better investment. Both tools are actively maintained, well-documented, and widely used in the bioinformatics community.

The most important thing is to use either one consistently rather than maintaining a collection of shell scripts. Even a basic Snakemake workflow with three rules is more reproducible, more auditable, and easier to extend than a script that chains commands together in a single file. Start small, add rules as you add pipeline steps, and version control your workflow alongside your analysis code.

Knowledge Check & Assessment

1. Concept Verification

What problem does a directed acyclic workflow graph solve that a sequence of manually run shell commands does not?

2. Practical Execution

Build and run a two-step test workflow that transforms an input file and records the result in a separate output directory. Pass Criteria: Record the command or analysis choice, keep the output, and explain why it answers the stated task.

3. Troubleshooting

If a workflow reruns unexpectedly or cannot find an output, how will you inspect rule inputs, wildcards, file timestamps, and the execution graph?

Reviewed: June 2026

All commands and outputs were verified with the software versions listed in this tutorial. If you encounter reproducibility issues, please report them through the Contact page.

Author: Nasir Mahmood Abbasi, PhD · Category: Workflow Management and Containerization

Continue Learning

Course Sequence