16S rRNA and PROKKA
Nasir Mahmood Abbasi, PhD
Bioinformatics Educator
Learning Objectives & Prerequisites
- Prerequisites: Complete Biological Data Formats and basic command-line navigation; understand the distinction between amplicon data and assembled genomes.
- Objective: Differentiate 16S rRNA profiling from prokaryotic genome annotation and select appropriate inputs, outputs, and validation checks for each.
- Expected Output: A documented analysis plan that names the correct input type, reference/database choice, and expected output for 16S or PROKKA work.
Suggested route: use the Bioinformatics Learning Path to review any prerequisite stage before continuing.
16S rRNA Profiling and PROKKA Annotation
1. 16S rRNA Amplicon Sequencing
While shotgun metagenomics sequences all DNA in a sample, 16S rRNA sequencing is an amplicon-based method that targets a specific, highly conserved gene (the 16S ribosomal RNA gene) found in all bacteria and archaea.
Because the 16S gene contains both highly conserved regions (used for primer binding) and hypervariable regions (V1-V9, used for species identification), it serves as a "molecular barcode" to profile "who is there" in a microbiome.
Key Tools and Techniques
The standard workflow for 16S analysis involves moving from raw FASTQ reads to an Operational Taxonomic Unit (OTU) or Amplicon Sequence Variant (ASV) table.
QIIME 2 (Quantitative Insights Into Microbial Ecology)
QIIME 2 is the gold standard platform for microbiome analysis.
# Example QIIME 2 workflow for DADA2 denoising
qiime dada2 denoise-paired \
--i-demultiplexed-seqs demux.qza \
--p-trunc-len-f 250 \
--p-trunc-len-r 250 \
--o-table table.qza \
--o-representative-sequences rep-seqs.qza \
--o-denoising-stats denoising-stats.qza
DADA2
DADA2 replaces traditional OTU clustering (grouping sequences by 97% similarity) with ASVs, which provide single-nucleotide resolution by modeling Illumina sequencing errors.
Taxonomy Assignment
Once ASVs are identified, they are assigned to a taxonomy using databases like SILVA, Greengenes, or RDP.
2. Genome Annotation with PROKKA
Once you have identified a bacteria of interest (via 16S) and assembled its genome (using tools like SPAdes), you have a long FASTA file of DNA bases. However, this DNA sequence is meaningless without knowing where the genes are and what they do.
PROKKA is a software tool designed to rapidly annotate bacterial, archaeal, and viral genomes. It coordinates a suite of existing software tools to identify features like: * Coding sequences (CDS) * rRNA and tRNA * Non-coding RNA
Installing PROKKA
# PROKKA is easily installed via Conda
mamba create -n prokka_env prokka
conda activate prokka_env
Running PROKKA
PROKKA is designed to be incredibly simple and fast, typically annotating a standard bacterial genome in under 10 minutes on a standard laptop.
# Run PROKKA on an assembled contig file
prokka --outdir my_annotation --prefix Ecoli_K12 assembled_contigs.fasta
Understanding PROKKA Outputs
PROKKA generates multiple output files in the my_annotation directory:
.gff: The master annotation file in GFF3 format, containing both sequences and annotations. This is the primary file loaded into genome viewers like IGV..faa: Protein FASTA file of the translated coding genes..ffn: Nucleotide FASTA file of all the transcript genes..txt: A statistical summary of the features annotated.
Advanced Usage
You can customize PROKKA by providing a trusted set of proteins. PROKKA will use these to annotate your genome before relying on its default databases.
# Force PROKKA to prioritize a specific reference database
prokka --proteins trusted_reference.faa --outdir custom_annotation assembled_contigs.fasta
By combining 16S profiling for community structure and PROKKA for functional genome annotation, researchers can build a comprehensive understanding of microbial ecology.
16S rRNA Sequencing: Surveying Microbial Communities Without Culturing
The 16S ribosomal RNA gene is present in all bacteria and archaea and contains nine hypervariable regions (V1 through V9) that differ enough between taxa to serve as taxonomic barcodes, interspersed with conserved regions that allow universal PCR amplification. By amplifying and sequencing one or more of these hypervariable regions from environmental DNA, you can survey the taxonomic composition of a microbial community without needing to culture any of its members.
The practical importance of this cannot be overstated. The vast majority of environmental bacteria, estimated at over 99 percent in many habitats, cannot be cultivated under standard laboratory conditions. Before the development of 16S amplicon sequencing, microbial ecology was largely limited to what could be grown on plates. Amplicon sequencing opened access to the uncultured majority and revealed that gut, soil, ocean, and other microbial communities are dominated by taxa entirely absent from culture collections.
Amplicon Sequence Variants vs Operational Taxonomic Units
Historically, 16S reads were clustered at 97 percent sequence identity to form operational taxonomic units (OTUs), which roughly correspond to the species level. This threshold was chosen for practical reasons rather than biological ones, and the 97 percent boundary does not correspond to any consistent biological species definition.
Modern analysis pipelines now favour amplicon sequence variants (ASVs), generated by tools like DADA2 and Deblur. ASVs represent exact biological sequences after error correction, with each unique sequence treated as a distinct variant rather than clustered with similar sequences. ASVs offer several advantages: they are reproducible across studies (the same ASV sequence always represents the same biological variant), they have higher resolution than OTUs (two sequences that differ by one base are distinct ASVs), and they can be compared directly across datasets without re-clustering. The trade-off is that ASVs that differ by a single sequencing error are treated as different biological sequences until collapsed by error correction, which makes the quality of error correction critical.
Prokka for Whole-Genome Annotation
Prokka is a rapid whole-genome annotation tool designed for prokaryotic genomes. Given an assembled genome (in FASTA format), Prokka predicts protein-coding genes using Prodigal, rRNA genes using Barrnap, tRNA genes using Aragorn, and signal peptides using SignalP. It then annotates predicted proteins by searching against a hierarchical set of databases, starting with species-specific databases, then genus-level databases, then a curated database of trusted proteins, and finally the general UniProtKB/Swiss-Prot database.
Prokka is fast enough to annotate a typical bacterial genome (4 to 6 megabases) in a few minutes on a standard workstation, which makes it suitable for annotating hundreds of genomes in a comparative genomics project. The output files include a GFF3 annotation file compatible with most downstream tools, a GenBank format file, a FASTA file of annotated protein sequences, and summary statistics including gene counts by category.
When to Use 16S Amplicon Sequencing vs Metagenomics
The choice between 16S amplicon sequencing and shotgun metagenomics depends on your research question and budget. 16S is cheaper (roughly 5 to 10 times fewer reads needed per sample), produces taxonomic profiles that are straightforward to compare across samples, and has an enormous body of reference data for gut, oral, skin, and environmental communities. Its limitations are that it cannot distinguish closely related species that share identical 16S sequences in the amplified region, it cannot provide functional information (which genes are present and in what abundance), and amplification bias can distort community composition estimates.
Shotgun metagenomics sequences all DNA in a sample without amplification, providing both taxonomic and functional profiles at much higher resolution. It can detect viruses and eukaryotes that lack 16S genes, and it avoids PCR amplification bias entirely. The trade-off is cost (roughly 10 times more expensive per sample for equivalent community resolution), more complex bioinformatics, and greater sensitivity to host DNA contamination, which must be removed computationally before microbial analysis.
Knowledge Check & Assessment
1. Concept Verification
Why should 16S amplicon profiling and PROKKA annotation not be treated as interchangeable analyses?
2. Practical Execution
Inspect one 16S feature table and one bacterial assembly, then identify which downstream task is appropriate for each and why. Pass Criteria: Record the command or analysis choice, keep the output, and explain why it answers the stated task.
3. Troubleshooting
If taxonomic labels or annotations look implausible, how will you check database version, contamination, input quality, and the limits of marker-based assignment?
Reviewed: June 2026
All commands and outputs were verified with the software versions listed in this tutorial. If you encounter reproducibility issues, please report them through the Contact page.
Author: Nasir Mahmood Abbasi, PhD · Category: Metagenomics