Advanced Shell Scripting
Nasir Mahmood Abbasi, PhD
Bioinformatics Educator
Learning Objectives & Prerequisites
- Prerequisites: Complete Basic Navigation and Text Processing; work in a disposable test directory.
- Objective: Write robust shell commands with awk, redirection, permissions, process control, compression, and documented pipeline steps.
- Expected Output: A commented Bash script that summarizes a small FASTQ or annotation file and writes a reproducible output table.
Suggested route: use the Bioinformatics Learning Path to review any prerequisite stage before continuing.
Advanced Text Processing with awk
awk is a powerful programming language built into Unix systems:
Basic awk Patterns
awk '{print $1}' file.txt # Print first column
awk '{print $1, $3}' file.txt # Print columns 1 and 3
awk '{print NF}' file.txt # Print number of fields
awk '{print NR, $0}' file.txt # Print line numbers
Conditional Processing
awk '$3 > 50' data.txt # Print lines where column 3 > 50
awk '$1 == "gene"' data.txt # Print lines where column 1 equals "gene"
awk 'length($0) > 80' file.txt # Print lines longer than 80 characters
Mathematical Operations
awk '{sum += $2} END {print sum}' numbers.txt # Sum column 2
awk '{print $1, $2*2}' data.txt # Multiply column 2 by 2
awk '{avg = ($2+$3)/2; print $1, avg}' data.txt # Calculate average
Pipes and Redirection: Connecting the Pieces
Pipes (|)
Pipes connect the output of one command to the input of another:
cat sequences.fasta | grep ">" | wc -l # Count sequences
ls -l | grep "\.fastq" | wc -l # Count FASTQ files
sort data.txt | uniq -c | sort -nr # Sort, count, sort by count
Redirection
Output redirection (> and >>)
ls > file_list.txt # Write output to file (overwrite)
ls >> file_list.txt # Append output to file
grep "error" log.txt > errors.txt # Save errors to file
Input redirection (<)
sort < unsorted.txt # Use file as input
wc -l < sequences.fasta # Count lines from file
Combining Commands
# Complex bioinformatics pipeline
cat *.fastq | \
grep -A 1 "^@" | \
grep -v "^@" | \
grep -v "^--" | \
awk 'length($0) > 50' | \
wc -l
File Permissions and Ownership
Understanding Permissions
When you run ls -l, you see something like:
-rw-r--r-- 1 user group 1024 Jan 15 10:30 file.txt
This breaks down as:
- -: File type (- for file, d for directory)
- rw-r--r--: Permissions (owner, group, others)
- 1: Number of links
- user: Owner
- group: Group
- 1024: File size
- Jan 15 10:30: Last modified
- file.txt: Filename
Permission Types
r(read): Can view file contentsw(write): Can modify filex(execute): Can run file as program
Changing Permissions (chmod)
chmod +x script.sh # Make script executable
chmod 755 script.sh # rwxr-xr-x
chmod 644 data.txt # rw-r--r--
chmod -R 755 directory/ # Apply to directory recursively
Process Management
Viewing Running Processes
ps # Show your processes
ps aux # Show all processes
top # Interactive process viewer
htop # Better interactive viewer (if installed)
Background Processes
long_command & # Run in background
nohup long_command & # Run in background, ignore hangup
jobs # Show background jobs
fg %1 # Bring job 1 to foreground
bg %1 # Send job 1 to background
Killing Processes
kill PID # Kill process by ID
kill -9 PID # Force kill process
killall process_name # Kill all processes by name
Working with Compressed Files
Compression and Decompression
gzip file.txt # Compress file
gunzip file.txt.gz # Decompress file
tar -czf archive.tar.gz files/ # Create compressed archive
tar -xzf archive.tar.gz # Extract compressed archive
Working with Compressed Files Directly
zcat file.txt.gz | head # View compressed file
zgrep "pattern" file.txt.gz # Search in compressed file
zless file.txt.gz # View compressed file interactively
Environment Variables and PATH
Viewing Environment Variables
echo $HOME # Show home directory
echo $PATH # Show executable search path
env # Show all environment variables
Setting Environment Variables
export MYVAR="value" # Set variable for session
echo 'export MYVAR="value"' >> ~/.bashrc # Set permanently
Command History and Shortcuts
History Commands
history # Show command history
!123 # Run command 123 from history
!! # Run last command
!grep # Run last command starting with grep
Keyboard Shortcuts
Ctrl+C: Cancel current commandCtrl+Z: Suspend current commandCtrl+A: Go to beginning of lineCtrl+E: Go to end of lineCtrl+U: Clear line before cursorCtrl+K: Clear line after cursorTab: Auto-complete↑/↓: Navigate command history
Best Practices for Bioinformatics
1. Organize Your Files
project/
├── data/
│ ├── raw/
│ └── processed/
├── scripts/
├── results/
└── docs/
2. Use Descriptive Filenames
# Good
sample_01_quality_filtered.fastq
alignment_results_2024_01_15.sam
# Bad
file1.txt
output.txt
3. Document Your Commands
# Keep a log of important commands
echo "$(date): Started quality control" >> analysis.log
fastqc *.fastq >> analysis.log 2>&1
4. Test Commands on Small Datasets
# Test on first 1000 lines
head -n 1000 large_file.fastq | your_command
5. Use Version Control
git init # Initialize repository
git add script.sh # Add file to staging
git commit -m "Added QC script" # Commit changes
Troubleshooting Common Issues
Command Not Found
which command_name # Check if command exists
echo $PATH # Check search path
Permission Denied
ls -l file.txt # Check permissions
chmod +x script.sh # Make executable
File Not Found
ls -la # Check if file exists
pwd # Verify current directory
Out of Disk Space
df -h # Check disk usage
du -sh * # Check directory sizes
Building Your First Bioinformatics Pipeline
Let's put it all together with a simple quality control pipeline:
#!/usr/bin/env bash
set -euo pipefail
# Quality-control summary for uncompressed, four-line FASTQ files.
# Usage: ./qc_pipeline.sh input_directory output_directory
if [[ $# -ne 2 ]]; then
echo "Usage: $0 input_directory output_directory" >&2
exit 1
fi
INPUT_DIR=$1
OUTPUT_DIR=$2
if [[ ! -d "$INPUT_DIR" ]]; then
echo "Input directory does not exist: $INPUT_DIR" >&2
exit 1
fi
mkdir -p "$OUTPUT_DIR"
summary="$OUTPUT_DIR/summary.tsv"
printf "file\treads\tadapter_sequence_lines\tmean_read_length\n" > "$summary"
shopt -s nullglob
fastq_files=("$INPUT_DIR"/*.fastq)
if (( ${#fastq_files[@]} == 0 )); then
echo "No .fastq files found in: $INPUT_DIR" >&2
exit 1
fi
for file in "${fastq_files[@]}"; do
filename=$(basename "$file" .fastq)
line_count=$(wc -l < "$file")
if (( line_count == 0 || line_count % 4 != 0 )); then
echo "Skipping malformed FASTQ file: $file" >&2
continue
fi
read_count=$(( line_count / 4 ))
adapter_count=$(grep -c "AGATCGGAAGAG" "$file" || true)
avg_length=$(awk 'NR % 4 == 2 {sum += length($0); count++} END {if (count) printf "%.1f", sum/count; else print "NA"}' "$file")
printf "%s\t%s\t%s\t%s\n" "$filename" "$read_count" "$adapter_count" "$avg_length" >> "$summary"
done
echo "Quality control complete. Review: $summary"
This is a small teaching script, not a replacement for FastQC or MultiQC. A successful run creates summary.tsv with one row per valid FASTQ file; inspect any skipped-file warning before interpreting results.
Advanced Topics to Explore Next
Once you're comfortable with these basics, consider learning:
- Regular expressions: Pattern matching on steroids
- Shell scripting: Automating complex workflows
- SSH and remote computing: Working on clusters
- Package managers: Installing bioinformatics software
- Workflow managers: Snakemake, Nextflow, CWL
Conclusion
Mastering the command line is like learning a new language: it takes practice, but once you're fluent, it opens up a world of possibilities. The commands and concepts covered in this tutorial form the foundation of computational biology work.
Remember:
- Practice regularly: Use the command line for daily tasks
- Start simple: Master basic commands before moving to complex pipelines
- Read the manual: Use man command_name to learn more about any command
- Don't be afraid to experiment: The best way to learn is by doing
The command line is your gateway to powerful bioinformatics analysis. With these skills, you're ready to tackle real biological datasets and start uncovering the secrets hidden in genomic data.
Next Steps
Ready to level up your bioinformatics skills? Check out our other tutorials:
- Package Management with Conda: Learn to install and manage bioinformatics software
- Single-cell RNA-seq Analysis: Apply your command line skills to cutting-edge analysis
- Introduction to Bioinformatics: Understand the bigger picture
Questions about command line usage? Need help with a specific bioinformatics task? Contact us: we're here to help you master computational biology!
Knowledge Check & Assessment
1. Concept Verification
Why are quoting, file permissions, exit codes, and test data essential when converting one-off commands into a script?
2. Practical Execution
Create and execute a Bash script that calculates read counts and average sequence length for small test FASTQ files. Pass Criteria: Record the command or analysis choice, keep the output, and explain why it answers the stated task.
3. Troubleshooting
If a script processes no files or produces an empty summary, how will you inspect glob expansion, input paths, line endings, and stderr logs?
Reviewed: February 2026
All commands and outputs were verified with the software versions listed in this tutorial. If you encounter reproducibility issues, please report them through the Contact page.
Author: Nasir Mahmood Abbasi, PhD · Category: Shell Command Basics