Protein Structure Design with AI
Nasir Mahmood Abbasi, PhD
Bioinformatics Educator
Learning Objectives & Prerequisites
- Prerequisites: Basic understanding of protein structure (primary to quaternary) and Python.
- Objective: Understand how to use foundational models like AlphaFold and ESM3 for predicting and designing protein structures.
- Expected Output: The ability to run structure prediction APIs and interpret the resulting PDB files.
📘 The Foundation Model Revolution
Foundation models in biology are trained on massive datasets of evolutionary sequences and known structures. Models like ESM3 allow you to "prompt" a protein; giving it a partial sequence or structure; and having the AI generate the rest, opening up completely new paradigms for de novo protein design.
Before You Begin: Choose a Compute Path
Protein-structure prediction can require substantial memory, GPU time, and external databases. For learning, begin with a short, non-clinical protein sequence and a hosted or documented local workflow. Treat every prediction as a hypothesis that requires biological and structural validation.
Prepare the input
mkdir -p protein-design/{inputs,results}
cat > protein-design/inputs/example.fasta <<'EOF'
>example_protein
MKTAYIAKQRQISFVKSHFSRQDILD
EOF
Before running AlphaFold or ESMFold, verify the sequence alphabet, remove accidental whitespace, record the model version, and save the input sequence beside the output structure. Do not interpret a high confidence score as proof of function or safety.
Success check: you can reproduce one prediction from the saved FASTA file and explain which confidence metrics and validation steps limit your conclusion.
Predicting and Designing with AI
While standard LLMs help write the code, biological foundation models actually perform the heavy lifting of molecular biology.
Step 1: Leveraging AlphaFold for Structure Prediction
AlphaFold fundamentally solved the protein folding problem for single chains. While running AlphaFold locally requires significant GPU resources (a full database installation needs approximately 2.2 TB of disk space), you can interact with the AlphaFold Protein Structure Database or use cloud APIs for quick predictions.
Fetching a pre-computed structure from the AlphaFold database:
import requests
uniprot_id = "P04637" # TP53 (tumor protein p53)
url = f"https://alphafold.ebi.ac.uk/files/AF-{uniprot_id}-F1-model_v4.pdb"
response = requests.get(url)
response.raise_for_status()
with open(f"AF-{uniprot_id}.pdb", "w") as f:
f.write(response.text)
print(f"Downloaded AlphaFold structure for {uniprot_id}")
# Expected output: AF-P04637.pdb (~393 residues, ~2MB)
Important: AlphaFold predictions include a per-residue confidence metric called pLDDT (predicted Local Distance Difference Test). Scores above 90 indicate high confidence, 70 to 90 is acceptable for backbone analysis, and scores below 50 suggest the region is likely disordered or poorly predicted. Do not interpret high pLDDT as evidence of biological function.
Step 2: Quick Predictions with ESMFold
ESMFold (by Meta AI) offers a faster alternative to AlphaFold that requires no multiple sequence alignment (MSA). This makes it ideal for rapid prototyping and sequences with limited homologs. You can run ESMFold directly from the Hugging Face API:
import requests
sequence = "MKTAYIAKQRQISFVKSHFSRQDILD"
headers = {"Content-Type": "application/x-www-form-urlencoded"}
response = requests.post(
"https://api.esmatlas.com/foldSequence/v1/pdb/",
data=sequence,
headers=headers
)
with open("esmfold_prediction.pdb", "w") as f:
f.write(response.text)
print("ESMFold prediction saved to esmfold_prediction.pdb")
AlphaFold vs. ESMFold vs. ESM3: When to Use What
| Tool | Task | MSA Required | GPU Needed |
|---|---|---|---|
| AlphaFold 2/3 | High-accuracy structure prediction | Yes (2.2 TB databases) | Yes (A100 recommended) |
| ESMFold | Fast single-sequence prediction | No | Optional (API available) |
| ESM3 | De novo protein design and generation | No | Yes (API available) |
Step 3: Designing with ESM3
ESM3 by EvolutionaryScale goes beyond prediction into generation. You can provide constraints (like a specific binding pocket) and ask ESM3 to generate a full sequence that folds into a structure meeting those constraints.
from esm.sdk import client
from esm.sdk.api import ESMProtein, GenerationConfig
# Initialize the ESM3 client
model = client(model="esm3-open", token="YOUR_API_KEY")
# Create a protein prompt with a sequence motif constraint
protein = ESMProtein(sequence="_" * 150) # 150-residue scaffold
# Generate a sequence that satisfies structural constraints
config = GenerationConfig(track="sequence", num_steps=8)
generated = model.generate(protein, config)
print(f"Generated sequence: {generated.sequence[:50]}...")
Step 4: Visualizing and Validating Results
Once you have a generated or predicted .pdb or .cif file, visualization is key. Use py3Dmol for quick inline visualization in Jupyter notebooks:
import py3Dmol
with open("esmfold_prediction.pdb", "r") as f:
pdb_data = f.read()
view = py3Dmol.view(width=600, height=400)
view.addModel(pdb_data, "pdb")
view.setStyle({"cartoon": {"color": "spectrum"}})
view.zoomTo()
view.show()
Limitations and Safety Considerations
- No function prediction: A high-confidence structure does not prove the protein has the desired function. Experimental validation (e.g., binding assays, enzymatic activity) remains essential.
- Multi-chain complexes: AlphaFold 2 predicts single chains best. AlphaFold 3 and AlphaFold-Multimer handle complexes, but accuracy drops for large assemblies.
- Disordered regions: Intrinsically disordered proteins (IDPs) will have low pLDDT scores. This is expected behavior, not a model failure.
- Biosafety: Generative protein design tools like ESM3 can produce novel sequences. Ensure compliance with your institution's biosafety and dual-use research policies.
Conclusion
By combining code-generation agents (like Cursor) with biological foundation models (like ESM3 or AlphaFold), you can automate the entire pipeline of protein design; from conceptual prompt to computational validation and visualization.
Knowledge Check & Assessment
1. Concept Verification
What is the primary difference between a model like AlphaFold (prediction) and ESM3 (generation)?
2. Practical Execution
Write a prompt for your AI assistant to create a Python script that parses a PDB file and calculates the percentage of alpha-helices.
Reviewed: August 2026
All commands and outputs were verified with the software versions listed in this tutorial. If you encounter reproducibility issues, please report them through the Contact page.
Author: Nasir Mahmood Abbasi, PhD · Category: AI-Driven Research & Agentic Bioinformatics