AI-Driven Research & Agentic Bioinformatics•2026-08-24

Protein Structure Design with AI

NM

Nasir Mahmood Abbasi, PhD

Bioinformatics Educator

Protein Structure Design
Tested on: ESM3, AlphaFold3, PyMOL
Last Review: 2026-09-08

Learning Objectives & Prerequisites

  • Prerequisites: Basic understanding of protein structure (primary to quaternary) and Python.
  • Objective: Understand how to use foundational models like AlphaFold and ESM3 for predicting and designing protein structures.
  • Expected Output: The ability to run structure prediction APIs and interpret the resulting PDB files.

📘 The Foundation Model Revolution

Foundation models in biology are trained on massive datasets of evolutionary sequences and known structures. Models like ESM3 allow you to "prompt" a protein; giving it a partial sequence or structure; and having the AI generate the rest, opening up completely new paradigms for de novo protein design.

Before You Begin: Choose a Compute Path

Protein-structure prediction can require substantial memory, GPU time, and external databases. For learning, begin with a short, non-clinical protein sequence and a hosted or documented local workflow. Treat every prediction as a hypothesis that requires biological and structural validation.

Prepare the input

mkdir -p protein-design/{inputs,results}
cat > protein-design/inputs/example.fasta <<'EOF'
>example_protein
MKTAYIAKQRQISFVKSHFSRQDILD
EOF

Before running AlphaFold or ESMFold, verify the sequence alphabet, remove accidental whitespace, record the model version, and save the input sequence beside the output structure. Do not interpret a high confidence score as proof of function or safety.

Success check: you can reproduce one prediction from the saved FASTA file and explain which confidence metrics and validation steps limit your conclusion.

Predicting and Designing with AI

While standard LLMs help write the code, biological foundation models actually perform the heavy lifting of molecular biology.

Step 1: Leveraging AlphaFold for Structure Prediction

AlphaFold fundamentally solved the protein folding problem for single chains. While running AlphaFold locally requires significant GPU resources (a full database installation needs approximately 2.2 TB of disk space), you can interact with the AlphaFold Protein Structure Database or use cloud APIs for quick predictions.

Fetching a pre-computed structure from the AlphaFold database:

import requests

uniprot_id = "P04637"  # TP53 (tumor protein p53)
url = f"https://alphafold.ebi.ac.uk/files/AF-{uniprot_id}-F1-model_v4.pdb"

response = requests.get(url)
response.raise_for_status()

with open(f"AF-{uniprot_id}.pdb", "w") as f:
    f.write(response.text)

print(f"Downloaded AlphaFold structure for {uniprot_id}")
# Expected output: AF-P04637.pdb (~393 residues, ~2MB)

Important: AlphaFold predictions include a per-residue confidence metric called pLDDT (predicted Local Distance Difference Test). Scores above 90 indicate high confidence, 70 to 90 is acceptable for backbone analysis, and scores below 50 suggest the region is likely disordered or poorly predicted. Do not interpret high pLDDT as evidence of biological function.

Step 2: Quick Predictions with ESMFold

ESMFold (by Meta AI) offers a faster alternative to AlphaFold that requires no multiple sequence alignment (MSA). This makes it ideal for rapid prototyping and sequences with limited homologs. You can run ESMFold directly from the Hugging Face API:

import requests

sequence = "MKTAYIAKQRQISFVKSHFSRQDILD"
headers = {"Content-Type": "application/x-www-form-urlencoded"}

response = requests.post(
    "https://api.esmatlas.com/foldSequence/v1/pdb/",
    data=sequence,
    headers=headers
)

with open("esmfold_prediction.pdb", "w") as f:
    f.write(response.text)

print("ESMFold prediction saved to esmfold_prediction.pdb")

AlphaFold vs. ESMFold vs. ESM3: When to Use What

ToolTaskMSA RequiredGPU Needed
AlphaFold 2/3High-accuracy structure predictionYes (2.2 TB databases)Yes (A100 recommended)
ESMFoldFast single-sequence predictionNoOptional (API available)
ESM3De novo protein design and generationNoYes (API available)

Step 3: Designing with ESM3

ESM3 by EvolutionaryScale goes beyond prediction into generation. You can provide constraints (like a specific binding pocket) and ask ESM3 to generate a full sequence that folds into a structure meeting those constraints.

from esm.sdk import client
from esm.sdk.api import ESMProtein, GenerationConfig

# Initialize the ESM3 client
model = client(model="esm3-open", token="YOUR_API_KEY")

# Create a protein prompt with a sequence motif constraint
protein = ESMProtein(sequence="_" * 150)  # 150-residue scaffold

# Generate a sequence that satisfies structural constraints
config = GenerationConfig(track="sequence", num_steps=8)
generated = model.generate(protein, config)
print(f"Generated sequence: {generated.sequence[:50]}...")

Step 4: Visualizing and Validating Results

Once you have a generated or predicted .pdb or .cif file, visualization is key. Use py3Dmol for quick inline visualization in Jupyter notebooks:

import py3Dmol

with open("esmfold_prediction.pdb", "r") as f:
    pdb_data = f.read()

view = py3Dmol.view(width=600, height=400)
view.addModel(pdb_data, "pdb")
view.setStyle({"cartoon": {"color": "spectrum"}})
view.zoomTo()
view.show()

Limitations and Safety Considerations

  • No function prediction: A high-confidence structure does not prove the protein has the desired function. Experimental validation (e.g., binding assays, enzymatic activity) remains essential.
  • Multi-chain complexes: AlphaFold 2 predicts single chains best. AlphaFold 3 and AlphaFold-Multimer handle complexes, but accuracy drops for large assemblies.
  • Disordered regions: Intrinsically disordered proteins (IDPs) will have low pLDDT scores. This is expected behavior, not a model failure.
  • Biosafety: Generative protein design tools like ESM3 can produce novel sequences. Ensure compliance with your institution's biosafety and dual-use research policies.

Conclusion

By combining code-generation agents (like Cursor) with biological foundation models (like ESM3 or AlphaFold), you can automate the entire pipeline of protein design; from conceptual prompt to computational validation and visualization.

Knowledge Check & Assessment

1. Concept Verification

What is the primary difference between a model like AlphaFold (prediction) and ESM3 (generation)?

2. Practical Execution

Write a prompt for your AI assistant to create a Python script that parses a PDB file and calculates the percentage of alpha-helices.

Reviewed: August 2026

All commands and outputs were verified with the software versions listed in this tutorial. If you encounter reproducibility issues, please report them through the Contact page.

Author: Nasir Mahmood Abbasi, PhD · Category: AI-Driven Research & Agentic Bioinformatics

Continue Learning

Course Sequence