Design protein sequences with ProteinMPNN
ProteinMPNN (Dauparas et al., Science 2022) is a structure-conditioned protein sequence design model, released under the MIT license. Given a protein backbone structure, it designs amino-acid sequences predicted to fold into that structure.
Status: beta. Install and run both work today. The registry keeps this at beta (rather than available) until execution is verified across more environments.
Install
$ moldesk install proteinmpnnSee CLI usage for reinstalling, uninstalling, and other commands.
Example
Download a sample structure — ubiquitin (1UBQ) from the RCSB PDB — and design new sequences for it:
$ curl -O https://files.rcsb.org/download/1UBQ.pdb$ moldesk run proteinmpnn 1UBQ.pdb --output ./resultsMoleculeDesk prints the run's status and the path to each output file as it completes. --output ./results also copies the designed sequences into ./results as .fa (FASTA) files you can open directly.
What ProteinMPNN is used for
ProteinMPNN is widely used upstream for:
- Fixed-backbone sequence design — generating new sequences for a known structure.
- Enzyme redesign — designing sequence variants around a target active site or fold.
- Binder design — as a component in pipelines that design proteins to bind a specific target.
Requirements
| Model version | v_48_020 (default upstream checkpoint) |
| License | MIT |
| Python | 3.11, via uv |
| Dependencies | torch 2.2.1, numpy 1.26.4 |
| Platforms | darwin-arm64, linux-x64 |
| Input format | .pdb |
| Output | sequences (seqs/*.fa) |
MoleculeDesk pins ProteinMPNN to a specific upstream commit and verifies the model checkpoint (~6.7 MB) by SHA-256 before use, so every install is reproducible.
ProteinMPNN benchmarks: speed, VRAM and cost per design batch
On an NVIDIA GeForce RTX 3090, ProteinMPNN takes about 8 s per design batch once warm (16 s on the first run), roughly 450 design batchs per GPU-hour, or about <$0.01 per design batch.
| Metric | NVIDIA GeForce RTX 3090 |
|---|---|
| Time per run, first (cold) | 16 s |
| Time per run, warm | 8 s |
| Design batchs per GPU-hour | 450 |
| Cost per run, first (cold) | <$0.01 |
| Cost per run, warm | <$0.01 |
| Peak GPU memory | 0.4 GiB |
| Peak GPU utilization | 9 % |
| Peak GPU power | 108 W |
| Install time | 6.1 min |
| Disk per install | 6 GiB |
Workload: One 20-residue chain (examples/diffdock/protein.pdb), default sampling, 1 sequence. Median of 3 runs.
Machine: RunPod GPU pod, Linux x64, 125 GiB RAM; AMD EPYC 7H12 (32 vCPU allocated); driver 580.126.20.
Cost: at $0.50/GPU-hour (RunPod Secure Cloud, EU-CZ-1, compute only; storage adds about $0.02/hr); excludes storage and idle time.
Warm time is the mean of runs 2 and 3 (7.9 s and 8.1 s). Install time includes downloading PyTorch with an empty package cache. Peak GPU power is near the idle-to-light-load range because the job is tiny.
FAQ
Can I run ProteinMPNN through MoleculeDesk today?
Yes — both moldesk install proteinmpnn and moldesk run proteinmpnn <input>.pdb work today (see the example above). See Model compatibility & registry for what beta means.
Does this require a GPU?
No specific GPU requirement is declared for ProteinMPNN's manifest, though torch will use one if available. Run moldesk doctor to check your machine.
What license is ProteinMPNN under? MIT.