ProteinMPNN is now available as a one-click web tool at tools.ranomics.com/tools/mpnn. Upload a backbone structure, configure sampling temperature and the number of sequences, and download a FASTA. No GPU, no conda environment, no command line required.
This post covers what the tool does, when to use it, how to use it, and what it costs.
What ProteinMPNN does (two sentences)
ProteinMPNN solves the inverse folding problem: given a backbone structure, it predicts which amino acid sequences are geometrically and thermodynamically compatible with that geometry. It outputs a probability distribution over the 20 amino acids at each position conditioned on the full backbone, from which sequences are sampled at a user-specified temperature.
For a full treatment of the model architecture, the temperature parameter, and how fixed-position constraints work, see the companion article: ProteinMPNN sequence design on RFdiffusion backbones.
When to use this tool vs. structure prediction tools
ProteinMPNN and AlphaFold/ColabFold/Boltz-2 operate on opposite directions of the sequence-structure map. They are not alternatives — they are sequential steps in the same pipeline. In a full AI protein binder design campaign, the order is backbone generation, ProteinMPNN sequence design, structure-prediction validation, then wet-lab screening.
Use ProteinMPNN immediately after a backbone generative step (RFdiffusion, BindCraft backbone output, or any designed PDB) to produce candidate sequences. ProteinMPNN takes backbone coordinates as input and returns sequences. It does not evaluate binding.
Use structure prediction (AlphaFold, ESMFold, ColabFold, Boltz-2) after ProteinMPNN to validate that the designed sequences actually fold into the intended backbone. The standard self-consistency check is: fold the designed sequence with ESMFold or ColabFold, align the result to the input backbone, and filter on RMSD. Candidates above ~2 Å are discarded.
Use complex structure prediction (AlphaFold 3, Boltz-2 in complex mode) as a separate downstream step to evaluate whether the designed binder actually engages the target. ProteinMPNN does not model the binding partner. A sequence that passes self-consistency may still produce a poor binding interface — complex prediction is the filter for that.
The practical pipeline order: RFdiffusion → ProteinMPNN → self-consistency filter → complex prediction → synthesis.
Tool walkthrough
The interface at tools.ranomics.com/tools/mpnn requires a free account.
Step 1: Upload a backbone. Accepted formats are PDB, CIF and mmCIF. The backbone must contain at least one protein chain with full backbone atom records (N, CA, C, O). Side-chain coordinates are not used. The chains-to-design field defaults to chain A, so on a multi-chain file such as a binder-target complex, set it deliberately rather than assuming the whole complex gets redesigned.
Step 2: Set sampling temperature. The default is 0.1. For most binder design workflows, running at two temperatures (0.1 and 0.5) and pooling the outputs provides conservative high-confidence predictions alongside a diversity buffer. The tool accepts any value between 0.01 and 1.0.
Step 3: Set the number of sequences. The default is 50 per backbone, and a single run accepts anywhere from 1 to 1000. That default is already generous for a production run feeding a downstream filter step. Lower it only when you are designing across a large backbone pool and want the per-backbone count small.
Step 4: Download output. Results are returned as a FASTA file with one sequence per record, annotated with the backbone filename, temperature, and ProteinMPNN’s per-sequence score. The score is a log-probability averaged over all positions — lower (more negative) is better. Sort by score before passing sequences to the next filter step.
A concrete example from binder design
In a typical RFdiffusion binder design run, 10,000–30,000 backbone structures are generated against a target hotspot. After initial geometry filtering (diffusion confidence, clash score), a subset of 5,000–10,000 backbones advances to sequence design. Running ProteinMPNN at T = 0.2 and T = 0.5 on each backbone, with 8 sequences per backbone per temperature, produces a pool of 80,000–160,000 candidate sequences. That pool is then filtered by self-consistency (ESMFold RMSD), pLDDT on the designed chain, and interface quality from complex prediction, converging to a synthesis set of 500–2,000 sequences.
The ProteinMPNN step in this pipeline takes roughly thirty to sixty seconds of GPU time per run. It is not the computational bottleneck. Running more sequences per backbone is cheap relative to the structure prediction steps that follow.
Pricing
ProteinMPNN runs on the Ranomics tools hub, funded from a USD wallet and billed by the second of GPU time, which at roughly thirty to sixty seconds per run makes it one of the cheapest steps on the platform. A single run is capped at 1000 sequences. New accounts start with $5 in the wallet, no credit card required, which is plenty to design sequences across a full backbone pool.
Fund your wallet in any amount from $20 and pay only for the compute delivered. See tools.ranomics.com for the current model.
Run ProteinMPNN now: tools.ranomics.com/tools/mpnn. Start with $5 in your wallet on signup, no credit card.
Related Ranomics services
- ProteinMPNN tool: One-click sequence design on the Ranomics tools hub. Funded from your wallet, $5 to start on signup.
- Binder Pilot: Full binder design campaign — RFdiffusion + ProteinMPNN + experimental validation. Target structure in, ranked hit list out.
- Epitope Scout: Free surface epitope identification — identify which patches on your target are worth designing against before running any sequence design.
Frequently asked questions
What sampling temperature should I run?
The form defaults to 0.1 and accepts 0.01 to 1.0. For binder design the useful pattern is running two temperatures, say 0.1 and 0.5, and pooling the outputs, which gives conservative high-confidence sequences alongside a diversity buffer for the downstream filter to work on.
How many sequences should I generate per backbone?
The form defaults to 50, and the hosted tool accepts 1 to 1000 per run. The structure prediction steps that follow are the expensive part, so extra sequences here are cheap, and the default is already generous for a production run feeding a downstream filter. Design itself takes roughly thirty to sixty seconds of GPU time, so it is not the bottleneck in any realistic pipeline.
Does a good ProteinMPNN score mean the design binds?
It does not, and the model never sees the binding partner. A score says the sequence is compatible with the backbone you uploaded. Whether it folds that way is a separate question, answered by folding the design with ESMFold or ColabFold and discarding anything above roughly 2 Å RMSD from the input backbone. Whether it engages the target is a third question, and that one needs complex prediction.
Can I hold part of my structure fixed?
Yes. Upload a PDB, CIF or mmCIF carrying backbone atom records, then name the chains to design. The field defaults to chain A rather than to every chain in the file, so on a binder-target complex set it deliberately instead of assuming the whole complex gets redesigned. Side-chain coordinates in the upload are ignored either way.
How do I rank the sequences that come back?
Sort by the per-sequence score in the output FASTA before anything advances. The score is a log-probability averaged over all positions, so lower and more negative is better, and each record is annotated with the backbone filename and the temperature it came from. That ordering is what decides which sequences are worth paying to fold.
Do I need a GPU or a credit card to run it?
Neither. The interface at tools.ranomics.com needs a free account, and new accounts start with $5 in the wallet with no card required, which is plenty to design sequences across a full backbone pool. Sequence design is among the cheapest things on the platform, roughly thirty to sixty seconds of GPU time per run, but it is billed by the second like everything else, the cost scales with how many sequences you ask for, and a single run is capped at 1000 sequences. If you were about to build a conda environment for this one step, do not.