Take a model that predicts structure from sequence and run it backwards. Hold the target fixed, carry a sequence that has not yet committed to specific amino acids, and push it downhill until the fold the model predicts is one that binds. That is ESMFold2 design, and in scFv mode it aims that loop at all six CDRs, heavy and light together, on a framework borrowed from a clinical antibody.
It is the only tool in our self-serve catalog that designs both chains of an scFv at once. Here is how it works, and the ways a first run gets wasted.
ESMFold2 design has no separate sequence model, and that is the whole idea
Built at the Chan Zuckerberg Biohub in 2026 on the ESMC protein language model, ESMFold2 design inverts the folding step. A normal prediction run reads a sequence and returns coordinates. This one holds the target fixed, keeps a soft sequence representation, and runs gradient descent through the folding network itself, so sequence and predicted binding pose are optimised together in one loop.
As binder design tools go that makes it the odd one out in a useful way. Nothing is handed off. Compare it to how RFdiffusion works, where a diffusion model proposes a backbone and ProteinMPNN writes a sequence onto it afterwards.
The same machinery drives both modes. Minibinder mode builds a free scaffold of 60 to 200 residues behind an isoelectric point filter. scFv mode locks the framework backbone and sequence and lets only the six CDR regions move. The framework is trastuzumab, atezolizumab, or ocankitug, all three humanized IgG1 with clinical precedent. Trastuzumab and atezolizumab are the backbones of approved drugs, and that is where the developability head start comes from.
There is no structure to prepare. The loop is sequence only: one chain of 30 to 800 residues, or one of five presets cropped to the relevant ectodomain. No PDB, no docking setup, no alignment. A minibinder design takes about ten minutes on a warm container, an scFv about twelve.
The published work behind it was tested at the bench against PDGFRB, EGFR, PD-L1, CD45 and CTLA4, in both modalities, with BLI-measured affinities from 68 pM to 70 nM. Reported scFv hit rates run 15 to 29 percent in the higher-compute cohorts, against 36 to 88 percent for minibinders. Read those as a ceiling, not a forecast: they come from testing the top 84 designs out of roughly 117,000 scFv candidates, ranked by a 19-critic ensemble at thousands of GPU-hours per target. A first run on a handful of seeds is a different experiment. Those five targets ship as presets, so they are the cheapest way to baseline your own wet-lab setup against a known result before spending a campaign on a novel target.
Design a fresh scFv when there is nothing to graft
The case is narrower than the capability makes it sound.
Grafting CDRs onto a framework assumes you already own a binder worth moving. Affinity maturation assumes you already own a weak one worth improving. Both are correct when the assumption holds, and neither needs a design tool.
Designed scFvs are for the case where neither holds. Nothing binds your target, and the default alternative is a discovery campaign followed by humanization and reformatting. Starting on a locked clinical framework inverts that order: the framework half of the developability question begins from a humanized IgG1 backbone with clinical precedent, rather than from whatever framework a discovery campaign happens to land on. That is the argument, and it is the only one this tool wins on its own.
One boundary, stated plainly: this is paired heavy and light only. For a single-domain nanobody use RFantibody instead, and the scFv versus nanobody decision covers how to choose the format before you choose the tool.
Four ways a first run gets wasted
Every one of these makes the tool look broken when it is behaving correctly.
Running a single design. One design is one sample from a stochastic process, and it is entirely normal for it to come back short of the bar. Batch size defaults to three for that reason. For a first pass at a new target, use six. Drop to one only when you already know the target gives clean hits.
Pasting the whole protein. People paste a full-length sequence when one domain is what matters. The five built-in presets are cropped to the relevant ectodomain deliberately. A full-length paste asks a harder and different question than the one you meant.
Reading the confidence score as an affinity. The score that ranks designs is a structural confidence number. It says the model is confident about the predicted pose. It is not a binding constant, it does not rank like one, and buying the top of the table as though it were is the expensive mistake in this workflow.
Running one seed and stopping. Different seeds initialise different soft sequences and produce genuinely different designs, so judging the method on one initialisation is judging noise. You can run 1 to 64 seeds in a single sweep, each on its own GPU in parallel, and the results merge into one ranked table. A sixteen-seed sweep costs sixteen times the money and almost no extra waiting, which is an unusually good trade.
The output is a ranked table, not a winner
Each design returns its sequence, its confidence scores and, where the run produced one, its predicted complex structure. The ranked table carries no verdict column, which is deliberate: a pass-or-fail word stamped during the run is only ever true against the bar that was live when it was written. Instead the best-design panel judges its own pick at the moment the page renders, and says something only when there is something to say: a pick that clears the bar carries no label at all, one that falls short is marked borderline or rejected, and a pick the tool could not measure is reported as not checked against the bar rather than as a failure. Multi-seed sweeps also carry the seed that produced each row. Isoelectric point is reported for minibinders only; in scFv mode the column is not shown at all, because the pI gate is a minibinder filter.
Nothing is withheld. Ranking is a sort, not a gate, so a design that falls short of the bar still comes back with everything a design that cleared it would have. The distinction matters when you are deciding what to buy, because a design the tool measured and found wanting and a design it could not score at all mean completely different things, which is why the panel names which of the two it is rather than calling both a failure.
Quick reference
| ESMFold2 design, scFv mode | ESMFold2 design, minibinder mode | RFantibody | |
|---|---|---|---|
| Output | paired heavy and light scFv | free scaffold, 60 to 200 aa | single-domain VHH |
| What moves | six CDR loops only | the whole scaffold | CDR loops on a VHH framework |
| Framework | trastuzumab, atezolizumab or ocankitug | none | fixed nanobody framework |
| Input | target sequence, 30 to 800 aa | target sequence, 30 to 800 aa | target structure |
| Typical runtime | about 12 min per design | about 10 min per design | see the RFantibody guide |
| Watch out | iPTM ranks, the CDR proxy gates | pI gate rejects undisplayable scaffolds | interface PAE is not a binding measurement |
Practical recommendations
- Start at batch six and at least eight seeds. Breadth is nearly free in wall-clock here because seeds run side by side. Spending it on a first look at a new target is the highest-value setting change available.
- Baseline on a preset first. Run one of the five validated targets before your novel one, so that when results look wrong you know whether the problem is the target or the setup.
- Take the shortlist to a developability screen, then an assay. A predicted complex is a hypothesis. The confidence score ranks plausibility, not binding.
- Run it alongside another design method, not instead of one. The designs worth screening are the ones your other tools did not propose.
Summary
ESMFold2 design is a folding model driven in reverse, and in scFv mode it designs all six CDRs of a paired fragment on a clinical antibody framework in about twelve minutes from a sequence alone. Reach for it when nothing binds your target yet and the alternative is discovery followed by humanization, because starting on a humanized backbone with clinical precedent inverts that order. Treat it as one method among several rather than the answer, and remember that the ranking column measures structural confidence and not affinity. The scFv is a hypothesis until an assay says otherwise.
Related Ranomics services
- ESMFold2 design technology page: how the inversion stack runs on the Ranomics platform, and what each mode returns.
- Yeast display antibody discovery: the experimental side, where designed scFvs get measured rather than predicted.
- Binder Pilot: take a shortlist of designed fragments to a ranked, wet-lab-validated hit list on a fixed-scope campaign.
Frequently asked questions
What does ESMFold2 design actually do?
It runs a structure prediction model in reverse. Rather than reading a sequence and returning a fold, it holds your target fixed and uses gradient descent on a soft sequence representation, backpropagating through the folding network, until the sequence it carries folds into something that binds. There is no separate sequence model downstream, because sequence and pose are optimised in the same loop.
When should I design a new scFv instead of grafting CDRs I already have?
Grafting assumes you already own a binder worth moving, and affinity maturation assumes you own a weak one worth improving. ESMFold2 design is for the case where neither assumption holds and nothing binds your target yet. Its advantage over discovering a binder and reformatting it later is the starting framework: these are humanized IgG1 backbones with clinical precedent, and two of the three are the backbones of approved drugs.
Can ESMFold2 design produce nanobodies or VHH fragments?
No. It covers paired heavy and light scFvs and free minibinder scaffolds only. Single-domain formats are RFantibody's job, and the two tools complement rather than compete. Choosing the fragment format is a separate decision from choosing the tool that designs it.
Why does my scFv run label every design as a reject?
It should not, and there is no setting to change. Every scFv run scores its designs on the CDR distogram proxy as a matter of course. If the best-design panel marks its pick rejected, that design was measured and genuinely fell short of the bar, and the answer is breadth rather than a checkbox: raise the batch size and run more seeds. A design the tool could not score at all is reported on that panel as not checked against the bar rather than as a rejection, so the two cases stay distinguishable. Every design is returned either way, with its sequence and, where the run produced one, its predicted complex.
Why did my run return no designs at all?
An empty result is a different problem from a rejected one. Zero designs usually means the target sequence itself failed to fold as a single chain, which happens with very short sequences, very long ones, and low-complexity stretches. Check the target before you change any design setting.