AI NewsWords 1673Read time5 min

Claude Designed Wet-Lab-Validated Protein Binders for 14 of 15 Targets

Claude produced 354 confirmed protein binders across 14 targets, though Anthropic's study lacked a matched human control.

Anthropic has reported that Claude designed experimentally confirmed protein binders for 14 of 15 targets with interpretable laboratory results. Across 1,320 tested designs, 354 bound to their intended targets, an overall hit rate of 26.8%.

The result goes beyond asking a language model to propose amino-acid sequences. Claude researched each target, selected binding sites, installed specialized protein-design software, generated and optimized candidates, and ranked 30 designs per target. Contract research organizations then produced the proteins and measured their binding in physical assays.

Individual hit rates ranged from 22.6% to 35.1%, depending on the model and campaign format. Anthropic compares this with a 10%–15% rate calculated from recent de novo protein-design campaigns. The comparison is useful but not definitive: the experiment had no matched human team operating the same software with the same compute budget, and Anthropic's technical report has not yet undergone independent peer review.

1. What Anthropic tested

A protein binder is a molecule engineered to attach to a particular surface on another protein. Binding can block a protein, activate it, or provide a way to deliver another molecule. This makes binder discovery relevant to therapeutics, diagnostics, and biological research, but a binder is not itself a drug.

Anthropic selected 16 targets, including EGFR, IL-7Rα, PD-L1, RBX1, TNFα, TREM2, TrkA, VEGF-A, and two forms of GDF-8. Mature GDF-8 aggregated under the assay conditions and produced nonspecific measurements at both testing organizations, so its 120 designs were excluded. That left 1,320 designs across 15 targets with interpretable results.

Claude Opus 4.8 and the research model Mythos Preview were tested in two configurations. In multi-target campaigns, each model worked on 14 targets during a single 48-hour session. In the single-target configuration, Mythos Preview received a separate 24-hour session for each target. Opus 4.8 was also run separately on three selected targets.

Each campaign began with a roughly 16,000-word protocol prompt—about 30,000 tokens—describing the stages of protein-binder design, relevant tools, screening criteria, operational rules, and a scientific reading list. It did not specify the epitope, scaffold, or sequence Claude should use for any individual target.

Human researchers selected the targets, prepared the protocol, supplied cloud accounts and reference material, approved access requests, handled infrastructure interruptions, and ordered the final sequences. Anthropic says they did not intervene in Claude's scientific design decisions after a campaign began.

2. Claude produced 354 confirmed binders

Opus 4.8 generated 88 binders from 390 designs in its multi-target campaign, a 22.6% hit rate. Mythos Preview produced 104 binders from 390 designs in the same format, reaching 26.7%.

When Mythos Preview received an independent campaign for each target, 158 of 450 designs bound, or 35.1%. On the 13 targets shared with its multi-target campaign, the single-target format produced 143 binders compared with 104. That improvement came with substantially more compute per target, so the experiment does not isolate whether the difference resulted from greater focus, additional compute, or both.

Claude's ordering also contained useful information. Among the candidates it ranked first for each target and campaign, 49% bound. The technical report found that its rankings separated binders from non-binders better than chance on 12 of 13 targets where that comparison was possible.

Performance varied sharply by target. Claude found 72 binders among 90 TREM2 designs, 54 among 90 VEGF-A designs, and 49 among 90 IL-7Rα designs. At the other end, it found three weak binders among 90 designs for the synthetic β-barrel BBF-14, one among 30 for 15-PGDH, and none among 90 for maltose-binding protein.

The complete failure on maltose-binding protein is particularly informative. Claude's computational confidence scores rated those candidates almost as highly as designs for productive targets. The pipeline therefore remained unable to recognize some target-level failure modes before laboratory testing.

3. The RBX1 and TNFα results provide the strongest comparisons

RBX1 offered a direct comparison with an earlier open competition conducted through Adaptyv Bio. Nine of 245 de novo competition entries bound to RBX1, compared with 28 of Claude's 90 designs.

Anthropic also had the competition winner reproduced and tested alongside Claude's candidates on the same assay plate. Claude's best design recorded a dissociation constant, or K_D, of 3.9 nanomolar in that test, while the competition winner measured 45 nanomolar. A lower K_D indicates tighter binding, making Claude's result approximately 11 times tighter under those assay conditions.

This is stronger evidence than comparing measurements from separate laboratories, but it still does not establish that Claude outperformed a controlled human team. Competition participants had different constraints, compute budgets, tools, and numbers of submissions. Results from the RBX1 competition were also included in Claude's reading material.

TNFα presented a different test. Several published computational methods had previously reported no binders for this target. Claude produced 12 binders from 150 candidates, all generated by Opus 4.8 rather than Mythos Preview.

Those 12 sequences came from only four distinct structural backbones, so they should not be treated as 12 completely independent solutions. Some nevertheless bound human, cynomolgus-monkey, and mouse TNFα. Across the broader study, 130 of 233 binders tested against a mouse version of their target also bound it, potentially reducing the redesign required before animal experiments.

4. Claude orchestrated specialized models rather than replacing them

Claude did not generate the proteins using its language-model weights alone. It acted as an agent controlling established, task-specific protein software.

The campaigns used ten structure-generation systems, including PXDesign, RFdiffusion3, Genie 3, FreeBindCraft, BoltzGen, RFdiffusion, and Proteina-Complexa. Most amino-acid sequences were produced with SolubleMPNN, a variant of ProteinMPNN. ESMFold2, ESMFold2-Fast, and Protenix v2 supplied the main structural confidence scores used to filter and rank candidates.

Claude selected and installed these tools, chose target regions and epitopes, allocated its compute budget, removed redundant or implausible sequences, and ran additional optimization when it judged that worthwhile. The delivered proteins represented 24 combinations of structure-generation and sequence-design methods.

The multi-target campaigns were allowed up to 12,500 Nvidia H100 GPU-hours over 48 hours. Each single-target Mythos campaign received up to 2,500 H100 hours over 24 hours. The technical report describes corresponding cloud-compute budgets of $50,000 for a multi-target campaign and $10,000 for a single-target campaign.

These resource levels matter when interpreting the claimed automation. Claude compressed the human orchestration work into one or two days, but it did not eliminate the computational cost, the expert knowledge encoded in the protocol, or the weeks required for external synthesis and wet-lab validation.

5. How the laboratory validation worked

Adaptyv Bio received all 1,320 designs, expressed them individually through cell-free synthesis, and primarily tested binding with surface plasmon resonance. It returned measurements for 1,296 designs, including 61 that failed to express.

Twist Bioscience received 1,260 designs, excluding those for latent GDF-8. It expressed them as human IgG1 Fc fusions in HEK293 cells and measured binding on a separate high-throughput surface-plasmon-resonance platform.

The organizations used different protein formats and classification procedures. Anthropic combined their measurements using a documented rule after blinded review of the response traces. Among 1,235 designs tested by both organizations, their initial classifications agreed in 1,099 cases, or 89%, with a reported Cohen's kappa of 0.71.

This is genuine physical validation of binding, not merely a computational prediction. It does not, however, establish the proteins' atomic structures, biological effects, specificity across a broad off-target panel, stability in an organism, toxicity, pharmacokinetics, or therapeutic efficacy. No experimental structure was determined for any design or binder-target complex.

6. What the result changes—and what it does not

The practical advance is the orchestration layer. Protein-generation and structure-prediction models already existed, but using them effectively required experts to choose target surfaces, install and combine rapidly changing software, manage large GPU jobs, filter failures, and decide which sequences deserved costly laboratory testing.

Anthropic's experiment shows that a general-purpose AI agent can execute much of that workflow across many targets and return an experimentally productive ranking. For laboratories already equipped to evaluate proteins, such an agent could increase the number of targets attempted without requiring a computational specialist to supervise every software step.

The study does not show that Claude can develop a drug autonomously. A high-affinity minibinder is an early candidate, not evidence of safety or efficacy. Anthropic itself notes that minibinders are not a standard therapeutic modality and that subsequent engineering, functional testing, toxicology, manufacturing work, and clinical trials remain necessary.

The exact capability is also not broadly available as a turnkey Claude feature. Anthropic says some advanced biological tasks remain restricted because of dual-use risks and that it intends to establish a scientific access program. It has nevertheless released the campaign prompts, predicted structures, provenance records, and experimental measurements under a CC BY 4.0 license, allowing researchers to audit the analysis and use the 1,320-design dataset as a benchmark.

Frequently Asked Questions

Did Claude design a drug?

No. Claude designed small proteins that bound to laboratory targets. Binding is an early drug-development step, but the candidates were not tested as safe or effective medicines.

How many of Claude's designs worked?

Laboratory testing classified 354 of 1,320 designs as binders, an overall hit rate of 26.8%. At least one binder was found for 14 of 15 targets with interpretable measurements.

Which Claude models were used?

The binder campaigns used Claude Opus 4.8 and Mythos Preview. Mythos achieved 26.7% in multi-target mode and 35.1% in separate single-target campaigns; Opus achieved 22.6% in multi-target mode.

Were the proteins tested in a real laboratory?

Yes. Adaptyv Bio and Twist Bioscience produced and measured the designs using physical binding assays. The study did not determine experimental three-dimensional structures or therapeutic effects.

Is the study peer-reviewed?

No. Anthropic published a technical report and the underlying dataset, but independent peer review was still pending at publication time.

Sources

Share

Share this article