Claude’s Protein Design Hit Rate Was 26.8%. One Target Returned 0 for 90.

AimostAll news brief curated from Towards AI.

Source details

Original source
Towards AI
Published
2026-09-01
Primary topic
AI Agents

Why it matters

Agent products, browser agents, autonomous workflows, operator systems, and orchestration tools. Use the original source for the full report, then use the directory shortcuts below to compare the products and workflows the story points toward.

What happened

Author(s): Decoding AI Originally published on Towards AI. We refit Anthropic’s published per-target protein design counts on Aug 25, 2026. Ninety designs went into the wet lab against maltose-binding protein. Ninety came back with nothing. That is the part of the story that stopped us. On August 18, 2026, Anthropic published the results of an autonomous protein design campaign, and the number everyone repeated was the Claude protein design hit rate: 354 binders out of 1,320 designs, 26.8%, against an industry norm of 10–15%. It is a real result, independently validated in two contract labs. But sitting inside that 26.8% is a target where the agent produced ninety designs, every one of which was synthesized, expressed, and measured — and not one of them bound. Both facts came out of the same campaign, the same prompt, the same model. So we spent August 25 doing the arithmetic nobody in the coverage did: 26.8% is a portfolio average across sixteen targets, not the probability that your target works. Run the published per-target counts through a binomial test and a single shared hit rate is rejected at p ≈ 1.8 × 10⁻⁶³. Fit a model that allows targets to differ and the picture inverts: for a new target, the chance of landing below the 10% industry floor is about 38%. Key takeaways Anthropic’s campaign (published 2026–08–18) produced 354 binders from 1,320 designs — a pooled hit rate of 26.8% — but per-target rates ran from 72/90 on TREM2 to 0/90 on maltose-binding protein. Under a single 26.8% rate, MBP’s 0-for-90 result has probability 6.3 × 10⁻¹³ — roughly one in 1.6 trillion. The pooled rate is a mixture, not a per-target expectation. Our beta-binomial fit to the eight published per-target counts (run 2026–08–25) puts the 90% predictive interval for a new target at 0.1% to 84.7%, with a median of 18.6% — well below the 26.8% headline. On that fit, a 30-design batch against an unseen target has a 20.1% chance of returning zero binders, versus 0.0086% if you assume one shared rate — a factor of roughly 2,300. The in-silico confidence scores did not distinguish the failures: per Anthropic’s report, designs against MBP and BBF-14 scored about the same as designs against targets that worked. TL;DR: An autonomous agent orchestrated a dozen open-source protein design tools and beat expert human hit rates on most targets. That happened, and it matters. But the headline number is a portfolio statistic, and almost nobody deploys a portfolio — they deploy against one target, one ticket, one customer. When we modeled the published spread instead of the mean, the expected experience of a single new target looked dramatically worse and dramatically noisier than 26.8% implies. And the pipeline’s own confidence scores could not tell in advance which regime it was in. That combination — high average, enormous variance, uncalibrated self-assessment — is the shape of nearly every agentic AI benchmark result you will read this year. What did Claude’s protein design campaign actually measure? Anthropic gave Claude (Mythos Preview and Opus 4.8) a set of 16 protein targets and asked for 30 minibinders per target per design arm — small proteins engineered to latch onto a target, the mechanism behind a large share of modern biologic drugs. With three arms running, most targets ended up with 90 designs in the lab. Fifteen targets produced usable measurements; one, mature GDF-8, was dropped because the target aggregated in the assay. The agent did not invent a protein model. It installed and ran existing open-source tools: backbones came from PXDesign (358 designs), RFdiffusion3 (267), Genie 3 (185), FreeBindCraft (135), BoltzGen (134), RFdiffusion (118) and Proteina-Complexa (100), among others, with sequences mostly from SolubleMPNN and filtering through ESMFold2 and Protenix v2. What was new was the layer above: choosing the epitope, installing the software, combining it across 24 distinct workflows, and ranking the output with no human touching a design decision. Validation was independent. Adaptyv Bio’s wet-lab case study, published 2026–08–19, reports that the designs arrived anonymized — the lab did not know which model produced which sequence — and were run on surface plasmon resonance at five target concentrations in duplicate. 95% of designs expressed. 354 of 1,320 bound. That is a genuinely strong result and we are not going to shave it down. Against RBX1, an open design competition had produced 9 binders from 245 entries (3.7%); Claude produced 28 from 90. Anthropic had the competition’s winning design rebuilt and measured on the same assay plate: it bound at 45 nM, while Claude’s best bound at 3.9 nM, roughly ten times tighter. Why doesn’t Claude’s protein design hit rate apply to your target? Because the targets are not interchangeable, and the published per-target counts make that impossible to ignore. Here are the eight per-target results disclosed across Anthropic’s post, Adaptyv Bio’s case study and The Decoder’s technical breakdown of the report: Binders per target, from the published counts: TREM2–72 of 90 designs bound (80.0%) VEGF-A — 54 of 90 (60.0%) IL-7Rα — 49 of 90 (54.4%) RBX1–28 of 90 (31.1%) TNFα — 12 of 150 (8.0%) BBF-14–3 of 90 (3.3%) 15-PGDH — 1 of 30 (3.3%) MBP — 0 of 90 (0.0%) Pooled, all 15 targets — 354 of 1,320 (26.8%) Now ask the question the coverage skipped: if every design really had a 26.8% chance of binding, how surprised should we be by the bottom row? The answer is that we should be about as surprised as it is possible to be. The probability of drawing zero successes in 90 independent trials at p = 0.268 is 0.732 raised to the 90th power, or 6.3 × 10⁻¹³ — one in roughly 1.6 trillion. TREM2 is equally impossible from the other direction: getting 72 or more hits in 90 trials at that rate has probability 1.1 × 10⁻²⁵. Neither of those is a fluke you explain away. They are the model being wrong. We ran a formal likelihood-ratio test comparing one shared rate against a model […]

What to do next

Move into automation and workflow tools next so you can evaluate whether the agent story is actionable or still mostly experimental.

Author(s): Decoding AI Originally published on Towards AI. We refit Anthropic’s published per-target protein design counts on Aug 25, 2026. Ninety designs went into the wet lab against maltose-binding protein. Ninety came back with nothing. That is the part of the story that stopped us. On August 18, 2026, Anthropic published the results of an autonomous protein design campaign, and the number everyone repeated was the Claude protein design hit rate: 354 binders out of 1,320 designs, 26.8%, against an industry norm of 10–15%. It is a real result, independently validated in two contract labs. But sitting inside that 26.8% is a target where the agent produced ninety designs, every one of which was synthesized, expressed, and measured — and not one of them bound. Both facts came out of the same campaign, the same prompt, the same model. So we spent August 25 doing the arithmetic nobody in the coverage did: 26.8% is a portfolio average across sixteen targets, not the probability that your target works. Run the published per-target counts through a binomial test and a single shared hit rate is rejected at p ≈ 1.8 × 10⁻⁶³. Fit a model that allows targets to differ and the picture inverts: for a new target, the chance of landing below the 10% industry floor is about 38%. Key takeaways Anthropic’s campaign (published 2026–08–18) produced 354 binders from 1,320 designs — a pooled hit rate of 26.8% — but per-target rates ran from 72/90 on TREM2 to 0/90 on maltose-binding protein. Under a single 26.8% rate, MBP’s 0-for-90 result has probability 6.3 × 10⁻¹³ — roughly one in 1.6 trillion. The pooled rate is a mixture, not a per-target expectation. Our beta-binomial fit to the eight published per-target counts (run 2026–08–25) puts the 90% predictive interval for a new target at 0.1% to 84.7%, with a median of 18.6% — well below the 26.8% headline. On that fit, a 30-design batch against an unseen target has a 20.1% chance of returning zero binders, versus 0.0086% if you assume one shared rate — a factor of roughly 2,300. The in-silico confidence scores did not distinguish the failures: per Anthropic’s report, designs against MBP and BBF-14 scored about the same as designs against targets that worked. TL;DR: An autonomous agent orchestrated a dozen open-source protein design tools and beat expert human hit rates on most targets. That happened, and it matters. But the headline number is a portfolio statistic, and almost nobody deploys a portfolio — they deploy against one target, one ticket, one customer. When we modeled the published spread instead of the mean, the expected experience of a single new target looked dramatically worse and dramatically noisier than 26.8% implies. And the pipeline’s own confidence scores could not tell in advance which regime it was in. That combination — high average, enormous variance, uncalibrated self-assessment — is the shape of nearly every agentic AI benchmark result you will read this year. What did Claude’s protein design campaign actually measure? Anthropic gave Claude (Mythos Preview and Opus 4.8) a set of 16 protein targets and asked for 30 minibinders per target per design arm — small proteins engineered to latch onto a target, the mechanism behind a large share of modern biologic drugs. With three arms running, most targets ended up with 90 designs in the lab. Fifteen targets produced usable measurements; one, mature GDF-8, was dropped because the target aggregated in the assay. The agent did not invent a protein model. It installed and ran existing open-source tools: backbones came from PXDesign (358 designs), RFdiffusion3 (267), Genie 3 (185), FreeBindCraft (135), BoltzGen (134), RFdiffusion (118) and Proteina-Complexa (100), among others, with sequences mostly from SolubleMPNN and filtering through ESMFold2 and Protenix v2. What was new was the layer above: choosing the epitope, installing the software, combining it across 24 distinct workflows, and ranking the output with no human touching a design decision. Validation was independent. Adaptyv Bio’s wet-lab case study, published 2026–08–19, reports that the designs arrived anonymized — the lab did not know which model produced which sequence — and were run on surface plasmon resonance at five target concentrations in duplicate. 95% of designs expressed. 354 of 1,320 bound. That is a genuinely strong result and we are not going to shave it down. Against RBX1, an open design competition had produced 9 binders from 245 entries (3.7%); Claude produced 28 from 90. Anthropic had the competition’s winning design rebuilt and measured on the same assay plate: it bound at 45 nM, while Claude’s best bound at 3.9 nM, roughly ten times tighter. Why doesn’t Claude’s protein design hit rate apply to your target? Because the targets are not interchangeable, and the published per-target counts make that impossible to ignore. Here are the eight per-target results disclosed across Anthropic’s post, Adaptyv Bio’s case study and The Decoder’s technical breakdown of the report: Binders per target, from the published counts: TREM2–72 of 90 designs bound (80.0%) VEGF-A — 54 of 90 (60.0%) IL-7Rα — 49 of 90 (54.4%) RBX1–28 of 90 (31.1%) TNFα — 12 of 150 (8.0%) BBF-14–3 of 90 (3.3%) 15-PGDH — 1 of 30 (3.3%) MBP — 0 of 90 (0.0%) Pooled, all 15 targets — 354 of 1,320 (26.8%) Now ask the question the coverage skipped: if every design really had a 26.8% chance of binding, how surprised should we be by the bottom row? The answer is that we should be about as surprised as it is possible to be. The probability of drawing zero successes in 90 independent trials at p = 0.268 is 0.732 raised to the 90th power, or 6.3 × 10⁻¹³ — one in roughly 1.6 trillion. TREM2 is equally impossible from the other direction: getting 72 or more hits in 90 trials at that rate has probability 1.1 × 10⁻²⁵. Neither of those is a fluke you explain away. They are the model being wrong. We ran a formal likelihood-ratio test comparing one shared rate against a model […]

This AimostAll brief summarizes the linked source so readers can scan AI developments quickly and jump to the original reporting when needed.

Read original source More agents news Anthropic page

Directory context

Tools, models, and guides to go deeper

Move from the headline to product evaluation with topic-matched tool pages, model references, and buyer guides.

Related coverage

More from this topic