The Funnel Is Bigger. Is It Better?

When a drug fails a clinical trial, nobody blames the chemist who designed it. The molecule simply joins the 90% of candidates that never reach patients. But when AI-designed drugs fail, the headline becomes The AI Drug Discovery Lie.

That asymmetry is unfair, but it points to a real question. AI is genuinely transforming the top of the drug discovery funnel: Insilico Medicine, Recursion, and Isomorphic Labs are each compressing years of work into months - through generative chemistry, automated phenomics, and structural prediction. The intuition is that more candidates should mean more successful drugs. History suggests otherwise.

Drug development has long been governed by Eroom's Law - the observation that drug development has become less efficient over time, with fewer new drugs approved per dollar spent despite major technological advances. Take combinatorial chemistry in the 1980s, which enabled simultaneous synthesis of thousands of compounds, or high-throughput screening in the early 1990s, which allowed robotic testing at industrial scale. Both expanded the top of the discovery funnel and accelerated early-stage research, but neither fundamentally changed downstream success rates. 

To be fair, Eroom’s Law is not simply evidence that drug discovery technology has failed. Biology itself has become harder. Today’s pipelines increasingly target complex diseases, rare patient populations, and mechanisms that were inaccessible decades ago. Even if AI dramatically improves molecular generation, however, it does not eliminate the central uncertainty of drug development: how a molecule actually behaves in a human body. If anything, increasingly difficult biology raises the value of technologies that generate more predictive human evidence before clinical trials begin.

Today’s drug discovery models are trained on large, heterogeneous datasets spanning molecular chemistry, structural biology, omics, imaging, and literature-derived data, much of it generated through decades of preclinical research. That lets models optimize compounds against experimentally defined endpoints. It does not close the gap between those endpoints and how a molecule actually behaves in a human body.

The track record so far reflects that gap. BenevolentAI’s BEN-2293, Recursion’s REC-994, and Verge Genomics’ VRG50635 all reached human testing. None has shown enough clinical efficacy to become an approved therapy. These outcomes should not be interpreted as evidence that AI has failed. Rather, they illustrate that AI has primarily improved hypothesis generation while the hardest problem - predicting human outcomes - remains downstream.

Insilico Medicine's rentosertib is the strongest signal yet that this can work. The AI system identified both the biological target and the molecule for idiopathic pulmonary fibrosis - a disease with no cure and a median survival of three to four years. In a Phase IIa study, patients on rentosertib showed a mean improvement in lung function of +98.4 mL, while placebo patients declined. In a disease that typically only gets worse, that's a reversal, not just a slowdown. Biomarker analysis further supported the predicted target, reinforcing that the underlying biological hypothesis - not just the chemistry - was correct. Insilico has now advanced rentosertib into Phase III; if it succeeds, it would be one of the first proof points that AI can improve not just the speed of drug discovery, but the quality of what comes out of it.

But one data point doesn’t remove the central constraint in drug development: even a perfect molecule-design engine still has to survive the slow, capital-intensive process of proving it works in humans.

Formation Bio argues the real constraint isn’t discovery at all - it’s clinical development, where promising drugs compete for scarce trial capacity. Under this view, faster discovery doesn’t remove the bottleneck. It just shifts it forward and lengthens the queue waiting to get into the clinic.

That reframing is why some of the more interesting companies today are attacking the bottom of the funnel instead of the top. Formation Bio acquires stalled or clinical-stage drugs and accelerates them with an AI-native development platform - it fast-tracked and out-licensed gusacitinib to Sanofi in just 2.5 years. Trially matches and enrolls patients into trials directly from unstructured medical records. Parallel Bio pushes even further upstream, using human immune organoids to generate preclinical evidence before a drug ever reaches a patient. Owkin approaches the bottleneck from another angle, mining multimodal patient data - histopathology, genomics, and clinical records - to identify the biomarkers that determine which patients a trial should even be built around.

The next competitive frontier in AI drug development may not be molecular generation at all. As designing molecules becomes abundant, the scarce resource shifts toward confidence: knowing which therapies deserve the enormous cost of clinical validation. Not all confidence is equal. Failed-trial data - dosing regimens, biomarker profiles, and adverse-event logs locked in pharma vaults - is valuable, but retrospective. It explains where past programs failed, not how a new therapy is likely to behave.

The deeper moat is prospective: longitudinal, patient-derived multi-omics data paired with deep phenotyping that links biological mechanisms to clinical outcomes in humans rather than mice. Combined with tools that identify the patients most likely to respond, this evidence turns clinical development from a wide net into a scalpel.

Owkin's collaboration with Servier offers an early illustration of this approach. It has evolved into an active effort to evaluate AI-identified, biomarker-defined patient subgroups for a therapy already in clinical development against Phase II data, testing whether these signals can sharpen trial design and patient selection.

But linking a biomarker to a mechanism is harder to establish than it sounds - which is exactly what Servier's retrospective analysis will have to prove out. A biomarker's association with a clinical outcome doesn't establish the underlying biological mechanism. Confounders, disease heterogeneity, and small effect sizes make causation difficult to distinguish from correlation. That's exactly why the moat holds: it isn't a dataset anyone can license or replicate overnight. It's built over years, at enormous capital cost, by generating proprietary longitudinal human data rather than borrowing someone else's.

What's changed is that generating and interpreting this kind of data is finally tractable. Sequencing and single-cell profiling costs have fallen orders of magnitude in a decade. Lab automation now lets organoid and iPSC experiments run at a throughput no manual wet lab could match. And models trained across modalities - genomic, transcriptomic, imaging - can find structure in high-dimensional biological noise that would have been statistically invisible before. AI didn't just make designing molecules faster; it made this deeper layer of human data usable. The moat was always theoretically valuable. It's only now buildable.

The moat also compounds. Every validated experiment and clinical outcome expands a proprietary feedback loop between biological prediction and human reality. Molecule generation becomes increasingly commoditized as more players adopt similar tools. Human evidence doesn't. That suggests the most valuable biotech AI platforms of the next decade may resemble learning systems more than traditional software: each patient treated, trial completed, or organoid experiment run becomes new proprietary evidence that compounds into a durable competitive advantage.

Formation Bio's work on sprifermin, a cartilage-regenerating drug acquired from Merck, shows the pattern. An earlier 500-patient study found that no patient on the highest dose needed a knee replacement within five years - a striking signal, but one that would normally take a five-to-ten-year trial to confirm, since structural joint changes are slow to prove. Instead, Formation Bio used an AI model trained on MRI imaging to identify patients likely to progress toward knee replacement, then enrolled the trial around that population - compressing a decade-long endpoint into one that could show up faster, with a clearer signal.

None of this predicts efficacy outright, and it isn't meant to. In more structured fields like toxicity, cheaper data and higher-throughput systems translate fairly directly into better models, because the failure modes are known and the benchmarks are clean. Efficacy doesn't get that same lift. What sprifermin shows instead is a narrower, more achievable win: not knowing whether a drug works before the trial, but knowing who to run the trial on - turning a wide, slow bet into a sharper, faster one.

So the funnel is bigger. Whether it's better depends on where the industry chooses to point its models. Aimed at the top, generating more candidates, AI runs into the same wall combinatorial chemistry and high-throughput screening hit decades ago. Aimed at the bottom - at who to trial, and how confidently - it targets the one constraint that's held for fifty years: not a shortage of molecules, but a shortage of evidence about how they behave in a human body.