AI in Drug R&D: At Last, the End of the Beginning

David Shaywitz
With the emergence of AI as an increasingly dominant force in industry (and society), much of the early discussion around AI in biopharma has tended to treat it largely as an abstraction. Advocates and investors insist AI is a transformative superhuman technology poised to domesticate R&D, while cynics, including many grizzled drug hunters, tend to assume it’s just the latest in a long series of overhyped bright and shiny objects, destined first to distract, then to disappoint.
What makes the conversation more interesting now is that we are beginning to get past the abstraction and caricatures, and head toward something more useful: a clearer capability profile of what AI actually does well, under what conditions, and what has to be built around it before technical possibility becomes a useful R&D tool.

Andreas Bender
Two recent publications help sharpen that picture: a Nature Reviews Drug Discovery Perspective I co-authored, led by Andreas Bender (and including his exceptional team as well as highly engaged collaborators such as Jack “Eroom’s Law” Scannell), and an essay (and associated references) by the brilliant Daphne Koller (more on her shortly).
What kind of problems is AI good at?
AI tends to perform especially well when several conditions line up: there are enough observations that genuinely bear on the question; the objective is specified well enough to optimize against; and feedback is sufficiently fast, cheap and trustworthy to learn from.
Protein structure prediction offered an unusually favorable version of this environment. Decades of experimental work supplied ground truth, and protein folding is sufficiently conserved that information can be pooled across enormous evolutionary distances. AlphaFold was a magnificent achievement, but it was also an achievement in a domain whose informational structure was unusually well suited to AI.
That distinction can get lost when AlphaFold is invoked as proof of what AI will do next. Most of the problems that determine whether a drug ultimately helps a patient have a much less generous information environment.
Our argument: start with what matters
Channeling Robert Solow, we ask in our Perspective why AI is visible everywhere in biopharma except in the productivity statistics. Evidence of clinically relevant impact is, as we put it, disappointingly limited.
There is also a sobering historical backdrop to the question. Drug R&D has absorbed one genuinely important technical advance after another, from recombinant proteins and genomics to high-throughput technologies, without an obvious corresponding improvement in aggregate R&D productivity. As Scannell points out, if AI actually bends that trajectory, it will have accomplished something that other extraordinarily important innovations did not. That sets a considerably higher bar for “transformation” than demonstrating impressive technical performance.
One place we start is by asking where an improvement would actually matter most. When we model equivalent improvements in speed, cost or probability of success across drug development, improving clinical success, especially around Phase II, has the greatest leverage. Yet much AI work remains concentrated considerably earlier, particularly around hit and ligand discovery. The reason is understandable: that is where labeled data are plentiful and the modeling problems are more tractable. But it risks becoming the familiar search for the keys under the streetlight, focusing on what we can readily model rather than what most needs improving.
Much of the problem is that both the things we’re measuring and the models we’re deriving sit pretty far from the things that matter most to drug developers and ultimately to patients. We describe part of this as the epistemic opacity of biology: we often don’t understand nearly well enough the relationship between the proxy we can readily measure and the clinical outcome we actually care about.
Biological data are also highly conditional. What happens in one cell type, at one concentration, in one experimental system may tell us much less than we would like about what happens in a particular patient. Sometimes even ostensibly the same experiment, using the same cell type and concentration, can give a materially different answer in another laboratory because variables that look incidental turn out to matter. This is also why simply pooling datasets can be deceptive: measurements generated in different assays, biological systems and experimental conditions may look comparable while capturing importantly different things.
This is why saying that biomedicine has enormous amounts of data can mislead. The problem is not simply quantity or technical quality. What matters, as Scannell has emphasized, is predictive validity: whether the thing we can measure actually tells us something useful about the outcome we care about. Generating exquisitely reproducible data around a weak proxy mostly allows us to characterize the wrong thing with greater precision. Our paper argues that the relationship between available data and useful outcomes needs to be established rather than assumed.
Brown & Goldstein (inspired by Magritte) observed that a gene sequence is not a drug; a ligand isn’t either. Finding a molecule that binds a target is useful, but binding is only one requirement. A drug also needs workable physicochemical and pharmacokinetic properties, sufficient selectivity, adequate exposure in the right tissue, an acceptable safety profile, and ultimately efficacy in people.
The numerical contrast is striking. Our paper notes that databases such as ChEMBL and PubChem contain more than a million known bioactive ligands, while the number of marketed drugs is only on the order of thousands. AI can become extraordinarily good at generating or identifying the former without necessarily producing correspondingly more of the latter. That ligand-drug distinction is a useful example of the larger problem: generating better ligands faster may be a genuine technical achievement without doing much to improve a team’s chances of delivering the drug it actually needs.
The same issue arises in how models themselves are evaluated. A key idea in our paper is the distinction between model validation and process validation. Two models can look nearly identical on a standard overall performance measure and behave completely differently in use: one might be much better at surfacing the few compounds worth advancing, while the other might be better at rejecting compounds with liabilities. There is no context-free definition of a good model; the relevant question is what the model is good for.
A model operates inside an actual project and is supposed to inform an actual decision. Does it change which experiment gets run, which compound advances, or which weak program gets killed? Without those connections to project context and downstream decision-making, better model performance need not translate into better drug discovery.

Suchi Saria, John C. Malone Associate Professor,
Johns Hopkins University
Tej Azad, Harlan Krumholz, and Suchi Saria make a strikingly parallel argument about clinical AI: readiness is a property of a model-task-context pairing rather than of a model. Their wonderfully compact prescription, “study form = use form,” could almost summarize ours: evaluate the tool with the users, information, constraints and consequences it will actually encounter.
Hence our call (familiar to TR readers) for less technology push and more science pull: start with the decisions and clinical outcomes we most need to improve, then work backward to the measurements and experimental systems that can actually inform them. A gorgeous model may represent a succès d’estime and a prestige publication. What matters is whether it improves consequential R&D decisions.
Koller: the missing biology and the slow scorecard
Koller puts even greater emphasis on what may be missing upstream: the relevant human biology itself. She is a particularly interesting person to hear make this case because she has spent her career at the AI frontier. I often joke that she’s probably taught AI to most of the Stanford grads now working at Google.

Daphne Koller, founder and CEO, insitro
Koller is also the founder of insitro, which I (unsuccessfully) pushed my pharma VC team to invest in at its founding (the corporate buy-in wasn’t there), and she was a 2017 guest on the Tech Tonics podcast Lisa Suennen and I hosted for six years.
Koller writes that she fully expects AI eventually to transform human health. But the “magic wand” vision, she argues, rests on a seductive assumption: that we already understand human biology well enough for a sufficiently clever AI to find the cures hidden in what we know. Her answer is wonderfully direct: “We don’t.” Aimed at biology we have only begun to measure and barely understand, she writes, AI may mostly help us generate failures faster.
This conclusion mirrors Scannell’s recent observation that feeding AI with data from poor biological models simply increases the number of wrong answers you can generate per second.
Koller divides drug discovery into three broad tasks: disease-to-mechanism, mechanism-to-drug, and drug-to-patient. Much of AI’s most visible progress has occurred in the middle. Given a biological mechanism we want to hit, increasingly powerful methods can now design proteins, small molecules and other therapeutic agents against it. Her concern is that getting dramatically better at making molecules may leave the harder biological question largely untouched. We are getting pretty good at manufacturing keys, as she memorably puts it, but they are often keys for the wrong locks.
To be sure, clinical efficacy failure is not synonymous with choosing the wrong mechanism. Koller’s own evidence review acknowledges that inadequate target engagement or tissue exposure, as well as problems with trial design or execution, can also sink a program. But her broader point is well supported: binding chemistry is much less often the limiting problem than it once was, while stronger evidence that a target genuinely drives disease substantially improves the odds of clinical success.
Koller sees one consequence in the remarkable crowding around biology where industry already has confidence. About a quarter of drug-target pairs in the 2024 pipeline were concentrated around only 37 targets, each with more than 50 drugs directed against it. This occurred even as the total pipeline nearly doubled over the preceding decade, while the annual number of novel targets entering the pipeline fell sharply.
There is also a deeper data problem. Huge biological datasets can contain remarkably little coverage of the biological space we actually need to understand. Koller’s companion analysis makes the distinction nicely: bytes and coverage do not move together. Hundreds of millions of single cells sound enormous, but those observations sample relatively few perturbations across relatively few cellular contexts, before adding disease state, dose, time or combinations of perturbations.
But (to borrow a phrase immortalized by Commentary’s Abe Greenwald), it’s worse than that. The biology that matters most can also be precisely the least transferable. Protein folding generalizes across species unusually well; neurodegenerative disease does not. The diseases where progress has been slowest often sit toward the human-specific end of biology, where the measurements are most expensive, least available and most ethically constrained to obtain. There is too much context and idiosyncrasy, Koller argues, simply to reason it all out in the abstract. You have to measure it.
There is a silver lining, of sorts: Koller does not conclude that we first need to map every possible state of human biology. “We don’t need a universal causal model in order to make medicines,” she writes. Instead, we can build targeted models that make particular regions of biological space navigable toward a desired outcome.
Of course, while mechanistic understanding is a powerful route to better prediction, but it isn’t the only possible route. A strong phenotypic model, human genetics, a validated biomarker, or potentially an AI-discovered empirical pattern (as Obermeyer and colleagues recently described) might all provide useful predictive information even when our causal explanation remains incomplete. The broader requirement is a sufficiently reliable connection between what we can observe now and what will eventually happen in people.
Koller’s other especially useful contribution is the scorecard. AI systems advance rapidly when they can iterate against feedback that is fast, accurate and cheap. Code has a compiler, she notes. Many molecular-design problems have rapid quantitative assays. Give a sufficiently capable optimization system a tight feedback loop and it can improve remarkably quickly.
Drug development tends to operate in almost the reverse environment. The ultimate scorecard is whether an intervention actually benefits a patient, and the definitive answer may require a human clinical trial that takes years, costs millions and is constrained by the pace of living biology. No amount of compute can make a neurodegenerative disease progress faster simply so we can learn sooner whether a drug worked (nor would we wish that clinical course on any patient or subject).
This makes the uneven progress of AI considerably less mysterious. Computational effort naturally gravitates toward the places where data and scorecards already exist. It also gives some of these efforts a certain self-satisfying, almost onanistic quality, pursuing what is available rather than what is actually most useful. AI seems likely to improve patient identification, site selection, regulatory drafting and data management, and those gains would be valuable and worth pursuing. But the pace of clinical development is limited by biology: patients still have to be treated and followed until the outcome matures, on a timetable set by the disease process rather than computation.
Can we create new pockets?
In thinking about opportunities for AI in R&D, there’s a useful phrase I’ve borrowed, somewhat loosely, from Stephen Wolfram: “pockets of reducibility.” I use it to describe bounded problems where the available information and feedback are good enough for computational approaches to gain meaningful traction. Protein structure is one obvious pocket; portions of molecular design and some operational aspects of clinical development are others.
But the opportunity is not limited to exploiting the pockets that already exist. Better science and better measurement can potentially create new ones, and this is where the prescriptions in our Perspective and Koller’s essay become especially interesting.
Our paper argues for starting with consequential R&D decisions and clinically relevant outcomes, then working backward to ask what measurements and experimental systems could actually inform them. That often means deliberately generating fit-for-purpose data rather than simply applying AI to whatever happens to be available. The critical step is establishing that an experimental proxy really predicts the in vivo or clinical outcome we care about; only then does it make sense to generate the data and models at scale. Over time, those preclinical models should also learn from what actually happens in the clinic.
Koller places greater emphasis on improving our understanding of disease biology itself. Her prescription is to measure more relevant human biology, particularly how biological systems respond to perturbation, and use AI to build targeted rather than universal causal models. Better biology, she argues, could in turn generate better clinical readouts: helping identify patients most likely to respond, confirming target engagement, and potentially detecting meaningful changes in disease biology earlier. Those readouts could allow trials to enroll better, read out sooner and catch failures earlier, rather than merely making their administration more efficient.
The common idea is appealing: if today’s scorecards are too slow or poorly connected to the outcomes we care about, perhaps better science can create better ones. But doing so is difficult, as I learned in my first industry role, working in Experimental Medicine at Merck nearly two decades ago. As a concept, biomarkers were seductive. But the amount of validation work required before a team or an R&D organization was willing to make a consequential (and often distasteful) internal decision based on one, potentially killing or materially redirecting a program because of a biomarker result, was substantial. Persuading a regulator to accept a biomarker as an endpoint for a clinical trial, the sort of thing that might actually permit a shorter study, is a far higher bar still.
A proxy is not a shortcut simply because it can be measured sooner. The work lies in establishing that the earlier signal actually predicts the later outcome we care about. That generally requires accumulating experience with both, which takes time, money and, in many cases, collaboration across programs or organizations.

Jack Scannell
There is an economic challenge as well. As Scannell has emphasized, a company that invents a proprietary molecule can generally capture much of the value it creates. A better translational model, assay or validated biomarker also may benefit competitors. Developing that validated proxy can therefore be enormously valuable to patients and the field while harder for the organization that paid to create it to monetize. This suggests a potentially important role for disease foundations, public funders and carefully designed precompetitive consortia. Our Perspective acknowledges that creating appropriately designed, harmonized datasets may require efforts beyond the capacity or incentives of any single organization.
From abstraction to capability
Read together, these analyses suggest a more useful way to think about AI in drug R&D. The relevant question is not whether “AI” works in some global sense, but where the technology has the information, problem structure and feedback required to help now, and where better science can create those conditions.

Subha Madhavan
There is also a distinct organizational dimension. I’ve argued elsewhere, with Pfizer’s Subha Madhavan, that realizing the value of a new technology generally requires more than substituting a new tool for an old one; over time, it changes how the work itself gets done. New capabilities can enable different ways of organizing information, collaborating and making decisions. But first, AI has to help scientists make better consequential decisions.
I find this evolution in the discussion encouraging. AI is increasingly being considered less as a higher power that will bring either salvation or ruin to drug R&D, and more as a distinctly earthbound capability that drug developers are beginning to develop a working intuition about: where it excels, where it struggles, and what it takes to use it well.



