16
Aug
2026

The Narrative Machine: What AI Really Adds to Personal Health Data

David Shaywitz

For years, the Quantified Self movement had a stubborn problem: our devices became increasingly good at generating data but remained remarkably poor at telling us what any of it meant. Generative AI now dangles a tantalizing possibility: that it can tease genuinely useful personal insights from all those measurements and hand them back to us in a form we can immediately understand and readily apply.

A recent Wall Street Journal article introduced us to what may be the latest iteration of the movement: “datamaxxers.” These are health-conscious technophiles who are feeding ever-larger portions of their lives into AI, seeking highly curated and customized insight in return.

The inputs can include the familiar streams from Garmin, Oura, WHOOP and Strava, including workouts, sleep, heart rate and heart-rate variability, along with calendars, meeting transcripts, emails, clinical records and even text messages. The hope is that AI can integrate the data, connect the dots, and just like that deliver exquisitely personalized health advice, one of digital medicine’s eternal promises.

The Journal introduces us to a datamaxxing runner named Julian Flieller. Before AI, his fitness tracker might tell him that his heart-rate variability was 60, a number that generated, as he put it, “more stress than clarity.” Now he supplies Claude with his workout, sleep and other biometric information and asks what it means for his training. Instead of receiving another isolated score, he gets an explanation –- a story.

Others have gone further. One runner built an AI dashboard that integrates Garmin, Strava and Oura data and helps her connect a sluggish run with poor sleep or the onset of her menstrual cycle. An engineer in India combined his Google Calendar with heart-rate data and asked Claude to determine which colleague was causing him stress. Sure enough, the AI fingered a specific co-worker as a “prime suspect.”

Once More, With AI

The original Quantified Self movement emerged around 2007, propelled by the growing ability of smartphones and wearables to capture information that previously would have been difficult or impossible to gather continuously. Its appealing premise was “self-knowledge through self-tracking.”

The first iteration largely fizzled. Avid self-trackers found themselves drowning in data but lacking in insight. After years spent tracking his activity, work and sleep, former Wired editor Chris Anderson famously concluded in 2016 that the effort had proved essentially pointless, yielding “no non-obvious lessons or incentives.”

I wrote for TR about the arrival of what I called “Quantified Self 2.0” in 2021. By then, the original Fitbits and Jawbones had given way to far more sophisticated technologies offering continuous glucose measurements, heart-rate variability, sleep staging and other physiological assessments. Yet the nagging question remained: were we actually learning anything more useful, or simply repackaging old approaches with a slicker interface and the nascent promise of AI, while continuing to preach the dubious familiar gospel that more data inevitably leads to better insight?

Five years later, generative AI appears to offer the latest wrinkle: a conversational superintelligence capable of looking across all our data and patiently explaining exactly what they mean.

The Great Technology Plan for Wellness, as TR readers will recall, is basically: inhale data Þ metabolize data with AI Þ deliver personalized coaching that improves as the data grow and the models sharpen. It sounds plausible, rather like the perpetually aspirational “learning health system”: a deeply appealing idea that has spent decades seeming both obviously right and frustratingly difficult to realize.

I also understand the appeal. I tried using WHOOP at least two separate times over the years, repeatedly hoping it would eventually tell me something sufficiently interesting or useful to justify my investment. Each time I gave up, because despite the volume of information, and most recently the eagerness of the associated AI to offer interpretation and advice, I was learning remarkably little that meaningfully changed what I did.

Lots of Data, Limited Insight

Kevin Hall, the distinguished former NIH metabolism researcher and author of Food Intelligence, offers a useful reality check.

Continuous glucose monitors (CGMs) seem tailor-made for personalized nutrition. They generate exquisitely detailed information about how glucose changes throughout the day, including after particular meals. It seems logical that enough measurements, analyzed intelligently, could reveal precisely which foods are good or bad for you.

Researcher Kevin Hall

Yet Hall has been particularly skeptical of using CGMs to generate precision dietary advice in people without diabetes. The instruments unquestionably generate data. The issue is whether those data contain the stable, individualized signal needed to support the recommendations drawn from them. As TR readers will recall, the ability of CGMs to offer advice to non-diabetics that is both meaningfully personalized and scientifically valid remains remarkably limited.

This distinction is easy to lose because modern technology can produce such an astonishing volume of measurements. A wearable might collect hundreds or thousands of physiological observations every day. Add sleep, exercise, diet, glucose, blood testing, calendar information and subjective reports, and the system can possess an extraordinarily rich empirical description of an individual. Even so, these abundant data do not necessarily provide the information needed to reliably predict the outcomes we care about, much less how those outcomes would change under a different set of circumstances.

At the population level, many wearable measurements have established relationships to health. What is far less established is whether a particular constellation of measurements supports a highly specific inference about a particular person, such as whether a lower HRV (heart rate variability) number and a poor night’s sleep explain a slow run or mean you should skip today’s hard workout. The measurements are abundant; the outcomes needed to validate many of these individualized predictions are far less so, and the relevant feedback often arrives slowly. We can therefore become measurement-rich while remaining surprisingly evidence-poor.

Professor Nathan Price

Additional personal data can sometimes yield genuinely useful insight. In a recent npj Digital Medicine Comment with Nathan Price and Dan Sodickson (discussed in TR), we argued that richer longitudinal and multimodal context could make health screening substantially more informative, in part by allowing us to detect meaningful changes from an individual’s own baseline rather than repeatedly comparing each measurement with a population norm.

Dr. Daniel Sodickson

But more data are valuable only to the extent that they add information. Sometimes repeated measurements reveal a meaningful trajectory that a single measurement would miss; sometimes they add surprisingly little. Complementary biological, imaging and clinical signals may add still more. AI is most valuable when it can integrate this context in ways that genuinely sharpen the inference.

Researcher Andreas Bender

There is also a deeper problem, familiar from drug R&D. In a recent Nature Reviews Drug Discovery Perspective (also reviewed in TR), Andreas Bender and colleagues, including Jack Scannell and me, discussed how biological datasets can be enormous while the outcomes we actually care about, such as clinical trial results or FDA approval, remain comparatively sparse, slow to emerge and highly context-dependent. Daphne Koller has described a related “data chasm”: huge datasets may sample remarkably little of the biological space we need to understand. Bytes and coverage do not necessarily move together.  (This was reviewed in TR as well.)

Daphne Koller, founder and CEO, insitro

And learning from even extraordinarily rich measurements can be stymied by the challenge (discussed in our NRDD paper) of epistemic opacity: intrinsic biological complexity can make it difficult to reliably infer the outcomes we care about from the things we can readily measure. AlphaFold can predict how a protein folds with extraordinary accuracy, for example, but structure alone does not tell us what role that protein plays in human disease or whether altering it will make a patient better. In many settings, the limiting factor may therefore be less the sophistication of the AI than the information content of the data and the distance between what we can measure and what we are asking the system to infer.

At the moment, much of what passes for precision wellness still seems like precision health theater. The specificity of the recommendation can wildly exceed what the underlying evidence can credibly support.

However, there may be another opportunity here: even when AI isn’t uncovering a consequential personalized insight, it can be remarkably good at extracting a coherent narrative from, or perhaps imposing one on, a pile of disparate data.

The Narrative Machine

Flieller’s HRV of 60 offers a prismatic example. The fitness tracker gave him a number that, in isolation, meant very little to him. The AI can place that number alongside his sleep, training load, recent performance and work schedule, then offer an account of how those observations might fit together. Perhaps his HRV has fallen while his sleep has deteriorated, his recent training load has risen, and his work calendar has been unusually demanding. Perhaps accumulated fatigue contributed to his sluggish run this morning, and today would be a sensible day to recover.

That account might be exactly right. But generative AI has also moved, almost imperceptibly, from aggregating observations to interpreting them, and from interpreting them to confidently explaining what happened.

This, I suspect, is the true superpower of current LLMs. For all the talk about godlike intelligence and discovering patterns beyond human comprehension, they are demonstrably exceptional narrative machines. Give them a messy collection of disparate observations and they can select what seems relevant, connect the pieces, incorporate outside knowledge and fashion a coherent, highly plausible account of what might be going on.

We tend to find coherent accounts enormously compelling. Daniel Kahneman spent much of his career documenting our susceptibility to stories that make available facts fit together, and how readily the feeling of coherence can become a feeling of understanding. Missing information bothers us less when the information we possess coalesces into a satisfying explanation.

Health AI adds another potent ingredient because the story is about you. Generic advice to sleep more, exercise regularly and eat thoughtfully is sensible but hardly riveting. An explanation that references your declining HRV, your shortened sleep, your increased mileage and your unusually stressful Tuesday feels entirely different. It appears not merely sensible but customized and insightful.

Sometimes, the LLM may genuinely do what the datamaxxing vision promises: integrate many disparate signals and identify a consequential individualized insight sufficiently well supported to change what you should do. In other cases, the AI may simply have taken several plausible facts and done what LLMs do exceptionally well: construct a remarkably plausible story.

The “prime suspect” identification is entertaining for a different reason. Perhaps the AI got it exactly right. But if you already know that a particular colleague drives you crazy, do you really need a wearable, a calendar integration and Claude to tell you so? This recalls an enduring temptation in digital health: using increasingly sophisticated technology to measure with exquisite precision something we could already determine well enough by simpler means. The question isn’t merely whether the new measurement or inference is correct, but whether it provides enough incremental information to improve an actual decision.

Early digital-health researchers, for example, explored accelerometers to quantify how much hospitalized patients were up and walking, information potentially relevant to recovery and discharge. While useful in some settings, the exercise could also amount to measuring with digital precision something a clinician could already learn by simply poking her head into the room.

Why the Story May Matter Anyway

There is, however, another source of potential value here. Even when the AI has not discovered a consequential new insight, the personalized account it constructs may influence how someone makes sense of an experience and what she does next.

Professor Alia Crum

I’ve written previously about the work of Stanford psychologist Alia Crum and others studying mindset, essentially the beliefs and expectations we bring to experiences and the meaning we assign to them. These effects are not confined to vague affirmations. In one classic study, hotel workers who were informed, accurately, that the physical work they already performed constituted meaningful exercise subsequently showed changes in weight, body fat and blood pressure, despite apparently not becoming more active.

In another striking experiment, people wearing Apple Watches were shown manipulated information about their activity. Participants who were falsely told that they were walking substantially less than they actually were subsequently felt worse, ate worse and developed higher blood pressure and heart rates, even though the underlying amount of walking barely changed. The measurement hadn’t changed, but the meaning attached to it had, and that interpretation became part of the experience.

Crum’s more recent work is especially relevant because she does not argue that people should simply substitute a pleasing story for an unpleasant truth. Complex situations frequently support more than one plausible reading. Her “metacognitive” approach emphasizes recognizing that our first, often reflexive, reading of an experience is only one possibility; appreciating that beliefs have consequences; and getting better at considering alternatives. The point is to become more deliberate and skillful in how we make sense of the many experiences that are open to more than one interpretation.

Health and fitness provide abundant examples of just these sorts of experiences. A difficult workout after a long week can become evidence that you are getting old and declining, or it can be understood as a demonstration of your ability to show up and persist even when you’re not feeling your best. Those different readings do not require pretending the underlying facts are different, but they can shape what happens next.

There may also be a simpler, almost accidental way in which wearables add value, one that has relatively little to do with the precision of the insights they generate. Simply wearing a WHOOP strap or an Oura ring can reinforce an identity: I am someone who exercises, trains seriously or pays attention to my health. The device can function as a kind of commitment device, making that health-oriented identity more salient and giving it visible expression.

AI offers an additional opportunity: to build deliberately on this narrative dimension rather than merely benefiting from it incidentally. It might help people notice and interpret experiences in ways that strengthen agency, our capacity to act intentionally and influence what happens next, and self-efficacy, our confidence that our actions can actually make a difference. A completed workout is not merely 300 calories expended or another ring closed on your Apple Watch; it can also become evidence that you’re someone who follows through, someone who can recover after missing a week, or simply someone capable of doing a little more than you thought.

Consider the difference between an AI that confidently asserts, “Your slow run was caused by elevated work stress, which suppressed your recovery,” and one that uses the same information to help you put the experience in perspective. Rather than simply supplying a more reassuring explanation, it might help you consider alternative interpretations and recognize evidence of persistence or progress that you might otherwise overlook, with the goal of helping you gradually become better at doing this for yourself.

Interpreter, Not Oracle

The original Quantified Self movement assumed that enough measurement would yield self-knowledge. Often it didn’t. The numbers piled up while the promised insights proved frustratingly elusive. Generative AI has now supplied something the movement conspicuously lacked: a remarkably fluent interpreter.

What is unusual about generative AI is that it does not have to wait until the evidence is strong enough to justify an inference before offering an explanation. This is both its attraction and its danger. It can transform an otherwise bewildering pile of numbers into a coherent, easily digestible account of your sleep, exercise, stress and physiology. But if you increasingly require an AI to tell you what every difficult workout, bad night’s sleep or stressful day means, you may be outsourcing precisely the capacity for interpretation that these tools should ideally help you develop. A technology intended to augment self-knowledge could instead leave you increasingly dependent on an external narrator.

At its best, AI could offer something considerably more valuable. Where the data genuinely support a non-obvious and consequential individualized insight, AI could surface it. Even where they do not, AI could help us consider alternative interpretations of our experiences and recognize evidence of progress and capability that we might otherwise miss. Over time, this could help cultivate the judgment, confidence and habits of interpretation that underlie agency and self-efficacy. (The opportunity to cultivate agency in this way is also part of what attracted me to my current work at Lore.)

Follow this strand forward and you arrive at what feels to me like a more thoughtful aspiration for AI in personalized health: AI that delivers real insight when real insight is available, while helping us become ever more capable interpreters and authors of our own lives rather than unhealthily dependent on the narrative machine.

15
Aug
2026

The Other Side of Bayes: Rethinking False Positives in Health Screening

David Shaywitz

Imagine screening 10,000 apparently healthy people for a disease that affects one person in a thousand. Let’s further assume the test is fairly good: 99% sensitive and 99% specific. (In case it’s been a while, this means the test detects 99% of true cases and correctly exonerates 99% of people without disease.)

Ten of those 10,000 people will actually have the disease –- that’s just the prevalence of the condition in the population.  The test will identify essentially all of them. That’s the good news. But among the 9,990 people who don’t have it, about 100 will nevertheless test positive. The screening program will therefore return roughly 110 positive results, of which only about 10 represent actual disease.

In other words, even with a test that is 99% sensitive and 99% specific, about nine out of ten positive results will be wrong.

This is the classic screening problem every physician learns in medical-school biostatistics and reliably encounters on every board exam. Bayes’ Theorem explains what’s going on. 

In essence, Bayes tells us how much a new piece of evidence should change what we believed beforehand. For a screening test, the interpretation therefore depends not only on the intrinsic characteristics of the test, but also on the probability that the person had the disease before the test was performed. When that “prior probability,” as it’s called, is very low, even an excellent test can deliver misleading results.

Comprehensive testing compounds the difficulty. Run enough measurements and the chance that something will be flagged rises quickly. This concern was memorably anticipated nearly two decades ago by Isaac Kohane, Daniel Masys and Russ Altman, who coined the term “incidentalome” to describe the proliferation of unexpected findings that could accompany genome-scale testing. They worried that clinicians would feel compelled to chase these abnormalities, subjecting patients to additional testing and morbidity while driving up costs, potentially with little corresponding benefit.

Their statistical illustration was striking. Suppose you evaluate someone with a genetic panel that effectively represents 10,000 independent (for the sake of argument) tests, each of which has a false-positive rate of just 0.01% — extraordinarily good individual performance. Even so, when you work through the math (or maths, if you’re in the U.K.), it turns out that more than 60% of the people tested would nevertheless wind up with at least one false-positive result.

This is why the proliferation of ever-more comprehensive testing, largely driven by consumer interest in disease pre-emption, appropriately concerns many doctors, including Eric Topol, and including me (and has for years).

At the same time, the appeal of such testing is readily understandable.  Most diseases of aging develop gradually over years, potentially offering an opportunity to identify trouble at an earlier and more actionable stage. But looking harder at apparently healthy people also risks incidental findings, anxiety, and unnecessary workups, occasionally (if inadvertently) leading to actual harm.

Rethinking Bayes

A few years ago, however, I came across an intriguing paper by Dr. Jacques Balayla at McGill that made me think about this problem differently.

Dr. Jacques Balayla

Balayla starts with the familiar dilemma. Most diseases amenable to screening are uncommon in the general population, meaning a significant proportion of positive screening results can be false positives that carry, as he notes, medical, psychological, social and practical consequences. He then asks whether anything can be done about what appears to be an intrinsic Bayesian limitation.

His answer is embedded in Bayes itself.

Bayesian reasoning is not static. New evidence updates what you believe, which is kind of the point. The resulting probability becomes the starting point for interpreting the next piece of evidence. Balayla shows mathematically how repeated testing, under appropriate assumptions, can progressively increase the positive predictive value of a screening result.

That insight led directly to a Comment that Nathan Price, Dan Sodickson and I recently published in npj Digital Medicine.

Professor Nathan Price

Nathan, a professor at the Buck Institute for Research on Aging, the CSO of Thorne, and co-author with systems-biology pioneer Lee Hood (see Luke Timmerman’s excellent biography of Hood, here) of The Age of Scientific Wellness, has spent more than a decade exploring what can be learned from dense longitudinal measurements of people beginning in health rather than disease.

Dan, a distinguished physician-scientist and MRI researcher at NYU, my former Harvard-MIT Health Sciences and Technology classmate, and author of The Future of Seeing, has pursued an analogous vision (so to speak) in imaging. (Dan is now Chief Medical Scientist of Function Health, though our paper was written before he joined.)

Dr. Daniel Sodickson

What Dan, Nathan, and I ask is fairly simple: as it becomes easier to collect serial molecular, imaging and clinical data, and as AI becomes increasingly capable of integrating them, can screening make greater use of individual context rather than evaluating each new result largely in isolation against a population reference range?

Insights from Imaging

Dan’s work provides an especially striking illustration.  In a terrific recent Big Brains podcast interview, he puts the central idea particularly well: “False positive rates aren’t fixed.” With enough prior context, Dan explains, the question becomes whether something represents a meaningful new change or simply that individual’s baseline.

In a retrospective analysis from his group involving nearly 30,000 patients followed for about a decade, AI models were trained to predict current and future risk of clinically significant prostate cancer. As the models were given more context in the form of prior imaging and clinical information, false positives fell dramatically while sensitivity remained around 90%. In the analysis reproduced in our paper, specificity rises from less than 30% with little prior information to more than 90% as richer prior information is incorporated. The work comes from a single center and requires external validation, but the underlying point is compelling: what today’s image means can depend enormously on what we already know about the person being imaged.

There is another potentially important aspect of Dan’s work. His group has demonstrated that prior imaging can help reconstruct a subsequent MR image from radically less newly acquired data. In the Big Brains podcast, Dan describes a broader “Everywhere Scanner” vision: cheaper, more accessible imaging that could fill the gaps between conventional scans, with AI focused on detecting meaningful change from an individual’s prior image.

If the approach generalizes, interval scans could eventually become quicker, cheaper and feasible on less specialized equipment. That creates the intriguing possibility of a virtuous cycle: easier imaging makes serial assessment more practical; serial assessment generates more individual context; and that context can make subsequent imaging both easier to obtain and more informative.

Insights from Molecular Measurement

Nathan has approached the problem from the molecular side. More than a decade ago, working with Hood and colleagues at the Institute for Systems Biology (ISB) in Seattle, he began exploring what they called “personal, dense, dynamic data clouds,” initially following 108 generally healthy people with repeated clinical and multi-omic measurements, whole-genome sequencing and activity tracking, work they reported in a 2017 Nature Biotechnology paper. They subsequently scaled the approach through Arivale, an ISB spinout focused on “scientific wellness.”

Across the five year program, nearly 5,000 participants contributed longitudinal data including genomic information, more than 1,000 blood-based measurements, microbiome analyses and Fitbit data, as described in a 2021 Nature Communications paper. The goal was to characterize individualized health trajectories that might reveal not simply whether someone falls outside a population norm, but whether that particular person is beginning to deviate from his or her own healthy state.

In a 2025 talk at the Buck, Nathan showed an evocative, though decidedly preliminary, illustration. Stored samples from three participants who subsequently developed metastatic cancer were retrospectively analyzed for the cancer-associated protein CEACAM5. Two of the participants had elevated values from the start, but in a third case, the value remained within the conventional normal range at an early time point, yet had jumped substantially from the individual’s previous measurement; it rose further before breast cancer was diagnosed. Nathan emphasized that there were only three cases, the measurements were retrospective, and this was emphatically not a validated screening test. But the example makes the intuition easy to grasp: a measurement — like the CEACAM5 value in the woman who developed breast cancer — can look normal compared with the population and still be distinctly abnormal for you.

Importantly, the broader principle does not rest on intriguing anecdotes or futuristic AI.

Consider prostate-specific antigen, or PSA. An elevated PSA can trigger MRI and biopsy, yet many such evaluations will not uncover clinically important cancer. Investigators have therefore developed reflex tests that bring additional biological information to bear. In an analysis from the Göteborg-2 prostate cancer screening trial cited in our paper, adding the 4Kscore, which combines several kallikrein measurements with clinical information, after an elevated PSA and before MRI would have reduced MRI use by 41%, biopsies by 28%, and diagnoses of low-grade cancers by 23%, while delaying detection of 4% of clinically significant cancers.

There is no AI magic in that example. The underlying model uses conventional statistics. The useful ingredient is additional, complementary information.

And “complementary” is important, though the relevant distinction is more subtle than simply serial versus different measurements. A trajectory of the same biomarker can itself be highly informative, as Nathan’s CEACAM5 example suggests. But that has to be demonstrated rather than assumed. There’s a useful counterexample from prostate screening: investigators hoped that PSA velocity, essentially the rate at which PSA changes over time, would improve prediction beyond the PSA level itself. In practice, it added surprisingly little independent predictive value. The lesson isn’t that trajectories don’t matter; it’s that some trajectories contain meaningful new information and others largely do not. Combining longitudinal change with genuinely complementary biological, imaging and clinical signals may offer still more opportunity to distinguish signal from noise.

AI and the Future of Screening

This is where AI becomes especially interesting.

Clinicians have always used context. We compare today’s scan with the old one. We interpret a laboratory result in light of earlier values, symptoms, medications and other findings. But there are obvious limits to how much disparate information a person can integrate. If an individual’s relevant context eventually includes years of imaging, hundreds or thousands of molecular measurements, physiological signals and clinical history, making sense of the whole may become as much a computational problem as a clinical one. AI could allow this sort of contextual integration to occur systematically and at scale.

None of this gives comprehensive screening a free pass. Earlier detection matters only if it leads to earlier useful intervention and ultimately better health, rather than simply more diagnoses, anxiety and procedures. Cost, overdiagnosis, liability and equity, as we discuss, remain substantial concerns. Context-aware screening will ultimately have to clear the same bar as any other medical intervention: prospective evidence that using it leaves patients better off.

But it’s difficult not to feel more hopeful. For years, I regarded Bayes primarily as the mathematical explanation for why broad screening of apparently healthy people was so fraught. Balayla helped me see the other side of the theorem. When disease is rare, a single positive result may tell us surprisingly little; but as additional context accumulates, sometimes through a revealing trajectory over time, sometimes through complementary signals, the picture can become much clearer.

The mathematics suggests that richer longitudinal context could make screening substantially more informative. The task now is to determine whether it can make patients healthier as well.

 

14
Aug
2026

AI in Drug R&D: At Last, the End of the Beginning

David Shaywitz

With the emergence of AI as an increasingly dominant force in industry (and society), much of the early discussion around AI in biopharma has tended to treat it largely as an abstraction. Advocates and investors insist AI is a transformative superhuman technology poised to domesticate R&D, while cynics, including many grizzled drug hunters, tend to assume it’s just the latest in a long series of overhyped bright and shiny objects, destined first to distract, then to disappoint.

What makes the conversation more interesting now is that we are beginning to get past the abstraction and caricatures, and head toward something more useful: a clearer capability profile of what AI actually does well, under what conditions, and what has to be built around it before technical possibility becomes a useful R&D tool.

Andreas Bender

Two recent publications help sharpen that picture: a Nature Reviews Drug Discovery Perspective I co-authored, led by Andreas Bender (and including his exceptional team as well as highly engaged collaborators such as Jack “Eroom’s Law” Scannell), and an essay (and associated references) by the brilliant Daphne Koller (more on her shortly).

What kind of problems is AI good at?

AI tends to perform especially well when several conditions line up: there are enough observations that genuinely bear on the question; the objective is specified well enough to optimize against; and feedback is sufficiently fast, cheap and trustworthy to learn from.

Protein structure prediction offered an unusually favorable version of this environment. Decades of experimental work supplied ground truth, and protein folding is sufficiently conserved that information can be pooled across enormous evolutionary distances. AlphaFold was a magnificent achievement, but it was also an achievement in a domain whose informational structure was unusually well suited to AI.

That distinction can get lost when AlphaFold is invoked as proof of what AI will do next. Most of the problems that determine whether a drug ultimately helps a patient have a much less generous information environment.

Our argument: start with what matters

Channeling Robert Solow, we ask in our Perspective why AI is visible everywhere in biopharma except in the productivity statistics. Evidence of clinically relevant impact is, as we put it, disappointingly limited.

There is also a sobering historical backdrop to the question. Drug R&D has absorbed one genuinely important technical advance after another, from recombinant proteins and genomics to high-throughput technologies, without an obvious corresponding improvement in aggregate R&D productivity. As Scannell points out, if AI actually bends that trajectory, it will have accomplished something that other extraordinarily important innovations did not. That sets a considerably higher bar for “transformation” than demonstrating impressive technical performance.

One place we start is by asking where an improvement would actually matter most. When we model equivalent improvements in speed, cost or probability of success across drug development, improving clinical success, especially around Phase II, has the greatest leverage. Yet much AI work remains concentrated considerably earlier, particularly around hit and ligand discovery. The reason is understandable: that is where labeled data are plentiful and the modeling problems are more tractable. But it risks becoming the familiar search for the keys under the streetlight, focusing on what we can readily model rather than what most needs improving.

Much of the problem is that both the things we’re measuring and the models we’re deriving sit pretty far from the things that matter most to drug developers and ultimately to patients. We describe part of this as the epistemic opacity of biology: we often don’t understand nearly well enough the relationship between the proxy we can readily measure and the clinical outcome we actually care about.

Biological data are also highly conditional. What happens in one cell type, at one concentration, in one experimental system may tell us much less than we would like about what happens in a particular patient. Sometimes even ostensibly the same experiment, using the same cell type and concentration, can give a materially different answer in another laboratory because variables that look incidental turn out to matter. This is also why simply pooling datasets can be deceptive: measurements generated in different assays, biological systems and experimental conditions may look comparable while capturing importantly different things.

This is why saying that biomedicine has enormous amounts of data can mislead. The problem is not simply quantity or technical quality. What matters, as Scannell has emphasized, is predictive validity: whether the thing we can measure actually tells us something useful about the outcome we care about. Generating exquisitely reproducible data around a weak proxy mostly allows us to characterize the wrong thing with greater precision. Our paper argues that the relationship between available data and useful outcomes needs to be established rather than assumed.

Brown & Goldstein (inspired by Magritte) observed that a gene sequence is not a drug; a ligand isn’t either. Finding a molecule that binds a target is useful, but binding is only one requirement. A drug also needs workable physicochemical and pharmacokinetic properties, sufficient selectivity, adequate exposure in the right tissue, an acceptable safety profile, and ultimately efficacy in people.

The numerical contrast is striking. Our paper notes that databases such as ChEMBL and PubChem contain more than a million known bioactive ligands, while the number of marketed drugs is only on the order of thousands. AI can become extraordinarily good at generating or identifying the former without necessarily producing correspondingly more of the latter. That ligand-drug distinction is a useful example of the larger problem: generating better ligands faster may be a genuine technical achievement without doing much to improve a team’s chances of delivering the drug it actually needs.

The same issue arises in how models themselves are evaluated. A key idea in our paper is the distinction between model validation and process validation. Two models can look nearly identical on a standard overall performance measure and behave completely differently in use: one might be much better at surfacing the few compounds worth advancing, while the other might be better at rejecting compounds with liabilities. There is no context-free definition of a good model; the relevant question is what the model is good for.

A model operates inside an actual project and is supposed to inform an actual decision. Does it change which experiment gets run, which compound advances, or which weak program gets killed? Without those connections to project context and downstream decision-making, better model performance need not translate into better drug discovery.

Suchi Saria, John C. Malone Associate Professor,
Johns Hopkins University

Tej Azad, Harlan Krumholz, and Suchi Saria make a strikingly parallel argument about clinical AI: readiness is a property of a model-task-context pairing rather than of a model. Their wonderfully compact prescription, “study form = use form,” could almost summarize ours: evaluate the tool with the users, information, constraints and consequences it will actually encounter.

Hence our call (familiar to TR readers) for less technology push and more science pull: start with the decisions and clinical outcomes we most need to improve, then work backward to the measurements and experimental systems that can actually inform them. A gorgeous model may represent a succès d’estime and a prestige publication. What matters is whether it improves consequential R&D decisions.

Koller: the missing biology and the slow scorecard

Koller puts even greater emphasis on what may be missing upstream: the relevant human biology itself. She is a particularly interesting person to hear make this case because she has spent her career at the AI frontier. I often joke that she’s probably taught AI to most of the Stanford grads now working at Google.

Daphne Koller, founder and CEO, insitro

Koller is also the founder of insitro, which I (unsuccessfully) pushed my pharma VC team to invest in at its founding (the corporate buy-in wasn’t there), and she was a 2017 guest on the Tech Tonics podcast Lisa Suennen and I hosted for six years.

Koller writes that she fully expects AI eventually to transform human health. But the “magic wand” vision, she argues, rests on a seductive assumption: that we already understand human biology well enough for a sufficiently clever AI to find the cures hidden in what we know. Her answer is wonderfully direct: “We don’t.” Aimed at biology we have only begun to measure and barely understand, she writes, AI may mostly help us generate failures faster.

This conclusion mirrors Scannell’s recent observation that feeding AI with data from poor biological models simply increases the number of wrong answers you can generate per second.

Koller divides drug discovery into three broad tasks: disease-to-mechanism, mechanism-to-drug, and drug-to-patient. Much of AI’s most visible progress has occurred in the middle. Given a biological mechanism we want to hit, increasingly powerful methods can now design proteins, small molecules and other therapeutic agents against it. Her concern is that getting dramatically better at making molecules may leave the harder biological question largely untouched. We are getting pretty good at manufacturing keys, as she memorably puts it, but they are often keys for the wrong locks.

To be sure, clinical efficacy failure is not synonymous with choosing the wrong mechanism. Koller’s own evidence review acknowledges that inadequate target engagement or tissue exposure, as well as problems with trial design or execution, can also sink a program. But her broader point is well supported: binding chemistry is much less often the limiting problem than it once was, while stronger evidence that a target genuinely drives disease substantially improves the odds of clinical success.

Koller sees one consequence in the remarkable crowding around biology where industry already has confidence. About a quarter of drug-target pairs in the 2024 pipeline were concentrated around only 37 targets, each with more than 50 drugs directed against it. This occurred even as the total pipeline nearly doubled over the preceding decade, while the annual number of novel targets entering the pipeline fell sharply.

There is also a deeper data problem. Huge biological datasets can contain remarkably little coverage of the biological space we actually need to understand. Koller’s companion analysis makes the distinction nicely: bytes and coverage do not move together. Hundreds of millions of single cells sound enormous, but those observations sample relatively few perturbations across relatively few cellular contexts, before adding disease state, dose, time or combinations of perturbations.

But (to borrow a phrase immortalized by Commentary’s Abe Greenwald), it’s worse than that. The biology that matters most can also be precisely the least transferable. Protein folding generalizes across species unusually well; neurodegenerative disease does not. The diseases where progress has been slowest often sit toward the human-specific end of biology, where the measurements are most expensive, least available and most ethically constrained to obtain. There is too much context and idiosyncrasy, Koller argues, simply to reason it all out in the abstract. You have to measure it.

There is a silver lining, of sorts: Koller does not conclude that we first need to map every possible state of human biology. “We don’t need a universal causal model in order to make medicines,” she writes. Instead, we can build targeted models that make particular regions of biological space navigable toward a desired outcome.

Of course, while mechanistic understanding is a powerful route to better prediction, but it isn’t the only possible route. A strong phenotypic model, human genetics, a validated biomarker, or potentially an AI-discovered empirical pattern (as Obermeyer and colleagues recently described) might all provide useful predictive information even when our causal explanation remains incomplete. The broader requirement is a sufficiently reliable connection between what we can observe now and what will eventually happen in people.

Koller’s other especially useful contribution is the scorecard. AI systems advance rapidly when they can iterate against feedback that is fast, accurate and cheap. Code has a compiler, she notes. Many molecular-design problems have rapid quantitative assays. Give a sufficiently capable optimization system a tight feedback loop and it can improve remarkably quickly.

Drug development tends to operate in almost the reverse environment. The ultimate scorecard is whether an intervention actually benefits a patient, and the definitive answer may require a human clinical trial that takes years, costs millions and is constrained by the pace of living biology. No amount of compute can make a neurodegenerative disease progress faster simply so we can learn sooner whether a drug worked (nor would we wish that clinical course on any patient or subject).

This makes the uneven progress of AI considerably less mysterious. Computational effort naturally gravitates toward the places where data and scorecards already exist. It also gives some of these efforts a certain self-satisfying, almost onanistic quality, pursuing what is available rather than what is actually most useful. AI seems likely to improve patient identification, site selection, regulatory drafting and data management, and those gains would be valuable and worth pursuing. But the pace of clinical development is limited by biology: patients still have to be treated and followed until the outcome matures, on a timetable set by the disease process rather than computation.

Can we create new pockets?

In thinking about opportunities for AI in R&D, there’s a useful phrase I’ve borrowed, somewhat loosely, from Stephen Wolfram: “pockets of reducibility.” I use it to describe bounded problems where the available information and feedback are good enough for computational approaches to gain meaningful traction. Protein structure is one obvious pocket; portions of molecular design and some operational aspects of clinical development are others.

But the opportunity is not limited to exploiting the pockets that already exist. Better science and better measurement can potentially create new ones, and this is where the prescriptions in our Perspective and Koller’s essay become especially interesting.

Our paper argues for starting with consequential R&D decisions and clinically relevant outcomes, then working backward to ask what measurements and experimental systems could actually inform them. That often means deliberately generating fit-for-purpose data rather than simply applying AI to whatever happens to be available. The critical step is establishing that an experimental proxy really predicts the in vivo or clinical outcome we care about; only then does it make sense to generate the data and models at scale. Over time, those preclinical models should also learn from what actually happens in the clinic.

Koller places greater emphasis on improving our understanding of disease biology itself. Her prescription is to measure more relevant human biology, particularly how biological systems respond to perturbation, and use AI to build targeted rather than universal causal models. Better biology, she argues, could in turn generate better clinical readouts: helping identify patients most likely to respond, confirming target engagement, and potentially detecting meaningful changes in disease biology earlier. Those readouts could allow trials to enroll better, read out sooner and catch failures earlier, rather than merely making their administration more efficient.

The common idea is appealing: if today’s scorecards are too slow or poorly connected to the outcomes we care about, perhaps better science can create better ones. But doing so is difficult, as I learned in my first industry role, working in Experimental Medicine at Merck nearly two decades ago. As a concept, biomarkers were seductive. But the amount of validation work required before a team or an R&D organization was willing to make a consequential (and often distasteful) internal decision based on one, potentially killing or materially redirecting a program because of a biomarker result, was substantial. Persuading a regulator to accept a biomarker as an endpoint for a clinical trial, the sort of thing that might actually permit a shorter study, is a far higher bar still.

A proxy is not a shortcut simply because it can be measured sooner. The work lies in establishing that the earlier signal actually predicts the later outcome we care about. That generally requires accumulating experience with both, which takes time, money and, in many cases, collaboration across programs or organizations.

Jack Scannell

There is an economic challenge as well. As Scannell has emphasized, a company that invents a proprietary molecule can generally capture much of the value it creates. A better translational model, assay or validated biomarker also may benefit competitors. Developing that validated proxy can therefore be enormously valuable to patients and the field while harder for the organization that paid to create it to monetize. This suggests a potentially important role for disease foundations, public funders and carefully designed precompetitive consortia. Our Perspective acknowledges that creating appropriately designed, harmonized datasets may require efforts beyond the capacity or incentives of any single organization.

From abstraction to capability

Read together, these analyses suggest a more useful way to think about AI in drug R&D. The relevant question is not whether “AI” works in some global sense, but where the technology has the information, problem structure and feedback required to help now, and where better science can create those conditions.

Subha Madhavan

There is also a distinct organizational dimension. I’ve argued elsewhere, with Pfizer’s Subha Madhavan, that realizing the value of a new technology generally requires more than substituting a new tool for an old one; over time, it changes how the work itself gets done. New capabilities can enable different ways of organizing information, collaborating and making decisions. But first, AI has to help scientists make better consequential decisions.

I find this evolution in the discussion encouraging. AI is increasingly being considered less as a higher power that will bring either salvation or ruin to drug R&D, and more as a distinctly earthbound capability that drug developers are beginning to develop a working intuition about: where it excels, where it struggles, and what it takes to use it well.

11
Aug
2026

Genetic Medicines For Kidney Diseases: Jason Coloma on The Long Run

Jason Coloma is today’s guest on The Long Run.

Jason is the CEO of South San Francisco-based Maze Therapeutics. The company is developing small molecule drugs to halt or potentially reverse kidney diseases. The drug discovery work at Maze is grounded in human genetics, looking at large pools of data to see how variants of one kind or another can make people sick, or protect them from falling ill.

Jason Coloma, CEO, Maze Therapeutics

Maze went public in early 2025, after surviving several lean years of biotech financing. An important partnership unraveled after the federal government sought to block it on antitrust grounds. Maze found a way to bounce back, and enticed a new partner to support that program. Those battles have strengthened the company culture. Maze is now in position to find out in the next year or so whether its two lead programs – an APOL1 directed kidney drug and another aimed at SLC6a19 – are good enough to pass Phase II clinical trial scrutiny.

Like many biotech industry leaders, Jason didn’t take a linear path into this moment of possibility in medicine. That’s part of what makes his story, and Maze’s story, so interesting.

The Long Run is sponsored by

 

 

So far 2026 has been a year of transition for healthcare and life sciences. Over the first half of this year, the AlphaSense healthcare research team analyzed the individual stories moving the needle across the life sciences ecosystem as part of the Sector Spotlight series. Download the eight spotlights here. They featuring the high-stakes battles for new drug markets (like oral GLP-1s and targeted radioligand therapies), a resurgent biotech IPO market, and more.

Follow the sector’s key debates during 2H26 and beyond, with the latest Sector Spotlights directly in AlphaSense.

 

Quick question: when was the last time a CRO told you what something costs before you signed an NDA, sat through two scoping calls, and waited a week for a Scope of Work? In bioanalysis, pricing gets treated like a closely-guarded secret. Dash Bio thinks that’s backward. It publishes pricing publicly, on its website, and even has a calculator that shows how the price changes if you redesign your study. You can see what your study costs before you talk to a human being. No surprise line items, no “it depends”, no games. Dash was built to be the bioanalysis company that’s actually straightforward to work with: transparent pricing, guaranteed timelines, and data in days, not months.

If you’re tired of pricing that feels like a negotiation every single time, run your bioanalysis with Dash.

Visit dash.bio/pricing and see exactly how much your next bioanalysis study will cost.

Now please enjoy this episode of The Long Run with Jason Coloma.

28
Jul
2026

Inhalable Medicines for Severe Lung Diseases: Lyn Baranowski on The Long Run

Lyn Baranowski is today’s guest on The Long Run.

Lyn is the CEO of Boston-based Avalyn Pharma. The company is developing inhalable drugs for rare, severe lung diseases, such as pulmonary fibrosis and other interstitial lung diseases.

Lyn Baranowski, CEO, Avalyn Pharma

Avalyn has taken a couple of small molecules that are given orally, and remade them into inhalable drugs so they can be more effectively delivered where they need to go, deep in the lungs. By concentrating delivery, the idea is that Avalyn will be able to lower the doses needed and make treatments that are easier for patients to tolerate over the long term.

In this episode you’ll hear Lyn’s enthusiasm for developing better treatments for respiratory diseases. It’s one of the therapeutic areas that doesn’t capture a lot of news coverage, but lung diseases of one kind of another affect many people.

The Long Run is sponsored by:

Under unprecedented time and cost pressures, many life sciences teams have reached for off-the-shelf AI tools to stay afloat. Yet they have discovered that generic AI tools introduce new risks, such as hallucination. Tools that cannot withstand scientific or regulatory scrutiny are a liability in MLR review, portfolio committees, and investment decisions where every claim must be traceable to a source.

More than isolated tools, teams need a purpose-built intelligence layer designed for the rigor that the industry demands. Across life sciences functions, platforms like AlphaSense are streamlining the entire decision cycle. Download this report that covers the top AI use cases that have transformed how the industry’s leading teams reach defensible answers.

 

Quick question: when was the last time a CRO told you what something costs before you signed an NDA, sat through two scoping calls, and waited a week for a Scope of Work? In bioanalysis, pricing gets treated like a closely-guarded secret. Dash Bio thinsk that’s backward. It publishes pricing publicly, on its website, and even has a calculator that shows how the price changes if you redesign your study. You can see what your study costs before you talk to a human being. No surprise line items, no “it depends”, no games. Dash was built to be the bioanalysis company that’s actually straightforward to work with: transparent pricing, guaranteed timelines, and data in days, not months.

If you’re tired of pricing that feels like a negotiation every single time, run your bioanalysis with Dash. Visit dash.bio/pricing and see exactly how much your next bioanalysis study will cost.

 

Now please enjoy this conversation with Lyn Baranowski on The Long Run.

14
Jul
2026

A Schizophrenia Drug Patients Can Stick With: Heather Turner on The Long Run

Heather Turner is today’s guest on The Long Run.

Heather is the CEO of New York-based LB Pharmaceuticals. It’s developing treatments for serious mental health disorders such as schizophrenia, depression, and bipolar disorder.

Heather Turner, CEO, LB Pharmaceuticals

The company is running late-stage clinical trials for its lead drug candidate, LB102. It’s a methylated derivative of amisulpride, a benzamide antipsychotic. It’s approved in more than 50 countries, but never made it to the market in the US before going generic.

The LB modification is supposed to make the compound more efficiently cross the blood-brain barrier, so it can be given in lower doses, and once a day instead of twice a day like with amisulpride. It should be easier for patients to adhere to. That’s always a challenge, and it’s especially important for mental health, where many patients struggle to stay on their meds consistently.

Heather comes to this opportunity in her second go-round as a CEO. The first was with Carmot Therapeutics, a company that developed GLP-1 receptor agonist based treatments for weight loss. Carmot was acquired by Roche for $2.7 billion upfront, plus $400 million in milestone payments, in 2024.

Heather came to the biotech industry from a career in law. Her path to become a successful biotech CEO wasn’t obvious from the early days, which is part of what makes her story interesting.

The Long Run is sponsored by AlphaSense

 

Under unprecedented time and cost pressures, many life sciences teams have reached for off-the-shelf AI tools to stay afloat. Yet they have discovered that generic AI tools introduce new risks, such as hallucination. Tools that cannot withstand scientific or regulatory scrutiny are a liability in MLR review, portfolio committees, and investment decisions where every claim must be traceable to a source.

More than isolated tools, teams need a purpose-built intelligence layer designed for the rigor that the industry demands. Across life sciences functions, platforms like AlphaSense are streamlining the entire decision cycle. Download this report that covers the top AI use cases that have transformed how the industry’s leading teams reach defensible answers.

AND

Quick question: when was the last time a CRO told you what something costs before you signed an NDA, sat through two scoping calls, and waited a week for a Scope of Work? In bioanalysis, pricing gets treated like a closely-guarded secret. Dash Bio thinsk that’s backward. It publishes pricing publicly, on its website, and even has a calculator that shows how the price changes if you redesign your study. You can see what your study costs before you talk to a human being. No surprise line items, no “it depends”, no games. Dash was built to be the bioanalysis company that’s actually straightforward to work with: transparent pricing, guaranteed timelines, and data in days, not months.

If you’re tired of pricing that feels like a negotiation every single time, run your bioanalysis with Dash. Visit dash.bio/pricing and see exactly how much your next bioanalysis study will cost.

Now please enjoy this conversation with Heather Turner on The Long Run.

11
Jul
2026

Useful Fictions or Distorting Frames? How Models Shape Drug R&D

David Shaywitz

We rely on models — deliberate simplifications — to navigate, make sense of, and engage productively with an ever-more complicated world.

Patients understand their illnesses through what Dr. Arthur Kleinman called “explanatory models.” Physicians use models such as the pathophysiologic model or the biopsychosocial model to integrate their patients’ symptoms and fashion treatments. Biopharma R&D would be impossible without models as well. Every program rests on a series of them, at every stage of the R&D process.

The catch, as Alfred Korzybski observed nearly a century ago, is that the map is not the territory. Unless we recognize the assumptions built into the models we use, we are liable to be led astray — especially because we tend to reach for models that are cognitively accessible, experimentally tractable, or organizationally useful, whether or not they are especially predictive.

Enter Jack Scannell

In a recent interview with Pablo Lubroth of Decibio, Jack Scannell offers an eloquent and deeply resonant account of the role of models across R&D, highlighting not only their limitations but, more interestingly, why we continue to rely on them. He conveys a deep familiarity with the lived experience of R&D — the critical tacit knowledge that other biopharma veterans will immediately recognize.

Jack Scannell

Scannell, a longtime analyst of drug R&D productivity who now runs the early-stage biotech Etheros, is perhaps best known for coining “Eroom’s Law”: the depressing observation that, from 1950 to roughly 2010, the number of new drugs approved per inflation-adjusted billion dollars of R&D spending fell by about half every nine years. The name is “Moore” spelled backward — a pointed contrast with Moore’s Law, the emblem of exponential progress in computing.

Within R&D, Scannell highlights problems with both the financial and biological models on which the industry relies. I’ll focus mostly on the biological side and simply note that his assertion that “standard drug and biotech valuation models are blind to decision quality,” making companies less sensitive to the value of better decisions, especially early in development.

Reductive Models in Biology

The key biological question Scannell — and really everyone in drug R&D — wrestles with is how to think about something as complex as human biology in a way that allows you to develop a medicine, introduce it into the midst of all this complexity, and have reasonable confidence that it will improve a particular condition without doing harm elsewhere.

It takes a prodigious amount of chutzpah to believe we can do this. The many remarkable medicines that have been developed deserve to be celebrated.

In the words of Dr. Roger Perlmutter, former head of R&D at Merck, “we have no idea what we’re doing.” He continues, “It’s a bloody miracle if you ever make a drug that works.”

To manifest such miracles more reliably, biopharma — following the lead of medical science — has leaned heavily into reductionism: the idea that we can best understand life by breaking complex processes down and studying their individual components.

This has worked especially well for monogenic diseases, conditions that result from a single catastrophic glitch that can be targeted and, ideally, remediated.

It has also encouraged a focus on individual targets, often receptors, and on developing the most potent possible binders. These assays are especially attractive because they are amenable to scale: you can evaluate a huge number of candidates in a high-throughput fashion.

This approach can be effective. But it can also generate premature confidence, since identifying a great ligand is not the same as finding a promising drug candidate. Tight binding to an individual target may be useful to select for, while still failing to predict physiological promise in the organism as a whole.

When Clean Stories Meet Messy Biology

Highly reductionist approaches tend to struggle with complex diseases and psychiatric conditions, Scannell explains, noting “drugs don’t know they’re meant to bind a target,” and often turn out to bind multiple targets at once. Some of the most effective drugs we have, he says, may be more like “magic shotguns” than “magic bullets.”

In psychopharmacology, for instance, clozapine — the only drug approved specifically for treatment-resistant schizophrenia — binds dopamine, serotonin, histamine, muscarinic, and adrenergic receptors. After more than fifty years of research, its mechanism of action remains unclear and debated. Its clinical efficacy may depend on this promiscuity across receptor systems rather than any single interaction. 

Clozapine is also a reminder of what highly reductionist approaches can miss. A target-first assay tends to reward exceptional activity against a specified interaction. A phenotypic readout may capture a pattern of modest effects across many interactions — a profile that might look unimpressive receptor by receptor, yet prove meaningful in the organism.

The broader point is not confined to clozapine and psychopharmacology. Some TKIs developed against specific kinase targets turned out to inhibit panels of kinases, and their clinical activity may depend on this broader activity. “In large parts of psychopharmacology,” Scannell explains, “there’s no easy relationship between pattern of binding across receptors and phenotypic impact on people.” He adds that we can confuse “neat stories about targets, creation myths, with how drugs actually work.”

He also points out that a number of drugs approved before the target-centric era — medicines like metformin, the first antipsychotics, and the first antidepressants — have messy or contested mechanisms. Metformin has been attributed to AMPK activation, mitochondrial complex I inhibition, and effects on gut microbiota, with little consensus around a single unifying mechanism. Chlorpromazine and the first antidepressants were discovered empirically, with mechanistic stories built around them later.

Even targeted drugs with seemingly clear mechanisms can become more complicated over time. Anti-VEGF therapy, for example, emerged from Judah Folkman’s elegant insight that tumors depend on angiogenesis. But subsequent work, especially by Rakesh Jain, suggested an additional mechanism: anti-VEGF agents may transiently normalize chaotic tumor vasculature, improving delivery of chemotherapy. The original story was not so much wrong as incomplete — a reminder that even elegant mechanistic stories tend to get more complicated as we learn more.

A Little More Predictive Validity Goes a Long Way

Scannell has particularly emphasized the importance of improving disease models — increasing the “predictive validity” of the systems in which drug candidates are studied. He notes that while everyone recognizes that good models are better than bad models, the less obvious point is that “marginally better models are much better than marginally worse models.” A model that improves predictive ability even slightly can be more valuable than running ever more compounds through a less predictive system.

By contrast, he warns that “feeding AI with data from poor biological models simply increases the number of wrong answers you can generate per second.” He worries that some AI drug-development companies may industrialize a poor biological model because it is easiest to fit “into an AI-led design-make-test loop” — leading at best, he says, to virtual ligands optimized against a good in silico representation of an irrelevant biological system. Other companies, he encouragingly notes, appear more focused on using AI to develop better biological models.

Why Flawed Models Endure

Given the obvious value of improved models, why are they so hard to come by? Asked differently, what makes existing models so sticky?

The answer, Scannell suggests, comes down largely to incentives.

Academic researchers, he explains, are not rewarded for the slow work of model validation. You “do not get Nature papers by doing long, tedious, gritty work to evaluate whether a particular mouse model of disease X recapitulates the human pathology and then responds the same way to drugs that have been in people,” he says.

Industry researchers, for their part, desperately want better models but are not always incentivized to invest in developing them. A validated model may quickly become useful to everyone, eroding any competitive advantage. Moreover, Scannell explains, while industry scientists may privately acknowledge the limitations of their models, admitting this too directly could threaten the programs built on them.

At the portfolio level, simpler explanations often prevail over more nuanced accounts. Scannell shares a thoroughly relatable story about a colleague preparing to present a project to a portfolio management committee at a mid-sized pharma. The colleague had slides with “really complicated maps and models of biological pathways.” Before the meeting, he was told to come back with something much simpler: “three or four boxes and a couple of arrows,” otherwise “the board would never approve it.”

Context Matters

Perhaps Scannell’s most important observation involves a distinction he borrows from innovation scholar Fred Steward: the difference between “technoscientific” problems and “sociotechnical” problems.

A technoscientific problem, Scannell says, is “like the classic Apollo moonshot” — a problem that can be solved with enough engineers and enough funding.

Sociotechnical problems, by contrast, involve “a complicated mix of economic, regulatory, and process changes happening simultaneously over time.” He cites the building of modern healthcare systems as an example and believes AI in R&D likely falls into this category.

On the process side, he explicitly cites the classic work from Stanford scholar Paul David that I’ve discussed before, on how the arrival of electricity itself did not immediately improve factory productivity. The benefits came later, when electricity allowed factory work to be reorganized. The new technology mattered, but the impact was felt only after the structure of work changed around it.

“I suspect that no one really knows what the reconfigured factory is going to look like when it comes to AI in drug R&D,” Scannell wisely observes — a statement that could apply just as well to the application of contemporary AI across healthcare.

It also speaks to the remarkable opportunity now before us, from biopharma R&D to disease prevention: to use emerging technologies such as AI to reconfigure the work itself, improve the models on which we rely, and turn better models into better decisions — decisions that help us stay healthy longer and, when we become patients, receive treatments more likely to work.

30
Jun
2026

Precision Cancer Drugs and Combos: Troy Wilson on The Long Run

Troy Wilson is today’s guest on The Long Run.

Troy is the co-founder and CEO of San Diego-based Kura Oncology. The company, founded in 2014, has gone all the way through R&D to develop its own novel cancer drug. It’s called ziftomenib (brand name Komzifti). It’s an oral pill that inhibits a target called menin. Ziftomenib was approved by the FDA in November 2025, initially for a small set of patients with acute myeloid leukemia who have specific genetic alterations such as NPM1 mutations or KMT2A rearrangements.

Troy Wilson, CEO, Kura Oncology

This is the start. Kura hopes to demonstrate the drug can help a broader group of patients with AML – a notoriously tough to treat cancer.

Not many biotech startup companies make it all the way through the R&D gauntlet to develop an FDA approved medicine. Kura is also hopeful lightning will hit twice, with another drug called darlifarnib. It’s a farnesyl transferase inhibitor (FTI) inhibitor. Scientists at the company believe this entire class of FTI medicines – once written off as a pharma graveyard – is due for a comeback, especially as it logically ought to complement the new breakthrough RAS inhibitor, daraxonrasib, from Revolution Medicines.

Troy arrives at this auspicious moment in cancer R&D after a long and distinguished career. He co-founded three previous companies that ended up being successfully acquired – Intellikine, a PI3kinase inhibitor company acquired by Takeda; Ambrx, a developer of antibody-drug conjugates acquired by Johnson & Johnson; and Avidity Biosciences, the developer of antibody oligonucleotide conjugates for rare muscular diseases acquired by Novartis.

The Long Run is sponsored by:

AI has advanced from pilot programs into real biopharma workflows. But has it lived up to the hype? Drawing from an expert call with Dr. Paul Agapow, former director of data science at GSK, this report dissects whether AI is on track to make drug development faster, cheaper, and more successful. Learn how AI is reshaping drug development in real time, turning AI execution into a key competitive differentiator.

 

Quick question: when was the last time a CRO told you what something costs before you signed an NDA, sat through two scoping calls, and waited a week for a Scope of Work? In bioanalysis, pricing gets treated like a closely-guarded secret. Dash Bio thinsk that’s backward. It publishes pricing publicly, on its website, and even has a calculator that shows how the price changes if you redesign your study. You can see what your study costs before you talk to a human being. No surprise line items, no “it depends”, no games. Dash was built to be the bioanalysis company that’s actually straightforward to work with: transparent pricing, guaranteed timelines, and data in days, not months.

If you’re tired of pricing that feels like a negotiation every single time, run your bioanalysis with Dash. Visit dash.bio/pricing and see exactly how much your next bioanalysis study will cost.

Now please enjoy this conversation with Troy Wilson on The Long Run.

29
Jun
2026

The Rise, Fall, and Rise (and Fall…?) of Cell and Gene Therapy

Carl Schoellhammer, partner, DeciBio

More than $14 billion in announced acquisition value has flowed into in vivo cell therapy companies over the last 18 months. Lilly acquired Kelonia, AbbVie bought Capstan, BMS acquired Orbital, and AstraZeneca bought EsoBiotec, to name a few. 

These acquisitions demonstrate that we’re getting closer to the day when single injections that reprogram cells can fight disease without complex shipping logistics, waiting times and toxic preconditioning regimens of first-generation ex vivo cell therapies.

Over the last decade, the cell therapy field has experienced extraordinary breakthroughs and painful corrections. And yet, just as it has emerged from its first major boom-and-bust cycle with hard-won discipline, in vivo cell therapy is generating the kind of excitement that should, by now, feel familiar.

The first modern wave of cell and gene therapy was built on scientific breakthroughs. CAR-T therapies demonstrated that living medicines could deliver long-lasting remissions for cancer patients who were facing a death sentence. Early AAV therapies showed that a single treatment could meaningfully alter the course of genetic disease. It was heady stuff. Investors responded accordingly.

During the pandemic-era biotech boom, capital flooded the sector in search for the next company like Moderna, with a platform technology with broad potential for multiple diseases, like messenger RNA vaccines. Every company seemed to need a cell and gene therapy strategy.

Looking across DeciBio’s TheraTrack database, there was a period where the average indication being addressed (across oncology, rare disease) attracted more than four gene therapy assets and more than 10 cell therapy assets in development. Multiple companies chased the same diseases with limited differentiation. Capital poured in faster than market realities could justify, and scientific possibility increasingly became confused with commercial inevitability.

The sector delivered effective, generally durable treatments, not the one-and-done “cures” the market was promised. The resulting correction was severe, but necessary.

Today, several approved cell and gene therapy products have reached blockbuster status (Table 1). Sectors written off by some investors as dead in 2022 have steadily continued to attract smaller, but still meaningful investments and generate significant transaction value (Figure 1). The manufacturing challenges that once seemed existential are being converted, steadily, from scientific uncertainties into engineering problems.

The correction hurt. It also forced companies to step up their game.

Table 1: Blockbuster cell and gene therapy products

Figure 1: Investment and financings into AAV-based biotech companies

Which brings us to today’s newest source of excitement: in vivo cell therapy. Rather than extracting cells from a patient, engineering them in a manufacturing facility, and reinfusing them weeks later, why not engineer those cells directly inside the body? In theory, this approach could simplify treatment, reduce costs, shorten timelines, and dramatically expand access. 

Many of today’s talking points from companies echo the promises made during the first cell and gene therapy boom. Manufacturing will be simplified, costs will fall, and platform technologies will unlock dozens of indications. Investors and companies have flocked to the sector. But who is left to buy the next wave of companies?

The most logical strategic buyers have largely committed. That leaves a meaningfully thinner pool of potential acquirers for the companies now raising capital, building platforms, and conducting early-stage trials. Exit paths matter, and the number of natural buyers for the next cohort may be more limited than the current pace of investment implies.

Manufacturing Challenges Ahead 

One of the most common narratives surrounding in vivo cell therapy is that it eliminates manufacturing complexity. It is more complicated than that. In many cases, we may simply be trading one manufacturing challenge for another.

Consider a targeted LNP-based system. Production involves at least three distinct steps, each requiring purification and each introducing yield losses: synthesis of the mRNA cargo, formulation and encapsulation within the lipid nanoparticle, and conjugation of the targeting ligand to the LNP surface. Losses compound at every stage.

Beyond yield, the chemistry itself is not trivial. Conjugating antibodies or peptides to lipid nanoparticle surfaces introduces questions around conjugation efficiency, batch-to-batch consistency, and analytical characterization that the field is still working through.

Then there is the CDMO problem. Targeted LNP products sit at the intersection of lipid chemistry, biologics manufacturing, nanoparticle formulation, and conjugation science. While each capability exists independently, relatively few organizations have deep expertise across all of them within a single integrated process. As these programs scale, developers will need to navigate manufacturing execution and increasingly complex analytical, comparability, and regulatory expectations.

Longer Term Questions 

Clinical reality will also eventually arrive. As programs move into larger patient populations and longer follow-up periods, new questions will emerge around safety, durability, repeat dosing, and immune responses.

Cell and gene therapy has never been a story of failure. It has been a story of overestimation followed, more quietly, by real progress.

The danger today is that we once again mistake early promise for inevitability. If the field applies the lessons learned from AAV and CAR-T, the next decade could be transformative. If not, we may simply be boarding the next rollercoaster before the last one has finished its climb.

16
Jun
2026

Small Molecules To Correct Disease: Sri Kosuri on The Long Run

Sri Kosuri is today’s guest on The Long Run.

Sri is the co-founder and CEO of Emeryville, California-based Octant Bio. The company is developing oral small molecule drugs that are designed to correct protein misfolding and mistrafficking. Quite a few rare diseases, cancers, and metabolic disorders are thought to be amenable to this strategy.

Sri Kosuri, co-founder and CEO, Octant Bio

Octant’s work starts with a platform designed to capture relevant information about these cell processes that go awry. But like any platform company, it needs to demonstrate its value in the form of new drug candidates.

Octant’s lead drug candidate is for rhodopsin-associated autosomal dominant Retinitis Pigmentosa. It’s caused by a gene mutation which causes misfolded proteins to accumulate in the retina. That leads to progressive vision loss over time. The Octant drug candidate, which entered its first clinical trial in March, is intended to help improve low-light vision and halt disease progression.

Sri has a wide range of experiences in academia and in industry. He and his colleagues are seeking to build on recent advances in technology to discover oral small molecules that not only make a big difference for patients, but which ought to be widely accessible to the vast majority of people in need.

The Long Run is sponsored by AlphaSense

AI has advanced from pilot programs into real biopharma workflows. But has it lived up to the hype? Drawing from an expert call with Dr. Paul Agapow, former director of data science at GSK, this report dissects whether AI is on track to make drug development faster, cheaper, and more successful. Learn how AI is reshaping drug development in real time, turning AI execution into a key competitive differentiator.

 

 

Bioanalysis should not rely on which scientists at the CRO you get. These are assays, not consulting. Your results should be the same, regardless of who does the work.

“We love working with our CRO… when we get the right PI” is a red flag, not a compliment. And if the quality of your bioanalysis depends on which principal investigator happens to be assigned to your study, you don’t have a CRO partnership, you have a lottery ticket.

Dash was built around automation and standardized workflows specifically to eliminate that variability. Every study gets the same rigor and the same rapid turnaround, days not weeks, across ELISA/MSD, LC-MS, and PCR. GLP-compliant, transparent pricing, guaranteed outcomes. You shouldn’t have to ask who’s running your study. You should just get great data.

Visit www.dash.bio

Now please enjoy this conversation with Sri Kosuri on The Long Run.

1 2 3 86