The Other Side of Bayes: Rethinking False Positives in Health Screening

David Shaywitz
Imagine screening 10,000 apparently healthy people for a disease that affects one person in a thousand. Let’s further assume the test is fairly good: 99% sensitive and 99% specific. (In case it’s been a while, this means the test detects 99% of true cases and correctly exonerates 99% of people without disease.)
Ten of those 10,000 people will actually have the disease –- that’s just the prevalence of the condition in the population. The test will identify essentially all of them. That’s the good news. But among the 9,990 people who don’t have it, about 100 will nevertheless test positive. The screening program will therefore return roughly 110 positive results, of which only about 10 represent actual disease.
In other words, even with a test that is 99% sensitive and 99% specific, about nine out of ten positive results will be wrong.
This is the classic screening problem every physician learns in medical-school biostatistics and reliably encounters on every board exam. Bayes’ Theorem explains what’s going on.
In essence, Bayes tells us how much a new piece of evidence should change what we believed beforehand. For a screening test, the interpretation therefore depends not only on the intrinsic characteristics of the test, but also on the probability that the person had the disease before the test was performed. When that “prior probability,” as it’s called, is very low, even an excellent test can deliver misleading results.
Comprehensive testing compounds the difficulty. Run enough measurements and the chance that something will be flagged rises quickly. This concern was memorably anticipated nearly two decades ago by Isaac Kohane, Daniel Masys and Russ Altman, who coined the term “incidentalome” to describe the proliferation of unexpected findings that could accompany genome-scale testing. They worried that clinicians would feel compelled to chase these abnormalities, subjecting patients to additional testing and morbidity while driving up costs, potentially with little corresponding benefit.
Their statistical illustration was striking. Suppose you evaluate someone with a genetic panel that effectively represents 10,000 independent (for the sake of argument) tests, each of which has a false-positive rate of just 0.01% — extraordinarily good individual performance. Even so, when you work through the math (or maths, if you’re in the U.K.), it turns out that more than 60% of the people tested would nevertheless wind up with at least one false-positive result.
This is why the proliferation of ever-more comprehensive testing, largely driven by consumer interest in disease pre-emption, appropriately concerns many doctors, including Eric Topol, and including me (and has for years).
At the same time, the appeal of such testing is readily understandable. Most diseases of aging develop gradually over years, potentially offering an opportunity to identify trouble at an earlier and more actionable stage. But looking harder at apparently healthy people also risks incidental findings, anxiety, and unnecessary workups, occasionally (if inadvertently) leading to actual harm.
Rethinking Bayes
A few years ago, however, I came across an intriguing paper by Dr. Jacques Balayla at McGill that made me think about this problem differently.

Dr. Jacques Balayla
Balayla starts with the familiar dilemma. Most diseases amenable to screening are uncommon in the general population, meaning a significant proportion of positive screening results can be false positives that carry, as he notes, medical, psychological, social and practical consequences. He then asks whether anything can be done about what appears to be an intrinsic Bayesian limitation.
His answer is embedded in Bayes itself.
Bayesian reasoning is not static. New evidence updates what you believe, which is kind of the point. The resulting probability becomes the starting point for interpreting the next piece of evidence. Balayla shows mathematically how repeated testing, under appropriate assumptions, can progressively increase the positive predictive value of a screening result.
That insight led directly to a Comment that Nathan Price, Dan Sodickson and I recently published in npj Digital Medicine.

Professor Nathan Price
Nathan, a professor at the Buck Institute for Research on Aging, the CSO of Thorne, and co-author with systems-biology pioneer Lee Hood (see Luke Timmerman’s excellent biography of Hood, here) of The Age of Scientific Wellness, has spent more than a decade exploring what can be learned from dense longitudinal measurements of people beginning in health rather than disease.
Dan, a distinguished physician-scientist and MRI researcher at NYU, my former Harvard-MIT Health Sciences and Technology classmate, and author of The Future of Seeing, has pursued an analogous vision (so to speak) in imaging. (Dan is now Chief Medical Scientist of Function Health, though our paper was written before he joined.)

Dr. Daniel Sodickson
What Dan, Nathan, and I ask is fairly simple: as it becomes easier to collect serial molecular, imaging and clinical data, and as AI becomes increasingly capable of integrating them, can screening make greater use of individual context rather than evaluating each new result largely in isolation against a population reference range?
Insights from Imaging
Dan’s work provides an especially striking illustration. In a terrific recent Big Brains podcast interview, he puts the central idea particularly well: “False positive rates aren’t fixed.” With enough prior context, Dan explains, the question becomes whether something represents a meaningful new change or simply that individual’s baseline.
In a retrospective analysis from his group involving nearly 30,000 patients followed for about a decade, AI models were trained to predict current and future risk of clinically significant prostate cancer. As the models were given more context in the form of prior imaging and clinical information, false positives fell dramatically while sensitivity remained around 90%. In the analysis reproduced in our paper, specificity rises from less than 30% with little prior information to more than 90% as richer prior information is incorporated. The work comes from a single center and requires external validation, but the underlying point is compelling: what today’s image means can depend enormously on what we already know about the person being imaged.
There is another potentially important aspect of Dan’s work. His group has demonstrated that prior imaging can help reconstruct a subsequent MR image from radically less newly acquired data. In the Big Brains podcast, Dan describes a broader “Everywhere Scanner” vision: cheaper, more accessible imaging that could fill the gaps between conventional scans, with AI focused on detecting meaningful change from an individual’s prior image.
If the approach generalizes, interval scans could eventually become quicker, cheaper and feasible on less specialized equipment. That creates the intriguing possibility of a virtuous cycle: easier imaging makes serial assessment more practical; serial assessment generates more individual context; and that context can make subsequent imaging both easier to obtain and more informative.
Insights from Molecular Measurement
Nathan has approached the problem from the molecular side. More than a decade ago, working with Hood and colleagues at the Institute for Systems Biology (ISB) in Seattle, he began exploring what they called “personal, dense, dynamic data clouds,” initially following 108 generally healthy people with repeated clinical and multi-omic measurements, whole-genome sequencing and activity tracking, work they reported in a 2017 Nature Biotechnology paper. They subsequently scaled the approach through Arivale, an ISB spinout focused on “scientific wellness.”
Across the five year program, nearly 5,000 participants contributed longitudinal data including genomic information, more than 1,000 blood-based measurements, microbiome analyses and Fitbit data, as described in a 2021 Nature Communications paper. The goal was to characterize individualized health trajectories that might reveal not simply whether someone falls outside a population norm, but whether that particular person is beginning to deviate from his or her own healthy state.
In a 2025 talk at the Buck, Nathan showed an evocative, though decidedly preliminary, illustration. Stored samples from three participants who subsequently developed metastatic cancer were retrospectively analyzed for the cancer-associated protein CEACAM5. Two of the participants had elevated values from the start, but in a third case, the value remained within the conventional normal range at an early time point, yet had jumped substantially from the individual’s previous measurement; it rose further before breast cancer was diagnosed. Nathan emphasized that there were only three cases, the measurements were retrospective, and this was emphatically not a validated screening test. But the example makes the intuition easy to grasp: a measurement — like the CEACAM5 value in the woman who developed breast cancer — can look normal compared with the population and still be distinctly abnormal for you.
Importantly, the broader principle does not rest on intriguing anecdotes or futuristic AI.
Consider prostate-specific antigen, or PSA. An elevated PSA can trigger MRI and biopsy, yet many such evaluations will not uncover clinically important cancer. Investigators have therefore developed reflex tests that bring additional biological information to bear. In an analysis from the Göteborg-2 prostate cancer screening trial cited in our paper, adding the 4Kscore, which combines several kallikrein measurements with clinical information, after an elevated PSA and before MRI would have reduced MRI use by 41%, biopsies by 28%, and diagnoses of low-grade cancers by 23%, while delaying detection of 4% of clinically significant cancers.
There is no AI magic in that example. The underlying model uses conventional statistics. The useful ingredient is additional, complementary information.
And “complementary” is important, though the relevant distinction is more subtle than simply serial versus different measurements. A trajectory of the same biomarker can itself be highly informative, as Nathan’s CEACAM5 example suggests. But that has to be demonstrated rather than assumed. There’s a useful counterexample from prostate screening: investigators hoped that PSA velocity, essentially the rate at which PSA changes over time, would improve prediction beyond the PSA level itself. In practice, it added surprisingly little independent predictive value. The lesson isn’t that trajectories don’t matter; it’s that some trajectories contain meaningful new information and others largely do not. Combining longitudinal change with genuinely complementary biological, imaging and clinical signals may offer still more opportunity to distinguish signal from noise.
AI and the Future of Screening
This is where AI becomes especially interesting.
Clinicians have always used context. We compare today’s scan with the old one. We interpret a laboratory result in light of earlier values, symptoms, medications and other findings. But there are obvious limits to how much disparate information a person can integrate. If an individual’s relevant context eventually includes years of imaging, hundreds or thousands of molecular measurements, physiological signals and clinical history, making sense of the whole may become as much a computational problem as a clinical one. AI could allow this sort of contextual integration to occur systematically and at scale.
None of this gives comprehensive screening a free pass. Earlier detection matters only if it leads to earlier useful intervention and ultimately better health, rather than simply more diagnoses, anxiety and procedures. Cost, overdiagnosis, liability and equity, as we discuss, remain substantial concerns. Context-aware screening will ultimately have to clear the same bar as any other medical intervention: prospective evidence that using it leaves patients better off.
But it’s difficult not to feel more hopeful. For years, I regarded Bayes primarily as the mathematical explanation for why broad screening of apparently healthy people was so fraught. Balayla helped me see the other side of the theorem. When disease is rare, a single positive result may tell us surprisingly little; but as additional context accumulates, sometimes through a revealing trajectory over time, sometimes through complementary signals, the picture can become much clearer.
The mathematics suggests that richer longitudinal context could make screening substantially more informative. The task now is to determine whether it can make patients healthier as well.



