The Narrative Machine: What AI Really Adds to Personal Health Data

David Shaywitz
For years, the Quantified Self movement had a stubborn problem: our devices became increasingly good at generating data but remained remarkably poor at telling us what any of it meant. Generative AI now dangles a tantalizing possibility: that it can tease genuinely useful personal insights from all those measurements and hand them back to us in a form we can immediately understand and readily apply.
A recent Wall Street Journal article introduced us to what may be the latest iteration of the movement: “datamaxxers.” These are health-conscious technophiles who are feeding ever-larger portions of their lives into AI, seeking highly curated and customized insight in return.
The inputs can include the familiar streams from Garmin, Oura, WHOOP and Strava, including workouts, sleep, heart rate and heart-rate variability, along with calendars, meeting transcripts, emails, clinical records and even text messages. The hope is that AI can integrate the data, connect the dots, and just like that deliver exquisitely personalized health advice, one of digital medicine’s eternal promises.
The Journal introduces us to a datamaxxing runner named Julian Flieller. Before AI, his fitness tracker might tell him that his heart-rate variability was 60, a number that generated, as he put it, “more stress than clarity.” Now he supplies Claude with his workout, sleep and other biometric information and asks what it means for his training. Instead of receiving another isolated score, he gets an explanation –- a story.
Others have gone further. One runner built an AI dashboard that integrates Garmin, Strava and Oura data and helps her connect a sluggish run with poor sleep or the onset of her menstrual cycle. An engineer in India combined his Google Calendar with heart-rate data and asked Claude to determine which colleague was causing him stress. Sure enough, the AI fingered a specific co-worker as a “prime suspect.”
Once More, With AI
The original Quantified Self movement emerged around 2007, propelled by the growing ability of smartphones and wearables to capture information that previously would have been difficult or impossible to gather continuously. Its appealing premise was “self-knowledge through self-tracking.”
The first iteration largely fizzled. Avid self-trackers found themselves drowning in data but lacking in insight. After years spent tracking his activity, work and sleep, former Wired editor Chris Anderson famously concluded in 2016 that the effort had proved essentially pointless, yielding “no non-obvious lessons or incentives.”
I wrote for TR about the arrival of what I called “Quantified Self 2.0” in 2021. By then, the original Fitbits and Jawbones had given way to far more sophisticated technologies offering continuous glucose measurements, heart-rate variability, sleep staging and other physiological assessments. Yet the nagging question remained: were we actually learning anything more useful, or simply repackaging old approaches with a slicker interface and the nascent promise of AI, while continuing to preach the dubious familiar gospel that more data inevitably leads to better insight?
Five years later, generative AI appears to offer the latest wrinkle: a conversational superintelligence capable of looking across all our data and patiently explaining exactly what they mean.
The Great Technology Plan for Wellness, as TR readers will recall, is basically: inhale data Þ metabolize data with AI Þ deliver personalized coaching that improves as the data grow and the models sharpen. It sounds plausible, rather like the perpetually aspirational “learning health system”: a deeply appealing idea that has spent decades seeming both obviously right and frustratingly difficult to realize.
I also understand the appeal. I tried using WHOOP at least two separate times over the years, repeatedly hoping it would eventually tell me something sufficiently interesting or useful to justify my investment. Each time I gave up, because despite the volume of information, and most recently the eagerness of the associated AI to offer interpretation and advice, I was learning remarkably little that meaningfully changed what I did.
Lots of Data, Limited Insight
Kevin Hall, the distinguished former NIH metabolism researcher and author of Food Intelligence, offers a useful reality check.
Continuous glucose monitors (CGMs) seem tailor-made for personalized nutrition. They generate exquisitely detailed information about how glucose changes throughout the day, including after particular meals. It seems logical that enough measurements, analyzed intelligently, could reveal precisely which foods are good or bad for you.

Researcher Kevin Hall
Yet Hall has been particularly skeptical of using CGMs to generate precision dietary advice in people without diabetes. The instruments unquestionably generate data. The issue is whether those data contain the stable, individualized signal needed to support the recommendations drawn from them. As TR readers will recall, the ability of CGMs to offer advice to non-diabetics that is both meaningfully personalized and scientifically valid remains remarkably limited.
This distinction is easy to lose because modern technology can produce such an astonishing volume of measurements. A wearable might collect hundreds or thousands of physiological observations every day. Add sleep, exercise, diet, glucose, blood testing, calendar information and subjective reports, and the system can possess an extraordinarily rich empirical description of an individual. Even so, these abundant data do not necessarily provide the information needed to reliably predict the outcomes we care about, much less how those outcomes would change under a different set of circumstances.
At the population level, many wearable measurements have established relationships to health. What is far less established is whether a particular constellation of measurements supports a highly specific inference about a particular person, such as whether a lower HRV (heart rate variability) number and a poor night’s sleep explain a slow run or mean you should skip today’s hard workout. The measurements are abundant; the outcomes needed to validate many of these individualized predictions are far less so, and the relevant feedback often arrives slowly. We can therefore become measurement-rich while remaining surprisingly evidence-poor.

Professor Nathan Price
Additional personal data can sometimes yield genuinely useful insight. In a recent npj Digital Medicine Comment with Nathan Price and Dan Sodickson (discussed in TR), we argued that richer longitudinal and multimodal context could make health screening substantially more informative, in part by allowing us to detect meaningful changes from an individual’s own baseline rather than repeatedly comparing each measurement with a population norm.

Dr. Daniel Sodickson
But more data are valuable only to the extent that they add information. Sometimes repeated measurements reveal a meaningful trajectory that a single measurement would miss; sometimes they add surprisingly little. Complementary biological, imaging and clinical signals may add still more. AI is most valuable when it can integrate this context in ways that genuinely sharpen the inference.

Researcher Andreas Bender
There is also a deeper problem, familiar from drug R&D. In a recent Nature Reviews Drug Discovery Perspective (also reviewed in TR), Andreas Bender and colleagues, including Jack Scannell and me, discussed how biological datasets can be enormous while the outcomes we actually care about, such as clinical trial results or FDA approval, remain comparatively sparse, slow to emerge and highly context-dependent. Daphne Koller has described a related “data chasm”: huge datasets may sample remarkably little of the biological space we need to understand. Bytes and coverage do not necessarily move together. (This was reviewed in TR as well.)

Daphne Koller, founder and CEO, insitro
And learning from even extraordinarily rich measurements can be stymied by the challenge (discussed in our NRDD paper) of epistemic opacity: intrinsic biological complexity can make it difficult to reliably infer the outcomes we care about from the things we can readily measure. AlphaFold can predict how a protein folds with extraordinary accuracy, for example, but structure alone does not tell us what role that protein plays in human disease or whether altering it will make a patient better. In many settings, the limiting factor may therefore be less the sophistication of the AI than the information content of the data and the distance between what we can measure and what we are asking the system to infer.
At the moment, much of what passes for precision wellness still seems like precision health theater. The specificity of the recommendation can wildly exceed what the underlying evidence can credibly support.
However, there may be another opportunity here: even when AI isn’t uncovering a consequential personalized insight, it can be remarkably good at extracting a coherent narrative from, or perhaps imposing one on, a pile of disparate data.
The Narrative Machine
Flieller’s HRV of 60 offers a prismatic example. The fitness tracker gave him a number that, in isolation, meant very little to him. The AI can place that number alongside his sleep, training load, recent performance and work schedule, then offer an account of how those observations might fit together. Perhaps his HRV has fallen while his sleep has deteriorated, his recent training load has risen, and his work calendar has been unusually demanding. Perhaps accumulated fatigue contributed to his sluggish run this morning, and today would be a sensible day to recover.
That account might be exactly right. But generative AI has also moved, almost imperceptibly, from aggregating observations to interpreting them, and from interpreting them to confidently explaining what happened.
This, I suspect, is the true superpower of current LLMs. For all the talk about godlike intelligence and discovering patterns beyond human comprehension, they are demonstrably exceptional narrative machines. Give them a messy collection of disparate observations and they can select what seems relevant, connect the pieces, incorporate outside knowledge and fashion a coherent, highly plausible account of what might be going on.
We tend to find coherent accounts enormously compelling. Daniel Kahneman spent much of his career documenting our susceptibility to stories that make available facts fit together, and how readily the feeling of coherence can become a feeling of understanding. Missing information bothers us less when the information we possess coalesces into a satisfying explanation.
Health AI adds another potent ingredient because the story is about you. Generic advice to sleep more, exercise regularly and eat thoughtfully is sensible but hardly riveting. An explanation that references your declining HRV, your shortened sleep, your increased mileage and your unusually stressful Tuesday feels entirely different. It appears not merely sensible but customized and insightful.
Sometimes, the LLM may genuinely do what the datamaxxing vision promises: integrate many disparate signals and identify a consequential individualized insight sufficiently well supported to change what you should do. In other cases, the AI may simply have taken several plausible facts and done what LLMs do exceptionally well: construct a remarkably plausible story.
The “prime suspect” identification is entertaining for a different reason. Perhaps the AI got it exactly right. But if you already know that a particular colleague drives you crazy, do you really need a wearable, a calendar integration and Claude to tell you so? This recalls an enduring temptation in digital health: using increasingly sophisticated technology to measure with exquisite precision something we could already determine well enough by simpler means. The question isn’t merely whether the new measurement or inference is correct, but whether it provides enough incremental information to improve an actual decision.
Early digital-health researchers, for example, explored accelerometers to quantify how much hospitalized patients were up and walking, information potentially relevant to recovery and discharge. While useful in some settings, the exercise could also amount to measuring with digital precision something a clinician could already learn by simply poking her head into the room.
Why the Story May Matter Anyway
There is, however, another source of potential value here. Even when the AI has not discovered a consequential new insight, the personalized account it constructs may influence how someone makes sense of an experience and what she does next.

Professor Alia Crum
I’ve written previously about the work of Stanford psychologist Alia Crum and others studying mindset, essentially the beliefs and expectations we bring to experiences and the meaning we assign to them. These effects are not confined to vague affirmations. In one classic study, hotel workers who were informed, accurately, that the physical work they already performed constituted meaningful exercise subsequently showed changes in weight, body fat and blood pressure, despite apparently not becoming more active.
In another striking experiment, people wearing Apple Watches were shown manipulated information about their activity. Participants who were falsely told that they were walking substantially less than they actually were subsequently felt worse, ate worse and developed higher blood pressure and heart rates, even though the underlying amount of walking barely changed. The measurement hadn’t changed, but the meaning attached to it had, and that interpretation became part of the experience.
Crum’s more recent work is especially relevant because she does not argue that people should simply substitute a pleasing story for an unpleasant truth. Complex situations frequently support more than one plausible reading. Her “metacognitive” approach emphasizes recognizing that our first, often reflexive, reading of an experience is only one possibility; appreciating that beliefs have consequences; and getting better at considering alternatives. The point is to become more deliberate and skillful in how we make sense of the many experiences that are open to more than one interpretation.
Health and fitness provide abundant examples of just these sorts of experiences. A difficult workout after a long week can become evidence that you are getting old and declining, or it can be understood as a demonstration of your ability to show up and persist even when you’re not feeling your best. Those different readings do not require pretending the underlying facts are different, but they can shape what happens next.
There may also be a simpler, almost accidental way in which wearables add value, one that has relatively little to do with the precision of the insights they generate. Simply wearing a WHOOP strap or an Oura ring can reinforce an identity: I am someone who exercises, trains seriously or pays attention to my health. The device can function as a kind of commitment device, making that health-oriented identity more salient and giving it visible expression.
AI offers an additional opportunity: to build deliberately on this narrative dimension rather than merely benefiting from it incidentally. It might help people notice and interpret experiences in ways that strengthen agency, our capacity to act intentionally and influence what happens next, and self-efficacy, our confidence that our actions can actually make a difference. A completed workout is not merely 300 calories expended or another ring closed on your Apple Watch; it can also become evidence that you’re someone who follows through, someone who can recover after missing a week, or simply someone capable of doing a little more than you thought.
Consider the difference between an AI that confidently asserts, “Your slow run was caused by elevated work stress, which suppressed your recovery,” and one that uses the same information to help you put the experience in perspective. Rather than simply supplying a more reassuring explanation, it might help you consider alternative interpretations and recognize evidence of persistence or progress that you might otherwise overlook, with the goal of helping you gradually become better at doing this for yourself.
Interpreter, Not Oracle
The original Quantified Self movement assumed that enough measurement would yield self-knowledge. Often it didn’t. The numbers piled up while the promised insights proved frustratingly elusive. Generative AI has now supplied something the movement conspicuously lacked: a remarkably fluent interpreter.
What is unusual about generative AI is that it does not have to wait until the evidence is strong enough to justify an inference before offering an explanation. This is both its attraction and its danger. It can transform an otherwise bewildering pile of numbers into a coherent, easily digestible account of your sleep, exercise, stress and physiology. But if you increasingly require an AI to tell you what every difficult workout, bad night’s sleep or stressful day means, you may be outsourcing precisely the capacity for interpretation that these tools should ideally help you develop. A technology intended to augment self-knowledge could instead leave you increasingly dependent on an external narrator.
At its best, AI could offer something considerably more valuable. Where the data genuinely support a non-obvious and consequential individualized insight, AI could surface it. Even where they do not, AI could help us consider alternative interpretations of our experiences and recognize evidence of progress and capability that we might otherwise miss. Over time, this could help cultivate the judgment, confidence and habits of interpretation that underlie agency and self-efficacy. (The opportunity to cultivate agency in this way is also part of what attracted me to my current work at Lore.)
Follow this strand forward and you arrive at what feels to me like a more thoughtful aspiration for AI in personalized health: AI that delivers real insight when real insight is available, while helping us become ever more capable interpreters and authors of our own lives rather than unhealthily dependent on the narrative machine.



