17
Aug
2026

Harell Data Debuts With $15M To Change Incentives, Speed Up AI Drug Discovery

The Protein Data Bank was the unsung hero that made AlphaFold possible. For decades, the structural biologists, biochemists and others who put carefully curated protein structures in the public domain laid down the conditions for AI to eventually do something useful – predicting protein structures in this case.

Those scientists never got a nickel. This is just one of many examples in which AI model builders hoovered up public data for free and now are making their fortunes.

Harlan Robins, founder and president, Harell Data

Harlan Robins thinks there has to be a way for data generators to share in the value of discoveries made based on the data. So he started a company, Bellevue, WA-based Harell Data, to create a more sustainable business for data generators, as well as AI model builders seeking to discover new medicines.

“Data itself, especially for training ML models, isn’t getting appropriately valued,” Robins said. “Everyone gives it lip service, how the important the data is. But all models generate huge value, whether it’s ChatGPT, Claude, or others, from training on public data.”

Robins, the scientific founder of Seattle-based Adaptive Biotechnologies, a public company worth $3.9 billion, is announcing today that his new startup has raised $15 million from Fuse, Cercano Management, and others. The idea is to create a more sustainable and mutually beneficial cloud computing service for data generators and AI model builders.

Here’s how it’s supposed to work. Harell has lined up secure cloud computing services through CoreWeave Cloud, with the industry standard Nvidia graphics processing unit technology, Robins said.

Those computers are being initially populated with two well-curated, proprietary datasets. One dataset is from Seattle-based A-Alpha Bio, which is contributing some of its data on protein-protein binding affinities. The other is from Robins’ company, Adaptive, which is chipping in some of its proprietary data on T cell receptor binding.

The companies that provide that data to the scientific community for drug discovery will receive a cut of the revenue Harell receives from AI model builders, Robins said. The AI model builders, he says, will benefit by getting access to well-curated datasets that provide relevant information for drug discovery. The price to customers will be lower than what they currently pay to big cloud service providers like Amazon Web Services, Microsoft’s Azure, and Google Cloud.

Robins tried to circulate this idea among the big cloud computing companies. He offered to drive hundreds of millions in revenue to the big companies, if they would offer a cut of revenue to data generators who put in the quality data that is the essential grist for the mill of drug discovery.

Those talks went nowhere.

“I understand why,” Robins said. “They are selling every bit of compute they are getting already. They are supply limited, not limited on customers.”

Robins will now get real-life feedback from the market to see if Harell has hit upon a viable business model. To work, it will need to appeal to AI drug discovery customers, and provide fair compensation for data providers. Robins isn’t disclosing the percentage cut of revenue it offers data generators, but he said they will be paid upfront and will not try to reach through for downstream royalties on discoveries years after the fact.

Harell – an amalgam of Harlan’s first name and that of his 8-year-old son Ellis – has hired a team of 11 people. None have been brought over from Adaptive, where Robins remains a consultant.

Robins, the scientist, has been itching to see more medical discoveries come out of this deep, rich, and proprietary dataset. So much industry data lives in silos and isn’t accessible to AI model builders. What would happen if data generators and AI model builders could share in the prosperity of new medicines that emerge in coming years? Could it accelerate discovery?

The only way to find out is to try a different business model. Harlan Robins is about to get some real-world feedback on whether this is an idea whose time has come.

You may also like

Save the Date: TR 10th Anniversary
The Other Side of The Story: Vaccines Must Produce Both Antibodies and T cell Immunity
The Immune Sequencing Frontier: Harlan and Chad Robins on The Long Run
Microsoft Makes Nine-Figure Bet on Adaptive’s TCR-Antigen Map; BioNTech Hauls in $270M for 10-Drug Cancer Pipeline