Guest Column | July 27, 2026

Drug Discovery AI Has A Ground Truth Problem

By Elizabeth Hudson, Ph.D.

Digital DNA AI Medical Research Technology-GettyImages-2243217788

“I would estimate we might know 10% to 15% of human biology.”— Dave Ricks, CEO, Eli Lilly, Cheeky Pint podcast, November 11, 2025

Every process runs at the speed of its slowest, least reliable step. Chemists call it the rate-limiting step. Agronomists call it Liebig's law of the minimum. Supply chains call it the bottleneck.

AlphaFold is brilliant, but protein folding is not the rate-limiting step in drug development. Folding was an attractive AI problem because it was computationally interesting and had unusually clean, deep data. That made it solvable. It showed what is possible when a bounded problem has the right data. It did not show what is probable when we turn to other problems – understanding whole patients, or complex disease states, or biological networks – even with the right architectures and massive compute.

If the goal is more healthy years of life, the bottlenecks are target identification, target validation, and clinical testing: identify a (possible) disease-driving process, and test that hypothesis against reality.

Why is AI not transforming those steps? Because the data are not there. So far, the field has worked backward from where data exist to what we can do with that power. The next phase must run the other way: Name the problems that gate new medicines, define the data they require, and generate those data on purpose.

What Data You Need, And Why It Isn't There

We’re going to use target identification/validation as our example in this piece to illustrate the need for better data and the attributes that attempt to fill this data gap must possess.

Most drug programs crowd onto the same well-understood proteins because proving that a new target drives disease and is not load-bearing in healthy biology is slow, expensive, and usually fails. If biology were a map, we have surveyed a few major cities and left whole continents blank.

Crowding has a cost. The commercial opportunity gets split across companies chasing the same biology. The patient impact also shrinks, because so much target space remains unexplored. The real prize is the much, much larger set of targets we have yet to pursue: those we have never found — or identified as potentially interesting and able to be isolated — because we have never mapped the biology where they live. Can AI help you to find these?

The short answer is no. Better architectures and algorithms cannot extract insight from a map that does not exist. They can reason over measured biology. They cannot infer, with confidence, the biology we have never measured.

And it is tempting to assume the missing map is hiding somewhere and only needs assembling. It is not. To see why, let’s look at the places people expect to find it.

It Isn't In The Literature

The NIH alone spends nearly $48 billion a year on biomedical research.1 That money generates reams of raw, novel measurement. But those data never reach (in most cases) the public domain. What we get is a paper.

Think of it this way: A lab forms a hypothesis, runs experiments, analyzes the relevant slice, and publishes a narrative. What reaches the literature is a derivative of a derivative of a fraction of what was measured. The raw signal is often unavailable, let alone a faithful recording of all of the assumptions, decisions, and contextual information that you’d need to fully understand the data, translate it, and relate it to that produced by someone else, and reproduce it. 

A model trained on papers gets prose. It cannot recover what was measured but never reported. Billions of dollars’ worth of raw biological signal gets dropped on the floor because it did not fit the story, was never explored, or was never shared. And that is before we get into the bias against publishing null results.2

It Isn't In The Public Databases

Public reference projects are essential. They are flags in the ground and infinitely better than nothing. They are also not enough.

Start with speed. These projects take years to plan, fund, recruit, measure, and release. That means they will always lag the frontier. By the time a major cohort releases data, the measurement frontier may already have moved. They also skew toward what can be done (relatively) cheaply and at scale: genomes, blood, surveys, imaging, health records, and other noninvasive measurements. UK Biobank, for example, is a major achievement at half-million scale.3 But it is much easier to genotype hundreds of thousands of people than to deeply profile proteins across their organs, and yet, you need the latter. A genome-wide association study can point toward a target. It cannot tell you whether the protein is present in the disease tissue, if it is present in a vulnerable tissue, which cell type expresses it, or what else depends on it — all data a practitioner needs to know.

Public reference projects give the field leverage, standards, and scale. But they are just one ingredient of what must be many. They are necessary, but they are not sufficient.

It Isn't In Health Records

Electronic health records do not fill the gap. They are sparse in content, sparse over time, and built for a different purpose.

They give you a handful of basic measurements: height, weight, diagnoses, medications, procedures, standard blood panels, notes, and outcomes. That is useful, but it is not a map of the organism. It does not tell you what the liver, lung, kidney, brain, immune system, or tumor microenvironment is made of at the molecular level. They are also sparse over time. For much of a person’s life, the record may contain almost nothing. Then when the person gets sick, the record becomes dense, and the data arrive in a burst: labs, drugs, procedures, complications, responses, and sometimes death. That sequence is valuable for studying care and outcomes. It is not the same as measuring biology.

EHRs are messy because they were not built to describe a body. Medications may stay on a list long after a patient stops taking them. Diagnoses may reflect billing codes more than biology. Even mortality status can be incomplete or delayed. These records can support phenotyping, safety signals, and real-world evidence. They are not the ground truth layer.

It Isn't In The Vaults Of Pharma

Pharma does not lack data. It lacks the kind of data that has no immediate asset-level payoff.

Companies generate data to advance or kill programs. That means causal data: perturb a target, dose a molecule, run an assay, study toxicity, and watch what fails. With access to that estate, a model would learn faster. But it wouldn’t be done.

What is thinner is reference data. What does healthy biology look like across tissues, ages, disease histories, environments, and platforms? Where does a target live before disease, and what else depends on it? Pharma often has answers inside narrow program contexts. It rarely has the broad observational layer needed to reason across the organism.

It also often lacks the metadata needed to combine what it has. Some measurements come from contractors, collaborators, or older systems. If the record is missing the instrument, settings, protocol, sample handling, reference database, and processing history, the result is harder to trust and harder to pool.

And pharma usually does not build the translation scaffolding AI would thrive on: the same tissue across sequencing platforms, the same material across mass spectrometers, or animal and human tissues measured side by side. That work helps many programs but belongs to no single one, so it loses to the nearest deadline.

Open every pharma vault and you would learn a lot, but would not be done. The missing layer is still a broad biological reference: many tissues, many organisms, many platforms, full metadata, and enough repeated measurements to tell biology from artifact.

It Isn't In The Automated Labs, Either

Large-scale automated labs help. They lower the cost of repeated work and make perturbation data sets easier to generate. But they do not solve the ground-truth problem by themselves.

Automation works best when a small number of standardized steps run at high volume. It is weaker when work requires a long tail of protocols, instruments, sample types, tissues, and platforms. The economics of automation break down when you move beyond a few platform SKUs with high utilization for each. And biology needs that long tail of idiosyncratic functions.

Depth is not breadth. Perturbing the same cell line a thousand ways can teach you a great deal about that cell line. It does not tell you whether the result holds across tissues, people, disease states, or species. Automated labs can help build the reference layer, including cross-platform Rosetta stones. But only if that is part of their job.

How You'd Actually Get It (The Data You Need)

The field has to generate the missing data directly and design it to combine from the start. That starts with metadata. A result only means something with its input material, donor context, collection and preservation methods, protocol, instrument, reference database, settings, thresholds, processing steps, and quality-control history. That is how you tell true absence from a missed measurement. A missing protein may be biology, but it may also mean the method did not look for it, the instrument could not see it, sample prep lost it, or the database excluded the relevant isoform.4

Then build the reference at the right scale. Measure the same material across modalities — RNA, protein, imaging, pathology, and function — and across platforms, including old and new instruments. Measure animal and human tissues in comparable ways. Work across enough organisms, and enough tissues within each organism, to learn normal variation rather than assume it. “Healthy” and “diseased” are gradients, not clean bins. And for many tissue-level and protein-level questions, power calculations require priors we do not yet have: effect size, variance, and correlation structure. That makes this a discovery mission, not a confirmation study. We build the reference partly to learn how big it needs to be.

Build This, And Today's Science Improves First

None of this waits on AI. A real biological reference improves ordinary drug discovery immediately: teams know the normal range, stop treating missed measurements as true absence, choose animal models based on where translation holds, compare platforms honestly, and see safety risk across the organism instead of one tissue at a time. You can be an AI skeptic and still want this built. It makes today’s science faster and less fragile. AI is the second payoff.

Return to AlphaFold. It showed that a model can produce useful new outputs when the problem is bounded and the ground truth exists. The model did not create the Protein Data Bank. Rather, it stood on it.

That is the lesson for drug discovery. The question is not whether a machine can discover biology in the abstract. The question is whether we have measured enough biology, broadly and deeply, with enough context, for anything to reason from it. For folding, the answer was yes. For the next generation of targets, the answer is still no.

Stop letting available data choose the problems we solve. Name the problems that gate new medicines, then build their ground truth on purpose. Build it, and today’s science improves first. Build it, and AI becomes more serious after that. Skip it, and the cleverest model in the world is still reasoning from fragments.

References

  1. https://www.nih.gov/about-nih/organization/budget
  2. Dwan K, Gamble C, Williamson PR, Kirkham JJ, Reporting Bias Group. “Systematic Review of the Empirical Evidence of Study Publication Bias and Outcome Reporting Bias — An Updated Review.” PLOS ONE. 2013;8(7):e66844.
  3. UK Biobank, “Health research data for the world,” UKBiobank.ac.uk.
  4. McGurk KA, et al. “The use of missing values in proteomic data-independent acquisition mass spectrometry to enable disease activity discrimination.” Bioinformatics. 2020;36(7):2217–2223.

About The Author

Elizabeth Hudson is the CEO and founder of Ambitious Bio, a life sciences AI infrastructure company generating foundational biological datasets for drug discovery in the AI era. Hudson’s professional path reflects her deep passion for math, science, education, and solving complex problems. Hudson earned her undergraduate degree from Yale University and a Ph.D. in neuroscience from the University of Geneva, and led large-scale robotics and software efforts at Symbotic. Elizabeth later served as CTO at a Boston-based biotech venture capital firm, and as vice chair of the advisory board for the Harvard & Smithsonian Center for Astrophysics, where she helped raise funds to support basic scientific research. Elizabeth is an intelligence officer in the United States Navy Reserve and is in her second term on the Cambridge Public Schools Committee. For more, follow her on X @LizzieHudson