Mon 27 Jul 2026 / 11:15 ET
Kernel
AI 4 min read

AI drug discovery data gap slows autonomous lab ambitions, Cytiva says

Cytiva’s Paul Belcher says AI can speed early drug design, but weak datasets and isolated lab systems still limit the payoff.

Felix Aranda

By Felix Aranda / Silicon Editor

AI drug discovery data gap slows autonomous lab ambitions, Cytiva says
img: MIT Technology Review

AI drug discovery data is becoming the awkward bottleneck in pharma’s favorite efficiency story. Drugmakers are using models to propose and rank new compounds earlier in research, but Paul Belcher, director of protein research strategy at Cytiva, says those candidates still need lab validation, better data, and systems that can actually talk to each other.

The stakes are not small. The OECD describes the long-running decline in biopharmaceutical research productivity as Eroom’s Law, under which the cost of developing new medicines has roughly doubled every nine years since the 1950s. Cytiva has cited estimates that a new drug typically takes 10 to 15 years to reach market and costs $1 billion to $2.5 billion, with failure rates above 90%.

Belcher says the clinical phase remains the main cost center, so the practical goal for AI is to reduce the chance that weak candidates reach expensive trials. In his view, AI may shorten discovery timelines and improve the quality of molecules that move toward the clinic, but it does not remove the need for wet-lab work.

What data does AI drug discovery need?

Drug discovery models need structured, diverse, trustworthy data that includes both successful and failed experiments. Belcher says many models trained on public datasets are hitting a wall because they often draw from the same published material, which skews toward positive results and leaves out compounds that did not bind or experiments that failed.

That missing negative data matters because machine-learning systems learn from the boundaries in their training sets. If the record mostly contains winners, the model has less information about what failure looks like, which can make its predictions less reliable.

Data integrity is another problem. Belcher points to research by Dutch microbiologist Elisabeth Bik, who found in 2016 that nearly 4% of biomedical papers contained duplicated or manipulated images. He argues that generative AI makes fabrication easier and raises the risk that bad scientific data will be used to train later models.

Cytiva is pitching tools for that problem, including its Image Integrity Checker, which Belcher says uses secure hash algorithms to detect whether scientific images have been altered. He says publishers have shown interest in adopting such checks as a way to screen material before it enters the literature.

How AI changes hit identification

One early use case is hit identification, the stage where researchers look for molecules that bind to a disease-related target such as a protein. A hit is a starting point for more testing and optimization, not a finished drug.

Belcher says the field is moving from large empirical screens toward predictive design. Instead of testing vast libraries first, companies can use AI to design possible compounds and forecast how they may interact with targets before spending R&D resources on physical testing.

That shift creates a second bottleneck. Traditional screening workflows were built for scale and often produce low-detail yes-or-no readouts. AI can generate more candidate hits, and potentially better ones, which increases demand for higher-throughput lab methods that can characterize, purify, and validate them in more detail.

Belcher says AI still cannot reliably predict kinetics or developability for new compounds. In plain English: models may suggest a molecule that looks promising against a target, but researchers still need experiments to learn how it behaves and whether it can become a practical drug candidate.

Has an AI-designed drug been approved by the FDA?

Belcher says no drug discovered primarily through AI-driven design has received full FDA approval yet. He expects that milestone could arrive within two to three years, though that is a forecast, not a regulatory timetable.

The more ambitious vision is the autonomous lab: instruments running around the clock, testing model-generated ideas, feeding results back into the model, and repeating the cycle. Belcher says that requires interoperable instruments, comprehensive datasets, and FAIR data, meaning data that is findable, accessible, interoperable, and reusable.

Most labs are not built that way today, according to Belcher. Many instruments remain standalone systems, and closed data environments limit the value of even strong experimental tools. The pitch for autonomous discovery depends on closing that loop between computational design and physical lab testing.

Cost could still bite. Stanford research has found that training frontier AI models has more than doubled in cost each year since 2016. Belcher says AI remains useful if the cost of compute does not exceed the cost and risk of clinical development. That is a sensible caveat in a sector already famous for expensive failures.

This story draws on original reporting from MIT Technology Review.

More AI/

view all ↗