Medicine

Medicine and AI, Part 4: Drug Discovery Still Has to Become Human Evidence

AI can rank molecules, predict protein structures and sharpen trial design, but a model score is not a medicine until chemistry, safety, dosing and human outcomes hold up.

Marco Linden ·

Medicine and AI, Part 4: Drug Discovery Still Has to Become Human Evidence

Artificial intelligence enters drug discovery at a tempting point in the story: before the expensive failures are visible. A model can compare protein structures, gene activity, disease pathways, chemical libraries and earlier assay results, then rank targets or propose molecules that look worth testing. That is genuinely useful. It can also create the illusion that a hard medical problem has become a search-engine problem. The fourth part of this medicine-and-AI series is about that gap: the distance between a promising model output and a medicine that helps patients.

![Original EBK diagram listing drug-development tasks where AI models can support scientists without replacing clinical proof. Credit: EveryBunnyKnows, CC BY 4.0](https://images.ctfassets.net/80ca4ljo2d4c/7epqziRO4icMKRhWXhER1i/b7da7cdbc5f18a61a0d8c70946a3e157/ai-models-drug-discovery.svg)

The mechanism starts with narrowing possibility. DeepMind’s AlphaFold changed expectations by predicting structures for many proteins, giving scientists better maps of shapes that influence binding and function. Other systems search chemical space, predict toxicity signals, suggest new uses for old compounds or cluster patients whose disease biology may respond differently. Companies such as Insilico Medicine, Recursion and Exscientia have built businesses around pieces of this pipeline, while academic groups use machine learning to search for antibiotics, protein designs and trial endpoints. The best version is not a robot pharmacist. It is a filter that helps human teams choose which experiments deserve scarce time and money.

Chemistry is where the first optimism test begins. A molecule that looks elegant in a model may be hard to synthesize, unstable in blood, poorly absorbed, rapidly metabolized by the liver or active against an unintended target. Cell assays and animal studies can remove many false leads, but they add their own limits because a dish, mouse or organoid is not a whole patient with age, pregnancy, kidney disease, other medicines and unequal access to care. AI can help choose experiments; it cannot make those experiments unnecessary.

![Original EBK graphic explaining why AI-designed candidates still pass through phased human testing and surveillance. Credit: EveryBunnyKnows, CC BY 4.0](https://images.ctfassets.net/80ca4ljo2d4c/7s168zdRgpa3AIn8D1qFFD/1f85f093670838a188cf7ad1a37038cf/drug-trial-evidence-ladder.svg)

The clinical road is deliberately slow. A phase 1 study usually asks about safety, tolerability and dose in dozens of volunteers or patients. Phase 2 looks for a signal of efficacy, dose selection and endpoint behavior. Phase 3 may involve hundreds or thousands of people across many sites to test whether benefits outweigh risks in a population close enough to real use. Regulators such as the U.S. Food and Drug Administration and the European Medicines Agency evaluate patient outcomes, manufacturing quality, protocol integrity and adverse events, not only how impressive a model seemed during discovery.

AI can still improve clinical trials in practical ways. It may identify patients whose biomarkers fit a mechanism, reduce screen failures, flag adverse-event patterns, simulate parts of trial logistics or help select endpoints that are sensitive enough to measure change. It can also make trials worse if training data overrepresent wealthy health systems, if a subgroup is excluded because it is statistically inconvenient, or if a black-box score hides a fragile assumption. In medicine, a cleaner dataset is not the same as a fairer or safer trial.

The limits are ethical as well as technical. A model trained on yesterday’s published biology may miss an unmeasured immune pathway. A model optimized for binding may ignore formulation or manufacturing. A retrospective success can vanish in a prospective trial. Even when AI helps create a candidate quickly, patents, pricing, trial access and post-market surveillance decide whether patients benefit. The useful future is therefore less theatrical than the slogan “AI discovers drugs.” It is a disciplined collaboration: models generate hypotheses faster, laboratories test them more intelligently, clinicians protect patients during trials, and regulators insist that the final claim belongs to evidence in humans.