Cireşan, D. C., Meier, U., Gambardella, L. M. & Schmidhuber, J. Deep, big, simple neural nets for handwritten digit recognition. Neural Comput. 22, 3207–3220 (2010).
Krizhevsky, A., Sutskever, I. & Hinton, G. E. ImageNet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems 25 (NIPS 2012) 1097–1105 (Curran Associates, Inc., 2012).
Hochreiter, S. & Schmidhuber, J. Long short-term memory. Neural Comput. 9, 1735–1780 (1997).
Walters, W. P. & Barzilay, R. Critical assessment of AI in drug discovery. Expert. Opin. Drug Discov. 16, 937–947 (2021).
Rodriguez, A. et al. Unlocking the Potential of AI in Drug Discovery (Boston Consulting Group Report, 2023).
Jayatunga, M. K. P., Xie, W., Ruder, L., Schulze, U. & Meier, C. AI in small-molecule drug discovery: a coming wave? Nat. Rev. Drug Discov. 21, 175–176 (2022).
Mullard, A. 2024 FDA approvals. Nat. Rev. Drug Discov. 24, 75–82 (2025).
Phares, S., Phillip, K. & Trusheim, M. Clinical development success rates for durable cell and gene therapies. Nat. Rev. Drug Discov. 24, 329–330 (2025).
Wouters, O. J. et al. Differential legal protections for biologics vs small-molecule drugs in the US. J. Am. Med. Assoc. 332, 2101–2108 (2024).
Jayatunga, M., Ayers, M., Bruens, L., Jayanth, D. & Meier, C. How successful are AI-discovered drugs in clinical trials? A first analysis and emerging lessons. Drug Discov. Today 29, 104009 (2024).
Lowe, D. AI drugs so far. In the Pipeline https://www.science.org/content/blog-post/ai-drugs-so-far (2024).
Pitt, W. R. et al. Real-world applications and experiences of AI/ML deployment for drug discovery. J. Med. Chem. 68, 851–859 (2025).
Marshall, S. et al. Model-Informed drug discovery and development: current industry good practice and regulatory expectations and future perspectives. CPT Pharmacometrics Syst. Pharmacol. 8, 87–96 (2019).
Madhavan, S. & Shaywitz, D. A. AI: an essential tool for managing the burgeoning complexity of clinical development in pharmaceutical R&D. Drug Discov. Today 30, 104271 (2025).
Wodak, S. J., Vajda, S., Lensink, M. F., Kozakov, D. & Bates, P. A. Critical assessment of methods for predicting the 3D structure of proteins and protein complexes. Annu. Rev. Biophys. 52, 183–206 (2023).
Vázquez Torres, S. et al. De novo designed proteins neutralize lethal snake venom toxins. Nature 639, 225–231 (2025).
Du, H. et al. Targeting peptide antigens using a multiallelic MHC I-binding system. Nat. Biotechnol. 43, 1683–1693 (2024).
De Jonghe, J. et al. A community effort to track commercial single-cell and spatial ’omic technologies and business trends. Nat. Biotechnol. 42, 1017–1023 (2024).
Göller, A. H. et al. Bayer’s in silico ADMET platform: a journey of machine learning over the past two decades. Drug Discov. Today 25, 1702–1709 (2020).
Scannell, J. W. et al. Predictive validity in drug discovery: what it is, why it matters and how to improve it. Nat. Rev. Drug Discov. 21, 915–931 (2022).
Clark, D. E. What has computer-aided molecular design ever done for drug discovery? Expert Opin. Drug Discov. 1, 103–110 (2006).
Bender, A. & Cortés-Ciriano, I. Artificial intelligence in drug discovery: what is realistic, what are illusions? Part 1: ways to make an impact, and why we are not there yet. Drug Discov. Today 26, 511–524 (2021).
Paul, S. et al. How to improve R&D productivity: the pharmaceutical industry’s grand challenge. Nat. Rev. Drug Discov. 9, 203–214 (2010).
Dreiman, G. H. S. et al. Changing the HTS Paradigm: AI-driven iterative screening for hit finding. SLAS Discov. 26, 257–262 (2021).
Morgan, P. et al. Impact of a five-dimensional framework on R&D productivity at AstraZeneca. Nat. Rev. Drug Discov. 17, 167–181 (2018).
Fernando, K. et al. Achieving end-to-end success in the clinic: Pfizer’s learnings on R&D productivity. Drug Discov. Today 27, 697–704 (2022).
Wong, C. H., Siah, K. W. & Lo, A. W. Estimation of clinical trial success rates and related parameters. Biostatistics 20, 273–286 (2019).
Minikel, E. V. et al. Refining the impact of genetic evidence on clinical success. Nature 629, 624–629 (2024).
Bender, A. & Cortes-Ciriano, I. Artificial intelligence in drug discovery: what is realistic, what are illusions? Part 2: a discussion of chemical and biological data. Drug Discov. Today 26, 1040–1052 (2021).
Ortiz de Montellano, P. R. Cytochrome P450-activated prodrugs. Future Med. Chem. 5, 213–228 (2013).
Rodriguez-Antona, C. & Ingelman-Sundberg, M. Cytochrome P450 pharmacogenetics and cancer. Oncogene 25, 1679–1691 (2006).
Zevin, S. & Benowitz, N. L. Drug interactions with tobacco smoking. An update. Clin. Pharmacokinet. 36, 425–438 (1999).
Zimmermann, M. et al. Separating host and microbiome contributions to drug pharmacokinetics and toxicity. Science 363, eaat9931 (2019).
Safikhani, Z. et al. Revisiting inconsistency in large pharmacogenomic studies. F1000Research 5, 2333 (2016).
Samad, S. S., Schwartz, J. M. & Francavilla, C. Functional selectivity of receptor tyrosine kinases regulates distinct cellular outputs. Front. Cell Dev. Biol. 11, 1348056 (2024).
Zhang, Z., Pan, Q., Lu, M. & Zhao, B. Intermediate endpoints as surrogates for outcomes in cancer immunotherapy: a systematic review and meta-analysis of phase 3 trials. eClinicalMedicine 63, 102156 (2023).
Nogales, C. et al. Network medicine-based unbiased disease modules for drug and diagnostic target identification in ROSopathies. Handb. Exp. Pharmacol. 264, 49–68 (2021).
Grotzinger, A. D. et al. Mapping the genetic landscape across 14 psychiatric disorders. Nature https://doi.org/10.1038/s41586-025-09820-3 (2025).
Polishchuk, P. G., Madzhidov, T. I. & Varnek, A. Estimation of the size of drug-like chemical space based on GDB-17 data. J. Comput. Aided Mol. Des. 27, 675–679 (2013).
Balestriero, R. et al. Learning in high dimension always amounts to extrapolation. Preprint at https://arxiv.org/abs/2110.09485 (2021).
Sheridan, R. P. Time-split cross-validation as a method for estimating the goodness of prospective prediction. J. Chem. Inf. Model. 53, 783–790 (2013).
Lee, K., Moldagulov, G. & Grzybowski, B. A. The fragility of bioactivity prediction: rigorous dataset splits expose the illusion of ML accuracy. Chem. Eur. J. 15, e71208 (2026).
Kapoor, S. & Narayanan, A. Leakage and the reproducibility crisis in machine-learning-based science. Patterns 4, 100804 (2023).
Rodgers, S., Glen, R. C. & Bender, A. Characterizing bitterness: identification of key structural features and development of a classification model. J. Chem. Inf. Model. 46, 569–576 (2006).
Fierro, F., Giorgetti, A., Carloni, P., Meyerhof, W. & Alfonso-Prieto, M. Dual binding mode of “bitter sugars” to their human bitter taste receptor target. Sci. Rep. 9, 8437 (2019).
Huang, G., Lv, M., Hu, J., Huang, K. & Xu, H. Glycosylation and activities of natural products. Mini Rev. Med. Chem. 16, 1013–1016 (2016).
D’Amour, A. et al. Underspecification presents challenges for credibility in modern machine learning. J. Mach. Learn. Res. 23, 1–61 (2022).
Bender, A. et al. Evaluation guidelines for machine learning tools in the chemical sciences. Nat. Rev. Chem. 6, 428–442 (2022).
Church, K. & Kordoni, V. Emerging trends: SOTA-chasing. Nat. Lang. Eng. 28, 249–269 (2020).
Peters, B., Brenner, S. E., Wang, E., Slonim, D. & Kann, M. G. Putting benchmarks in their rightful place: the heart of computational biology. PLoS Comput. Biol. 14, e1006494 (2018).
Moult, J., Pedersen, J. T., Judson, R. & Fidelis, K. A large-scale experiment to assess protein structure prediction methods. Proteins 23, 2–5 (1995).
Newman, J. et al. Practical aspects of the SAMPL challenge: providing an extensive experimental data set for the modeling community. J. Biomol. Screen. 14, 1245–1250 (2009).
Ackloo, S. et al. CACHE (Critical Assessment of Computational Hit-finding Experiments): a public–private partnership benchmarking initiative to enable the development of computational methods for hit-finding. Nat. Rev. Chem. 6, 287–295 (2022).
Lander, E. S. et al. Initial sequencing and analysis of the human genome. Nature 409, 860–921 (2001).
Venter, J. C. et al. The sequence of the human genome. Science 291, 1304–1351 (2001).
Reiss, T. Drug discovery of the future: the implications of the human genome project. Trends Biotechnol. 19, 496–499 (2001).
Scannell, J. W. & Bosley, J. When quality beats quantity: decision theory, drug discovery, and the reproducibility crisis. PLoS ONE 11, e0147215 (2016).
Heyndrickx, W. et al. MELLODDY: cross-pharma federated learning at unprecedented scale unlocks benefits in QSAR without compromising proprietary information. J. Chem. Inf. Model. 64, 2331–2344 (2024).
Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021).
Liebeschuetz, J., Hennemann, J., Olsson, T. & Groom, C. R. The good, the bad and the twisted: a survey of ligand geometry in protein crystal structures. J. Comput. Aided Mol. Des. 26, 169–183 (2012).
Markowetz, F. All models are wrong and yours are useless: making clinical prediction models impactful for patients. npj Precis. Oncol. 8, 54 (2024).
Arnold, C. Inside the nascent industry of AI-designed drugs. Nat. Med. 29, 1292–1295 (2023).
Ackoff, R. L. The future of operational research is past. J. Oper. Res. Soc. 30, 93–104 (1979).
Kearnes, S. Pursuing a prospective perspective. Trends Chem. 3, 77–79 (2021).
Wellnitz, J. et al. One size does not fit all: revising traditional paradigms for assessing accuracy of QSAR models used for virtual screening. J. Cheminform. 17, 7 (2025).
Seal, S. et al. Machine learning for toxicity prediction using chemical structures: pillars for success in the real world. Chem. Res. Toxicol. 38, 759–807 (2025).
Obrezanova, O. et al. Prediction of in vivo pharmacokinetic parameters and time-exposure curves in rats using machine learning from the chemical structure. Mol. Pharm. 19, 1488–1504 (2022).
Miljković, F. et al. Machine learning models for human in vivo pharmacokinetic parameters with in-house validation. Mol. Pharm. 18, 4520–4530 (2021).
Walters, P. We need better benchmarks for machine learning in drug discovery. Practical Cheminformatics http://practicalcheminformatics.blogspot.com/2023/08/we-need-better-benchmarks-for-machine.html (2023).
Steiner, J. SOTA seeking — a knife fight in a phone booth. Techbio<>Biotech https://biotechbio.substack.com/p/sota-seeking-a-knife-fight-in-a-phone (2023).
Ott, S. et al. Mapping global dynamics of benchmark creation and saturation in artificial intelligence. Nat. Commun. 13, 6793 (2022).
Youssef, A. et al. External validation of AI models in health should be replaced with recurring local validation. Nat. Med. 29, 2686–2687 (2023).
Del Rosario, Z., Rupp, M., Kim, Y., Antono, E. & Ling, J. Assessing the frontier: active learning, model accuracy, and multi-objective candidate discovery and optimization. J. Chem. Phys. 153, 024112 (2020).
Kalliokoski, T., Kramer, C., Vulpetti, A. & Gedeck, P. Comparability of mixed IC50 data – a statistical analysis. PLoS ONE 8, e61007 (2013).
Landrum, G. A. & Riniker, S. Combining IC50 or Ki values from different sources is a source of significant noise. J. Chem. Inf. Model. 64, 1560–1567 (2024).
Cortés-Ciriano, I. & Bender, A. How consistent are publicly reported cytotoxicity data? Large-scale statistical analysis of the concordance of public independent cytotoxicity measurements. ChemMedChem 11, 57–71 (2016).
Brown, S. P., Muchmore, S. W. & Hajduk, P. J. Healthy skepticism: assessing realistic model performance. Drug Discov. Today 14, 420–427 (2009).
Crusius, D., Cipcigan, F. & Biggin, P. C. Are we fitting data or noise? Analysing the predictive power of commonly used datasets in drug-, materials-, and molecular-discovery. Faraday Discuss. 256, 304–321 (2025).
van Tilborg, D., Alenicheva, A. & Grisoni, F. Exposing the limitations of molecular machine learning with activity cliffs. J. Chem. Inf. Model. 62, 5938–5951 (2022).
Ahdritz, G. et al. OpenFold: retraining AlphaFold2 yields new insights into its learning mechanisms and capacity for generalization. Nat. Methods 21, 1514–1524 (2024).
Kryshtafovych, A., Schwede, T., Topf, M., Fidelis, K. & Moult, J. Critical assessment of methods of protein structure prediction (CASP)—round XIV. Proteins 89, 1607–1617 (2021).
Scardino, V., Di Filippo, J. I. & Cavasotto, C. N. How good are AlphaFold models for docking-based virtual screening? iScience 26, 105920 (2023).
Lyu, J. et al. AlphaFold2 structures guide prospective ligand discovery. Science 384, eadn6354 (2024).
Karelina, M., Noh, J. J. & Dror, R. O. How accurately can one predict drug binding modes using AlphaFold models? eLife 12, RP89386 (2023).
Masters, M. R., Mahmoud, A. H. & Lill, M. A. Investigating whether deep learning models for co-folding learn the physics of protein–ligand interactions. Nat. Commun. 16, 8854 (2025).
Star-studded, A. I. Biotech launch. Nat. Biotechnol. 42, 689 (2024).
Du, Y. et al. Machine learning-aided generative molecular design. Nat. Mach. Intell. 6, 589–604 (2024).
Thomas, M. et al. Identification of nanomolar adenosine A2A receptor ligands using reinforcement learning and structure-based drug design. Nat. Commun. 16, 5485 (2025).
Atomwise AIMS Program AI is a viable alternative to high throughput screening: a 318-target study. Sci. Rep. 14, 7526 (2024).
Xu, Z. et al. A generative AI-discovered TNIK inhibitor for idiopathic pulmonary fibrosis: a randomized phase 2a trial. Nat. Med. 31, 2602–2610 (2025).
Ren, F. et al. A small-molecule TNIK inhibitor targets fibrosis in preclinical and clinical models. Nat. Biotechnol. 43, 63–75 (2025).
Lenselink, E. B. et al. Beyond the hype: deep neural networks outperform established methods using a ChEMBL bioactivity benchmark set. J. Cheminform. 9, 45 (2017).
Mayr, A. et al. Large-scale comparison of machine learning methods for drug target prediction on ChEMBL. Chem. Sci. 9, 5441–5451 (2018).
Janela, T. & Bajorath, J. Simple nearest-neighbour analysis meets the accuracy of compound potency predictions using complex machine learning models. Nat. Mach. Intell. 4, 1246–1255 (2022).
Robinson, M. C., Glen, R. C. & Lee, A. A. Validating the validation: reanalyzing a large-scale comparison of deep learning and machine learning models for bioactivity prediction. J. Comput. Aided Mol. Des. 34, 717–730 (2020).
Lane, T. R. et al. Bioactivity comparison across multiple machine learning algorithms using over 5000 datasets for drug discovery. Mol. Pharm. 18, 403–415 (2021).
Ferreira, F. J. & Carneiro, A. S. AI-driven drug discovery: a comprehensive review. ACS Omega 10, 23889–23903 (2025).
Eken, B. et al. A multivocal review of MLOps practices, challenges and open issues. ACM Comput. Surv. 58, 1–35 (2025).
Bray, M. A. et al. Cell Painting, a high-content image-based assay for morphological profiling using multiplexed fluorescent dyes. Nat. Protoc. 11, 1757–1774 (2016).
Seal, S. et al. Cell Painting: a decade of discovery and innovation in cellular imaging. Nat. Methods 22, 254–268 (2025).
Pruteanu, L. L. & Bender, A. Using transcriptomics and cell morphology data in drug discovery: the long road to practice. ACS Med. Chem. Lett. 14, 386–395 (2023).
Chandrasekaran, S. N. et al. Morphological map of under- and overexpression of genes in human cells. Nat. Methods 22, 1742–1752 (2025).
Celik, S. et al. Building, benchmarking, and exploring perturbative maps of transcriptional and morphological data. PLoS Comput. Biol. 20, e1012463 (2024).
Chen, D. et al. A combined AI and cell biology approach surfaces targets and mechanistically distinct Inflammasome inhibitors. iScience 27, 111404 (2024).
Trapotsi, M. A. et al. Cell morphological profiling enables high-throughput screening for PROteolysis TArgeting Chimera (PROTAC) phenotypic signature. ACS. Chem. Biol. 17, 1733–1744 (2022).
Keenan, A. B. et al. Connectivity mapping: methods and applications. Annu. Rev. Biomed. Data Sci. 2, 69–92 (2019).
Ewald, J. D. et al. Cell Painting for cytotoxicity and mode-of-action analysis in primary human hepatocytes. Cell Syst. 17, 101566 (2026).
Platani, M. et al. Screening for variable drug responses using human iPSC cohorts. PLoS ONE 20, e0323953 (2025).
Roohani, Y. H. et al. Virtual cell challenge: toward a turing test for the virtual cell. Cell 188, 3370–3374 (2025).
Fu, H., Hardy, J. & Duff, K. E. Selective vulnerability in neurodegenerative diseases. Nat. Neurosci. 21, 1350–1358 (2018).
Summers, R. A. et al. Novel human iPSC models of neuroinflammation in neurodegenerative disease and regenerative medicine. Trends Immunol. 45, 799–813 (2024).
Okano, H. & Morimoto, S. iPSC-based disease modeling and drug discovery in cardinal neurodegenerative disorders. Cell Stem Cell 29, 189–208 (2022).
Morimoto, S. et al. Phase 1/2a clinical trial in ALS with ropinirole, a drug candidate identified by iPSC drug discovery. Cell Stem Cell 30, 766–780 (2023).
Atmaramani, R. et al. Deep learning analysis on images of iPSC-derived Motor Neurons Carrying fALS-genetics Reveals Disease-Relevant Phenotypes. Preprint at bioRxiv https://doi.org/10.1101/2024.01.04.574270 (2024).
Zushin, P. H., Mukherjee, S. & Wu, J. C. FDA Modernization Act 2.0: transitioning beyond animal models with human cells, organoids, and AI/ML-based approaches. J. Clin. Invest. 133, e175824 (2023).
Yu, S. et al. Integrating inflammatory biomarker analysis and artificial-intelligence-enabled image-based profiling to identify drug targets for intestinal fibrosis. Cell Chem. Biol. 30, 1169–1182 (2023).
Ingber, D. E. Human organs-on-chips for disease modelling, drug development and personalized medicine. Nat. Rev. Genet. 23, 467–491 (2022).
Chen, B., Du, C., Wang, M., Guo, J. & Liu, X. Organoids as preclinical models of human disease: progress and applications. Med. Rev. 4, 129–153 (2024).
Calandrini, C. & Drost, J. Normal and tumor-derived organoids as a drug screening platform for tumor-specific drug vulnerabilities. STAR Protoc. 3, 101079 (2022).
Scalia, G. et al. Deep-learning-based virtual screening of antibacterial compounds. Nat. Biotechnol. https://doi.org/10.1038/s41587-025-02814-6 (2025).
Chungyoun, M. & Gray, J. Fitness landscape for antibodies 2: benchmarking reveals that protein AI models cannot yet consistently predict developability properties. Preprint at bioRxiv https://doi.org/10.64898/2025.12.27.696706 (2025).
Lim, J. et al. Advances in single-cell omics and multiomics for high-resolution molecular profiling. Exp. Mol. Med. 56, 515–526 (2024).
Ota, M. et al. Causal modelling of gene effects from regulators to programs to traits. Nature 650, 399–408 (2025).
Sivanandan, S. et al. A pooled Cell Painting CRISPR screening platform enables de novo inference of gene function by self-supervised deep learning. Nat. Commun. 17, 77 (2025).
Ghiandoni, G. M. et al. Augmenting DMTA using predictive AI modelling at AstraZeneca. Drug Discov. Today 29, 103945 (2024).
Medcalf, M. et al. Overcoming DMTA cycle challenges: a unified AI-driven system for efficient drug design. Preprint at https://doi.org/10.26434/chemrxiv-2024-0z7g6-v2 (2025).
Ma, S. et al. Evolving drug discovery using AI, automation, and ASMS through an integrated D-preMTA-MTA strategy for target-focused library exploration Preprint at. Research Square https://doi.org/10.21203/rs.3.rs-3753964/v1 (2023).
Tang, Q. et al. AI-driven robotics laboratory identifies pharmacological TNIK inhibition as a potent senomorphic agent. Aging Dis. 17, 432–451 (2025).
Alves, V. M. et al. Curated data in — trustworthy in silico models out: the impact of data quality on the reliability of artificial intelligence models as alternatives to animal testing. Altern. Lab. Anim. 49, 73–82 (2021).
Saha, U. S. et al. Step forward cross validation for bioactivity prediction: out of distribution validation in drug discovery. Preprint at bioRxiv https://doi.org/10.1101/2024.07.02.601740 (2024).
Brumfield, M. A. et al. Delivering regulatory impact from consortium-based projects. Nat. Rev. Drug Discov. 24, 889–890 (2025).
Scannell, J. et al. Diagnosing the decline in pharmaceutical R&D efficiency. Nat. Rev. Drug Discov. 11, 191–200 (2012).
Truebel, H. & Seidler, M. Mitigating bias in pharmaceutical R&D decision-making. Nat. Rev. Drug Discov. 21, 874–875 (2022).
Shaywitz, D. You have chosen poorly: why drug developers make bad decisions. Will technology help? Timmermann Report https://timmermanreport.com/2022/10/you-have-chosen-poorly-why-drug-developers-make-bad-decisions/ (2022).
Otamendi, F. J. & Sutil Martín, D. L. The emotional effectiveness of advertisement. Front. Psychol. 11, 2088 (2020).
Bieske, L., Zinner, M., Dahlhausen, F. & Trübel, H. Trends, challenges, and success factors in pharmaceutical portfolio management: cognitive biases in decision-making and their mitigating measures. Drug Discov. Today 28, 103734 (2023).
Peplow, M. Robot chemist sparks row with claim it created new materials. Nature https://doi.org/10.1038/d41586-023-03956-w (2023).
Cokol, M., Iossifov, I., Rodriguez-Esteban, R. & Rzhetsky, A. How many scientific papers should be retracted? EMBO Rep. 8, 422–423 (2007).
Lazebnik, Y. Can a biologist fix a radio? Or, what I learned while studying apoptosis. Cancer Cell 2, 179–182 (2002).
Anderson, P. W.More is different. Science 177, 393–396 (1972).
Frangogiannis, N. G. Why animal model studies are lost in translation. J. Cardiovasc. Aging 2, 22 (2022).
Kitano, H. Systems biology: a brief overview. Science 295, 1662–1664 (2002).
Hasin, Y., Seldin, M. & Lusis, A. Multi-omics approaches to disease. Genome Biol. 18, 83 (2017).
Ravarani, C. N. J. et al. Retrospective evaluation of human genetic evidence for clinical trial success using Mendelian randomization and machine learning. Preprint at medRxiv https://doi.org/10.64898/2026.02.19.26346536 (2026).
Lagunin, A., Stepanchikova, A., Filimonov, D. & Poroikov, V. PASS: prediction of activity spectra for biologically active substances. Bioinformatics 16, 747–748 (2000).
Lounkine, E. et al. Large-scale prediction and testing of drug activity on side-effect targets. Nature 486, 361–367 (2012).
Bender, A. et al. Analysis of pharmacology data and the prediction of adverse drug reactions and off-target effects from chemical structure. ChemMedChem 2, 861–873 (2007).
Kirchmair, J. et al. Predicting drug metabolism: experiment and/or computation? Nat. Rev. Drug Discov. 14, 387–404 (2015).
Davies, M. et al. Improving the accuracy of predicted human pharmacokinetics: lessons learned from the astrazeneca drug pipeline over two decades. Trends Pharmacol. Sci. 41, 390–408 (2020).
Schneckener, S. et al. Prediction of oral bioavailability in rats: transferring insights from in vitro correlations to (deep) machine learning models using in silico model outputs and chemical structure parameters. J. Chem. Inf. Model. 59, 4893–4905 (2019).
Handa, K. et al. Prediction of compound plasma concentration-time profiles in mice using random forest. Mol. Pharm. 20, 3060–3072 (2023).
Andrews-Morger, A., Reutlinger, M., Parrott, N. & Olivares-Morales, A. A machine learning framework to improve rat clearance predictions and inform physiologically based pharmacokinetic modeling. Mol. Pharm. 20, 5052–5065 (2023).
Obach, R. S., Lombardo, F. & Waters, N. J. Trend analysis of a database of intravenous pharmacokinetic parameters in humans for 670 drug compounds. Drug Metab. Dispos. 36, 1385–1405 (2008).
Bauer, J. et al. Ritonavir: an extraordinary example of conformational polymorphism. Pharm. Res. 18, 859–866 (2001).
Yang, Z. Y., He, J. H., Lu, A. P., Hou, T. J. & Cao, D. S. Frequent hitters: nuisance artifacts in high-throughput screening. Drug Discov. Today 25, 657–667 (2020).
Fraser, J. S. & Murcko, M. A. Structure is beauty, but not always truth. Cell 187, 517–520 (2024).
Segler, M. H. S., Preuss, M. & Waller, M. P. Planning chemical syntheses with deep neural networks and symbolic AI. Nature 555, 604–610 (2018).
Genheden, S. et al. AiZynthFinder: a fast, robust and flexible open-source software for retrosynthetic planning. J. Cheminform. 12, 70 (2020).
Voinarovska, V., Kabeshov, M., Dudenko, D., Genheden, S. & Tetko, I. V. When yield prediction does not yield prediction: an overview of the current challenges. J. Chem. Inf. Model. 64, 42–56 (2024).
Gupta, A. How unfair is the coin? https://ankitg.me/blog/2025/01/06/unfair-coins.html (2025).
Bilodeau, C. et al. Generative models for molecular discovery: recent advances and challenges. WIREs Comput. Mol. Sci. 12, e1608 (2022).
Gómez-Bombarelli, R. et al. Automatic chemical design using a data-driven continuous representation of molecules. ACS Cent. Sci. 4, 268–276 (2018).
Baillif, B., Cole, J., McCabe, P. & Bender, A. Deep generative models for 3D molecular structure. Curr. Opin. Struct. Biol. 80, 102566 (2023).
Cretu, M. et al. SynFlowNet: design of diverse and novel molecules with synthesis constraints. Preprint at https://arxiv.org/abs/2405.01155 (2025).
Buttenschoen, M., Morris, G. M. & Deane, C. M. PoseBusters: AI-based docking methods fail to generate physically valid poses or generalise to novel sequences. Chem. Sci. 15, 3130–3139 (2023).
Handa, K. et al. On the difficulty of validating molecular generative models realistically: a case study on public and proprietary data. J. Cheminform. 15, 112 (2023).
Liu, H. et al. How good are current pocket-based 3D generative models?: the benchmark set and evaluation of protein pocket-based 3D molecular generative models. J. Chem. Inf. Model. 64, 9260–9275 (2024).
Nie, D. et al. Durian: a comprehensive benchmark for structure-Based 3D molecular generation. J. Chem. Inf. Model. 65, 173–186 (2025).
Baillif, B. et al. Benchmarking structure-based three-dimensional molecular generative models using GenBench3D: ligand conformation quality matters. Preprint at https://arxiv.org/abs/2407.04424 (2024).
Renz, P. et al. On failure modes in molecule generation and optimization. Drug Discov. Today 32-33, 55–63 (2019).
Langevin, M., Vuilleumier, R. & Bianciotto, M. Explaining and avoiding failure modes in goal-directed generation of small molecules. J. Cheminform. 14, 20 (2022).
Senior, A. W. et al. Improved protein structure prediction using potentials from deep learning. Nature 577, 706–710 (2020).
Abramson, J. et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630, 493–500 (2024).
Pereira, J. et al. High-accuracy protein structure prediction in CASP14. Proteins 89, 1687–1699 (2021).
Momin, A. A. et al. PYK2 senses calcium through a disordered dimerization and calmodulin-binding element. Commun. Biol. 5, 800 (2022).
Jänes, J. & Beltrao, P. Deep learning for protein structure prediction and design-progress and applications. Mol. Syst. Biol. 20, 162–169 (2024).
Watson, J. L. et al. De novo design of protein structure and function with RFdiffusion. Nature 620, 1089–1100 (2023).
Dauparas, J. et al. Robust deep learning–based protein sequence design using ProteinMPNN. Science 378, 49–56 (2022).
Goverde, C. A. et al. Computational design of soluble and functional membrane protein analogues. Nature 631, 449–458 (2024).
Li, Q., Vlachos, E. N. M. & Bryant, P. Design of linear and cyclic peptide binders from protein sequence information. Commun. Chem. 8, 211 (2025).
Lauko, A. et al. Computational design of serine hydrolases. Science 388, eadu2454 (2025).
Desautels, T. A. et al. Computationally restoring the potency of a clinical antibody against Omicron. Nature 629, 878–885 (2024).
Stark, H. et al. BoltzGen: toward universal binder design. Preprint at bioRxiv https://doi.org/10.1101/2025.11.20.689494 (2025).
Chai Discovery Team et al. Zero-shot antibody design in a 24-well plate. Preprint at bioRxiv https://doi.org/10.1101/2025.07.05.663018 (2025).
Boretti, A. Improving chimeric antigen receptor T-cell therapies by using artificial intelligence and internet of things technologies: a narrative review. Eur. J. Pharmacol. 974, 176618 (2024).
Bujak, J. et al. Creating an innovative artificial intelligence-based technology (TCRact) for designing and optimizing T cell receptors for use in cancer immunotherapies: protocol for an observational trial. JMIR Res. Protoc. 12, e45872 (2023).
Klontzas, M. E. et al. Machine learning and metabolomics predict mesenchymal stem cell oteogenic differentiation in 2D and 3D cultures. J. Funct. Biomater. 15, 367 (2024).
Bäckel, N. et al. Elaborating the potential of artificial intelligence in automated CAR-T cell manufacturing. Front. Mol. Med. 3, 1250508 (2023).
Duran, I. et al. Detection of senescence using machine learning algorithms based on nuclear features. Nat. Commun. 15, 1–20 (2024).
Bock, C. et al. High-content CRISPR screening. Nat. Rev. Methods Primers 2, 8 (2022).
Schmidhuber, J. 2011: DanNet triggers deep CNN revolution. https://people.idsia.ch/~juergen/DanNet-triggers-deep-CNN-revolution-2011.html (2021).
Rumelhart, D. E. & McClelland, J. L. in Parallel Distributed Processing: Explorations in the Microstructure of Cognition (eds Feldman, J. A. et al.) 318–362 (MIT Press, 1987).
Archit, A. et al. Segment anything for microscopy. Nat. Methods 22, 579–591 (2025).
Kenyon-Dean, K. et al. ViTally consistent: scaling biological representation learning for cell microscopy. In Proc. 42nd International Conference on Machine Learning (eds Singh, A. et al.) Vol. 267, 29735–29752 (PMLR, 2025).
Brown, T. B. et al. Language models are few-shot learners. In Proc.34th International Conference on Neural Information Processing Systems (eds Larochelle H. et al.) article no. 159, 1877–1901 (Curran Associates, Inc., 2020).
Asher, N. et al. Limits for learning with language models. In Proc. 12th Joint Conference on Lexical and Computational Semantics (eds Palmer, A. & Camacho-Collados, J.) 236–248 (Association for Computational Linguistics, 2023).
Grisoni, F. Chemical language models for de novo drug design: challenges and opportunities. Curr. Opin. Struc. Biol. 97, 102527 (2023).
Bran, M. et al. Augmenting large language models with chemistry tools. Nat. Mach. Intell. 6, 525–535 (2024).
Chaves, J. M. Z. et al. Tx-LLM: a large language model for therapeutics. Preprint at https://arxiv.org/abs/2406.06316 (2024).
Huynh, D. L. et al. AI agents in drug discovery: applications and case studies. Drug Discov. Today 31, 104650 (2026).
Hu, Q. et al. Machine learning to predict adverse drug events based on electronic health records: a systematic review and meta-analysis. J. Int. Med. Res. 52, 3000605241302304 (2024).
Singhal, P. et al. Opportunities and challenges for biomarker discovery using electronic health record data. Trends Mol. Med. 29, 765–776 (2023).
Moynihan, D. et al. Analysis and visualisation of electronic health records data to identify undiagnosed patients with rare genetic diseases. Sci. Rep. 14, 5056 (2024).
Yadav, P., Steinbach, M., Kumar, V. & Simon, G. Mining electronic health records (EHRs): a survey. ACM Comput. Surv. 50, 6 (2018).
Sarwar, T. et al. The secondary use of electronic health records for data mining: data characteristics and challenges. ACM Comput. Surv. https://doi.org/10.1145/3490234 (2023).
Gaber, F. et al. Evaluating large language model workflows in clinical decision support for triage and referral and diagnosis. npj Digit. Med. 8, 263 (2025).
Hager, P. et al. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nat. Med. 30, 2613–2622 (2024).
Ranji, S. R. Large language models — misdiagnosing diagnostic excellence? JAMA Netw. Open 7, e2440901 (2024).
Harrer, S., Shah, P., Antony, B. & Hu, J. Artificial intelligence for clinical trial design. Trends Pharmacol. Sci. 40, 577–591 (2019).
Deloitte Centre for Health Solutions. Intelligent clinical trials. Transforming through AI-enabled engagement. https://www2.deloitte.com/content/dam/insights/us/articles/22934_intelligent-clinical-trials/DI_Intelligent-clinical-trials.pdf (Deloitte University EMEA CVBA, 2024).
Razuvayevskaya, O. et al. Genetic factors associated with reasons for clinical trial stoppage. Nat. Genet. 56, 1862–1867 (2024).
Askin, S., Burkhalter, D., Calado, G. & El Dakrouni, S. Artificial intelligence applied to clinical trials: opportunities and challenges. Health Technol. 13, 203–213 (2023).
Ciray, F. & Doğan, T. Machine learning-based prediction of drug approvals using molecular, physicochemical, clinical trial, and patent-related features. Expert Opin. Drug Discov. 17, 1425–1441 (2022).
Aliper, A. et al. Prediction of clinical trials outcomes based on target choice and clinical trial design with multi-modal artificial intelligence. Clin. Pharmacol. Ther. 114, 972–980 (2023).
Zheng, W. et al. Multimodal clinical trial outcome prediction via large language models and mixture-of-experts. Findings Assoc. Comput. Linguistics EMNLP 2025, 7503–7517 (2025).
Proctor, W. R. et al. Utility of spherical human liver microtissues for prediction of clinical drug-induced liver injury. Arch. Toxicol. 91, 2849–2863 (2017).
Rudolf, A. F. et al. A comparison of protein kinases inhibitor screening methods using both enzymatic activity and binding affinity determination. PLoS ONE 10, e98800 (2014).
Huang, K. et al. Therapeutics data commons: machine learning datasets and tasks for drug discovery and development. Preprint at https://arxiv.org/abs/2102.09548 (2021).
Goldstein, J. L. & Brown, M. S. The clinician-investigator: bewitched, bothered, and bewildered — but still beloved. J. Clin. Invest. 99, 2803–2812 (1997).
Hughes, J. P., Rees, S., Kalindjian, S. B. & Philpott, K. L. Principles of early drug discovery. Br. J. Pharmacol. 162, 1239–1249 (2011).
Hernández-Orozco, S., Kiani, N. A. & Zenil, H. Algorithmically probable mutations reproduce aspects of evolution, such as convergence rate, genetic memory and modularity. R. Soc. Open. Sci. 5, 180399 (2018).