From COVID-19 vaccines to anticancer drugs and gene therapies for rare diseases, drug discovery is a meticulous process with research and therapy-specific milestones. A well-established drug discovery pipeline can help pave the path to market.
What is the Drug Discovery Pipeline?
Understanding the Drug Discovery Pipeline
A drug discovery pipeline involves several stages, including preclinical and clinical studies, submission to regulatory authorities and post-market surveillance.
A typical pipeline starts with target identification, which involves characterizing the genomic and biochemical abnormalities underlying the disease. The clinical significance of a potential target must be validated by assessing the therapeutic value of correcting its activity.
After establishing the target, researchers begin with a hit-discovery process, screening libraries containing hundreds of thousands of compounds. Here, assay development helps characterize the compounds' impact on the target(s) and disease phenotypes while revealing potential off-target interactions. Top-ranking compounds are modified to optimize potency and derisk adverse interactions in the hit-to-lead phase.
The hits are continuously optimized until the lead compound is selected to proceed to preclinical development, which comprises animal studies to delineate efficacy, safety, pharmacokinetics (PK, what the drug does to the body) and pharmacodynamics (PD, what the body does to the drug).
In sequence, these early-stage steps are:
- Target identification and validation: identifying a disease-relevant target and confirming its therapeutic potential
- Hit discovery: screening compounds to identify activity against the target
- Hit-to-lead: improving promising hits for potency, selectivity and other drug-like properties
- Lead optimization: refining one or more leads before candidate selection and preclinical development
Following preclinical development, the sponsor submits an Investigational New Drug (IND) application that includes preclinical findings, manufacturing information and proposed clinical protocols. If the FDA allows the study to proceed, clinical development generally advances through the phases summarized below.
*Advancement-rate figures sourced from IntuitionLabs' Drug Development Pipeline Guide funnel analysis (https://intuitionlabs.ai/articles/drug-development-pipeline-guide); included for context, not page-specific data.
After the drug development process, a new drug application (NDA), comprising preclinical and clinical reports, is submitted to the authorities.
Even after approval, the pharmaceutical company continues bilateral communication with the regulatory authorities. At this stage, a phase 4 study must be conducted, including post-marketing surveillance, to analyze the drug’s long-term safety and efficacy.
When companies adhere to the standard drug discovery pipeline format, they can deliver life-saving drugs to market in a timely and effective manner. The pipeline ensures that patients can benefit from novel drugs without the risk of adverse events.
Key Takeaways
- The pipeline follows a fixed sequence: target ID, hit discovery, hit-to-lead, lead optimization, preclinical testing, IND filing, Phases 1–4 and NDA/BLA submission, with no stages skipped.
- “Drug discovery" covers target-to-lead work, while "drug development" spans preclinical to post-market; both terms have similar search volume and are relevant.
- Attrition is high at each stage: from thousands screened, only a small fraction reach Phase 1 and roughly 1 in 15 get approval.
- AI and high-throughput screening now accelerate target identification and hit discovery, rather than being optional.
- Regulatory checkpoints (IND, NDA, BLA) require proof of safety and efficacy before approval, not just formalities.
How is Target Identification Conducted
The Role of Biological Targets in Drug Discovery
A biological target is a gene, protein or enzyme that functions abnormally in a diseased state but can be restored to a healthy phenotype when treated with a drug candidate.
Given the immense time and financial resources required for drug discovery, development and clinical studies, the target must be correctly identified and validated.
Methods for Identifying Drug Targets
Target identification runs on genomic and proteomic levels. On a genomic scale, sequences in a genome can be probed with enhancers or suppressors to identify the gene or set of genes conferring the most therapeutic outcome. However, considering post-transcriptional and post-translational modifications from gene to protein, proteomics-based target discovery can be more informative for targeting structural and functional changes in proteins between healthy, diseased and drug-treated states.
Both methods can be significantly accelerated by utilizing high-throughput screening (HTS), which involves testing many molecules for their impact on the target’s activity. HTS workflows comprise automated equipment, which maximizes reproducibility and output.
Challenges in Target Identification
Several challenges lie between a target and its drug treatment. Firstly, diseases involve multiple pathways, making it challenging to pinpoint a single target. Even when a potential target is found, it may not be druggable because of its structural properties that shield potential drug-binding sites. Cell lines and animal models do not sufficiently recapitulate human physiology, which may lead to the identification of false targets or target-drug interactions. Finally, the analysis of large transcriptomics and proteomics datasets remains a challenge in understanding disease mechanisms.
- Disease complexity: multiple pathways are often involved, making a single target hard to isolate
- Druggability: a validated target may still lack a structural pocket that a drug molecule can bind to
- Model translation: cell lines and animal models don't fully replicate human physiology, risking false targets
- Data complexity: large transcriptomic and proteomic datasets remain difficult to interpret at scale
The Importance of Assay Development in Drug Discovery
A key milestone in a drug discovery pipeline is the in vitro assessment of drug candidates for efficacy and safety using human cell and tissue models.
Types of Assays Used in Drug Discovery
Biochemical assays inspect the direct binding between the drug and its target to determine binding affinity and the drug's subsequent activating or inhibitory effects.
Cell-based or in vitro assays test the drug molecule in cell cultures to assess its effects on cell viability, cell membrane integrity, proliferation, morphological features, gene expression, ion channel function (for neuronal and cardiac cell models) and metabolic activity.
Computation-based or in silico assays are performed using computer simulation software to predict binding scores, drug-target interactions, the free energy of binding and the overall impact of introducing the drug-target on the entire biological network.
How Assays Help Further the Drug Discovery Pipeline
The high rate of drug failures in clinical trials underscores the importance of assay development for documenting various aspects of a drug's mechanism of action. These assays ensure that the drug molecule is worth the years of investment and guide researchers in optimizing drug performance. From this perspective, assay development is key to time-and cost-efficient drug discovery pipelines.
How Do Databases Support Drug Discovery Pipelines?
Types of Databases in Drug Discovery
Several databases play a crucial role in drug discovery by storing vast amounts of data that can assist at all stages of drug discovery. Researchers can benefit from these databases to identify biologically relevant targets, optimize lead compounds and streamline preclinical and clinical studies. Thus, the balance of maximum efficacy and minimum adverse effects is obtained.
Applications of Databases in Drug Discovery
Novel databases are emerging to foster personalized medicine and AI-driven drug discovery.
UK Biobank is a patient-powered platform encompassing a broad range of objective (e.g., data on patient samples) and subjective (e.g., patient lifestyle information) information. With contributions from many UK-based participants, the database bolsters a deeper understanding of diseases at the individual level.
Another initiative is Chan Zuckerberg’s CELLxGENE database, which hosts a large pool of single-cell transcriptomics data. Researchers can visualize, compare and interpret large-scale single-cell RNA sequencing (scRNA-seq) datasets to achieve more accurate target identification and a better understanding of disease mechanisms and drug responses.
The diverse and high-quality datasets from these databases are invaluable inputs for AI-driven drug discovery platforms to identify novel drug targets and biomarkers, predict drug toxicity and support drug design. For example, the UK Biobank was used in research to identify targets that highly influence lipid levels in heart failure.1 Meanwhile, single-cell RNA-seq datasets of different cancer subtypes were compiled and uploaded to CELLxGENE to aid the construction of predictive models of cancer cell responses to immune checkpoint inhibitors.²
How AI Adoption Is Changing the Pipeline
AI is increasingly being applied across the drug discovery pipeline, including target identification, hit discovery, biomarker analysis and preclinical development. Industry analyses estimate that AI-enabled approaches could reduce time and costs during early research and preclinical development by 25–50%, although realized gains vary by program and have not yet been demonstrated consistently at industry scale. As of July 2026, no drug widely classified as AI-discovered or AI-designed had received full FDA approval, underscoring that faster candidate generation has not eliminated the biological, clinical and regulatory factors that drive pipeline attrition.
What are the Key Steps in Clinical Development and Approval?
Significant Milestones of Phase 1-4 Clinical Trials
Clinical trials are conducted in four phases to evaluate safety, efficacy, benefit-risk profile and longitudinal effects.
Phase 1 trials enroll 20-100 healthy volunteers to assess safety, side effects and the maximum tolerated dose.
Phase 2 tests drug efficacy and safety on 100-500 patients with the target condition. Here, optimal therapeutic dose and short-term side effects are determined.
Phase 3 is conducted across multiple sites with 1000-5000 patients to confirm long-term efficacy and safety and to record rare adverse events. Upon successful completion, submission of a New Drug Application (NDA) or Biologics License Application (BLA) to regulatory agencies (FDA, EMA) becomes possible.
Phase 4 involves post-market surveillance of the millions of patients worldwide who use the drug. It can provide insight into the long-term effects while creating room for repurposing for new indications.
Below are brief explanations of various applications associated with a drug discovery pipeline submitted to regulatory agencies.
Investigational New Drug (IND) applications are submitted to the FDA to allow the pharmaceutical company to commence clinical trials. The application must include preclinical data, manufacturing information and proposed clinical trial protocols.
A Biological License Application (BLA) is submitted to the FDA to approve vaccines, monoclonal antibodies, gene therapies and recombinant proteins for marketing. It is a comprehensive report that includes product characterization, manufacturing information, quality control, preclinical and clinical data and risk management plans.
A New Drug Application (NDA) is similar to a BLA in content but is submitted to the FDA for approval of pharmaceutical drugs (small molecules).
Why the Phased System Exists
The phase-by-phase structure reflects decades of evolving scientific and regulatory safeguards. Two pivotal U.S. milestones were the 1938 Federal Food, Drug and Cosmetic Act, enacted after more than 100 deaths from the untested Elixir Sulfanilamide formulation, which required manufacturers to provide evidence of a new drug’s safety before marketing. The 1962 Kefauver-Harris Amendments, influenced by the thalidomide tragedy, added the requirement for evidence of effectiveness from adequate and well-controlled studies. Together, these reforms helped establish the modern expectation that drug developers progressively characterize safety, dosage and efficacy as clinical testing expands from early studies to larger patient populations.
Featured Product
Exploring the Drug Discovery Pipeline: From Research to Development
Echo® MS+ Systems
Chosen because it is an acoustic-ejection mass spectrometry system that screens compound libraries at very high throughput without LC separation, tying directly to this page's subject of drug product pipeline.
Echo 525 Acoustic Liquid Handler
Chosen as a complementary second option because it is an acoustic droplet-ejection liquid handler for touchless, nanoliter-precision dispensing in HTS and library prep, also directly relevant to drug product pipeline.
See how Danaher Life Sciences can help
FAQs
What are the stages of drug discovery?
The drug discovery process includes target identification, hit discovery, lead optimization, preclinical testing and clinical trials (Phases 1–3). If successful, the drug undergoes regulatory approval (e.g., NDA/BLA submission) before post-marketing surveillance (Phase 4).
What is the role of artificial intelligence in drug discovery?
AI accelerates drug discovery by predicting drug-target interactions, analyzing biological data and optimizing molecule design. It improves efficiency in virtual screening, biomarker discovery and clinical trial optimization.
What are the main factors influencing the success rate of new drugs?
Key factors include drug efficacy, safety, target selection, trial design, regulatory submissions and market competition.
How long does the drug discovery process typically take?
On average, it takes 10–15 years from discovery to approval, with high costs (~$1–2 billion) and a low success rate (~10%) from Phase 1 to market.
What is the difference between drug discovery and drug development?
Drug discovery finds and refines the candidate, whereas drug development tests whether it can become a safe, effective and approvable medicine.
How many patients are typically involved in each phase of clinical trials?
Actual enrollment depends on the disease, treatment type, trial design and regulatory requirements. Typical enrollment ranges are:
- Phase 1: 20–100 healthy volunteers or people with the disease or condition
- Phase 2: A few dozen to about 300 patients with the disease or condition
- Phase 3: Several hundred to about 3,000 patients
- Phase 4: Varies by study because it occurs after approval in defined or routine-use populations
References
- Xiao J, Ji J, Zhang N, Yang X, Chen K, Chen L, et al. Association of genetically predicted lipid traits and lipid-modifying targets with heart failure. Eur J Prev Cardiol 2023;30(4):358-366.
- Gondal MN, Cieslik M, Chinnaiyan AM. Integrated cancer cell-specific single-cell RNA-seq datasets of immune checkpoint blockade-treated patients. Sci Data 2025;12(1):139.