Your task as a reader: decide what the trial actually demonstrated, what remains uncertain, and which next piece of evidence could change the conclusion. Do not begin with the p-value. Begin with the clinical question.

Drugnews reading rule: a “positive” readout can support a narrow claim while leaving the larger investment thesis unproven. The endpoint hierarchy, comparator, effect size, uncertainty, durability, and safety all determine how wide that claim can be.

The five-pass endpoint audit

  1. Define the question. Who was studied, what treatment was compared with what control, what outcome was measured, and over what time?
  2. Locate the endpoint in the hierarchy. Was it primary, multiplicity-controlled secondary, exploratory, or post hoc?
  3. Read the effect in absolute and relative terms. Pair the headline statistic with event rates, scale changes, medians, or response counts.
  4. Stress-test uncertainty. Check the confidence interval, missing data, censoring, analysis population, follow-up, and sensitivity analyses.
  5. Reconnect the result to patients and the program. Ask whether the outcome reflects how patients feel, function, or survive; whether the benefit is durable; and what toxicity or treatment burden accompanies it.

1. An endpoint is a question, not a floating number

ClinicalTrials.gov defines an outcome measure as a pre-specified measurement used to determine the effect of an experimental variable. A complete entry also identifies the outcome type, description, time frame, arms, analysis population, unit, and measure of precision.[1] If a company reports “a 40% response rate” without those elements, the reader still does not know enough to interpret the result.

The ICH E9(R1) framework adds an important question: which treatment effect is the trial trying to estimate? The answer depends on the population, endpoint variable, handling of events such as treatment discontinuation or rescue medication, and the population-level summary. These pieces form an estimand.[2]

This matters because two analyses can use the same endpoint label but answer different questions. An analysis that follows every randomized participant after discontinuation asks about the effect of assigning the treatment strategy. An analysis that censors participants at discontinuation may move closer to an “on-treatment” question and rely on stronger assumptions. Neither label is sufficient on its own; the protocol and statistical analysis plan determine the claim.

Primary, secondary, exploratory, and post-hoc results

The primary endpoint is the trial’s main pre-specified efficacy test. Secondary endpoints can support or extend the primary result, but only those protected by the statistical testing plan can generally carry confirmatory weight. Exploratory and post-hoc analyses are useful for generating hypotheses, not for silently replacing a failed primary endpoint.

Testing many endpoints increases the chance of a false-positive conclusion unless multiplicity is controlled. FDA’s 2022 guidance describes grouping, ordering, and statistical adjustment strategies for this problem.[3] A press release that lists several nominally significant secondary results without explaining the testing hierarchy deserves a narrower interpretation.

2. Ask whether the endpoint measures patient benefit—or predicts it

FDA describes a clinical outcome assessment as a measure that reflects how a patient feels, functions, or survives. It may be reported by the patient, a clinician, an observer, or a performance test.[4] These outcomes are not automatically perfect: a scale still needs a defined context of use, reliable administration, and an interpretable change threshold.

A surrogate endpoint is different. It is a biomarker, imaging measure, physical sign, or other marker used to predict clinical benefit rather than measure the benefit directly. FDA explicitly says the acceptability of a surrogate is context-dependent—disease, population, mechanism, and available therapy all matter—and a surrogate used in one program should not be assumed valid in another.[5]

Under accelerated approval, a drug for a serious condition may be approved on a surrogate or intermediate clinical endpoint that is reasonably likely to predict benefit. Confirmatory evidence is then required; failure to verify sufficient benefit can lead to a changed indication or withdrawal.[6] Therefore, “FDA has accepted this endpoint before” is not the same as “this endpoint guarantees traditional approval in this setting.”

Endpoint formWhat it can showWhat to inspect before believing the headline
Binary
Response / no response
The proportion meeting a defined thresholdDenominator, confidence interval, confirmation rule, blinded or independent review, missing scans, and duration of response
Time-to-event
PFS, OS, MACE
How event risk evolves over follow-upHazard ratio and confidence interval, absolute event rates, Kaplan–Meier curves, censoring, follow-up maturity, proportional-hazards assumption, and competing events
Continuous scale
Symptoms, function, biomarkers
Average change from baseline or difference between groupsBaseline balance, scale direction, clinically interpretable change, missing-data assumptions, rescue treatment, and responder analysis
Composite
Several events combined
Time to the first event among defined componentsWhether common but less serious components dominate, whether components move in the same direction, and whether the result is driven by one component

3. Translate the statistic into a statement a reader cannot misquote

Response rate: add depth, duration, and denominator

Objective response rate can reveal antitumor activity quickly, which is why it is useful in early oncology development. But response is not survival. FDA’s oncology endpoint guidance treats response rate together with response duration, assessment method, and disease context.[7] A 60% response rate in 20 highly selected patients with four months of follow-up is a signal; it is not yet a durable comparative-benefit claim.

Hazard ratio: do not translate it as “lived 20% longer”

A hazard ratio compares instantaneous event rates over the observed period under a statistical model. An HR of 0.80 is commonly described as a 20% relative reduction in hazard, not a 20% extension of survival and not a 20-percentage-point reduction in absolute risk. Pair it with absolute event rates, medians or restricted mean survival when available, the full curves, and the confidence interval.

P-value: ask how large and how precise

A p-value addresses compatibility with a null hypothesis under the specified analysis; it does not measure clinical importance, the probability that the drug “works,” or the chance that the result will replicate. The confidence interval is often more informative because it displays a range of effects compatible with the data. A narrow interval around a modest effect and a wide interval spanning trivial and transformative effects are not the same evidence package.

Subgroups: look for interaction, not isolated stars

A subgroup can appear “positive” and another “negative” because each contains fewer patients. The relevant question is whether the treatment effect differs between subgroups, supported by an interaction test and a plausible pre-specified rationale. Post-hoc slicing across many biomarkers, regions, ages, and treatment histories creates another multiplicity problem.

4. Worked example: reading SELECT beyond “20% risk reduction”

The SELECT trial provides a clean example of how a positive endpoint should be translated. More than 17,600 adults with established cardiovascular disease and overweight or obesity were randomized to semaglutide 2.4 mg or placebo on top of standard care. The primary endpoint was time to first major adverse cardiovascular event—a composite of cardiovascular death, nonfatal myocardial infarction, or nonfatal stroke.

FDA reported events in 6.5% of semaglutide participants and 8.0% of placebo participants. The prescribing information reports an HR of 0.80 with a 95% confidence interval of 0.72 to 0.90.[8] The two descriptions are complementary:

  • Relative view: the estimated hazard was 20% lower over follow-up.
  • Absolute view: the observed event proportions differed by 1.5 percentage points over the trial period.
  • Precision: the confidence interval excluded 1.00 for the primary endpoint.
  • Context: both groups received cardiovascular risk management and lifestyle counseling, so the comparison was against placebo plus current standard care—not against no care.

The label also reports a median follow-up of 41.8 months, 96.9% trial completion, vital-status ascertainment for 99.4%, and permanent study-drug discontinuation in 31% versus 27% of participants.[8] Those details help assess maturity and treatment exposure.

The FDA clinical review noted that the confirmatory secondary endpoint of cardiovascular death did not reach statistical significance: HR 0.85, 95% CI 0.71 to 1.01. The point estimate was favorable, but that narrower claim was not established by the same standard as the primary composite.[9]

Defensible conclusion: SELECT supported a reduction in the risk of first MACE in the studied population. It did not justify converting the HR into “20% longer life,” nor did it establish every component or subgroup as independently positive.

5. The due-diligence checklist

Before updating a clinical or company thesis, write one sentence for each line:

  1. Population: inclusion criteria, disease stage, biomarkers, prior treatment, geography, and baseline risk.
  2. Design: randomized or single-arm; blinded or open-label; control choice; stratification; sample size.
  3. Question: primary endpoint, time frame, estimand, and handling of discontinuation, rescue therapy, or death.
  4. Result: absolute values, relative effect, confidence interval, and pre-specified statistical threshold.
  5. Integrity: missing data, censoring, protocol deviations, interim looks, and sensitivity analyses.
  6. Durability: median follow-up, number still at risk, response duration, and curve behavior over time.
  7. Safety and burden: serious and grade ≥3 events, discontinuations, dose intensity, monitoring, administration, and competing risks.
  8. Regulatory fit: whether the endpoint and design can support the proposed indication and approval pathway in this exact context.
  9. Commercial fit: whether the measured benefit is differentiated enough to change practice, reimbursement, or market share.

Limits and uncertainty

Public readers rarely have the full protocol, statistical analysis plan, regulatory correspondence, adjudication charter, and patient-level data at first readout. Company releases may omit denominators, censoring detail, multiplicity rules, and negative secondary outcomes. When those materials are unavailable, state the missing evidence instead of assuming the most favorable analysis.

Endpoint acceptability is also context-specific and can change as standards of care and regulatory expectations evolve. This framework supports evidence review; it does not determine an individual patient’s treatment or replace medical, statistical, or regulatory advice.

Primary sources

  1. ClinicalTrials.gov — Results Data Element Definitions for Interventional and Observational Studies.
  2. FDA / ICH — E9(R1): Estimands and Sensitivity Analysis in Clinical Trials.
  3. FDA — Multiple Endpoints in Clinical Trials, final guidance.
  4. FDA — Clinical Outcome Assessment: Frequently Asked Questions.
  5. FDA — Table of Surrogate Endpoints That Were the Basis of Drug Approval or Licensure.
  6. FDA — Accelerated Approval.
  7. FDA — Clinical Trial Endpoints for the Approval of Cancer Drugs and Biologics.
  8. FDA — Wegovy prescribing information, SELECT clinical-study section.
  9. FDA — Clinical review for the Wegovy cardiovascular-risk-reduction supplement.
This guide is intended for industry research and education only. It does not constitute medical, investment, legal, fundraising, or individual securities advice.