Clinical trials
Clinical trials
Trial phases
- Phase 0 (microdosing): PK/PD signal; very small N.
- Phase 1: dose-finding, safety, PK, MTD/RP2D. Traditional 3+3 design, model-based (CRM), or model-assisted (BOIN, mTPI). Not designed to test efficacy, though activity is recorded.
- Phase 2: preliminary efficacy; ORR, PFS often primary. Single-arm (vs historical control) or randomized phase 2. Result is "preliminary" because it may be uncontrolled, use a surrogate endpoint, or be too small to exclude chance.
- Phase 3: pivotal; randomized, controlled, large enough to separate chance from true effect; efficacy vs SOC for FDA approval. Usually OS or PFS primary.
- Phase 4: post-marketing surveillance.
- Note: there is no governing body for trial nomenclature (phase Ib, II/III, etc.); the FDA mandates phase definitions for studies it regulates.
- Master protocols (multiple parallel drug subtrials under one protocol):
- Basket: single drug across multiple cancers with a shared biomarker (e.g., vemurafenib in BRAF V600 tumors).
- Umbrella: multiple drugs in one cancer based on biomarker subgroups (e.g., Lung-MAP).
- Platform: multiple drugs added/dropped over time (e.g., I-SPY2).
Phase I design in depth
- Goal: find the highest dose with acceptable toxicity, to carry to phase II. For cytotoxics, higher dose usually means more efficacy and more toxicity; for targeted agents, higher doses may add off-target toxicity without more benefit, so the aim may be an optimal biologic dose rather than an MTD.
- Key terms: MTD (maximum tolerated dose), RP2D (recommended phase 2 dose, may be below MTD), DLT (dose-limiting toxicity, typically grade ≥3 nonhematologic or grade ≥4 hematologic attributable to the agent, assessed usually over 1 cycle).
- Starting dose: often one-tenth of the rodent STD10 (severely toxic dose in 10%); if a nonrodent is the most relevant species, one-sixth of its highest nonseverely toxic dose (ICH S9). Dose escalation by decreasing fractions (100%, 67%, 50%, 40%, 33%) = modified Fibonacci.
- 3+3 design: cohorts of 3; escalate if 0/3 DLT; add 3 more if 1/3 DLT; MTD is the dose below the level where ≥2 patients have DLT. Widely used but considered outdated vs model-based designs.
- Alternatives: accelerated titration (single-patient cohorts until a grade ≥2 event); continual reassessment method (CRM, Bayesian model of the dose-toxicity curve); Bayesian optimal interval (BOIN), simple table-driven, treats fewer patients above MTD and better identifies the true MTD.
Study objectives & hypothesis testing
- Objective → outcome → endpoint: the objective is the research question; the outcome is the measured variable (e.g., tumor size); the endpoint is derived (e.g., ORR). The primary endpoint should be clinically relevant, objective, and sensitive to the intervention, and defined a priori.
- Null hypothesis (H0): no difference. Alternative hypothesis (Ha): reflects the scientific question (1-sided or 2-sided). Trials quantify evidence against H0; we "reject" or "do not reject" H0 (never "accept").
- Type I error (α): false positive, reject a true H0 (typically 0.05; phase II often 0.10; genomics much lower).
- Type II error (β): false negative, fail to reject a false H0 (typically 0.20). Power = 1 − β (commonly 0.80 to 0.90).
Endpoints
- Overall survival (OS): time from randomization to death; gold standard but slow and can be confounded by later therapies/crossover.
- Progression-free survival (PFS): time to progression or death; surrogate.
- Disease-free survival (DFS): relapse or death (adjuvant trials).
- Event-free survival (EFS): broader event set (neoadjuvant trials).
- Time to progression (TTP): to progression only.
- Objective response rate (ORR): RECIST v1.1, CR + PR. iRECIST for IO.
- Pathologic complete response (pCR): neoadjuvant surrogate.
- Patient-reported outcomes (PROs): QoL, symptoms; require fit-for-purpose validated instruments and prespecified collection/analysis. Blinding reduces bias but PROs are used in open-label trials, where reporting bias and missing data need explicit handling.
- Surrogate endpoints: PFS/DFS/ORR substitute for OS only if validated; accelerated approval may use a surrogate reasonably likely to predict benefit (confirmatory trial required), and a sufficiently large PFS/DFS gain can itself be a benefit. A valid surrogate must reliably capture the treatment effect on the clinical outcome in the relevant context; causal-pathway placement is neither sufficient (off-target effects) nor strictly required, and trial-level (not just patient-level) correlation is the key evidence; a marker merely correlated with outcome (not causal) can mislead (analogy: lowering a correlated-but-not-causal marker need not improve the outcome). Superiority in PFS/DFS can occur despite more toxic deaths or from a better safety profile without a true anticancer effect.
RECIST response categoriesTarget lesions, up to 5
| Response (RECIST) | Definition |
|---|---|
| Complete response (CR) | Disappearance of all non-nodal target lesions; target lymph nodes must shrink to short axis <10 mm |
| Partial response (PR) | ≥30% decrease in the sum of diameters (vs baseline) |
| Stable disease (SD) | Neither PR nor PD criteria met |
| Progressive disease (PD) | ≥20% increase in the sum of diameters from nadir (and ≥5 mm absolute) or new lesions |
Statistical concepts
- Hazard ratio (HR): relative instantaneous event rate. <1 = experimental better. A 95% CI excluding 1.0 = statistically significant. If event risk is not small, HR (and especially OR) diverges from a simple risk ratio.
- Odds ratio (OR): odds = p/(1−p); OR overstates effect when p is large (e.g., 90% vs 80% risk gives OR 2.25 though risk is far from doubled). Used in logistic regression.
- Type I error (α), Type II error (β), power (1−β) as above.
- Intention-to-treat (ITT): analyze all randomized patients as assigned; standard for the primary endpoint. Per-protocol/as-treated are secondary and interpreted cautiously.
- Crossover: can dilute the OS signal in 2L+ trials.
- Subgroup analysis: forest plots; exploratory unless pre-specified and adequately powered (trials rarely powered for subsets).
- Multiple comparisons: many tests inflate false positives; corrections include Bonferroni (divide α by number of tests, conservative) and false discovery rate (FDR, error among positive tests).
- Meta-analysis: pools studies; heterogeneity (I²).
- Bayesian methods: adaptive designs, prior probabilities (CRM, BOIN).
P value & confidence intervals
- P value: probability of observing the data (or more extreme) if H0 is true; it measures evidence against H0. Small P (<α) → reject H0.
- Common misinterpretations: (1) the P value is NOT the probability that H0 is true; (2) a large/non-significant P does NOT prove no difference (absence of evidence is not evidence of absence).
- Clinical vs statistical significance: a large trial can make a 1-day OS difference statistically significant (P = .001) yet clinically trivial; conversely a meaningful 6-month difference can be non-significant if underpowered.
- Confidence interval (CI): a range expected to contain the true value 95% of the time; conveys magnitude and uncertainty better than a P value. A 95% CI for an HR that crosses 1.0 implies no significant difference (equivalent to P > .05).
Sample size & power
- Sample size depends on three elements: the statistical factor (α and power), the noise (outcome variability), and the effect size (Δ, the smallest clinically meaningful difference).
- Sample size increases with higher power, smaller α, and greater variability; it decreases with a larger effect size and a more homogeneous population. Effect size dominates: halving Δ roughly quadruples the required N.
Phase II designs
- Single-arm vs historical control: compares observed response with a historical rate p0; p1 is the minimum meaningful improvement. Phase II typically uses α ~0.10 and power ~0.90 (higher error accepted because phase III confirms).
- Simon two-stage (optimal / minimax): interim stopping for futility if too few responses in stage 1, minimizing patients exposed to an ineffective drug.
- Pick-the-winner (selection) design: randomized among experimental variants; selects the best arm without formal type I error control (not powered for between-arm hypothesis testing).
- Prognostic vs predictive biomarkers: a prognostic marker predicts outcome regardless of treatment (e.g., metastatic status); a predictive marker identifies who benefits from a specific therapy (e.g., HER2 for trastuzumab). Predictive biomarkers require a treatment-by-biomarker interaction test (four groups: biomarker +/− by treatment +/−). Integral biomarkers guide treatment decisions in the trial (e.g., Oncotype DX in TAILORx); integrated biomarkers are analyzed for association.
- Adaptive randomization: shifts allocation toward better-performing arms as data accrue (I-SPY2); may raise equipoise concerns and can require a larger total N.
- Biases in phase II: single-arm trials are most prone, chiefly from lack of a contemporaneous randomized control (differing eligibility, supportive care, endpoint definitions, surrogate endpoints, random variation), producing false positives that later fail in phase III.
Phase III design principles
- Randomization: the key tool to minimize bias; balances known and unknown prognostic factors. Usually 1:1 (most efficient, smallest N, best equipoise); 2:1 needs a larger N.
- Stratified randomization & blocking: improve balance for strong prognostic factors and over time; most useful in smaller trials (largely superfluous in large trials). Dynamic allocation is an adaptive alternative.
- Blinding & placebo: single-blind (patient) or double-blind (patient + physician) reduce post-randomization bias; matters most for subjective endpoints (tumor response, PROs), less for mortality. Placebo controls are rare in oncology except adjuvant or add-on trials.
- ITT for the primary analysis in superiority trials; in non-inferiority/equivalence trials both ITT and per-protocol are prespecified (ITT can falsely favor non-inferiority), so neither is automatically secondary.
- Interim analysis / group sequential design: pre-specified looks with the option to stop for efficacy, harm, or futility; alpha-spending controls the overall type I error (interim thresholds may be very stringent, e.g., P < .001). Reviewed by an independent Data and Safety Monitoring Board (DSMB).
Non-inferiority & equivalence
- Rationale: a new therapy that is less toxic, less invasive, cheaper, or more convenient may be acceptable if not meaningfully less effective.
- Equivalence: shows the difference falls within a pre-specified margin in either direction; very large N (a 5% equivalence margin needs ~4× the N of a superiority trial powered for 10%).
- Non-inferiority: sets an acceptable upper limit of inferiority (the non-inferiority margin); usually a 1-sided test with a smaller N than equivalence. After establishing non-inferiority, one may then argue for less toxicity or better QoL.
Adaptive & Bayesian designs
- Adaptive designs: pre-specified modifications during conduct in one of four areas, randomization ratio (add/drop arms), sample-size re-estimation, enrichment (restrict to a benefiting subgroup at interim), or interim stopping. Blinded adaptations (e.g., increasing N for a lower-than-expected event rate) are lower-risk than unblinded ones.
- Seamless phase II/III: combines dose selection and confirmation, saving time and combining data for inference.
- Bayesian statistics: prior distribution + data → posterior distribution; allows direct statements like "90% probability arm A is better than arm B." Controversy centers on the choice of prior (a "flat"/non-informative prior often reproduces frequentist results). More common in phase I/II adaptive designs.
Survival (time-to-event) analysis
- Censoring: follow-up ends before the event; assumed non-informative (unrelated to later event risk).
- Kaplan-Meier estimator: step-function survival curve; each step multiplies the prior height by the fraction surviving at that event time. Median survival = time the curve first drops below 50%. CIs widen over time as fewer patients remain at risk.
- Log-rank test: compares curves via the relative number of events given those at risk (not the curve heights directly); the standard comparison for OS/PFS/DFS.
- Hazard function: underlying event rate among those at risk; the HR summarizes the between-group ratio (Cox model, below).
- Immortal time bias: grouping by a post-baseline event (e.g., responders vs not) from the start of therapy biases the comparison; the fix is a landmark analysis (start survival at a fixed later time).
- Competing risks: with multiple event types (e.g., relapse vs non-relapse death), Kaplan-Meier overestimates a single cause; use cumulative incidence functions (which sum, with EFS, to 100%). Effective therapy that prevents relapse can raise the apparent cumulative incidence of non-relapse death by leaving more patients at risk.
Regression models & biomarker evaluation
- Cox proportional hazards: relates covariates to a time-to-event outcome; effect summarized as an HR; allows fixed or time-dependent covariates. Fine-Gray subdistribution models handle competing risks.
- Logistic regression: models the odds of a binary outcome; effect summarized as an OR. Binary comparisons use χ² (or Fisher exact for small samples).
- Multivariable vs multivariate: multivariable = many predictors, one outcome (common); multivariate = multiple outcomes (rare; often misused term).
- Biomarker discrimination: ROC curve plots sensitivity (true-positive rate) against 1−specificity; AUC = 1 is perfect, 0.5 is no better than chance (useful prognostic markers often need AUC >0.80). Time-to-event discrimination uses Harrell's C-index.
Real-world evidence (RWE)
- Registries, EHR-derived datasets: Flatiron, SEER-Medicare.
- External controls: for rare cancers.
- FDA pathways increasingly accept RWE supplements.
Regulatory pathways (FDA)
- Standard / traditional approval: direct clinical benefit (OS, or clinically meaningful PFS, DFS/EFS, or PRO improvement) or a validated surrogate.
- Accelerated approval: surrogate endpoint reasonably likely to predict benefit; requires confirmatory trial.
- Breakthrough therapy: substantial improvement over existing.
- Priority review: 6 mo (vs 10).
- Project Orbis: international concurrent reviews.
- Project Frontrunner: encourages early-line studies.
- Project Optimus: dose optimization (avoiding MTD bias).
- Tumor-agnostic approvals: pembro for MSI-H/dMMR/TMB-high; larotrectinib/entrectinib/repotrectinib (NTRK); selpercatinib (RET); dabrafenib + trametinib (BRAF V600E); dostarlimab (dMMR); trastuzumab deruxtecan (HER2 IHC 3+).
Equity & access
- Diversity in clinical trials: underrepresentation of Black, Hispanic, AYA, elderly, rural.
- Decentralized trials and telehealth components.
Trial conduct & ethics
- IRB approval, informed consent, GCP.
- Equipoise: required for randomization.
- Data and Safety Monitoring Board (DSMB): interim analyses, futility / efficacy stopping.
- CONSORT, REMARK, STARD reporting standards; statistical analysis plan (SAP) finalized before unblinding.
2024-2026 trial-design updates
- Project Optimus in action: FDA increasingly requires randomized dose-comparison in registration studies, e.g., sotorasib 240 vs 960 mg (CodeBreaK 100 randomized dose-comparison cohort), adagrasib approved starting dose 600 mg BID (400 mg BID is the first dose-reduction level), cabozantinib in RCC dose-optimization; MTD no longer accepted as RP2D by default.
- Eligibility criteria modernization: ASCO-Friends of Cancer Research joint recs, expand for prior/concurrent malignancy, brain mets stable 2+ wk, HIV/HBV/HCV, organ dysfunction, ECOG 2 (Kim et al JCO 2024 update).
- Decentralized / hybrid trials: FDA guidance Sep 2024 (final) on decentralized clinical trials; telehealth visits, home nursing, digital PROs, expands access esp. rural/underrepresented.
- Real-world data / external controls: FDA final guidance Aug 2023 on RWE for regulatory decision-making; Flatiron-derived external controls in rare cancers (e.g., MSI-H, NRG1-fusion).
- Time-to-next-treatment (TTNT), quality-adjusted TTF: gaining traction as endpoints when crossover confounds OS.
- Platform / master protocols mature: I-SPY 2.2 sequential agents, ComboMATCH (NCI), Lung-MAP, TAPUR (ASCO basket study of approved targeted agents, cohort-level results).
High-yield trials pearls
- Phase 1 = safety, MTD/RP2D (DLT-driven; 3+3, CRM, BOIN).
- Phase 2 = efficacy signal, ORR (Simon 2-stage, single-arm vs historical control).
- Phase 3 = pivotal, OS/PFS, randomized.
- HR <1 = experimental better; CI crossing 1.0 = not significant.
- P value = evidence against H0, not the probability H0 is true; report CIs.
- Power = 1 − β; effect size dominates sample size.
- ITT analysis is standard for the primary endpoint.
- Surrogate endpoint valid only if it reliably captures the treatment effect on the clinical outcome (trial-level validation); causal-pathway placement or patient-level correlation alone is not enough.
- Kaplan-Meier curves + log-rank test; watch for immortal time bias (use landmark) and competing risks (use cumulative incidence).
- Non-inferiority needs a pre-specified margin; accelerated approval requires a confirmatory trial.
- Project Optimus shifts toward optimal dose vs MTD.
- Master protocols: basket (one drug, many cancers, shared biomarker), umbrella (one cancer, many drugs), platform (arms added/dropped).
- Prognostic (outcome regardless of treatment) vs predictive (benefit from a specific therapy) biomarkers.
Veli Bakalov MD, Board Review Notes 2026