Sample Size Calculator

Survival non-inferiority — Schoenfeld

Determines the required number of events and sample size to establish that the hazard ratio satisfies a prespecified non-inferiority criterion.

Key takeaways

  • What is calculated: Find the number of events at which the one-sided test excludes a hazard ratio worse than the margin with the target power, then the enrolment needed to accrue them.
  • Quantities you must supply: True; Non-inferiority margin on the HR scale; Overall probability of an event; Significance level; Power; Allocation ratio.
  • Type of calculation: Power-based: it answers how much data is needed to detect a stated effect. Power 0.80 at α = 0.05 is the usual target; confirmatory studies use 0.90.

Method used for Survival non-inferiority — Schoenfeld

  • Time-to-event non-inferiority — one-sided log hazard ratio margin Number of events, and from it the sample size, to show a hazard ratio does not exceed a pre-specified non-inferiority margin. Survival power is governed by events, not enrolment, so a non-inferiority claim on survival must be specified and powered in events before it is translated into a recruitment target.

When to use

Use when the primary objective is to determine whether the time-to-event outcome under a new treatment is not unacceptably worse than that under a reference treatment.

Sample size

The largest HR still considered non-inferior.

Type I error rate.

Probability of detecting the specified effect.

Set to 1 for equal group sizes.

Final n is inflated by 1/(1 − dropout).

Reset to defaults

Awaiting inputs

Set the parameters and calculate.

Methodology — Time-to-event non-inferiority — one-sided log hazard ratio margin

Formula used Time-to-event non-inferiority — one-sided log hazard ratio margin

d=(z1α+z1β)2q1q2(lnHRmarginlnHRtrue)2
Required events, one-sided

where —

HR_t
Hazard ratio actually expected; 1 is optimistic. > 0
HR_m
Largest hazard ratio still regarded as non-inferior. Must differ from HR_true
P(e)
Overall probability of experiencing the event. 0 – 1
α
One-sided significance level; regulatory convention 0.025. Conventionally 0.025 one-sided
1 − β
Probability of rejecting the null hypothesis when the specified alternative is true. 0.50 – 0.9999; conventionally 0.80 or 0.90
k
Allocation ratio n₂/n₁. 0.05 – 20; 1 gives equal groups
Expected proportion of enrolled units lost before analysis. 0 – 0.95

How this method works

The Schoenfeld relationship is applied on the log hazard scale with the hypothesis shifted by the margin: the required number of events is (z_{1−α} + z_{1−β})² divided by the squared distance between the assumed log hazard ratio and the log margin, scaled by the allocation. Sample size follows by dividing the required events by the expected event probability.

The null is reversed: rejecting it supports non-inferiority. The effective effect size is the gap between margin and truth on the log scale, so assuming exact equality is optimistic — if the experimental arm is genuinely slightly worse, power falls sharply.

The margin must preserve a defined fraction of the active control's established benefit over placebo. One chosen for feasibility is not defensible.

Null hypothesis. H0: HR ≥ HR_margin (the test regimen is inferior by at least the margin).

Alternative hypothesis. H1: HR < HR_margin. One-sided by construction.

Calculation procedure

  1. Justify the non-inferiority margin on the hazard ratio scale, HRmargin — for example 1.30.
  2. State the hazard ratio you actually expect, usually 1, and the probability that an enrolled subject has the event.
  3. The test is one-sided; compute the events required as d=(z1α+z1β)2q1q2[lnHRlnHRmargin]2.
  4. Convert events into enrolment: N=d/Pr(event).
  5. Round up, then inflate for dropout. Check proportional hazards first; a delayed effect invalidates the calculation.

Exact or approximate

Asymptotic. The Schoenfeld relationship is accurate once the required events run into the dozens; it is unreliable for very small event counts.

The formula shown always matches what was computed: when a one-sided test is selected the rendered quantile changes from z1−α/2 to z1−α, and where several methods exist, the formula follows the method selected in the calculator.

Advantages & limitations

Advantages

  • Reports events and enrolment separately, which is what drives survival power.
  • Allows follow-up to be extended instead of recruitment when accrual is limited.
  • Consistent with the Cox model used in the analysis.

Limitations

  • Fails if the hazard ratio varies over time; a delayed treatment effect breaks the calculation entirely.
  • The requirement is highly sensitive to the margin on the log scale.
  • Ignores staggered entry and administrative censoring unless these are reflected in the event probability.
  • Loss to follow-up biases towards concluding non-inferiority.

Assumptions

  • Proportional hazards over the whole follow-up period.
  • The margin is justified against the historical effect of the active control.
  • Censoring is non-informative and similar in both arms.
  • The assumed true hazard ratio is realistic, usually 1 (no true difference).
  • Follow-up is long enough to accrue the required events.

Applicable adjustments

Supports dropout inflation. Non-proportional hazards require an average hazard ratio or an RMST-based design instead.

Conclusion

State the margin, the assumed hazard ratio, the event probability, the accrual and follow-up periods, and the one-sided alpha. Check proportional hazards on any pilot or historical data before relying on this design.

References

  1. Jung S-H, Chow S-C. On sample size calculation for comparing survival curves under general hypothesis testing. J Biopharm Stat. 2012;22(3):485–495. Link
  2. Com-Nougue C, Rodary C, Patte C. How to establish equivalence when data are censored: a randomized trial of treatments for B non-Hodgkin lymphoma. Stat Med. 1993;12(14):1353–1364. Link
  3. European Medicines Agency. Guideline on the Choice of the Non-Inferiority Margin. EMEA/CPMP/EWP/2158/99; 2005. Link