Sample Size Calculator

Superiority trial

Determines the sample size required to demonstrate that one treatment is statistically superior to another treatment or control by a clinically relevant amount, with specified significance level and power.

Key takeaways

  • What is calculated: Find the smallest n per arm at which a two-group test of the specified difference attains the target power.
  • Quantities you must supply: Expected difference to detect; Standard deviation; Clinical superiority threshold; Significance level; Power; Allocation ratio.
  • Type of calculation: Power-based: it answers how much data is needed to detect a stated effect. Power 0.80 at α = 0.05 is the usual target; confirmatory studies use 0.90.

Method used for Superiority trial

  • Superiority trial — normal-theory two-group comparison Sizes the conventional randomised comparison: enough participants per arm that a difference of stated size would be detected, with a threshold available when the difference must also clear a clinically meaningful margin rather than merely differ from zero. It is the design for a trial whose claim is that the new treatment is better, and it carries the two-sided error rate regulators expect. Its limitation is worth knowing before the data arrive: a result that falls short of significance establishes nothing about the treatments being alike, so if comparability is the claim that may eventually be made, one of the other two designs has to be chosen at the outset.

When to use

Use when the primary objective is to determine whether a new treatment provides a clinically meaningful improvement over a control or standard treatment.

Sample size

In outcome units. The difference must exceed this margin, not merely differ from zero.

Type I error rate.

Probability of detecting the specified effect.

A two-sided alternative splits α between both tails; a one-sided alternative places all of α in one tail, needs a smaller n, and must be pre-specified.

Set to 1 for equal group sizes.

Final n is inflated by 1/(1 − dropout).

Reset to defaults

Awaiting inputs

Set the parameters and calculate.

Methodology — Superiority trial — normal-theory two-group comparison

Formula used Superiority trial — normal-theory two-group comparison

n=(z1α/2+z1β)2σ2(1+1/k)(ΔΔ0)2
Sample size per arm
  • z1α/2(two-sided)in place ofz1α(one-sided) Two-sided alternative selected: α is split between both tails, so the larger quantile is used and the requirement rises

where —

Δ
Difference to detect — the smallest clinically worthwhile effect. Must exceed the threshold
σ
Standard deviation of the outcome, assumed equal in both arms. > 0
Δ₀
Clinical threshold the difference must exceed; 0 for a conventional test. ≥ 0
α
Probability of rejecting a true null hypothesis (Type I error). 0.0001 – 0.5; conventionally 0.05, or 0.025 one-sided for regulatory non-inferiority
1 − β
Probability of rejecting the null hypothesis when the specified alternative is true. 0.50 – 0.9999; conventionally 0.80 or 0.90
k
Allocation ratio n₂/n₁. 0.05 – 20; 1 gives equal groups
Expected proportion of enrolled units lost before analysis. 0 – 0.95

How this method works

The standard two-group comparison written in trial terms: the required n per arm is (z_{1−α/2} + z_{1−β})² times the variance sum, divided by the square of the difference to be detected. Setting a non-zero superiority threshold subtracts it from the difference, which is the clinical superiority variant of the design.

The conventional two-group test, with an optional threshold so the trial is powered to show the difference exceeds a stated value rather than merely differing from zero — useful in large trials where a trivial difference is easy to make significant.

Enter the smallest clinically worthwhile effect. Powering on an optimistic effect size is the commonest reason trials of useful interventions return null results.

Null hypothesis. H0: mu_T − mu_C = 0 (or p_T − p_C = 0). With a clinical threshold, H0: the difference does not exceed the threshold.

Alternative hypothesis. H1: the difference is non-zero, or exceeds the threshold. Two-sided is the regulatory default; a one-sided test must be pre-specified and justified.

Calculation procedure

  1. State the difference Δ worth detecting and the standard deviation σ of the outcome.
  2. Set a clinical threshold Δ0 if the difference must exceed a margin rather than merely differ from zero; otherwise Δ0=0.
  3. Set the allocation ratio k and take z1α/2 and z1β.
  4. Evaluate n=(z1α/2+z1β)2σ2(1+1/k)(ΔΔ0)2 per arm.
  5. Round each arm up, then inflate for dropout. A non-significant result does not show the treatments are equivalent.

Exact or approximate

Normal approximation. At small n the true t-based requirement is one or two subjects higher per arm; at trial-scale sample sizes the difference is negligible.

The formula shown always matches what was computed: when a one-sided test is selected the rendered quantile changes from z1−α/2 to z1−α, and where several methods exist, the formula follows the method selected in the calculator.

Advantages & limitations

Advantages

  • The interpretation is unambiguous: a significant result establishes a difference in a stated direction.
  • Handles unequal allocation directly through the ratio.
  • Accommodates a clinical superiority margin without changing the analysis model.

Limitations

  • A non-significant result does not establish equivalence — absence of evidence is not evidence of absence.
  • Highly sensitive to the assumed difference: halving it roughly quadruples the required n.
  • The normal approximation understates the requirement slightly at small n.
  • Does not account for interim analyses, which require an alpha-spending adjustment.

Assumptions

  • Randomised, parallel-group allocation with independent subjects.
  • The outcome is approximately normal (continuous) or the normal approximation to the binomial is adequate (binary).
  • Equal variance in both arms for the continuous endpoint.
  • The difference specified is the smallest one worth detecting, not the one hoped for.

Applicable adjustments

Supports unequal allocation and dropout inflation. Cluster randomisation requires a design effect; interim analyses require an alpha-spending correction.

Conclusion

Report the difference, the variance assumption, alpha, sidedness and power together with n. State the source of the assumed effect size — pilot data, published trials or clinical judgement.

References

  1. Lachin JM. Introduction to sample size determination and power analysis for clinical trials. Control Clin Trials. 1981;2(2):93–113. Link
  2. Schulz KF, Altman DG, Moher D. CONSORT 2010 Statement: updated guidelines for reporting parallel group randomised trials. BMJ. 2010;340:c332. Link
  3. Chow S-C, Shao J, Wang H, Lokhnygina Y. Sample Size Calculations in Clinical Research. 3rd ed. Boca Raton: Chapman & Hall/CRC; 2017.