Sample Size Calculator

Equivalence trial

Determines the sample size required to demonstrate that the difference between two treatments lies within a prespecified equivalence margin, using the Two One-Sided Tests (TOST) procedure.

Key takeaways

  • What is calculated: Find the smallest n per arm at which both one-sided tests reject with the target overall power.
  • Quantities you must supply: True difference; Equivalence margin; Standard deviation; One-sided significance level; Power; Allocation ratio.
  • Type of calculation: Power-based: it answers how much data is needed to detect a stated effect. Power 0.80 at α = 0.05 is the usual target; confirmatory studies use 0.90.

Method used for Equivalence trial

  • Equivalence trial — two one-sided tests (TOST) Sizes two one-sided tests that must both succeed, so that the difference between treatments is shown to lie wholly inside a stated margin in either direction. It is the requirement where being unacceptably better matters as much as being worse — generic and biosimilar comparisons, alternative formulations, batch comparability. It needs more participants than a non-inferiority trial at the same margin because both bounds have to be excluded, and the requirement climbs sharply as the difference genuinely expected approaches the margin, which is where optimistic planning most often fails.

When to use

Use when the primary objective is to determine whether the difference between two treatments is sufficiently small to fall within prespecified clinically acceptable equivalence limits.

Sample size

Usually 0. The requirement rises steeply as this approaches the margin.

Each of the two one-sided tests is carried out at this level.

Probability of detecting the specified effect.

Set to 1 for equal group sizes.

Final n is inflated by 1/(1 − dropout).

Reset to defaults

Awaiting inputs

Set the parameters and calculate.

Methodology — Equivalence trial — two one-sided tests (TOST)

Formula used Equivalence trial — two one-sided tests (TOST)

n=(z1α+z1β/2)2σ2(1+1/k)(δ|Δ|)2
TOST requirement per arm

where —

Δ
True difference expected, usually 0. Strictly within ± δ
δ
Equivalence margin defining practical indistinguishability. > 0
σ
Standard deviation of the outcome, assumed equal in both arms. > 0
α
Probability of rejecting a true null hypothesis (Type I error). 0.0001 – 0.5; conventionally 0.05, or 0.025 one-sided for regulatory non-inferiority
1 − β
Probability of rejecting the null hypothesis when the specified alternative is true. 0.50 – 0.9999; conventionally 0.80 or 0.90
k
Allocation ratio n₂/n₁. 0.05 – 20; 1 gives equal groups
Expected proportion of enrolled units lost before analysis. 0 – 0.95

How this method works

Equivalence requires both bounds to be excluded, so the calculation uses a one-sided alpha with z_{1−β/2} rather than z_{1−β}: each of the two one-sided tests must succeed, and the power is apportioned between them. The effective effect size is δ − |Δ|, which shrinks sharply as the assumed true difference moves away from zero.

Two one-sided tests: the difference must exceed −δ and fall below +δ, so both must reject and the power term uses z₁₋β/₂.

Acutely sensitive to the assumed true difference — if Δ turns out to be δ/2, the requirement quadruples. Run a sensitivity analysis across plausible values. Failing to reject does not demonstrate a difference.

Null hypothesis. H0: |mu_T − mu_C| ≥ δ (the treatments differ by at least the margin in one direction or the other).

Alternative hypothesis. H1: −δ < mu_T − mu_C < δ. Sidedness is fixed by the design: two one-sided tests, each at α.

Calculation procedure

  1. Set the symmetric margin ±δ inside which the two treatments count as interchangeable.
  2. State the true difference Δ expected, usually 0, and the standard deviation σ. The effective margin is δ|Δ|.
  3. Both one-sided tests must reject, so the power quantile is z1β/2, not z1β.
  4. Evaluate n=(z1α+z1β/2)2σ2(1+1/k)(δ|Δ|)2 per arm.
  5. Round each arm up, then inflate for dropout. As |Δ| approaches δ the requirement diverges.

Exact or approximate

Normal approximation using the conservative power apportionment. Exact TOST power is marginally higher, so the reported n is slightly conservative.

The formula shown always matches what was computed: when a one-sided test is selected the rendered quantile changes from z1−α/2 to z1−α, and where several methods exist, the formula follows the method selected in the calculator.

Advantages & limitations

Advantages

  • Directly answers the interchangeability question in both directions.
  • Maps exactly onto the analysis: equivalence holds if the whole confidence interval lies within ± δ.
  • The type I error rate is controlled at α without adjustment despite the two tests, because both must reject.

Limitations

  • Demands appreciably more subjects than non-inferiority at the same margin.
  • Extremely sensitive to the assumed true difference: as |Δ| approaches δ, the requirement diverges.
  • The symmetric margin is often not clinically symmetric.
  • Assumes equal variances, which is questionable when the formulations genuinely differ.

Assumptions

  • Randomised allocation with independent subjects (or an appropriate crossover analysis for bioequivalence).
  • The equivalence margin δ is symmetric and clinically justified.
  • The assumed true difference lies inside the margin; if it does not, no sample size suffices.
  • Equal variance in both arms.

Applicable adjustments

Supports unequal allocation and dropout inflation. Crossover designs use the within-subject standard deviation, which is smaller and gives a lower n.

Conclusion

Report the margin, the assumed true difference and the power apportionment. Powering at |Δ| = 0 when a difference is genuinely expected is the most common reason equivalence trials fail.

References

  1. Schuirmann DJ. A comparison of the two one-sided tests procedure and the power approach for assessing the equivalence of average bioavailability. J Pharmacokinet Biopharm. 1987;15(6):657–680. Link
  2. Chow S-C, Liu J-P. Design and Analysis of Bioavailability and Bioequivalence Studies. 3rd ed. Boca Raton: Chapman & Hall/CRC; 2008.
  3. Walker E, Nowacki AS. Understanding equivalence and noninferiority testing. J Gen Intern Med. 2011;26(2):192–196. Link