Paired means
Determines the sample size required to detect a specified mean difference between paired or matched observations, such as before-after measurements on the same subjects.
Key takeaways
- What is calculated: Determine the number of pairs at which the paired test attains the specified power against the stated mean difference.
- Quantities you must supply: Variability specified as; Expected mean difference; SD at time 1; SD at time 2; Correlation between pairs; Significance level; Power.
- Type of calculation: Power-based: it answers how much data is needed to detect a stated effect. Power 0.80 at α = 0.05 is the usual target; confirmatory studies use 0.90.
Method used for Paired means
- Paired means — one-sample test on within-pair differences Sample size for a paired design, driven by the standard deviation of the differences. Chosen because pairing removes between-subject variability from the comparison. At ρ = 0.5 with equal measurement SDs, σ_d equals σ, so the paired design needs roughly half the observations of the independent-groups design.
When to use
Use when the primary objective is to determine whether there is a meaningful mean change between paired or matched observations, such as measurements taken before and after an intervention.
Sample size
Awaiting inputs
—
Set the parameters and calculate.
Methodology — Paired means — one-sample test on within-pair differences
Formula used Paired means — one-sample test on within-pair differences
- Two-sided alternative selected: α is split between both tails, so the larger quantile is used and the requirement rises
where —
- μ_d
- Expected mean of the within-pair differences. Non-zero
- σ₁
- SD at the first time point or condition. ≥ 0
- σ₂
- SD at the second time point or condition. ≥ 0
- ρ
- Correlation between paired measurements. −0.99 – 0.99; typically 0.4 – 0.8
- α
- Probability of rejecting a true null hypothesis (Type I error). 0.0001 – 0.5; conventionally 0.05, or 0.025 one-sided for regulatory non-inferiority
- 1 − β
- Probability of rejecting the null hypothesis when the specified alternative is true. 0.50 – 0.9999; conventionally 0.80 or 0.90
- ℓ
- Expected proportion of enrolled units lost before analysis. 0 – 0.95
How this method works
A paired design reduces to a one-sample problem on the differences D = X₁ − X₂. The relevant variability is σ_d, the standard deviation of those differences, which is smaller than either measurement SD whenever the pair members are positively correlated.
Each unit contributes two measurements, so the analysis reduces to a one-sample test on within-pair differences. The correlation ρ is what makes the design efficient: at ρ = 0.5 with equal variances it needs roughly half the observations of an independent-groups design.
Supplying σ_d directly is preferable when pilot data on differences exist, since it avoids compounding three estimates.
Null hypothesis. H₀: μ_d = 0, equivalently μ₁ = μ₂
Alternative hypothesis. Two-sided H₁: μ_d ≠ 0. One-sided H₁: μ_d > 0 or μ_d < 0.
Calculation procedure
- State the mean difference worth detecting — the average within-pair change, not the difference between two groups.
- Supply the standard deviation of the differences directly, or derive it from .
- Take and , and form the standardised effect .
- Evaluate — the number of pairs, not of measurements.
- Round up, then inflate for dropout. A higher correlation shrinks and so shrinks .
Exact or approximate
Both are reported: the normal approximation in closed form, and the exact noncentral t solution, which is the primary result.
The formula shown always matches what was computed: when a one-sided test is selected the rendered quantile changes from z1−α/2 to z1−α, and where several methods exist, the formula follows the method selected in the calculator.
Advantages & limitations
Advantages
- Substantially more efficient than an independent-groups design when pairing is effective.
- Controls for all stable between-subject characteristics by design.
Limitations
- Requires σ_d or a defensible correlation, and ρ is the parameter investigators most often guess badly.
- Vulnerable to carry-over and period effects in before-and-after designs.
- Loss of one pair member removes the whole pair from the analysis.
Assumptions
- Pairs are independent of one another, though members within a pair are not.
- The within-pair differences are approximately normally distributed.
- σ_d is correctly specified, either directly or derived from σ₁, σ₂ and ρ.
- Pairing is determined by design, not chosen after seeing the outcome.
Applicable adjustments
Supports dropout inflation, which should be specified at the pair level since losing either member removes the pair.
Conclusion
Report the number of pairs, not the number of measurements, and state the assumed correlation. The efficiency of the design rests entirely on that correlation being real.
References
- Chow S-C, Shao J, Wang H, Lokhnygina Y. Sample Size Calculations in Clinical Research. 3rd ed. Boca Raton: Chapman & Hall/CRC; 2017.
- Guenther WC. Sample size formulas for normal theory t tests. The American Statistician. 1981;35(4):243–244. Link
- Rosner B. Fundamentals of Biostatistics. 8th ed. Boston: Cengage; 2015.