Sample Size Calculator

Pearson correlation

Determines the sample size required to detect a specified population Pearson correlation against a null correlation with a given significance level and statistical power.

Key takeaways

  • What is calculated: Determine the smallest n at which the correlation can be distinguished from the null value with the target power.
  • Quantities you must supply: Expected correlation; Null correlation; Significance level; Power.
  • Type of calculation: Power-based: it answers how much data is needed to detect a stated effect. Power 0.80 at α = 0.05 is the usual target; confirmatory studies use 0.90.

Method used for Pearson correlation

  • Pearson correlation — Fisher's z transformation Estimates the sample size required to detect a specified correlation with a given significance level and statistical power. It can be applied to Pearson correlation, Spearman rank correlation (as an approximation), and partial correlation with appropriate adjustment for covariates.

When to use

Use when the primary objective is to determine whether a specified linear association between two continuous variables can be detected with adequate statistical power.

Sample size

Type I error rate.

Probability of detecting the specified effect.

A two-sided alternative splits α between both tails; a one-sided alternative places all of α in one tail, needs a smaller n, and must be pre-specified.

Final n is inflated by 1/(1 − dropout).

Reset to defaults

Awaiting inputs

Set the parameters and calculate.

Methodology — Pearson correlation — Fisher's z transformation

Formula used Pearson correlation — Fisher's z transformation

n=(z1α/2+z1βzrzr0)2+3
Required sample size
zr=12ln1+r1r
Fisher transformation
  • z1α/2(two-sided)in place ofz1α(one-sided) Two-sided alternative selected: α is split between both tails, so the larger quantile is used and the requirement rises

where —

r
Correlation to detect. Cohen: 0.10 small, 0.30 medium, 0.50 large
r₀
Correlation under the null; 0 for the conventional test. Must differ from r
α
Probability of rejecting a true null hypothesis (Type I error). 0.0001 – 0.5; conventionally 0.05, or 0.025 one-sided for regulatory non-inferiority
1 − β
Probability of rejecting the null hypothesis when the specified alternative is true. 0.50 – 0.9999; conventionally 0.80 or 0.90
Expected proportion of enrolled units lost before analysis. 0 – 0.95

How this method works

The sampling distribution of r is skewed and its variance depends on rho, so the test is conducted on Fisher's variance-stabilising transformation, which is approximately normal with variance 1/(n−3).

The sampling distribution of r is skewed and its variance depends on ρ, so the test is conducted on Fisher's variance-stabilising transformation, which is approximately normal with variance 1/(n − 3). The additive 3 restores the degrees of freedom lost.

r is highly sensitive to the sampled range: restricting it attenuates the correlation, so an estimate borrowed from a study with wider sampling overstates what is achievable.

Null hypothesis. H0: rho = rho0, usually zero.

Alternative hypothesis. Two-sided H1: rho != rho0. One-sided H1: rho > rho0 or rho < rho0.

Calculation procedure

  1. State the correlation r worth detecting, and the null value r0 (usually 0).
  2. Apply Fisher's transformation to both, zr=12ln1+r1r, which makes the sampling distribution approximately normal.
  3. The standard error of zr is 1/n3, which is why 3 is added back at the end.
  4. Evaluate n=(z1α/2+z1βzrzr0)2+3.
  5. Round up, then inflate for dropout. Weak correlations are expensive: r=0.2 needs roughly four times the sample of r=0.4.

Exact or approximate

Approximate — a normal approximation on the transformed scale, accurate for moderate n.

The formula shown always matches what was computed: when a one-sided test is selected the rendered quantile changes from z1−α/2 to z1−α, and where several methods exist, the formula follows the method selected in the calculator.

Advantages & limitations

Advantages

  • Closed form and widely tabulated.
  • Supports a non-zero null, which few calculators do.
  • Stable across the range of rho.

Limitations

  • Assumes bivariate normality.
  • r is highly sensitive to the sampled range: restricting it attenuates the correlation.
  • A correlation borrowed from a study with wider sampling will overstate what is achievable.

Assumptions

  • Pairs of observations are independent.
  • The two variables are jointly normally distributed.
  • The relationship is linear.
  • The sampled range of both variables matches that of the population of interest.

Applicable adjustments

Supports dropout inflation.

Conclusion

Report n with the assumed correlation and the null value tested. Confirm that the anticipated r comes from a population sampled over a comparable range.

References

  1. Fisher RA. On the "probable error" of a coefficient of correlation deduced from a small sample. Metron. 1921;1:3–32.
  2. Cohen J. Statistical Power Analysis for the Behavioral Sciences. 2nd ed. Hillsdale: Lawrence Erlbaum; 1988.
  3. Bonett DG, Wright TA. Sample size requirements for estimating Pearson, Kendall and Spearman correlations. Psychometrika. 2000;65(1):23–28. Link