Pearson correlation
Determines the sample size required to detect a specified population Pearson correlation against a null correlation with a given significance level and statistical power.
Key takeaways
- What is calculated: Determine the smallest n at which the correlation can be distinguished from the null value with the target power.
- Quantities you must supply: Expected correlation; Null correlation; Significance level; Power.
- Type of calculation: Power-based: it answers how much data is needed to detect a stated effect. Power 0.80 at α = 0.05 is the usual target; confirmatory studies use 0.90.
Method used for Pearson correlation
- Pearson correlation — Fisher's z transformation Estimates the sample size required to detect a specified correlation with a given significance level and statistical power. It can be applied to Pearson correlation, Spearman rank correlation (as an approximation), and partial correlation with appropriate adjustment for covariates.
When to use
Use when the primary objective is to determine whether a specified linear association between two continuous variables can be detected with adequate statistical power.
Sample size
Awaiting inputs
—
Set the parameters and calculate.
Methodology — Pearson correlation — Fisher's z transformation
Formula used Pearson correlation — Fisher's z transformation
- Two-sided alternative selected: α is split between both tails, so the larger quantile is used and the requirement rises
where —
- r
- Correlation to detect. Cohen: 0.10 small, 0.30 medium, 0.50 large
- r₀
- Correlation under the null; 0 for the conventional test. Must differ from r
- α
- Probability of rejecting a true null hypothesis (Type I error). 0.0001 – 0.5; conventionally 0.05, or 0.025 one-sided for regulatory non-inferiority
- 1 − β
- Probability of rejecting the null hypothesis when the specified alternative is true. 0.50 – 0.9999; conventionally 0.80 or 0.90
- ℓ
- Expected proportion of enrolled units lost before analysis. 0 – 0.95
How this method works
The sampling distribution of r is skewed and its variance depends on rho, so the test is conducted on Fisher's variance-stabilising transformation, which is approximately normal with variance 1/(n−3).
The sampling distribution of r is skewed and its variance depends on ρ, so the test is conducted on Fisher's variance-stabilising transformation, which is approximately normal with variance 1/(n − 3). The additive 3 restores the degrees of freedom lost.
r is highly sensitive to the sampled range: restricting it attenuates the correlation, so an estimate borrowed from a study with wider sampling overstates what is achievable.
Null hypothesis. H0: rho = rho0, usually zero.
Alternative hypothesis. Two-sided H1: rho != rho0. One-sided H1: rho > rho0 or rho < rho0.
Calculation procedure
- State the correlation worth detecting, and the null value (usually 0).
- Apply Fisher's transformation to both, , which makes the sampling distribution approximately normal.
- The standard error of is , which is why 3 is added back at the end.
- Evaluate .
- Round up, then inflate for dropout. Weak correlations are expensive: needs roughly four times the sample of .
Exact or approximate
Approximate — a normal approximation on the transformed scale, accurate for moderate n.
The formula shown always matches what was computed: when a one-sided test is selected the rendered quantile changes from z1−α/2 to z1−α, and where several methods exist, the formula follows the method selected in the calculator.
Advantages & limitations
Advantages
- Closed form and widely tabulated.
- Supports a non-zero null, which few calculators do.
- Stable across the range of rho.
Limitations
- Assumes bivariate normality.
- r is highly sensitive to the sampled range: restricting it attenuates the correlation.
- A correlation borrowed from a study with wider sampling will overstate what is achievable.
Assumptions
- Pairs of observations are independent.
- The two variables are jointly normally distributed.
- The relationship is linear.
- The sampled range of both variables matches that of the population of interest.
Applicable adjustments
Supports dropout inflation.
Conclusion
Report n with the assumed correlation and the null value tested. Confirm that the anticipated r comes from a population sampled over a comparable range.
References
- Fisher RA. On the "probable error" of a coefficient of correlation deduced from a small sample. Metron. 1921;1:3–32.
- Cohen J. Statistical Power Analysis for the Behavioral Sciences. 2nd ed. Hillsdale: Lawrence Erlbaum; 1988.
- Bonett DG, Wright TA. Sample size requirements for estimating Pearson, Kendall and Spearman correlations. Psychometrika. 2000;65(1):23–28. Link