One-way ANOVA
Determines the sample size required to detect a specified difference among the means of three or more independent groups using a one-way ANOVA framework.
Key takeaways
- What is calculated: Find the smallest per-group n at which the omnibus F test attains the target power.
- Quantities you must supply: Number of groups; Effect size; Significance level; Power.
- Type of calculation: Power-based: it answers how much data is needed to detect a stated effect. Power 0.80 at α = 0.05 is the usual target; confirmatory studies use 0.90.
Method used for One-way ANOVA
- One-way ANOVA — noncentral F with Cohen's f Omnibus F test comparing three or more group means, solved from the noncentral F distribution. A series of two-sample tests would inflate the type I error rate; the omnibus F controls it across all groups simultaneously.
When to use
Use when the primary objective is to determine whether the mean outcome differs among three or more independent groups.
Sample size
Awaiting inputs
—
Set the parameters and calculate.
Methodology — One-way ANOVA — noncentral F with Cohen's f
Formula used One-way ANOVA — noncentral F with Cohen's f
where —
- k
- Number of independent groups compared. 2 – 50
- f
- Cohen's f — SD of the group means over the within-group SD. 0.10 small, 0.25 medium, 0.40 large
- α
- Probability of rejecting a true null hypothesis (Type I error). 0.0001 – 0.5; conventionally 0.05, or 0.025 one-sided for regulatory non-inferiority
- 1 − β
- Probability of rejecting the null hypothesis when the specified alternative is true. 0.50 – 0.9999; conventionally 0.80 or 0.90
- ℓ
- Expected proportion of enrolled units lost before analysis. 0 – 0.95
How this method works
The F statistic follows a noncentral F distribution under the alternative, with noncentrality lambda = f²N on k−1 and N−k degrees of freedom. Cohen's f expresses the dispersion of the true group means relative to the within-group standard deviation.
The omnibus F test compares k group means at once. Under the alternative the statistic follows a non-central F distribution, evaluated here exactly as a Poisson mixture of incomplete beta functions rather than by the Patnaik approximation, which overstates power by one to two points.
A significant omnibus test establishes only that the means are not all equal. If specific contrasts matter, power those separately and handle multiplicity.
Null hypothesis. H0: mu1 = mu2 = ... = muk
Alternative hypothesis. H1: at least one group mean differs. The test is inherently non-directional, so no one-sided option applies.
Calculation procedure
- State the number of groups and the effect size — Cohen's benchmarks are 0.10 small, 0.25 medium, 0.40 large.
- For a candidate per group, the total is and the noncentrality is .
- The F statistic has and degrees of freedom; compute exact power from the noncentral F distribution.
- Search for the smallest whose power reaches the target, then multiply by for the total.
- Round up, then inflate for dropout. Any planned pairwise comparison must be powered separately at a corrected .
Exact or approximate
Exact. The noncentral F is evaluated as a Poisson mixture of incomplete beta functions, not by the Patnaik central-F approximation, which overstates power by one to two points.
The formula shown always matches what was computed: when a one-sided test is selected the rendered quantile changes from z1−α/2 to z1−α, and where several methods exist, the formula follows the method selected in the calculator.
Advantages & limitations
Advantages
- Controls the type I error across all groups at once.
- Exact rather than approximate.
- Agrees with G*Power.
Limitations
- Says only that the means are not all equal, not which differ.
- Assumes equal variances and balanced groups.
- Cohen's f is hard to elicit directly from investigators.
Assumptions
- Observations are independent within and between groups.
- The outcome is approximately normal in each group.
- Variances are equal across groups (homoscedasticity).
- Groups are of equal size (balanced design).
Applicable adjustments
Supports dropout inflation. Clustering requires a design-effect adjustment applied to the per-group figure.
Conclusion
Report the per-group and total sample sizes together with k and the assumed f. If pairwise comparisons are planned, power them separately and state the multiplicity adjustment.
References
- Cohen J. Statistical Power Analysis for the Behavioral Sciences. 2nd ed. Hillsdale: Lawrence Erlbaum; 1988.
- Faul F, Erdfelder E, Lang A-G, Buchner A. G*Power 3: A flexible statistical power analysis program for the social, behavioral, and biomedical sciences. Behav Res Methods. 2007;39(2):175–191. Link
- Chow S-C, Shao J, Wang H, Lokhnygina Y. Sample Size Calculations in Clinical Research. 3rd ed. Boca Raton: Chapman & Hall/CRC; 2017.