Repeated-measures ANOVA
Determines the sample size required to detect a specified within-subject change, between-group difference, or group-by-time interaction when measurements are repeatedly obtained from the same subjects.
Key takeaways
- What is calculated: Find the smallest number of subjects at which the specified effect attains the target power.
- Quantities you must supply: Number of groups; Effect size; Effect tested; Nonsphericity correction; Measurements per subject; Correlation among repeated measures; Significance level; Power.
- Type of calculation: Power-based: it answers how much data is needed to detect a stated effect. Power 0.80 at α = 0.05 is the usual target; confirmatory studies use 0.90.
- Calculation routes: 2 are available for this design; the one selected determines the formula, the quantities requested and the result.
Methods available for Repeated-measures ANOVA
- Omnibus F test (noncentral F, Cohen f) Sample size for within-subject, interaction or between-subject effects when each subject is measured several times. Treating repeated observations as independent would badly overstate the information available and underpower the study.
- Two-group comparison of means across m measurements Sample size for a two-group mean comparison where each subject is measured m times and the measurements are correlated. Averaging correlated measurements reduces the variance of each subject's value, so the same difference is detectable with fewer subjects. Ignoring the correlation and treating m measurements as m independent observations would badly overstate the information available.
When to use
Use when the primary objective is to determine whether the outcome changes over time, differs between groups, or shows a group-by-time interaction based on repeated measurements from the same participants.
Sample size
Awaiting inputs
—
Set the parameters and calculate.
Methodology — Repeated-measures ANOVA — noncentral F with sphericity correction
Formula used Repeated-measures ANOVA — noncentral F with sphericity correction
- Two-sided alternative selected: α is split between both tails, so the larger quantile is used and the requirement rises
where —
- g
- Between-subject groups. 1 – 20
- f
- Cohen's f for the effect tested. 0.10 small, 0.25 medium, 0.40 large
- ε
- Sphericity correction on the degrees of freedom. 0.2 – 1.0; plan at 0.6 – 0.8 when unknown
- m
- Repeated measurements per subject. 2 – 50
- ρ
- Average correlation between repeated measurements. 0 – 0.99; typically 0.4 – 0.7
- α
- Probability of rejecting a true null hypothesis (Type I error). 0.0001 – 0.5; conventionally 0.05, or 0.025 one-sided for regulatory non-inferiority
- 1 − β
- Probability of rejecting the null hypothesis when the specified alternative is true. 0.50 – 0.9999; conventionally 0.80 or 0.90
- ℓ
- Expected proportion of enrolled units lost before analysis. 0 – 0.95
Formulas for the 2 methods
| Calculation method | Formula |
|---|---|
| Omnibus F test (noncentral F, Cohen f) | |
| Two-group comparison of means across m measurements |
How this method works
The noncentrality parameter depends on which effect is tested. Correlation among repeated measures reduces the variance of within-subject contrasts but inflates it for between-subject comparisons, so the two forms differ.
The non-centrality depends on which effect is tested, and the contrast between the two forms below is the practical message. Correlation among repeated measures reduces the variance of within-subject contrasts but inflates it for between-subject comparisons, because m correlated observations carry less information than m independent ones.
Sphericity is rarely exact beyond two occasions; planning at ε = 1 when the analysis will apply a correction produces an underpowered study.
Null hypothesis. H0: no within-subject effect, no interaction, or no between-subject difference, according to the effect selected.
Alternative hypothesis. H1: the specified effect is non-zero. The omnibus F is non-directional.
Calculation procedure
- Choose which effect is being tested: within-subject, the group-by-time interaction, or between-subject.
- State the groups , the measurements per subject , the effect size , and the correlation between repeated measurements.
- Form the noncentrality. Within-subject and interaction effects use ; the between-subject effect uses .
- Compute exact power from the noncentral F with degrees of freedom corrected by , and search for the smallest .
- Round up, then inflate for dropout. Planning at when the analysis will correct for non-sphericity leaves the study underpowered.
Exact or approximate
Exact given the covariance assumptions, which are themselves a simplification of a general covariance structure.
The formula shown always matches what was computed: when a one-sided test is selected the rendered quantile changes from z1−α/2 to z1−α, and where several methods exist, the formula follows the method selected in the calculator.
Advantages & limitations
Advantages
- Efficient: each subject serves as their own control.
- Distinguishes within, interaction and between effects.
- Accounts for nonsphericity explicitly.
Limitations
- Assumes compound symmetry, corrected only approximately by epsilon.
- Sensitive to the assumed correlation, which is often guessed.
- Does not model dropout across occasions.
Assumptions
- Subjects are independent of one another.
- The outcome is approximately normal.
- The correlation between repeated measurements is constant (compound symmetry), corrected by epsilon where it is not.
- Missing measurements are not informative.
Applicable adjustments
Supports dropout inflation, which should be specified at the subject level.
Conclusion
Report the effect tested, the assumed correlation and epsilon alongside the sample size. Planning at epsilon = 1 when the analysis will apply a correction produces an underpowered study.
References
- Guo Y, Logan HL, Glueck DH, Muller KE. Selecting a sample size for studies with repeated measures. BMC Med Res Methodol. 2013;13:100. Link
- Cohen J. Statistical Power Analysis for the Behavioral Sciences. 2nd ed. Hillsdale: Lawrence Erlbaum; 1988.
- Vonesh EF, Schork MA. Sample sizes in the multivariate analysis of repeated measurements. Biometrics. 1986;42(3):601–610. Link