ROC area under the curve
Determines the sample size required to estimate or detect a specified area under the ROC curve with a desired precision, significance level, or statistical power.
Key takeaways
- What is calculated: Determine the diseased and non-diseased subjects required to detect the specified AUC.
- Quantities you must supply: Area under the ROC curve; Null hypothesis value; Ratio of sample sizes in negative / positive groups; Significance level; Power.
- Type of calculation: Power-based: it answers how much data is needed to detect a stated effect. Power 0.80 at α = 0.05 is the usual target; confirmatory studies use 0.90.
- Calculation routes: 4 are available for this design; the one selected determines the formula, the quantities requested and the result.
Methods available for ROC area under the curve
- Hypothesis test — Hanley & McNeil (1982) variance Searches for the number of positive cases at which the area under the ROC curve can be separated from a stated value, using the Hanley & McNeil standard error, which keeps the group sizes inside it and so has to be solved by search rather than in closed form. This is the form MedCalc implements, so the figure is the one a reviewer familiar with that tool will expect. The ratio of negative to positive cases is the design lever: where cases are scarce, recruiting more controls per case recovers useful precision, and stating the case mix as prevalence instead is the same design written the other way round.
- Hypothesis test — Obuchowski variance function Uses a closed-form variance function derived from the binormal model, so the number of positive cases follows directly from the two AUC values and the case mix without a search. It is the calculation behind most published ROC sample sizes and the widely used web calculators, which makes it the recognisable figure to quote. It returns fewer subjects than the Hanley & McNeil route at the same inputs — both are defensible and published, so whichever is used should be named in the protocol rather than left to be inferred from the number.
- Estimation to a stated precision Fixes the width of the confidence interval the study will report around its AUC, with no null value and no power term involved. It matches a validation study whose purpose is to characterise a test rather than to prove it beats a threshold. An AUC of 0.82 reported with an interval from 0.71 to 0.93 is a weak result however good the point estimate looks, and planning on precision is what prevents that; if a claim of superiority over chance or over an existing test is also intended, that has to be powered separately.
- Comparison of two AUCs on the same subjects Sizes a comparison of two tests applied to every subject, using the variance of the difference between their AUCs, which the correlation between them reduces. It is the usual diagnostic-accuracy design, because each subject supplies both results and one reference standard. That correlation is the whole efficiency of the arrangement: treating the two tests as though they came from separate samples inflates the variance and overstates the requirement, often by a wide margin, so the correlation should be taken from pilot data and the requirement shown across a plausible range of it.
When to use
Use when the primary objective is to determine how well a diagnostic test or biomarker discriminates between individuals with and without the disease across different decision thresholds.
Sample size
Awaiting inputs
—
Set the parameters and calculate.
Methodology — Area under the ROC curve — Hanley & McNeil (1982) variance
Formula used Area under the ROC curve — Hanley & McNeil (1982) variance
- Two-sided alternative selected: α is split between both tails, so the larger quantile is used and the requirement rises
where —
- A₁
- Expected area under the ROC curve. 0.7 – 0.8 acceptable, 0.8 – 0.9 excellent
- A₀
- AUC under the null; 0.5 means no discrimination. Must differ from A₁
- κ
- Ratio of sample sizes in the negative and positive groups. 0.05 – 50; gains flatten beyond 4
- α
- Probability of rejecting a true null hypothesis (Type I error). 0.0001 – 0.5; conventionally 0.05, or 0.025 one-sided for regulatory non-inferiority
- 1 − β
- Probability of rejecting the null hypothesis when the specified alternative is true. 0.50 – 0.9999; conventionally 0.80 or 0.90
- ℓ
- Expected proportion of enrolled units lost before analysis. 0 – 0.95
Formulas for the 4 methods
| Calculation method | Formula |
|---|---|
| Hypothesis test — Hanley & McNeil (1982) variance | |
| Hypothesis test — Obuchowski variance function | |
| Estimation to a stated precision | |
| Comparison of two AUCs on the same subjects |
How this method works
The area under the ROC curve is the probability that a random diseased subject scores higher than a random non-diseased one. Hanley and McNeil derived its variance under a binormal assumption; because the variance differs under the null and alternative, both appear.
Selected method - Hanley & McNeil (1982), the form MedCalc implements. Q1 and Q2 come from their exponential approximation to the binormal ROC curve, and the standard error keeps the group sizes inside it rather than factoring them out.
Because n appears on both sides, the requirement is found by searching for the smallest number of positive cases at which z_(1-alpha/2) SE(A0) + z_(1-beta) SE(A1) no longer exceeds the AUC difference. That is why this method reports standard errors rather than a variance.
The positive group carries the information; the negative group is however many the ratio kappa asks for. Specifying prevalence instead of a ratio is the same design stated the other way round, since kappa = (1 - p)/p: at a prevalence of 0.10, nine negative cases are screened for every positive one.
Null hypothesis. H0: AUC = A0, usually 0.5.
Alternative hypothesis. Two-sided H1: AUC != A0. One-sided H1: AUC > A0.
Calculation procedure
- State the AUC expected, , and the null value — usually 0.5, meaning no discrimination.
- State the case mix as the ratio of negative to positive cases, , or equivalently as the prevalence, .
- For a candidate number of positives, compute the Hanley & McNeil standard error at and at from and .
- Search for the smallest satisfying ; the negatives are .
- Round up, then inflate for indeterminate results. The search is needed because appears inside the standard error.
Exact or approximate
Approximate — an asymptotic result under a binormal assumption.
The formula shown always matches what was computed: when a one-sided test is selected the rendered quantile changes from z1−α/2 to z1−α, and where several methods exist, the formula follows the method selected in the calculator.
Advantages & limitations
Advantages
- Uses the whole score range rather than one threshold.
- Supports a non-zero benchmark null.
- The case-to-control ratio can be tuned when cases are scarce.
Limitations
- The binormal assumption may not hold for skewed scores.
- Gains from extra controls flatten beyond about four per case.
- AUC alone does not identify a usable threshold.
Assumptions
- Test scores follow a binormal distribution in the two groups.
- Disease status is determined by a reference standard applied to all subjects.
- Observations are independent.
- The ratio of non-diseased to diseased subjects is fixed by design.
Applicable adjustments
Supports dropout inflation.
Conclusion
Report both group sizes with the anticipated AUC and the ratio used. If a clinical threshold is the real objective, power sensitivity and specificity at that threshold instead.
References
- MedCalc Software Ltd. Sample size calculation: area under ROC curve. Link
- Zhou XH, Obuchowski NA, McClish DK. Statistical methods in diagnostic medicine. 2nd ed. Hoboken: Wiley; 2011. Equations 6.2, 6.6, 6.8. Link
- Hanley JA, McNeil BJ. The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology. 1982;143(1):29–36. Link
- Hanley JA, McNeil BJ. A method of comparing the areas under receiver operating characteristic curves derived from the same cases. Radiology. 1983;148(3):839–843. Link
- Obuchowski NA, McClish DK. Sample size determination for diagnostic accuracy studies involving binormal ROC curve indices. Stat Med. 1997;16(13):1529–1542. Link
- Arifin WN. Sample size calculator (web) — Area under ROC curve. Link