Sample Size Calculator

Specificity

Determines the sample size required to estimate the specificity of a diagnostic test with a specified precision and confidence level, based on the expected specificity.

Key takeaways

  • What is calculated: Determine the non-diseased subjects required, and hence the total enrolment.
  • Quantities you must supply: Objective; Expected specificity; Absolute precision; Minimum acceptable specificity — test mode; Disease prevalence in the sampled population; Significance level.
  • Type of calculation: Power-based: it answers how much data is needed to detect a stated effect. Power 0.80 at α = 0.05 is the usual target; confirmatory studies use 0.90.

Method used for Specificity

  • Specificity — Buderer prevalence-adjusted method Sizes the disease-free group — the number of confirmed negatives needed to pin down how often the test correctly clears them — and converts that into the total to screen using the prevalence. It is the mirror of the sensitivity calculation and is easier to satisfy when the condition is rare, because most of those screened contribute. Specificity usually needs tighter precision than sensitivity: across a large, mostly healthy population, a drop of one or two percentage points turns into a large number of false positives and the follow-up work that comes with them.

When to use

Use when the primary objective is to determine how accurately the study can estimate the probability that a diagnostic test correctly identifies individuals who do not have the disease.

Sample size

Type I error rate.

Final n is inflated by 1/(1 − dropout).

Reset to defaults

Awaiting inputs

Set the parameters and calculate.

Methodology — Specificity — Buderer prevalence-adjusted method

Formula used Specificity — Buderer prevalence-adjusted method

nwell=z1α/22Sp(1Sp)d2
Precision objective
N=nwell1φ
Total enrolment

where —

Sp
Expected specificity of the test. 0 – 1 exclusive
d
Absolute precision. Often 0.02 – 0.03 for screening tests
Sp₀
Minimum acceptable specificity, in test mode. Must differ from Sp
φ
Disease prevalence in the sampled population. 0 – 1 exclusive
α
Probability of rejecting a true null hypothesis (Type I error). 0.0001 – 0.5; conventionally 0.05, or 0.025 one-sided for regulatory non-inferiority
Expected proportion of enrolled units lost before analysis. 0 – 0.95

How this method works

Specificity is estimated among truly non-diseased subjects, so the subgroup formula applies there and total enrolment is obtained by dividing by one minus the prevalence.

The mirror image of the sensitivity calculation, applied among truly non-diseased subjects.

Non-diseased subjects are usually plentiful, so specificity is cheap to estimate — but it is typically required to tighter precision, because across a large healthy population 0.98 rather than 0.99 doubles the false-positive workload. Apply the reference standard to all subjects, or partial verification bias distorts both measures.

Null hypothesis. In test mode, H0: Sp = Sp0. In precision mode there is no hypothesis.

Alternative hypothesis. In test mode, H1: Sp > Sp0 (one-sided).

Calculation procedure

  1. State the specificity expected, Sp, and either the precision wanted or the minimum acceptable value.
  2. Only disease-free subjects inform specificity, so the calculation is for the number of true negatives nwell.
  3. Evaluate nwell=z1α/22Sp(1Sp)d2 for a precision target.
  4. Convert to the total to screen: N=nwell1φ.
  5. Round up, then inflate for indeterminate results. A low prevalence helps here, unlike for sensitivity.

Exact or approximate

Approximate — a normal approximation to the binomial within the non-diseased subgroup.

The formula shown always matches what was computed: when a one-sided test is selected the rendered quantile changes from z1−α/2 to z1−α, and where several methods exist, the formula follows the method selected in the calculator.

Advantages & limitations

Advantages

  • Non-diseased subjects are usually plentiful, so specificity is cheap to estimate.
  • Supports both objectives.
  • Reports the subgroup separately from total enrolment.

Limitations

  • Tighter precision is usually demanded than for sensitivity, offsetting the abundance of subjects.
  • Vulnerable to partial verification bias.
  • The normal approximation degrades as specificity approaches one.

Assumptions

  • The reference standard is applied to all subjects.
  • Disease status is classified without error.
  • Subjects are independent and representative.
  • The stated prevalence applies to the sampled population.

Applicable adjustments

Supports dropout inflation.

Conclusion

Report the non-diseased subgroup and total enrolment with the assumed prevalence. Consider the false-positive workload implied when choosing the precision target.

References

  1. Buderer NMF. Statistical methodology: I. Incorporating the prevalence of disease into the sample size calculation for sensitivity and specificity. Acad Emerg Med. 1996;3(9):895–900. Link
  2. Hajian-Tilaki K. Sample size estimation in diagnostic test studies of biomedical informatics. J Biomed Inform. 2014;48:193–204. Link
  3. Bossuyt PM, Reitsma JB, Bruns DE, et al. STARD 2015: An updated list of essential items for reporting diagnostic accuracy studies. BMJ. 2015;351:h5527. Link