Non-proportional hazards
Determines the sample size required for a time-to-event study when the treatment effect is expected to vary over time and the proportional-hazards assumption does not hold.
Key takeaways
- What is calculated: Determine the events and participants required under a time-varying hazard ratio.
- Quantities you must supply: Median survival — control; HR in the early period; HR in the late period; Time at which the HR changes; Accrual period; Additional follow-up; Significance level; Power.
- Type of calculation: Power-based: it answers how much data is needed to detect a stated effect. Power 0.80 at α = 0.05 is the usual target; confirmatory studies use 0.90.
Method used for Non-proportional hazards
- Non-proportional hazards — event-weighted average hazard ratio Piecewise-constant hazard ratio summarised as an event-weighted average on the log scale. Because most events occur early, an effect confined to the late period is heavily discounted, and the average sits much closer to one than the late value suggests.
When to use
Use when the primary objective is to evaluate a time-to-event treatment effect that is expected to vary over time, such that the proportional-hazards assumption may not hold.
Sample size
Awaiting inputs
—
Set the parameters and calculate.
Methodology — Non-proportional hazards — event-weighted average hazard ratio
Formula used Non-proportional hazards — event-weighted average hazard ratio
- Two-sided alternative selected: α is split between both tails, so the larger quantile is used and the requirement rises
where —
- m₁
- Median survival in the control arm. > 0
- HR_e
- Hazard ratio before the switch; 1 means no early benefit. > 0
- HR_l
- Hazard ratio after the switch. > 0
- t*
- Time at which the hazard ratio changes. ≥ 0
- a
- Recruitment period with uniform entry. ≥ 0
- f
- Additional follow-up; longer follow-up weights the late period more. ≥ 0
- α
- Probability of rejecting a true null hypothesis (Type I error). 0.0001 – 0.5; conventionally 0.05, or 0.025 one-sided for regulatory non-inferiority
- 1 − β
- Probability of rejecting the null hypothesis when the specified alternative is true. 0.50 – 0.9999; conventionally 0.80 or 0.90
- ℓ
- Expected proportion of enrolled units lost before analysis. 0 – 0.95
How this method works
When the hazard ratio changes over time, no single value describes the effect. This method models an early and a late hazard ratio and averages them weighted by the events expected in each period, reflecting how the log-rank statistic actually accumulates information.
A piecewise-constant hazard ratio summarised as an event-weighted average on the log scale, reflecting how the log-rank statistic actually accumulates information.
Because most events occur early, an effect confined to the late period is heavily discounted and the average sits much closer to 1. Where non-proportionality is expected, RMST or a weighted log-rank test is often the better primary analysis.
Null hypothesis. H0: the averaged hazard ratio equals one.
Alternative hypothesis. Two-sided H1: it differs from one.
Calculation procedure
- Split follow-up into an early and a late period and state the hazard ratio in each, and .
- State what share of events falls in the early period; the rest fall in the late one.
- Average on the log scale, weighted by events: .
- Apply the Schoenfeld relation to that average: , then .
- Round up, then inflate for dropout. A delayed effect always costs events relative to a constant hazard ratio.
Exact or approximate
Approximate. A weighted log-rank test would be a more powerful analysis than the standard log-rank assumed here.
The formula shown always matches what was computed: when a one-sided test is selected the rendered quantile changes from z1−α/2 to z1−α, and where several methods exist, the formula follows the method selected in the calculator.
Advantages & limitations
Advantages
- Quantifies the cost of a delayed effect, which proportional-hazards calculators hide.
- Reports the sample size under the late HR alone for comparison.
- Reflects how the log-rank statistic accumulates information.
Limitations
- Assumes a single switch point, which is a simplification.
- The standard log-rank test is not the optimal analysis under non-proportional hazards.
- RMST is often the better primary endpoint in this situation.
Assumptions
- The hazard ratio is constant within each of the two periods.
- Control survival is exponential.
- Accrual is uniform.
- Censoring is non-informative.
Applicable adjustments
Supports dropout inflation.
Conclusion
Report the averaged hazard ratio alongside the early and late values and the switch time. Consider RMST or a weighted log-rank test as the primary analysis.
References
- Lin RS, Lin J, Roychoudhury S, et al. Alternative analysis methods for time to event endpoints under nonproportional hazards: a comparative analysis. Stat Biopharm Res. 2020;12(2):187–198. Link
- Fine GD. Consequences of delayed treatment effects on analysis of time-to-event endpoints. Drug Inf J. 2007;41(4):535–539. Link
- Royston P, Parmar MKB. Restricted mean survival time: an alternative to the hazard ratio for the design and analysis of randomized trials with a time-to-event outcome. BMC Med Res Methodol. 2013;13:152. Link