Statistical Power: How Big a Study Do You Need?
A study can miss a real effect just because the sample was too small. Power is the probability a study actually detects an effect that's really there; drag the sample-size slider and watch it happen in real time.
Simulating one study at a time
The idea. Every simulated "study" draws a fresh baseline group and treatment group, then runs the same bootstrap test you've seen elsewhere on this site: resample each group (centered on its own median) with replacement, many times, and see how often the resampled gap is at least as big as the one actually observed. That fraction is the study's p-value. Do this for thousands of simulated studies at a given sample size, and the fraction landing at or below α = 0.01 — the shaded region below — is the power of a study that size. Everything else is the Type II error rate (β): real effects the study missed.
Power at a given sample size (using the median)
Drag n and watch power updatePurple bars fall at or below α = 0.01 — the study's test caught the effect. Blue bars missed it.
The power curve
The idea. Repeat the whole simulation above at every sample size from 5 to 50, and power traces out a curve: low and unreliable at small n, climbing steeply, then flattening out near 1.0 once the sample is big enough that chance alone can barely hide a real effect. The gold dot below tracks whatever n the slider above is set to.
Power vs. sample size (using the median)
n = 5–50Group measure: median. Each sample size n: 10,000 simulated studies at that sample size, each tested with a bootstrap of 5,000 inner resamples for p-value.
What if we used the mean instead?
The idea. Everything about the bootstrap test stays the same as above — same baseline, same 25%-higher treatment effect, same α = 0.01 — except each group is now summarized by its mean instead of its median. Watch one study at a time first, then see how the whole power curve changes: on this page's uniform baseline, the sample mean is a far more efficient (lower-variance) estimator than the median, so a test built on means detects the same effect reliably at much smaller sample sizes.
Power at a given sample size (using the mean)
Drag n and watch power updatePurple bars fall at or below α = 0.01 — the study's test caught the effect. Blue bars missed it.
The power curve, using the mean
The idea. Repeat the mean-based simulation above at every sample size from 5 to 50. The curve looks completely different from the median's: it climbs steeply and is essentially flat at 1.0 within a handful of n values, instead of climbing gradually all the way out to n = 50.
Power vs. sample size (using the mean)
n = 5–50Each sample size n: 10,000 simulated studies at that sample size, each tested with a bootstrap of 250 inner resamples for p-value.
So which should you use: mean or median?
The idea. There's no universal winner here — the right choice depends on the shape of the data and the conventions of the field, not a fixed rule of thumb.
Use the mean when the data is roughly symmetric and outlier-free, like this page's uniform(45–75) baseline. The mean uses every value's magnitude, not just its rank, so it estimates the group's center with less sample-to-sample noise — which is exactly why the mean-based test above reaches high power at much smaller n than the median-based one.
Use the median when the data is skewed or has outliers — income, hospital costs, reaction times, anything with a long tail. A single extreme value can drag the mean away from where most of the data actually sits; the median barely notices.
For survival time specifically, median is the field standard — for a reason beyond just skew. Survival times are almost always right-skewed (most events cluster early, a few subjects survive a very long time), so the median already better represents a "typical" outcome than the mean. But the sharper reason is censoring: in a real survival study, some subjects haven't experienced the event by the time the study ends, so their exact survival time is unknown. The mean can't be computed from censored data without extra assumptions, but the median can still be read directly off the data (or a Kaplan–Meier curve) as long as more than half the subjects have had the event. That's why "median survival time" is the number reported in almost every clinical trial — even though this page's simulated version (a clean, uncensored uniform distribution) doesn't have that problem, which is what lets the mean win on efficiency here instead.