← Back to Simulations
Setup: a baseline group drawn from a uniform range (45–75), and a treatment group that's genuinely 25% higher — the same simulated experiment as Figures 11.8–11.9. Each simulated study is tested with a bootstrap NHST (α = 0.01); power is the fraction of studies whose test actually catches the effect.

Simulating one study at a time

The idea. Every simulated "study" draws a fresh baseline group and treatment group, then runs the same bootstrap test you've seen elsewhere on this site: resample each group (centered on its own median) with replacement, many times, and see how often the resampled gap is at least as big as the one actually observed. That fraction is the study's p-value. Do this for thousands of simulated studies at a given sample size, and the fraction landing at or below α = 0.01 — the shaded region below — is the power of a study that size. Everything else is the Type II error rate (β): real effects the study missed.

Power at a given sample size (using the median)

Drag n and watch power update
n = 10

Purple bars fall at or below α = 0.01 — the study's test caught the effect. Blue bars missed it.

The power curve

The idea. Repeat the whole simulation above at every sample size from 5 to 50, and power traces out a curve: low and unreliable at small n, climbing steeply, then flattening out near 1.0 once the sample is big enough that chance alone can barely hide a real effect. The gold dot below tracks whatever n the slider above is set to.

Power vs. sample size (using the median)

n = 5–50

Group measure: median. Each sample size n: 10,000 simulated studies at that sample size, each tested with a bootstrap of 5,000 inner resamples for p-value.

Reading the result: at n = 10, a real 25%-higher effect gets missed more often than not — the study is underpowered. Push the slider to n = 25 and power jumps substantially, exactly as Figures 11.8 and 11.9 show side by side. Past roughly n = 35–40 here, power is already near 1.0 and more subjects buy little extra.

What if we used the mean instead?

The idea. Everything about the bootstrap test stays the same as above — same baseline, same 25%-higher treatment effect, same α = 0.01 — except each group is now summarized by its mean instead of its median. Watch one study at a time first, then see how the whole power curve changes: on this page's uniform baseline, the sample mean is a far more efficient (lower-variance) estimator than the median, so a test built on means detects the same effect reliably at much smaller sample sizes.

Power at a given sample size (using the mean)

Drag n and watch power update
n = 10

Purple bars fall at or below α = 0.01 — the study's test caught the effect. Blue bars missed it.

The power curve, using the mean

The idea. Repeat the mean-based simulation above at every sample size from 5 to 50. The curve looks completely different from the median's: it climbs steeply and is essentially flat at 1.0 within a handful of n values, instead of climbing gradually all the way out to n = 50.

Power vs. sample size (using the mean)

n = 5–50

Each sample size n: 10,000 simulated studies at that sample size, each tested with a bootstrap of 250 inner resamples for p-value.

Reading the result: power with the mean is already near 100% by n = 14 and stays there — compare that to n = 35–40 above for the median to reach the same point. The effect itself hasn't changed; only how each group is summarized has.

So which should you use: mean or median?

The idea. There's no universal winner here — the right choice depends on the shape of the data and the conventions of the field, not a fixed rule of thumb.

Use the mean when the data is roughly symmetric and outlier-free, like this page's uniform(45–75) baseline. The mean uses every value's magnitude, not just its rank, so it estimates the group's center with less sample-to-sample noise — which is exactly why the mean-based test above reaches high power at much smaller n than the median-based one.

Use the median when the data is skewed or has outliers — income, hospital costs, reaction times, anything with a long tail. A single extreme value can drag the mean away from where most of the data actually sits; the median barely notices.

For survival time specifically, median is the field standard — for a reason beyond just skew. Survival times are almost always right-skewed (most events cluster early, a few subjects survive a very long time), so the median already better represents a "typical" outcome than the mean. But the sharper reason is censoring: in a real survival study, some subjects haven't experienced the event by the time the study ends, so their exact survival time is unknown. The mean can't be computed from censored data without extra assumptions, but the median can still be read directly off the data (or a Kaplan–Meier curve) as long as more than half the subjects have had the event. That's why "median survival time" is the number reported in almost every clinical trial — even though this page's simulated version (a clean, uncensored uniform distribution) doesn't have that problem, which is what lets the mean win on efficiency here instead.

A side effect specific to the median: the power curve isn't smooth. Look closely at the median-based power curve above and you'll see a sawtooth — each even n sits noticeably higher than the odd n right after it (e.g., n = 17 vs. n = 18). That's not simulation noise; it's a real parity effect: with an odd-sized sample, the median is a single middle data point (higher variance), while with an even-sized sample, it's the average of the two middle points, which pulls the variance down. That lower variance at even n translates directly into higher power. This fluctuation shrinks as sample size increases — at small n (5–20) the jump between consecutive odd/even points is on the order of 0.04–0.13 in power; by n = 40–50 it's down to a few thousandths, since the odd/even effect is a fixed, one-data-point quirk while the median's overall sampling variance shrinks like 1/n. The mean has no such effect — its curve is smooth at every n, since averaging works the same regardless of whether n is odd or even.