← Back to Simulations
Data: T-cell count per unit area for three groups — Control (A, n = 55), Drug 1 (B, n = 58), and Drug 2 (C, n = 55).

Looking at the three groups

Each dot is one subject. Toggle the box below to see each group's mean, MAD (median absolute deviation), and median — B looks highest and most spread out, but with three groups there are now three gaps to explain, not just one.

Beeswarm comparison

Toggle box + stats

Is any group different? The omnibus test

The idea. Before asking which pair differs, ask whether there's a group effect at all. A median-based, ANOVA-like F statistic compares the spread of the three group medians around the grand median (between-group spread) to the spread of each group's own points around its own median (within-group spread). To find F's null distribution, first neutralize any real group differences: per the same rule the pairwise section uses, if A, B and C have roughly equal shape and variability, pool everyone into one shared big box and resample from it; otherwise center each group on zero — subtract its own median — and bootstrap-resample each group from itself. Either way, recompute the same F on the resample — each resample's F joins the null distribution below.

Omnibus resampling test

Original, prepared, and resample — all at once

Click the first button to prepare the data as described above, then Illustrate one draw to watch a single resample get built step by step below — some points get drawn more than once (extra pings on the same dot), some not at all — and watch its F join the null distribution. +100/+1,000/+10,000 add that many resamples at once; Reset clears everything and starts over.

Reading the result: the observed F sits far out past nearly every simulated draw under "no group effect" — a value this large essentially never happens by chance alone, so the omnibus test says at least one group really is different. That doesn't yet say which one.

Which groups differ? Pairwise post-hoc comparisons

The idea. Test each pair separately. We follow the rule of thumb that when one group's variability is more than double the other's, pooling them ("big box") would smear the more-variable group's spread onto the less-variable one, so each group is instead centered and resampled from itself ("re-centered 2-box"). Otherwise both are pooled into one box and resampled from that shared pool. The same three groups also give a 99% confidence interval for each pairwise gap, from a separate resampling distribution: each group resampled from itself with no centering, so the distribution sits near the observed gap instead of at zero.

Pairwise comparisons

p-value + 99% CI, all three pairs at once

Each row tests one pair: the two groups, a null distribution (for the p-value), and a sampling distribution with its 99% CI band.

A vs B
A vs C
B vs C
Reading the result: a pair is treated as significantly different when its p-value is small and its 99% CI excludes zero — the two views are two different resampling distributions of the same gap, so they occasionally disagree near the boundary, which is itself worth noticing rather than papering over.

Correcting for testing three pairs at once: Benjamini-Hochberg

The idea. Testing three pairs at once raises the odds that at least one "significant" result is just a false alarm. The Benjamini-Hochberg (BH) procedure controls that risk: rank the m = 3 p-values from smallest to largest (k = 1, 2, 3), compare each one to its own threshold α·km instead of the same flat α, and find the largest rank k^ whose p-value is still below that threshold. Every comparison at or below rank k^ is called significant; the rest are not — even ones that would have passed the flat, uncorrected α on their own.

The table and chart below re-rank the same three p-values from the pairwise section above and re-check them against the BH threshold, live, as more draws come in.

Benjamini-Hochberg correction

updates from the pairwise sim above
Comparisonkpk α·km pk < α? pk < α·km?
Run some draws on all three pairs above to see the correction.
BH-corrected confidence intervals. A confidence interval used for its own sake, to show precision, needs no correction. But using one as a stand-in significance test — declaring a difference whenever it excludes zero — is still multiple testing, and needs the same correction as the p-values above. Run some draws above to see the corrected interval.