← Back to Simulations
Data: 1,768 paired CO2/temperature measurements spanning roughly 800,000 years, built from NOAA's (2018) long-term CO2 concentration ice-core record and an Antarctic ice-core deuterium temperature record, interpolated onto the CO2 sampling years. Temperature is an anomaly (°C) relative to the average of the last 1000 years.

Looking at the relationship

Each point below is one moment in Earth's climate history: its CO2 concentration and the temperature at that time. The marginal histograms show each variable's own distribution. The two look tightly linked — the Pearson correlation coefficient is ρ = 0.929 — but with 1,768 points drawn from an autocorrelated climate record, could a correlation this strong show up even by chance?

CO2 concentration vs. temperature

n = 1,768 pairs
Hover a dot or histogram bar to highlight its matching points

Testing the correlation: shuffle the data

The idea. If CO2 and temperature were really unrelated, it shouldn't matter which CO2 value we attach to which temperature reading — the pairing is arbitrary noise. So: keep every temperature exactly where it is, and randomly shuffle the CO2 values among them. Recompute the correlation coefficient. Do that thousands of times, and you get the distribution of ρ you'd see from pure chance alone — using all 1,768 pairs every time, even though only a legible sample of dots is animated below.

Shuffle test (permutation NHST)

Temperature fixed, CO2 shuffled

The faint backdrop is all 1,768 pairs; the ~220 bold dots are a representative sample that animates so the shuffle stays legible — every ρ value dropped into the histogram below is computed from the full dataset.

How precise is our estimate? Resampling

The idea. This time we don't scramble anything — each (CO2, temperature) pair stays intact. Instead we ask: if we'd happened to sample a slightly different 1,768 moments (by resampling with replacement from the ones we have), how much would our correlation coefficient move around? That spread is our uncertainty about the true correlation. The confidence interval below uses the pivotal resample formula:

CIlower = 2 × ρactual − ρupper
CIupper = 2 × ρactual − ρlower

Resampling-based confidence interval

Pairs resampled with replacement

Each pick is its own small dot, offset sideways from its point — a pair drawn twice shows up as two dots side by side. Pairs left out of this resample appear as faint dashed outlines. This is shown only for the ~220-point representative sample; every ρ value is computed from all 1,768 pairs.

Same machinery, different question

Both methods recompute the correlation coefficient thousands of times and look at how it varies. It's tempting to think of them as the same thing — they aren't. The procedure is what tells them apart:

Shuffle test (NHST)
Breaks the CO2↔temperature pairing on purpose.
Asks: "if there were truly no relationship, how often would chance alone produce a correlation this strong?"
Answer is a single number — the p-value.
Resample CI
Preserves every CO2↔temperature pairing.
Asks: "given the relationship we found, how precisely have we pinned down its strength?"
Answer is a range — the confidence interval.
In short: shuffling destroys the relationship to build a null distribution for testing whether it exists at all. Resampling keeps the relationship intact to build a sampling distribution for measuring how precisely we've estimated it. With n = 1,768 and ρ = 0.929, both come out looking dramatic — the null distribution clusters tightly near zero while the observed ρ sits far outside it, and the resample CI is very narrow — which is exactly what you'd expect from such a strong, well-sampled relationship.

NHST and CI, side by side

Plotting both distributions on the same axis makes the contrast concrete: the null distribution sits centered near zero, while the resampling distribution sits centered near ρactual — worlds apart. The axis breaks between them because the two ranges don't overlap at all.

Null distribution vs. resampling distribution