← Back to Simulations

What are data?

We are mostly concerned here with data that is a number, or numbers, describing some feature we are measuring. The most typical situation is when we have a set of subjects \( 1, 2, 3, \dots, n \) and measure a quantity \(x\) for each. We call these \( x_1, x_2, \dots, x_n \) and speak of a data set \( \{x_1, x_2, \dots, x_n\} \) describing the group.

For example, we might have a group of people, and we measure a systolic blood pressure (sbp) reading for each of them. Blood pressures are measured in mmHg (millimeters of mercury). The data set might look like:

Subject No. Name Blood Pressure (mmHg)
1Sally E\(x_1 = 108\)
2Bob J\(x_2 = 126\)
⋮⋮⋮
nSteve W\(x_n = 118\)

The first thing to do with any data set is: look at it. This may seem obvious, but it is surprising how often this step is skipped, sometimes with serious consequences.

How to look at data

Dot plot. When the data set is small to middling — say \( n < 100 \) — the most important visualization is the dot plot: we simply plot the data values as points on a single axis (horizontal or vertical).

Horizontal / Vertical dot plot

Hover over table rows or data points

In this chapter we use the horizontal presentation. In later chapters, when comparing several groups, we switch to the vertical one.

When the data set gets larger than 20 or so, there is a real chance that data points might overlap, giving a misleading picture. Consider the four data sets below. For \(n = 10\) (first row) the dot plot is fine. By \(n = 20\) (second row) some points already overlap. By \(n = 50\) and \(n = 100\) the mis-representation is substantial.

Dot plots for data sets of increasing size

Set μ, σ, click Generate

Note the increasing overlap of data points in the larger sets.

Jitter. One solution is to add "jitter": since the data is one-dimensional, we can use the second dimension to spread points out. We create a fictitious second axis and add a small random quantity to each data point along it. This separates equal or close points so they don't overlap. The second axis is meaningless — it carries no information.

Dot plots with jitter

Set μ, σ, click Generate

Beeswarm. A variant of the jitter plot is the beeswarm plot. When data points would otherwise overlap, we use a fixed spacing in the fictitious axis to offset each point — producing a tidy, symmetric spread that still encodes density faithfully.

Dot plot → Beeswarm plot

Set DRAWS, μ, σ, click Generate

Looking at the three forms of presentation, the simple dot plot is fine for small data sets. For larger ones — say \(n > 20\) — the beeswarm gives the best picture of the data.

The most important first step in looking at a data set is to make a picture of its distribution.

What to look for in a data plot

The true picture of the data set is its distribution, and we must begin by looking at what it is trying to tell us. Some important features:

Bunching up (symmetry vs. skew). Are the data fairly uniformly distributed over their range, or do they bunch up? If they bunch up in the middle the distribution is fairly symmetric. If they bunch up on the left, the distribution is skewed to the right; if on the right, skewed to the left. Think of skew as meaning sticking out asymmetrically.

Skewed left / right

Set DRAWS and Skew Direction, click Generate

Outliers. Are all the data in the same general range, or are there some points "way off" from the others? These are often called outliers.

Outliers

Set sample size (DRAWS), click Generate

Subgroups. Are there noticeable gaps between bunches of data? Does the data set seem to group into distinct subgroups, like a low group and a high group? Or does it vary continuously with no noticeable gaps?

Subgroups

▶ Interactive
Notice that the language of this section has been very visual and qualitative. You are being asked to form a visual impression of some qualitative features of the data set. This absolutely critical step is often skipped, partly out of a prejudice that these impressions are "subjective." In fact, they are the most important features of the data set, and our judgments of the distribution are critical in deciding how to analyze it.

Later on, we learn methods for turning these visual impressions into precise mathematical concepts. But the visual impression is always the first step.