Visualizing and Presenting Data
The most important first step in any data analysis is to look at your data. This module introduces dot plots, jitter, and beeswarm plots — the tools for turning a table of numbers into a picture you can actually reason about.
What are data?
We are mostly concerned here with data that is a number, or numbers, describing some feature we are measuring. The most typical situation is when we have a set of subjects \( 1, 2, 3, \dots, n \) and measure a quantity \(x\) for each. We call these \( x_1, x_2, \dots, x_n \) and speak of a data set \( \{x_1, x_2, \dots, x_n\} \) describing the group.
For example, we might have a group of people, and we measure a systolic blood pressure (sbp) reading for each of them. Blood pressures are measured in mmHg (millimeters of mercury). The data set might look like:
| Subject No. | Name | Blood Pressure (mmHg) |
|---|---|---|
| 1 | Sally E | \(x_1 = 108\) |
| 2 | Bob J | \(x_2 = 126\) |
| ⋮ | ⋮ | ⋮ |
| n | Steve W | \(x_n = 118\) |
The first thing to do with any data set is: look at it. This may seem obvious, but it is surprising how often this step is skipped, sometimes with serious consequences.
How to look at data
Dot plot. When the data set is small to middling — say \( n < 100 \) — the most important visualization is the dot plot: we simply plot the data values as points on a single axis (horizontal or vertical).
Horizontal / Vertical dot plot
Hover over table rows or data pointsIn this chapter we use the horizontal presentation. In later chapters, when comparing several groups, we switch to the vertical one.
When the data set gets larger than 20 or so, there is a real chance that data points might overlap, giving a misleading picture. Consider the four data sets below. For \(n = 10\) (first row) the dot plot is fine. By \(n = 20\) (second row) some points already overlap. By \(n = 50\) and \(n = 100\) the mis-representation is substantial.
Dot plots for data sets of increasing size
Set μ, σ, click GenerateNote the increasing overlap of data points in the larger sets.
Jitter. One solution is to add "jitter": since the data is one-dimensional, we can use the second dimension to spread points out. We create a fictitious second axis and add a small random quantity to each data point along it. This separates equal or close points so they don't overlap. The second axis is meaningless — it carries no information.
Dot plots with jitter
Set μ, σ, click GenerateBeeswarm. A variant of the jitter plot is the beeswarm plot. When data points would otherwise overlap, we use a fixed spacing in the fictitious axis to offset each point — producing a tidy, symmetric spread that still encodes density faithfully.
Dot plot → Beeswarm plot
Set DRAWS, μ, σ, click GenerateLooking at the three forms of presentation, the simple dot plot is fine for small data sets. For larger ones — say \(n > 20\) — the beeswarm gives the best picture of the data.
What to look for in a data plot
The true picture of the data set is its distribution, and we must begin by looking at what it is trying to tell us. Some important features:
Bunching up (symmetry vs. skew). Are the data fairly uniformly distributed over their range, or do they bunch up? If they bunch up in the middle the distribution is fairly symmetric. If they bunch up on the left, the distribution is skewed to the right; if on the right, skewed to the left. Think of skew as meaning sticking out asymmetrically.
Skewed left / right
Set DRAWS and Skew Direction, click GenerateOutliers. Are all the data in the same general range, or are there some points "way off" from the others? These are often called outliers.
Outliers
Set sample size (DRAWS), click GenerateSubgroups. Are there noticeable gaps between bunches of data? Does the data set seem to group into distinct subgroups, like a low group and a high group? Or does it vary continuously with no noticeable gaps?
Subgroups
▶ InteractiveLater on, we learn methods for turning these visual impressions into precise mathematical concepts. But the visual impression is always the first step.