Sign in to save

Bookmark this page so you can find it later.

Sign in to save

Bookmark this page so you can find it later.

A histogram shows how often data values fall within intervals, making it one of the fastest ways to understand a data set. The overall shape can reveal patterns that are hidden in a list of numbers. Symmetric, skewed, uniform, and bimodal shapes each suggest different stories about the data.

Recognizing these shapes helps students choose better summaries and make better comparisons.

Understanding Statistics: Distribution Shape at a Glance

A distribution shape is built from individual observations, but the picture depends on choices made before the graph is drawn. The width of each interval matters. Very wide intervals can hide gaps, clusters, or two peaks.

Very narrow intervals can make ordinary random variation look important. A useful habit is to inspect the scale on the horizontal axis and count how many observations are in the data set.

With only a few observations, a shape can change sharply when one new value is added. Larger samples usually give a more stable view of the underlying pattern.

Skew often comes from a real limit or from rare extreme cases. Test scores may be left-skewed when a test is easy, because many students score near the highest possible mark and fewer receive low marks. Waiting times can be right-skewed because most people wait a short time, while a few face long delays.

Income data are commonly right-skewed since a small number of very high incomes stretch the upper end. In these cases, the tail points toward the unusual values. It does not point toward the side where most data are located.

The shape affects which summary gives a fair description. A mean uses every value, so one unusually large or small value can pull it away from the typical observation. The median is based on position after values are ordered, so it changes much less when an extreme value appears.

For strongly skewed data, the median and the interquartile range often describe a typical value and the middle spread more honestly than the mean and standard deviation. Students should not treat a single average as the whole story. Two classes can have the same mean score while one class has scores packed near that mean and the other has a wide spread.

Two peaks deserve careful investigation rather than a quick label. They can occur when data from separate groups have been combined. Heights from children and adults, travel times on weekdays and weekends, or results from two teaching methods may form separate clusters.

A second peak can sometimes be caused by rounding, a change in measuring equipment, or poorly chosen intervals. Check the original data, the data collection method, and any categories that were mixed together. Separating the groups may reveal a clearer pattern.

Relative frequency helps compare distributions with different sample sizes. It turns a count in an interval into the share of all observations in that interval. This makes it possible to compare, for instance, scores from a class of twenty students with scores from a year group of two hundred students.

When comparing graphs, use matching interval widths and the same axis range whenever possible. Notice outliers, empty intervals, uneven spread, and any ceiling or floor caused by limits in the measurement. These details often explain the shape better than its name alone.

Key Facts

  • A symmetric distribution has roughly matching left and right sides, and the mean is usually close to the median.
  • A right-skewed distribution has a long tail to the right, and often mean > median.
  • A left-skewed distribution has a long tail to the left, and often mean < median.
  • A uniform distribution has bars with nearly equal heights, meaning values occur at similar frequencies across the range.
  • A bimodal distribution has two clear peaks, which may suggest two different groups or processes in the data.
  • Relative frequency = class frequency / total number of observations.

Vocabulary

Histogram
A graph that displays the frequency of numerical data values within intervals called bins.
Skewness
Skewness describes how much a distribution has a longer tail on one side than the other.
Symmetric distribution
A symmetric distribution has a shape where the left and right sides are approximately mirror images.
Bimodal distribution
A bimodal distribution has two distinct peaks, showing that two value ranges occur especially often.
Outlier
An outlier is a data value that is far from most other values in the data set.

Common Mistakes to Avoid

  • Calling any tall bar a mode, which is wrong because a mode is a value or interval with especially high frequency compared with nearby intervals.
  • Ignoring the tail when naming skew, which is wrong because skew is named for the direction of the long tail, not the side with the tallest bars.
  • Using the mean alone for a strongly skewed distribution, which is wrong because extreme values can pull the mean away from the typical data value.
  • Assuming a bimodal histogram is just random noise, which is wrong because two peaks can point to two separate groups, conditions, or processes.

Practice Questions

  1. 1 A histogram of quiz scores has bin frequencies 2, 5, 9, 5, 2 from lowest to highest score bins. What shape is suggested, and should the mean be close to the median?
  2. 2 A data set has 40 observations. In one histogram bin, the frequency is 10. What is the relative frequency for that bin?
  3. 3 A histogram of commute times has many values between 10 and 25 minutes, but a few values stretch out to 80 minutes. Identify the likely shape and explain which measure of center, mean or median, better represents a typical commute.