Comparing two data sets helps you decide how groups are similar and different using evidence instead of guesses. In statistics, you usually compare center, spread, shape, and outliers. Comparative box plots are useful because they place the medians, quartiles, ranges, and possible outliers of two groups on the same scale.
Back-to-back stem plots are useful because they show the actual data values while making the two distributions easy to compare.
Understanding Statistics: Comparing Two Data Sets
A comparison starts with a fair setup. The two groups must measure the same variable in the same units. Comparing heights in centimetres with heights in inches can create a false difference.
The way data were collected matters too. A class survey, a random sample, and a set of volunteers can give very different pictures.
If one group has ten observations and another has one hundred, the smaller group may look more uneven just because each individual value has more effect. Before drawing a conclusion, check who was included, when measurements were taken, and whether the conditions were similar.
Center tells you what is typical, but the best measure depends on the data. The mean uses every value, so a few unusually large or small results can pull it away from where most values lie. The median is often more reliable for data such as income, house prices, or waiting times, where a long upper tail is common.
Imagine two groups with similar medians. One group could still have a higher mean because several people have very large values.
That difference is useful evidence about the distribution, not a mistake. Students should avoid saying one group is simply higher without naming the measure used.
Spread describes consistency. A group with a narrow middle half has results packed close together. A group with a wide middle half has more variation among typical results.
This matters in real situations. Two bus routes may have the same usual travel time, while one route has much less reliable arrival times. Two factories may produce parts with the same average length, while one factory produces lengths that vary too much for safe assembly.
The full range can be affected strongly by one extreme value. The interquartile range focuses on the middle half, so it often gives a clearer view of ordinary variation.
Shape adds details that a single number cannot show. A symmetric pattern has similar sides around its center. A skewed pattern has a longer tail on one side, often caused by a small number of unusual cases.
A distribution with two clear peaks may show that the data contain two different subgroups, such as weekday and weekend customers. Gaps can be meaningful too. They may show missing categories, different conditions, or a limit in how the data were recorded.
Possible outliers deserve investigation. They could be recording errors, genuine rare events, or clues that one observation came from a different population.
Good statistical writing states the evidence carefully. It describes differences in typical value and variability, then notes shape or unusual values instead of claiming that every member of one group differs from every member of the other.
Key Facts
- Median = the middle value when the data are ordered.
- Range = maximum value - minimum value.
- IQR = Q3 - Q1.
- Mean = sum of all data values / number of data values.
- A common outlier rule is: values less than Q1 - 1.5(IQR) or greater than Q3 + 1.5(IQR) may be outliers.
- When comparing two data sets, describe center, spread, shape, and outliers using the same scale and units.
Vocabulary
- Center
- The center of a data set is a typical or middle value, often described by the median or mean.
- Spread
- Spread describes how far apart the data values are, often measured by range or interquartile range.
- Median
- The median is the middle value of an ordered data set, or the average of the two middle values if there is an even number of values.
- Interquartile range
- The interquartile range, or IQR, is the distance between the first quartile and third quartile and describes the spread of the middle half of the data.
- Outlier
- An outlier is a data value that is unusually far from the rest of the data.
Common Mistakes to Avoid
- Comparing only the highest values is wrong because the maximum does not describe the typical value or the overall distribution.
- Using different number scales for the two graphs is wrong because it can make one data set look more spread out or shifted than it really is.
- Saying the data set with the larger range is always more variable is incomplete because one extreme outlier can make the range large while most values are still close together.
- Calling the mean the best center for every data set is wrong because outliers or strong skew can pull the mean away from a typical value.
Practice Questions
- 1 Data Set A: 4, 6, 7, 8, 10, 12, 13. Data Set B: 3, 5, 5, 9, 11, 14, 16. Find the median and range of each data set, then state which set has the larger center and which has the larger spread by range.
- 2 A box plot for Group 1 has Q1 = 20, median = 28, Q3 = 34, minimum = 16, and maximum = 42. A box plot for Group 2 has Q1 = 22, median = 25, Q3 = 31, minimum = 18, and maximum = 36. Find the IQR and range for each group, then compare center and spread.
- 3 Two classes took the same quiz. Class A has a higher median score, but Class B has a much larger IQR and one very low outlier. Explain what this means about typical performance and consistency in the two classes.