A distribution shows how often different values occur in a data set, and its shape tells you a lot before you do any calculations. Symmetric distributions are balanced around the center, while skewed distributions have a long tail on one side. Recognizing the shape helps you choose the best measure of center and describe typical values clearly.
This matters in science, economics, medicine, and any field where data can be pulled by extreme values.
In a symmetric distribution, the mean and median are usually close together because the data balance evenly on both sides. In a right-skewed distribution, a few large values pull the mean to the right, so the mean is usually greater than the median. In a left-skewed distribution, a few small values pull the mean to the left, so the mean is usually less than the median.
The median is often a better measure of a typical value when a distribution is strongly skewed.
Understanding Statistics: Skewed vs Symmetric Distributions
Start by sorting the values or placing them into equal-width intervals on a histogram. The choice of intervals can change how the shape looks. Very wide intervals can hide clusters or gaps.
Very narrow intervals can make random bumps look important. Use sensible intervals that fit the size and range of the data. Then inspect the whole pattern, not only the tallest bar.
Notice where most observations lie, how spread out they are, and whether values thin out gradually or stop suddenly. A tail is made from relatively few observations spread far from the main group. Its direction matters more than the location of the highest bar.
A skewed pattern often has a real-world cause. Household incomes commonly have a long upper tail because there is a practical lower limit to income, while a small number of people earn extremely large amounts. Travel times can be right-skewed because most journeys finish near the usual time, but traffic, accidents, or missed connections create a few much longer trips.
Scores on an easy test may be left-skewed when many students score near full marks and only a few receive very low scores. Knowing the situation helps explain the shape instead of treating it as a mysterious feature of a graph.
Outliers and skewness are related but not identical. One recording error, such as entering an extra zero in a measurement, can create a single extreme point without producing a genuinely skewed population. Check unusual values against the original source before using them.
A real extreme value should not automatically be removed. For example, an unusually expensive house is still part of a housing price study.
Report it, explain its effect, and choose summaries that do not let it dominate the description. The median and the interquartile range are resistant measures because a few distant values change them less than they change the mean and standard deviation.
Shape affects later statistical work. Many methods for comparing groups or estimating uncertainty work best when data are roughly balanced, especially with small samples. Strong skew can make averages unstable from one sample to another.
Students often meet this when comparing salaries, waiting times, rainfall totals, hospital stays, or social media follower counts. A useful habit is to graph the data before calculating anything. Compare the graph with the mean, median, range, and quartiles.
If the data are strongly skewed, state that clearly in a conclusion. A single average can be mathematically correct while giving a misleading picture of what is typical for most people.
Key Facts
- Symmetric distribution: mean ≈ median ≈ mode when the curve is single-peaked and balanced.
- Right-skewed distribution: mean > median because the long tail points toward larger values.
- Left-skewed distribution: mean < median because the long tail points toward smaller values.
- Mean formula: mean = (sum of all data values) / n.
- Median: the middle value when data are ordered, or the average of the two middle values if n is even.
- Skewness describes asymmetry: positive skew means right-skewed, and negative skew means left-skewed.
Vocabulary
- Distribution
- A distribution is the pattern showing how data values are spread across possible values.
- Symmetric distribution
- A symmetric distribution has roughly the same shape on both sides of its center.
- Right-skewed distribution
- A right-skewed distribution has a long tail extending toward larger values.
- Left-skewed distribution
- A left-skewed distribution has a long tail extending toward smaller values.
- Median
- The median is the middle value of an ordered data set and is resistant to extreme values.
Common Mistakes to Avoid
- Calling the tall side of the graph the direction of skew is wrong because skew is named for the long tail, not the peak.
- Assuming the mean is always the best center is wrong because extreme values can pull the mean away from a typical value in skewed data.
- Saying mean equals median in every bell-shaped graph is wrong because small asymmetries or outliers can separate them.
- Ignoring the scale of the horizontal axis is wrong because stretched or compressed axes can make a distribution look more or less skewed than it really is.
Practice Questions
- 1 For the data set 4, 5, 5, 6, 6, 7, 7, 8, 30, find the mean and median, then decide whether the distribution is likely right-skewed or left-skewed.
- 2 A left-skewed data set has a median of 72. Which is more likely for its mean: 65 or 80? Explain using the position of the tail.
- 3 A town reports household income using the mean, but most households earn less than the reported mean. Explain what distribution shape is likely present and why the median may be more informative.