A Pareto chart is a graph that helps you see which categories contribute the most to a total problem. It combines bars sorted from largest to smallest with a line showing the cumulative percentage. This makes it useful for finding the “vital few” causes that deserve the most attention.
In quality control, business, engineering, and school projects, Pareto charts help turn raw counts into clear priorities.
The bars show category frequencies, such as defects, delays, complaints, errors, returns, setup issues, and miscellaneous causes. The cumulative percentage line adds each category’s contribution as you move from left to right. If the first few categories account for most of the total, the chart supports the 80-20 idea, meaning a small number of causes may create a large share of the effects.
A Pareto chart does not prove cause and effect, but it helps teams choose where to investigate and act first.
Understanding Statistics: Pareto Charts
To build a reliable Pareto chart, start by defining one clear unit of observation. A factory might record each faulty product. A school office might record each late arrival.
A website team might record each reported problem. Every observation needs one category chosen using the same rules. If one person labels a problem as shipping delay while another calls it poor service, the totals become unreliable.
Create short category definitions before collecting data. Include a category for other, but review it carefully. If other becomes a large bar, the categories are too vague or important causes have been hidden.
The cumulative line is built from running totals. Imagine 100 recorded errors. If the largest category has 35 errors, its cumulative percentage is 35 percent.
If the next category has 25 errors, the running total becomes 60 errors, so the line reaches 60 percent. Continue adding each bar until all 100 errors are included. The line is useful because it shows the share covered by a group of categories, not just the size of one category.
A team may decide to study the first three categories because together they represent 70 percent of the recorded errors. This is often more practical than trying to fix every category at once.
Students meet this type of thinking in ordinary decisions. A class could sort reasons for missed homework, such as unclear instructions, absent students, lost work, lack of time, or technical problems. A sports team could track the kinds of mistakes that lead to lost points.
A student planning revision could count errors on practice tests by topic. The largest error category can guide the next study session. Still, frequency is not the only thing that matters.
A rare safety problem may deserve urgent action even if it appears near the right side of the chart. Cost, severity, time, fairness, and whether a cause can actually be changed all affect the final decision.
A Pareto chart describes what was recorded during a particular period. It can be misleading when the data collection period is too short, the sample is small, or conditions changed during collection. Ten complaints in one week may not represent a normal month.
Counts can mislead when categories have different chances to occur. For example, one machine may produce far more items than another, so compare defect rates per item when exposure differs. Do not treat the 80 percent idea as a law.
Some charts have one dominant cause, while others have many similar bars. After choosing a priority, investigate the reason behind it, make a change, then collect new data. A later Pareto chart can show whether the change reduced the problem or merely shifted it into another category.
Key Facts
- A Pareto chart uses bars sorted from highest frequency to lowest frequency.
- Cumulative count after category k = sum of counts from category 1 through category k.
- Cumulative percentage = cumulative count / total count x 100%.
- The cumulative percentage line usually starts above the first bar and ends at 100%.
- The 80-20 idea suggests that about 80% of effects may come from about 20% of causes.
- Pareto charts are best for categorical data, such as defect types, complaint reasons, or error sources.
Vocabulary
- Pareto chart
- A graph that displays categories in descending order with bars and adds a cumulative percentage line.
- Frequency
- The number of times a category or event occurs in a data set.
- Cumulative percentage
- The running percent of the total after adding each category from left to right.
- 80-20 rule
- The idea that a large share of outcomes often comes from a small share of causes.
- Category
- A group or label used to organize data, such as defects, delays, or returns.
Common Mistakes to Avoid
- Leaving the bars unsorted is wrong because a Pareto chart depends on ranking categories from largest to smallest to show priorities clearly.
- Using percentages without checking the total is wrong because each category’s percent must be based on the same overall total.
- Treating the 80-20 rule as exact is wrong because it is a guideline, not a law that must always produce exactly 80% and 20%.
- Assuming the biggest bar is the root cause is wrong because a Pareto chart shows frequency, not proof of why the problem happens.
Practice Questions
- 1 A company records defect counts of 45, 30, 15, 6, and 4 for five defect types already sorted from largest to smallest. What is the cumulative percentage after the first two categories?
- 2 A service desk has 50 delays, 25 complaints, 15 errors, 5 returns, and 5 setup issues. Make the cumulative percentage values for the categories in this order.
- 3 A Pareto chart shows that complaints are frequent but inexpensive, while returns are less frequent but very costly. Explain why a team might need more than the Pareto chart before choosing what to fix first.