Sign in to save

Bookmark this page so you can find it later.

Sign in to save

Bookmark this page so you can find it later.

Statistics begins with knowing what kind of data you have. This cheat sheet helps students tell the difference between categorical data, which describe groups or labels, and quantitative data, which measure numbers. Choosing the correct data type matters because it determines which graph, table, and summary measure makes sense.

Students can use this reference when organizing survey results, experiment data, or real-world measurements.

Categorical data are usually summarized with counts, relative frequencies, bar graphs, or two-way tables. Quantitative data are summarized with measures of center such as mean and median, measures of spread such as range, and displays such as dot plots, histograms, and box plots. A helpful rule is that arithmetic only makes sense for quantitative data, so values like jersey numbers and ZIP codes are not quantitative measurements.

Good statistical thinking starts by identifying the variable, classifying the data, and selecting an appropriate display.

Key Facts

  • Categorical data place individuals into groups or labels, such as eye color, favorite sport, or type of pet.
  • Quantitative data are numerical measurements or counts where arithmetic is meaningful, such as height, time, age, or number of siblings.
  • A frequency is the count in a category, and the total sample size is n=fn = \sum f.
  • Relative frequency is found with relative frequency=category frequencytotal frequency\text{relative frequency} = \frac{\text{category frequency}}{\text{total frequency}}.
  • A percent frequency is found with percent=partwhole×100%\text{percent} = \frac{\text{part}}{\text{whole}} \times 100\%.
  • The mean of quantitative data is xˉ=x1+x2++xnn\bar{x} = \frac{x_1 + x_2 + \cdots + x_n}{n}.
  • The range of quantitative data is range=maximumminimum\text{range} = \text{maximum} - \text{minimum}.
  • Use bar graphs for categorical data and dot plots, histograms, or box plots for quantitative data.

Vocabulary

Categorical Data
Data that describe qualities, groups, names, or labels rather than measurements.
Quantitative Data
Data made of numbers that represent counts or measurements where arithmetic is meaningful.
Variable
A characteristic being recorded or measured for each individual in a data set.
Frequency
The number of times a value or category appears in a data set.
Relative Frequency
The fraction or proportion of the total data that belongs to a category.
Distribution
The pattern of values in a data set, including where values cluster and how they spread out.

Common Mistakes to Avoid

  • Treating every number as quantitative is wrong because some numbers are labels, such as jersey numbers, room numbers, or ZIP codes.
  • Using a histogram for categorical data is wrong because histograms show intervals of numerical values, not separate labels or groups.
  • Finding the mean of categorical data is wrong because categories like colors or favorite foods cannot be added and divided meaningfully.
  • Confusing frequency with relative frequency is wrong because frequency is a count, while relative frequency is a fraction such as fn\frac{f}{n}.
  • Forgetting to include units for quantitative data is wrong because measurements like 1212 could mean 1212 seconds, 1212 meters, or 1212 students.

Practice Questions

  1. 1 Classify each variable as categorical or quantitative: favorite ice cream flavor, number of books read in a month, shoe size, and type of phone.
  2. 2 A survey of 4040 students found that 1212 chose soccer as their favorite sport. Find the relative frequency and percent for soccer.
  3. 3 The data set 4,7,7,9,134, 7, 7, 9, 13 gives the number of pets owned by different families. Find the mean and range.
  4. 4 A student says ZIP codes are quantitative because they are written with numbers. Explain why this reasoning is incorrect.

Understanding Categorical vs Quantitative Data

A variable is one recorded feature of each person, object, or event in a data set. Before organizing the data, decide what the recorded value means. A number can represent an amount, yet it can also be only a name tag.

A test score is quantitative because a difference of ten points has meaning. A bus route number is a label, even though it uses digits. Some data need extra care.

Letter grades and rating choices such as poor, fair, good, and excellent have an order. They are called ordinal categories.

Their gaps are not known to be equal, so finding an average rating can give a misleading result. Dates can be used to find elapsed time, but the written date itself often acts like a label.

Quantitative values can be discrete or continuous. Discrete values come from counting and usually jump from one whole value to the next. The number of books checked out is discrete.

Continuous values come from measuring and can fall anywhere within a range, depending on the precision of the tool. Height, mass, temperature, and travel time are continuous. Units matter because they explain what a value measures.

A class should not combine heights in centimeters with heights in inches until they are converted. Rounding matters too. If times are rounded to the nearest minute, two students with slightly different times may appear tied in the table.

When quantitative data are grouped into intervals for a histogram, the chosen interval width affects the picture. Very wide intervals can hide important clusters or gaps. Very narrow intervals can make a small sample look random and scattered.

Histogram bars touch because the intervals represent connected number ranges. Bar graph bars have spaces because categories are separate groups. The shape of a distribution gives useful evidence.

A long tail on the high end is right skewed. In a skewed set, a few unusually large values can pull the mean upward. The median is often a better description of a typical value in that case.

Outliers should be checked, not automatically deleted. They may show a recording error, or they may reveal something important.

Comparisons need a fair basis. Counts alone can be misleading when groups have different sizes. A club with forty members may have more students choosing soccer than a club with ten members, even if soccer is less common in the larger club.

Relative frequencies make that comparison fair by describing the share of each group. Two-way tables help students compare one category across another, such as preferred lunch option by grade level. A pattern in a table shows an association, not proof that one factor caused the other.

Results also depend on who was measured. A survey of one friend group cannot reliably describe an entire school. Clear labels, honest scales, and a representative sample are as important as the calculations.