Sign in to save

Bookmark this page so you can find it later.

Sign in to save

Bookmark this page so you can find it later.

Statistics helps us turn raw data into useful information. Descriptive statistics focuses on organizing and summarizing the data you actually have, such as a class set of test scores or a list of daily temperatures. Inferential statistics uses data from a sample to make reasonable conclusions about a larger population.

Knowing the difference matters because summaries describe what was observed, while inferences estimate what may be true beyond the observed data.

Descriptive statistics includes measures like mean, median, range, and standard deviation, often shown with tables, graphs, and charts. Inferential statistics depends on sampling, probability, confidence intervals, and hypothesis tests to handle uncertainty. A good inference requires a sample that represents the population well, not just a large sample.

In science, business, medicine, and public policy, inferential statistics helps people make decisions when measuring every individual is impossible.

Understanding Statistics: Descriptive vs Inferential Statistics

A summary can be accurate but still hide an important pattern. Imagine two groups of students with the same average score. One group may have scores tightly clustered near the average.

The other may include very high and very low scores. Looking at spread reveals this difference. A histogram can show whether values form one main cluster, have two separate clusters, or contain unusual values called outliers.

A box plot makes the middle half of the data easier to see. Students should choose graphs to match the data type.

Bar charts suit categories such as favourite sport. Histograms suit measured numerical values such as height or time.

The mean is useful because it uses every value, but it is pulled by extreme observations. If most weekly earnings are similar but one person earns far more, the mean can suggest that a typical person earns more than most people actually do. In this situation, the median is often a fairer description of a typical value.

The mode is useful for the most common category or value. There is no single best summary for every data set.

Good statistical work reports enough information for a reader to see the center, spread, shape, and any unusual values. It should never use one number to tell a misleading story.

Making an inference adds uncertainty because another random sample would not give exactly the same result. Random sampling helps each member of the population have a fair chance of selection. This reduces systematic bias, though it does not remove ordinary sample variation.

A larger random sample usually gives a more precise estimate because chance fluctuations tend to be smaller. It cannot fix a biased method.

For example, an online poll about school transport may miss students without internet access or students who do not choose to respond. Even thousands of responses can give a poor estimate if the respondents differ from the population in an important way.

Confidence intervals express how precise an estimate is. A narrow interval means the data give a more focused estimate. A wide interval means there is more uncertainty.

The stated confidence level describes the long run performance of the method, not a guarantee that one particular interval contains the true population value. Hypothesis tests are used when data are compared with a claim, such as whether a new revision method changes average test performance. A small result can occur by chance, especially when many comparisons are made.

Students should check the population, sampling method, sample size, graph, and wording of any claim. They should remember that association does not prove causation. Two trends can move together because of a third factor, such as temperature affecting both ice cream sales and swimming pool visits.

Key Facts

  • Descriptive statistics summarize the data collected, such as mean, median, mode, range, and standard deviation.
  • Inferential statistics use a sample to estimate or test claims about a larger population.
  • Mean = sum of all data values / number of data values.
  • Range = maximum value - minimum value.
  • Sample proportion: p-hat = x / n, where x is the number of successes and n is the sample size.
  • A confidence interval has the form estimate ± margin of error, showing a plausible range for a population value.

Vocabulary

Population
The entire group of individuals or items that a study wants to understand.
Sample
A smaller group selected from a population and used to collect data.
Descriptive statistics
Methods used to organize, display, and summarize the data that were actually measured.
Inferential statistics
Methods used to make estimates, predictions, or decisions about a population based on sample data.
Confidence interval
A range of values calculated from sample data that is likely to contain a true population value.

Common Mistakes to Avoid

  • Calling every graph inferential statistics, which is wrong because many graphs only describe the data that were collected.
  • Using a sample mean as if it must equal the population mean, which is wrong because samples vary and usually include sampling error.
  • Ignoring how the sample was chosen, which is wrong because a biased sample can lead to a misleading conclusion about the population.
  • Thinking a larger sample always fixes bias, which is wrong because a big biased sample can still represent the wrong group.

Practice Questions

  1. 1 A teacher records quiz scores for 8 students: 6, 7, 7, 8, 8, 9, 10, 10. Find the mean, median, and range of the data set.
  2. 2 In a random sample of 200 voters, 118 say they support a new school policy. Find the sample proportion p-hat and express it as a decimal and a percent.
  3. 3 A website surveys only its most active users and concludes that 92 percent of all users love a new design. Explain whether this is descriptive or inferential, and identify one reason the conclusion may be unreliable.