Sign in to save

Bookmark this page so you can find it later.

Sign in to save

Bookmark this page so you can find it later.

A census collects data from every member of a population, while a sample collects data from only part of the population. This distinction matters because real studies often face limits in time, money, access, and accuracy. A census can sound perfect because it includes everyone, but it can still be wrong if people are missed, counted twice, or measured poorly.

A sample can give excellent information when it is chosen carefully and represents the population well.

The main trade-off is between coverage and practicality. A census tries to remove sampling error by measuring all units, but it can increase cost, delay results, and create more chances for nonresponse or recording errors. A sample is faster and cheaper, and statistics such as sample means and proportions can estimate population values with measurable uncertainty.

In many situations, a well-designed random sample is more useful than a rushed or flawed census.

Understanding Statistics: Census vs Sample

Before collecting data, a researcher needs a clear list or method for reaching the group of interest. This is called a sampling frame. A school survey might use the current enrollment list.

A study of local shoppers might use customer records, street addresses, or people leaving stores. No frame is complete in every case. Students who are absent, families without stable addresses, or people who do not use a certain app may be left out.

This creates coverage error. A large survey cannot represent people who had no chance of being selected.

Random selection gives each member a known chance of entering the study. It does not mean choosing people who seem typical. It means using a process such as random numbers, a lottery, or computer selection.

Random selection protects against the researcher quietly choosing convenient answers. Sometimes researchers divide the population into important groups first, then randomly choose within each group. This is called stratified sampling.

For example, a school may sample students from each grade level rather than taking most responses from one large grade. This can make estimates more precise when groups differ in useful ways.

The result from a sample is an estimate, not an exact answer for the whole group. If repeated random samples were taken, their results would vary. This natural variation is why reports often include a margin of error.

A margin of error describes the likely amount of random sampling variation under stated conditions. It does not cover every possible mistake.

It cannot repair misleading wording, missing groups, dishonest responses, or a low response rate. If only a small fraction of selected people reply, the respondents may differ from the nonrespondents in ways that matter.

Censuses have special problems when the measurement is difficult or changes quickly. Counting a country's residents is possible, but people move, forms are misunderstood, and some households do not respond. In manufacturing, testing every item can even destroy the product.

A factory that makes safety matches cannot strike every match to test it. Instead, it tests selected matches and uses the findings to judge the batch. Medical studies, election polls, quality checks, and school feedback forms all depend on similar decisions about whom to measure and how to measure them.

When reading statistics, pay attention to the target group, the selection method, the number of responses, and the exact wording used. A survey of followers of one social media account says little about all teenagers. A question such as whether students support longer lunches may get different results if it mentions a shorter school day.

Bigger samples usually reduce random variation, but they do not remove systematic bias. Good statistics comes from careful planning before numbers are collected. The most important habit is to ask whether the data collection process gave a fair view of the group being studied.

Key Facts

  • Population = the entire group you want to study.
  • Census = data collected from every member of the population.
  • Sample = data collected from a subset of the population.
  • Sample proportion: p-hat = x/n, where x is the number with the trait and n is the sample size.
  • Sampling error usually decreases as sample size increases, roughly proportional to 1/sqrt(n).
  • A biased sample can give worse results than a smaller random sample because increasing n does not fix bias.

Vocabulary

Population
The complete group of individuals, objects, or measurements that a study is trying to describe.
Census
A data collection method that attempts to measure every member of the population.
Sample
A smaller group selected from the population to provide information about the whole population.
Bias
A systematic error that pushes results away from the true population value.
Random sample
A sample chosen by a chance process so that every member of the population has a known chance of selection.

Common Mistakes to Avoid

  • Assuming a census is always accurate is wrong because nonresponse, duplicate counting, and measurement mistakes can still distort the results.
  • Using a convenient sample and treating it as representative is wrong because easy-to-reach people may differ from the full population in important ways.
  • Thinking a larger sample automatically removes bias is wrong because bias comes from the selection or measurement method, not just from having too few people.
  • Confusing population size with sample quality is wrong because a small random sample can estimate a huge population well if the sampling method is sound.

Practice Questions

  1. 1 A school has 1,200 students. A survey randomly selects 150 students and finds that 96 prefer online grade reports. What is the sample proportion p-hat, and how many students in the whole school would you estimate prefer online grade reports?
  2. 2 A city wants to estimate support for a new park plan. A census would cost 4perhouseholdfor80,000households.Arandomsampleof1,000householdswouldcost4 per household for 80,000 households. A random sample of 1,000 households would cost 7 per household. How much money is saved by using the sample instead of the census?
  3. 3 A restaurant asks only customers who follow its social media page to rate a new menu item. Explain whether this is closer to a census or a sample, and identify one reason the results may be biased.