Sampling bias happens when the people, objects, or events chosen for a study do not represent the full population of interest. It matters because even a large sample can give a wrong answer if it is collected in a biased way. Good statistical conclusions depend more on how the sample is chosen than on sample size alone.
Recognizing common bias types helps researchers design fairer surveys, experiments, and observational studies.
Bias can enter through the sampling frame, the recruitment method, who chooses to respond, or which cases remain visible in the data. Selection bias, undercoverage, nonresponse bias, voluntary response bias, and survivorship bias each distort a sample in a different way. These problems affect estimates such as proportions, means, and comparisons between groups.
Careful random sampling, follow-up with nonresponders, and clear population definitions reduce the risk of misleading results.
Understanding Statistics: Sampling Bias Types
The first place bias can appear is before anyone receives a survey. A sampling frame is the list or system used to find possible participants. A school email list misses students without regular access to that account.
A landline phone list misses many younger households. This is undercoverage. Selection bias occurs when inclusion depends on a factor linked to the result being studied.
For example, measuring exercise habits by recruiting at a gym will tend to overstate how much exercise people do. The problem is not that gym members gave false answers. The group was built in a way that favored one kind of person.
Nonresponse bias develops after people have been selected. Some selected people reply while others do not. If nonresponders differ in a relevant way, the returned answers can lean in one direction.
A survey about homework time may receive more replies from students who are organized or strongly interested in school. Repeated reminders can help, but they do not guarantee a fair result.
Researchers should compare known features of responders and nonresponders, such as age group, grade level, or location. Large differences are warning signs that the final sample may not reflect the intended population.
Voluntary response samples are especially vulnerable because people choose themselves into the study. Online polls below news articles often attract people with strong feelings. A customer complaint form mainly reaches unhappy customers.
These sources can reveal opinions and experiences, but they cannot reliably measure how common those opinions are among everyone. Survivorship bias works differently. It focuses only on cases that remain visible.
Looking at successful businesses to find the secret of success ignores businesses that failed. Studying only repaired machines may miss machines that were thrown away. The missing cases can completely change the pattern.
Students should separate random error from systematic error. Random error causes results to vary by chance from one sample to another. A larger, well-chosen sample usually makes that variation smaller.
Systematic error pulls estimates in a similar direction each time because of the collection method. Taking thousands of votes from one biased online poll can produce a very precise estimate of the wrong quantity. When judging a study, identify the target population, the source used to reach it, the people excluded, and the reasons people may not respond.
Check whether the sample matches the population on important traits. A result deserves more trust when the method makes missing groups and self-selection less likely.
Key Facts
- Sampling bias occurs when a sample is systematically different from the population it is meant to represent.
- A random sample gives every member of the population a known chance of selection.
- Sample proportion: p-hat = x/n, where x is the number with the trait and n is the sample size.
- Bias of an estimator: bias = E(estimator) - true value.
- Increasing n reduces random sampling error, but it does not automatically remove sampling bias.
- Common sampling bias types include selection bias, undercoverage bias, nonresponse bias, voluntary response bias, and survivorship bias.
Vocabulary
- Population
- The entire group of people, objects, or events that a study wants to learn about.
- Sample
- A smaller group selected from the population to collect data from.
- Sampling frame
- The list or method used to identify members of the population who can be selected for the sample.
- Undercoverage bias
- A bias that occurs when some groups in the population are left out or are less likely to be included in the sample.
- Nonresponse bias
- A bias that occurs when people who do not respond differ in important ways from people who do respond.
Common Mistakes to Avoid
- Assuming a large sample is automatically unbiased. A large biased sample can still give a very precise but wrong estimate if the same groups are overrepresented or excluded.
- Confusing voluntary response with random sampling. People who choose to respond often have stronger opinions or different experiences than the overall population.
- Ignoring the sampling frame. If the frame leaves out people without phones, internet access, addresses, or membership in a list, the resulting sample can suffer from undercoverage.
- Treating nonresponse as harmless missing data. If nonresponders differ from responders on the topic being studied, estimates such as means and proportions can be distorted.
Practice Questions
- 1 A school has 1200 students, but a survey about cafeteria food is emailed only to the 800 students who joined the school app. If 160 students respond and 112 say they like the food, what is the sample proportion p-hat, and what type of sampling bias might occur?
- 2 A city survey calls 1000 randomly selected phone numbers about public transportation. Only 420 people answer, and 252 of them support increasing bus service. What percent of respondents support the plan, and what bias could occur if nonresponders have different commuting habits?
- 3 A magazine asks readers to vote online on whether homework should be banned. Explain why the result may not represent all students, even if 50,000 people respond.