Sign in to save

Bookmark this page so you can find it later.

Sign in to save

Bookmark this page so you can find it later.

Data collection is the first major decision in any statistical study because the quality of the conclusions depends on the quality of the data. A strong research question helps determine whether to use surveys, experiments, observation, or existing records. Each method answers different kinds of questions and has different strengths, limits, costs, and risks of bias.

Choosing the right method makes results more reliable and easier to interpret.

Understanding Statistics: Data Collection Methods

A study begins with a precise plan for what will be measured and who will be included. The population is the full group the study is about. A sample is the smaller group actually measured.

To make a sample represent the population, every member needs a known or fair chance of selection. Random selection helps prevent researchers from choosing people who are easiest to reach or most likely to agree. A school poll taken only from one lunch table may reflect that table well, but it cannot reliably describe the whole school.

The list used to select people matters too. A list that leaves out students who are absent, homeschooled, or lack internet access creates coverage bias.

Surveys can produce misleading results even when many people respond. Wording can push people toward a certain answer. A question that combines two ideas, such as whether students like homework and tests, cannot show which idea caused the response.

Response choices must fit the question and include reasonable options. People may forget details, guess, or give answers that sound socially acceptable. This is called response bias.

Nonresponse is another problem. If students who dislike a topic ignore a survey about it, the completed answers can be unrepresentative. Anonymous surveys can sometimes reduce pressure, especially for personal topics.

Experiments are the strongest method for testing whether one factor causes a change in another. The key feature is random assignment. Participants are placed into treatment groups by chance, so the groups tend to be similar before the treatment begins.

If one group uses a new study app and another uses the usual method, a later difference in test scores has stronger evidence of being caused by the app. A control group provides a comparison point. Researchers should keep other conditions as similar as possible.

Factors that change along with the treatment, such as different teachers or unequal study time, can confound the result. Experiments still have limits. Some treatments are unsafe, unfair, too expensive, or impossible to assign randomly.

Observational studies are useful when researchers cannot control what people do. They can reveal patterns in real settings, such as a link between sleep length and attendance. A link does not prove that one variable caused the other.

Stress, family schedules, illness, or many other factors might affect both sleep and attendance. Existing records can provide very large data sets across many years, but students should check who collected the records, why they were collected, and how consistently values were recorded. A missing entry may mean no measurement was taken rather than a value of zero.

When reading any result, pay attention to the population, the selection method, the definitions used for each variable, and the possible sources of error. Those details determine how far a conclusion can honestly reach.

Key Facts

  • Survey: collects self-reported answers from a sample using questions, polls, or interviews.
  • Experiment: changes one variable on purpose to study its effect on a response variable.
  • Observation: records behavior or measurements without assigning treatments.
  • Existing records: uses data already collected, such as medical files, school records, or government databases.
  • Sample proportion: p-hat = x/n, where x is the number with a trait and n is the sample size.
  • Larger random samples usually reduce sampling variability, but they do not automatically remove bias.

Vocabulary

Population
The entire group of individuals or items that a study wants to learn about.
Sample
A smaller group selected from the population to provide data for a study.
Bias
A systematic error that makes collected data or conclusions consistently favor one outcome.
Treatment
A condition or action assigned in an experiment to test its effect on a response.
Response variable
The outcome measured in a study to see how it changes or differs across groups.

Common Mistakes to Avoid

  • Using a survey when behavior is the real target is a mistake because people may misremember, exaggerate, or give answers they think sound acceptable.
  • Calling an observational study an experiment is wrong because an experiment requires the researcher to assign treatments or conditions.
  • Assuming a large sample is automatically unbiased is wrong because a huge sample chosen badly can still give misleading results.
  • Using existing records without checking how they were collected is a mistake because missing data, old definitions, or inconsistent measurement methods can distort conclusions.

Practice Questions

  1. 1 A school wants to estimate the proportion of 1,200 students who bike to school. A random sample of 80 students finds that 18 bike to school. Find p-hat and estimate the number of students in the whole school who bike.
  2. 2 A researcher wants to test whether a new fertilizer increases plant height. She randomly assigns 60 plants to two equal groups, one with the new fertilizer and one with the old fertilizer. How many plants are in each group, and what data collection method is being used?
  3. 3 A city wants to know whether adding more streetlights reduces nighttime accidents. Explain whether a survey, experiment, observation, or existing records would be most appropriate, and justify your choice.