Skip to main content
뒤로

Data Collection and Experimental Design in Introductory Statistics

스터디 가이드 - 스마트 노트

자료에 맞춘 맞춤형 노트, 핵심 정의, 예시, 맥락을 확장해 제공합니다.

Data Collection and Experimental Design

Introduction to Statistical Studies

Statistical studies aim to collect data and use it to make informed decisions. The reliability of these decisions depends on the quality of the data collection process. Flawed data collection leads to questionable results, so understanding the design and execution of statistical studies is essential for interpreting their outcomes.

  • Key Point: The process of data collection directly impacts the validity of statistical conclusions.

  • Key Point: Before interpreting results, assess the reliability of the study's methodology.

Steps in Designing a Statistical Study

Designing a statistical study involves several systematic steps to ensure meaningful and unbiased results.

  • Identify Variables and Population: Determine the focus of the study and the population to be examined.

  • Develop a Data Collection Plan: Ensure the sample is representative of the population.

  • Collect Data: Gather information according to the plan.

  • Describe Data: Use descriptive statistics to summarize the data.

  • Interpret Data: Apply inferential statistics to make decisions about the population.

  • Identify Errors: Recognize and account for possible errors in the study.

Types of Statistical Studies

Statistical studies are generally categorized as observational studies or experiments.

  • Observational Study: The researcher observes and measures characteristics without influencing responses or changing existing conditions.

  • Experiment: The researcher deliberately applies a treatment to part of the population (treatment group) and compares responses to a control group, which may receive a placebo.

  • Example: Observing health outcomes in a population versus testing a new drug by assigning subjects to treatment and control groups.

Methods of Data Collection

Data can be collected through various methods, each suited to different research contexts.

  • Simulation: Uses mathematical or physical models to replicate real-life conditions, often with computers. Useful for studying dangerous or impractical situations (e.g., crash tests).

  • Survey: Investigates characteristics of a population, typically by asking questions via interviews, Internet, phone, or mail. Question wording is crucial to avoid bias.

Elements of Experimental Design

To ensure valid results, experiments must be carefully designed. Three key elements are control, randomization, and replication.

  • Control: Keeping certain factors constant to compare results and minimize confounding influences.

  • Randomization: Randomly assigning participants to groups to reduce bias and establish cause-and-effect relationships.

  • Replication: Repeating experiments to confirm results and identify flaws.

Experimental Terminology

Understanding key terms is essential for interpreting experimental results.

  • Control Group: Subjects treated identically to the experimental group, except for the independent variable.

  • Control Variable: Factor kept constant throughout the experiment.

  • Dependent Variable: The measured outcome in the experiment.

  • Independent Variable: The manipulated factor in the experiment.

  • Variability: Degree of dispersion in data points. Example: Data sets (3, 5, 7) vs. (0, 5, 10) both have mean 5, but the second has greater variability.

  • Lurking Variable: Unmeasured variable affecting the response variable.

  • Confounding Variables: Variables whose effects on the response variable cannot be distinguished.

  • Treatment: Specific experimental condition applied to subjects.

Placebo and Nocebo Effects

Psychological effects can influence experimental outcomes.

  • Placebo Effect: Subjects respond favorably to a fake treatment, believing it is real.

  • Nocebo Effect: Subjects respond negatively to a fake treatment, believing it is real.

  • Example: Patients experiencing side effects from a placebo due to expectations.

Blocking and Matched Pair Design

These techniques help control variability and improve experimental validity.

  • Block: Group of individuals sharing common features (e.g., age, gender, health status).

  • Example of Blocking: Subjects divided by age group, then randomly assigned to treatment or control within each group.

  • Matched Pair Design: Subjects paired by similarity; one receives treatment, the other receives control.

Randomization and Blinding

Randomization and blinding are crucial for reducing bias and ensuring validity.

  • Complete Random Experiment: Random assignment of individuals to treatments.

  • Randomized Block Experiment: Individuals sorted into blocks, then randomly assigned treatments within blocks.

  • Blinding: Minimizes placebo effect. Single Blind: Subjects unaware of treatment. Double Blind: Both subjects and experimenters unaware.

  • Hawthorne Effect: Subjects alter behavior because they know they are being studied.

Validity, Sample Size, and Sampling Frame

Ensuring validity and appropriate sample size is fundamental for reliable results.

  • Validity: Accuracy and reliability of experimental results. Improved by replication.

  • Sample Size: Number of subjects; should be representative of the population.

  • Sampling Frame: List of population elements from which the sample is drawn.

  • Sample Design: Process of selecting sample elements from the sampling frame.

Sampling Techniques

Different sampling methods are used to collect data, each with advantages and limitations.

  • Census: Measures the entire population; complete but often impractical.

  • Sampling: Measures part of the population; more practical but must be representative.

  • Sampling Error: Difference between sample results and population results.

Bias in Sampling

Bias occurs when the sample is not representative of the population, leading to inaccurate results.

  • Biased Sample: Systematic tendency in data collection or analysis that skews results.

  • Sources of Bias: Data source, collection method, estimator, analysis methods.

  • Example: Sampling only college students to represent all young adults.

Random and Simple Random Samples

Random sampling minimizes bias and ensures each population member has an equal chance of selection.

  • Random Sample: No bias introduced; selection progresses without pattern.

  • Random Error: Measurement mistake due to chance.

  • Simple Random Sample: Each element has equal probability; every possible sample of the same size has equal chance.

  • Note: Simple random sample is a specific technique; random sample is the ideal.

Stratified Sampling

Stratified sampling divides the population into strata (groups with similar characteristics) and samples from each stratum.

  • Stratum: Group sharing a characteristic (e.g., age, gender).

  • Process: Assign elements to strata, then perform simple random sampling within each stratum.

  • Advantage: Ensures representation from each segment of the population.

Cluster Sampling

Cluster sampling divides the population into clusters (naturally occurring groups), then randomly selects clusters and includes all elements within them.

  • Cluster: Group with similar characteristics (e.g., zip codes, company branches).

  • Process: Randomly select clusters, include all members in sample.

Systematic Sampling

Systematic sampling selects every nth element from a list after a random start. This method is efficient but may introduce bias if there is a pattern in the list.

  • Example: Selecting every 10th person from a roster.

  • Advantage: Simplicity and speed.

  • Disadvantage: Potential bias if list is ordered in a way that correlates with the variable of interest.

  • Additional info: Systematic sampling is best used when the population list is randomly ordered.

Convenience Sampling

Convenience sampling selects subjects based on ease of access. It is generally unreliable and biased, but sometimes unavoidable.

  • Examples: Self-selection bias, sampling whoever is available.

  • Should be avoided: Results are often not representative.

Multi-stage Sampling

Multi-stage sampling involves dividing the population into primary groups, then subdividing and sampling at each stage.

  • Process: Divide into primary groups, sample, subdivide, sample again, repeat as needed.

  • Advantage: Useful for large, complex populations.

Sampling With or Without Replacement

Sampling can be done with or without replacement, affecting the probability of selection.

  • With Replacement: Same member can be selected more than once.

  • Without Replacement: Each member can be selected only once.

Biased and Unbiased Sampling Methods

Sampling methods can be classified as biased or unbiased based on their tendency to produce representative data.

  • Biased Sampling Method: Systematically differs from the population.

  • Unbiased Sampling Method: Does not introduce systematic bias.

  • Common Biased Methods: Convenience sample, volunteer sample.

Summary Table: Sampling Methods

The following table summarizes key sampling methods, their characteristics, and potential for bias.

Sampling Method

Description

Bias Potential

Simple Random Sample

Each element has equal probability; random selection

Low

Stratified Sample

Population divided into strata; random sample from each stratum

Low

Cluster Sample

Population divided into clusters; random clusters selected, all members included

Medium

Systematic Sample

Select every nth element after random start

Medium

Convenience Sample

Subjects selected based on ease of access

High

Volunteer Sample

Subjects self-select to participate

High

Key Formulas and Concepts

  • Sample Mean: The average value in a sample.

  • Sample Variance: Measures variability in a sample.

  • Sampling Error: Difference between sample statistic and population parameter.

Conclusion

Understanding data collection and experimental design is foundational for interpreting statistical studies. Proper sampling techniques, experimental controls, and awareness of bias ensure that statistical conclusions are valid and reliable.

Pearson Logo

스터디 프렙