Skip to main content
뒤로

Introductory Statistics: Data Collection, Sampling, and Study Design

스터디 가이드 - 스마트 노트

자료에 맞춘 맞춤형 노트, 핵심 정의, 예시, 맥락을 확장해 제공합니다.

Introduction to Statistics

What is the Important Goal of Statistics?

Statistics is the science of collecting, analyzing, interpreting, and presenting data. The primary goal is to make informed decisions and draw conclusions about populations based on sample data.

  • Key Point: Statistics helps us understand variability and uncertainty in data.

  • Example: Estimating the average height of college students using a sample.

Populations, Samples, and Census

Definitions and Differences

Understanding the basic units of statistical study is essential for proper data analysis.

  • Population: The entire group of individuals or items of interest.

  • Census: A study that collects data from every member of the population.

  • Sample: A subset of the population selected for analysis.

  • Example: Surveying 100 students (sample) from a university (population).

Parameters vs. Statistics

Key Differences

Parameters and statistics are measures that describe populations and samples, respectively.

  • Parameter: A numerical summary of a population (e.g., population mean ).

  • Statistic: A numerical summary of a sample (e.g., sample mean ).

  • Example: is the average height of all students; is the average height in the sample.

Descriptive vs. Inferential Statistics

Definitions and Applications

Statistics is divided into two main branches: descriptive and inferential.

  • Descriptive Statistics: Methods for summarizing and organizing data (e.g., tables, graphs, averages).

  • Inferential Statistics: Methods for making predictions or inferences about a population based on sample data.

  • Example: Calculating the mean and standard deviation (descriptive); estimating population parameters (inferential).

Statistical Thinking vs. Practical Thinking

Comparison

Statistical thinking involves understanding data variability and uncertainty, while practical thinking focuses on real-world application and feasibility.

  • Statistical Thinking: Emphasizes randomness, probability, and data-driven decisions.

  • Practical Thinking: Considers logistics, cost, and practicality of data collection.

Potential Pitfalls in Statistical Analysis

Common Issues

Several pitfalls can affect the validity of statistical analysis.

  • Sampling Bias: Systematic error in sample selection.

  • Measurement Error: Inaccurate data collection.

  • Confounding Variables: Factors that affect the outcome but are not accounted for.

Types of Data and Levels of Measurement

Classification of Data

Data can be classified by type and level of measurement, which affects analysis methods.

  • Types of Data:

    • Qualitative (Categorical): Describes qualities or categories (e.g., gender, color).

    • Quantitative (Numerical): Measures quantities (e.g., height, weight).

  • Levels of Measurement:

    • Nominal: Categories without order (e.g., colors).

    • Ordinal: Categories with order (e.g., rankings).

    • Interval: Ordered, equal intervals, no true zero (e.g., temperature in Celsius).

    • Ratio: Ordered, equal intervals, true zero (e.g., height, weight).

Techniques for Collecting Sample Data

Sampling Methods

Choosing the right sampling method is crucial for obtaining representative data.

  • Random Sampling: Every member has an equal chance of selection.

  • Simple Random Sampling: Randomly select individuals from the population.

  • Systematic Sampling: Select every nth individual from a list.

  • Cluster Sampling: Divide population into clusters, randomly select clusters, sample all within clusters.

  • Stratified Sampling: Divide population into strata, randomly sample from each stratum.

  • Convenience Sampling: Select individuals who are easiest to reach.

  • Volunteer Sampling: Individuals self-select to participate.

Sources of Sampling Bias

Types of Bias

Bias can distort results and lead to incorrect conclusions.

  • Undercoverage: Some groups are not represented in the sample.

  • Nonresponse: Selected individuals do not respond.

  • Response Bias: Respondents provide inaccurate answers.

Sampling Design

Importance and Structure

Sampling design refers to the plan for selecting a sample from the population.

  • Key Point: Good design minimizes bias and maximizes representativeness.

Variables in Statistical Studies

Response vs. Explanatory Variables

Variables are classified based on their role in the study.

  • Response Variable: The outcome or dependent variable.

  • Explanatory Variable: The factor or independent variable that may influence the response.

  • Example: In a study of medication, the response variable is improvement; the explanatory variable is dosage.

Types of Studies: Observational vs. Experimental

Study Designs

Statistical studies can be observational or experimental.

  • Observational Study: Researchers observe subjects without intervention.

  • Experimental Study: Researchers apply treatments and observe effects.

Experimental Units and Treatments

Definitions

In experiments, the subjects and interventions are clearly defined.

  • Experimental Unit: The subject or object receiving a treatment.

  • Treatment: The specific condition applied to experimental units.

Randomization in Experiments

Purpose and Methods

Randomization ensures unbiased assignment of treatments.

  • Key Point: Randomization reduces confounding and increases validity.

Experimental Design Concepts

Blind, Double Blind, Replication, Block

Several concepts improve the quality of experimental studies.

  • Blind Study: Subjects do not know which treatment they receive.

  • Double Blind Study: Neither subjects nor experimenters know treatment assignments.

  • Replication: Repeating the experiment to confirm results.

  • Block: Grouping subjects by similar characteristics.

Randomized Block Design (RBD) vs. Completely Randomized Design (CRD)

Comparison of Designs

Experimental designs differ in how subjects are assigned to treatments.

  • Randomized Block Design (RBD): Subjects are grouped into blocks, then randomly assigned treatments within each block.

  • Completely Randomized Design (CRD): Subjects are randomly assigned to treatments without blocking.

Types of Errors in Sampling

Error Classification

Errors can arise from various sources during sample collection.

  • Sampling Error: Difference between sample statistic and population parameter due to random chance.

  • Nonsampling Error: Errors not related to sampling, such as measurement or data entry mistakes.

  • Nonrandom Sampling Error: Errors due to nonrandom selection of samples.

Summary Table: Sampling Methods and Biases

The following table summarizes sampling methods and potential biases:

Sampling Method

Description

Potential Bias

Simple Random Sampling

Each member has equal chance

Low if properly conducted

Systematic Sampling

Select every nth member

Possible if list has patterns

Cluster Sampling

Randomly select clusters, sample all within

High if clusters are not representative

Stratified Sampling

Divide into strata, sample from each

Low if strata are well-defined

Convenience Sampling

Sample easiest to reach

High due to lack of randomness

Volunteer Sampling

Self-selected participants

High due to self-selection bias

Additional info: Academic context and examples were added to clarify definitions and concepts for self-contained study notes.

Pearson Logo

스터디 프렙