Skip to main content
뒤로

Chapter 1: Data Collection – Introductory Statistics Study Guide

스터디 가이드 - 스마트 노트

자료에 맞춘 맞춤형 노트, 핵심 정의, 예시, 맥락을 확장해 제공합니다.

Data Collection

Big Idea of Statistics

Statistics is the science of using data to answer questions, understand variability, and make conclusions with a measure of confidence. The process involves collecting, organizing, summarizing, and analyzing information to draw meaningful conclusions.

1.1 The Practice of Statistics

Key Terms and Definitions

  • Statistics: The science of collecting, organizing, summarizing, and analyzing information to draw conclusions or answer questions. It also measures the reliability of those conclusions.

  • Data: Facts or information used to draw a conclusion or make a decision. Data describe characteristics of individuals or objects.

  • Population: The entire group of individuals to be studied.

  • Individual: A single member of the population.

  • Sample: A subset of the population being studied.

  • Statistic: A numerical summary based on a sample.

  • Parameter: A numerical summary of a population.

  • Descriptive statistics: Methods for organizing and summarizing data using numerical summaries, tables, and graphs.

  • Inferential statistics: Methods that use sample results to make conclusions about the population and measure reliability.

Memory Trick: Parameter = Population; Statistic = Sample.

The Process of Statistics

  1. Identify the research objective: Clearly state the question and identify the population.

  2. Collect the data: Usually from a sample; poor data collection can invalidate conclusions.

  3. Describe the data: Use descriptive statistics to summarize and understand the data.

  4. Perform inference: Extend sample results to the population and report the reliability of the result.

Classifying Variables

  • Qualitative (Categorical) Variable: Classifies individuals based on an attribute or characteristic. Arithmetic operations are not meaningful (e.g., color, type, category).

  • Quantitative Variable: Provides numerical measures of individuals. Arithmetic operations are meaningful.

    • Discrete: Quantitative; finite or countable values (e.g., number of students).

    • Continuous: Quantitative; infinitely many possible values, measured on a continuum (e.g., height, time).

Decision Rule: Is the variable a category or a number? If a number, is it counted (discrete) or measured (continuous)?

Levels of Measurement

Level

What It Allows

Order Matters?

Zero Means "None"?

Example

Nominal

Names, labels, categories

No

Not applicable

Eye color

Ordinal

Categories can be ranked

Yes

Not necessarily

Class rank

Interval

Rank + meaningful differences

Yes

No

Temperature (°C)

Ratio

Interval properties + meaningful ratios

Yes

Yes

Height, weight

Mnemonic: N-O-I-R: Nominal → Ordinal → Interval → Ratio (each adds more information).

1.2 Observational Studies vs. Designed Experiments

Comparison Table

Feature

Observational Study

Designed Experiment

Researcher Action

Observes/measures without influencing variables

Randomly assigns groups, manipulates variables, controls others

Main Conclusion

Association only

Can support cause-and-effect (if well designed)

Common Concern

Lurking variables

Confounding variables

Explanatory vs. Response Variables

  • Explanatory Variable: Used to explain or predict changes in another variable.

  • Response Variable: The outcome being measured.

  • Example: In a study of music exposure and IQ, music exposure is explanatory, IQ is the response.

Confounding and Lurking Variables

  • Confounding Variable: Effects of two or more explanatory variables are not separated; individual effects cannot be distinguished (common in experiments).

  • Lurking Variable: An explanatory variable not considered in the study that affects the response and is typically related to a considered explanatory variable (common in observational studies).

Important: Observational studies cannot establish causation—only association.

Types of Observational Studies

Type

Time Direction

How It Works

Memory Cue

Cross-sectional

Now / short time

Collects information at a specific point in time

Snapshot

Case-control

Backward

Retrospective; compares people with and without a characteristic

Look back

Cohort

Forward

Follows a group over time, recording characteristics

Follow forward

Other Data Collection Terms

  • Census: A list of all individuals in a population with certain characteristics recorded for each.

  • Web scraping / Data mining: Extracting data from the Internet. Note: May raise ethical issues if done without permission.

1.3–1.4 Sampling Methods

Why Randomness?

Random sampling uses chance to select individuals, ensuring that the sample is representative of the population. Convenience samples can lead to meaningless results due to bias.

Sampling Methods Comparison

Method

How to Recognize

What Gets Randomly Selected?

Key Phrase

Simple Random

Every possible sample of size n is equally likely

Individuals

Pure chance

Stratified

Population split into homogeneous groups (strata); SRS from each stratum

Individuals within every stratum

Some from every group

Systematic

Choose every kth individual after a random start

Individuals at regular intervals

Every kth

Cluster

Divide into groups; randomly select some groups; survey all in selected groups

Whole groups

All from some groups

Convenience

Use individuals easiest to obtain; not random

Whoever is easy to reach

Easy, not random

Stratified vs. Cluster: Stratified = some people from every group; Cluster = everyone from some groups.

Systematic Sampling Formula

  • Approximate population size

  • Choose desired sample size

  • Compute and round down

  • Randomly choose from 1 to

  • Select individuals at positions

Sampling Method Decision Tree

  • If people are chosen because they are easiest to reach: Convenience

  • If every possible sample of size n is equally likely: Simple Random

  • If every kth individual is selected after a random start: Systematic

  • If an SRS is taken from each homogeneous group: Stratified

  • If some groups are randomly chosen and all people in those groups are surveyed: Cluster

1.5 Bias in Sampling

Types of Bias

Source

Definition

How to Spot / Fix

Sampling bias

Sampling technique favors one part of the population

Look for undercoverage; ensure all segments are represented

Nonresponse bias

Selected individuals who do not respond differ from those who do

Use callbacks or incentives to improve response

Response bias

Survey answers do not reflect true feelings

Check for interviewer error, misrepresented answers, question wording/order

Data-entry error

Incorrect input makes results unrepresentative

Carefully check data for accuracy

Sampling vs. Nonsampling Errors

  • Nonsampling errors: Sampling bias, nonresponse bias, response bias, and data-entry error. These can occur even in a census.

  • Sampling error: Occurs because a sample gives incomplete information about a population. It is the natural result of using a sample instead of the whole population.

Bias Recognition Chart

Situation

Likely Issue

Part of the population has too little representation

Undercoverage → sampling bias

Selected people choose not to answer and differ from responders

Nonresponse bias

Question wording or interviewer influences the answer

Response bias

A value is typed incorrectly

Data-entry error

Sample differs from population simply because only part of the population was observed

Sampling error

Chapter 1 One-Page Review

  • ALL individuals: Population

  • SOME individuals: Sample

  • Number describing sample: Statistic

  • Number describing population: Parameter

  • Category / label: Qualitative

  • Numerical measurement: Quantitative

  • Countable numerical values: Discrete

  • Any value on an interval: Continuous

  • Names only: Nominal

  • Ranked categories: Ordinal

  • Meaningful differences, zero ≠ none: Interval

  • True zero + meaningful ratios: Ratio

  • Observe only: Observational study → association

  • Manipulate + randomly assign: Designed experiment

  • Snapshot: Cross-sectional

  • Look backward: Case-control

  • Follow forward: Cohort

  • Some from every group: Stratified

  • Everyone from some groups: Cluster

  • Every kth: Systematic

  • Easiest people to reach: Convenience

Self-Check: Essential Skills

  • Identify population, sample, individual, statistic, and parameter.

  • Decide whether a variable is qualitative or quantitative.

  • Classify a quantitative variable as discrete or continuous.

  • Identify nominal, ordinal, interval, and ratio levels of measurement.

  • Distinguish an observational study from a designed experiment.

  • Identify explanatory and response variables.

  • Explain why observational studies show association rather than causation.

  • Distinguish cross-sectional, case-control, and cohort studies.

  • Recognize simple random, stratified, systematic, cluster, and convenience samples.

  • Identify sampling, nonresponse, and response bias, plus data-entry and sampling error.

Mini Practice Examples

  1. A college randomly selects 50 students from each class year (freshman, sophomore, junior, senior). Sampling method: Stratified

  2. A researcher surveys everyone in 8 randomly selected classrooms. Sampling method: Cluster

  3. A survey asks people to rate service as poor, fair, good, or excellent. Level of measurement: Ordinal

  4. A researcher records whether people who already exercise more tend to have lower resting heart rates, without assigning exercise plans. Study type: Observational study

  5. A survey about campus dining is emailed to 500 randomly selected students, but students who dislike the dining hall are much more likely to respond. Type of bias: Nonresponse bias

Pearson Logo

스터디 프렙