뒤로Chapter 1: Data Collection – Introductory Statistics Study Guide
스터디 가이드 - 스마트 노트
자료에 맞춘 맞춤형 노트, 핵심 정의, 예시, 맥락을 확장해 제공합니다.
Data Collection
Big Idea of Statistics
Statistics is the science of using data to answer questions, understand variability, and make conclusions with a measure of confidence. The process involves collecting, organizing, summarizing, and analyzing information to draw meaningful conclusions.
1.1 The Practice of Statistics
Key Terms and Definitions
Statistics: The science of collecting, organizing, summarizing, and analyzing information to draw conclusions or answer questions. It also measures the reliability of those conclusions.
Data: Facts or information used to draw a conclusion or make a decision. Data describe characteristics of individuals or objects.
Population: The entire group of individuals to be studied.
Individual: A single member of the population.
Sample: A subset of the population being studied.
Statistic: A numerical summary based on a sample.
Parameter: A numerical summary of a population.
Descriptive statistics: Methods for organizing and summarizing data using numerical summaries, tables, and graphs.
Inferential statistics: Methods that use sample results to make conclusions about the population and measure reliability.
Memory Trick: Parameter = Population; Statistic = Sample.
The Process of Statistics
Identify the research objective: Clearly state the question and identify the population.
Collect the data: Usually from a sample; poor data collection can invalidate conclusions.
Describe the data: Use descriptive statistics to summarize and understand the data.
Perform inference: Extend sample results to the population and report the reliability of the result.
Classifying Variables
Qualitative (Categorical) Variable: Classifies individuals based on an attribute or characteristic. Arithmetic operations are not meaningful (e.g., color, type, category).
Quantitative Variable: Provides numerical measures of individuals. Arithmetic operations are meaningful.
Discrete: Quantitative; finite or countable values (e.g., number of students).
Continuous: Quantitative; infinitely many possible values, measured on a continuum (e.g., height, time).
Decision Rule: Is the variable a category or a number? If a number, is it counted (discrete) or measured (continuous)?
Levels of Measurement
Level | What It Allows | Order Matters? | Zero Means "None"? | Example |
|---|---|---|---|---|
Nominal | Names, labels, categories | No | Not applicable | Eye color |
Ordinal | Categories can be ranked | Yes | Not necessarily | Class rank |
Interval | Rank + meaningful differences | Yes | No | Temperature (°C) |
Ratio | Interval properties + meaningful ratios | Yes | Yes | Height, weight |
Mnemonic: N-O-I-R: Nominal → Ordinal → Interval → Ratio (each adds more information).
1.2 Observational Studies vs. Designed Experiments
Comparison Table
Feature | Observational Study | Designed Experiment |
|---|---|---|
Researcher Action | Observes/measures without influencing variables | Randomly assigns groups, manipulates variables, controls others |
Main Conclusion | Association only | Can support cause-and-effect (if well designed) |
Common Concern | Lurking variables | Confounding variables |
Explanatory vs. Response Variables
Explanatory Variable: Used to explain or predict changes in another variable.
Response Variable: The outcome being measured.
Example: In a study of music exposure and IQ, music exposure is explanatory, IQ is the response.
Confounding and Lurking Variables
Confounding Variable: Effects of two or more explanatory variables are not separated; individual effects cannot be distinguished (common in experiments).
Lurking Variable: An explanatory variable not considered in the study that affects the response and is typically related to a considered explanatory variable (common in observational studies).
Important: Observational studies cannot establish causation—only association.
Types of Observational Studies
Type | Time Direction | How It Works | Memory Cue |
|---|---|---|---|
Cross-sectional | Now / short time | Collects information at a specific point in time | Snapshot |
Case-control | Backward | Retrospective; compares people with and without a characteristic | Look back |
Cohort | Forward | Follows a group over time, recording characteristics | Follow forward |
Other Data Collection Terms
Census: A list of all individuals in a population with certain characteristics recorded for each.
Web scraping / Data mining: Extracting data from the Internet. Note: May raise ethical issues if done without permission.
1.3–1.4 Sampling Methods
Why Randomness?
Random sampling uses chance to select individuals, ensuring that the sample is representative of the population. Convenience samples can lead to meaningless results due to bias.
Sampling Methods Comparison
Method | How to Recognize | What Gets Randomly Selected? | Key Phrase |
|---|---|---|---|
Simple Random | Every possible sample of size n is equally likely | Individuals | Pure chance |
Stratified | Population split into homogeneous groups (strata); SRS from each stratum | Individuals within every stratum | Some from every group |
Systematic | Choose every kth individual after a random start | Individuals at regular intervals | Every kth |
Cluster | Divide into groups; randomly select some groups; survey all in selected groups | Whole groups | All from some groups |
Convenience | Use individuals easiest to obtain; not random | Whoever is easy to reach | Easy, not random |
Stratified vs. Cluster: Stratified = some people from every group; Cluster = everyone from some groups.
Systematic Sampling Formula
Approximate population size
Choose desired sample size
Compute and round down
Randomly choose from 1 to
Select individuals at positions
Sampling Method Decision Tree
If people are chosen because they are easiest to reach: Convenience
If every possible sample of size n is equally likely: Simple Random
If every kth individual is selected after a random start: Systematic
If an SRS is taken from each homogeneous group: Stratified
If some groups are randomly chosen and all people in those groups are surveyed: Cluster
1.5 Bias in Sampling
Types of Bias
Source | Definition | How to Spot / Fix |
|---|---|---|
Sampling bias | Sampling technique favors one part of the population | Look for undercoverage; ensure all segments are represented |
Nonresponse bias | Selected individuals who do not respond differ from those who do | Use callbacks or incentives to improve response |
Response bias | Survey answers do not reflect true feelings | Check for interviewer error, misrepresented answers, question wording/order |
Data-entry error | Incorrect input makes results unrepresentative | Carefully check data for accuracy |
Sampling vs. Nonsampling Errors
Nonsampling errors: Sampling bias, nonresponse bias, response bias, and data-entry error. These can occur even in a census.
Sampling error: Occurs because a sample gives incomplete information about a population. It is the natural result of using a sample instead of the whole population.
Bias Recognition Chart
Situation | Likely Issue |
|---|---|
Part of the population has too little representation | Undercoverage → sampling bias |
Selected people choose not to answer and differ from responders | Nonresponse bias |
Question wording or interviewer influences the answer | Response bias |
A value is typed incorrectly | Data-entry error |
Sample differs from population simply because only part of the population was observed | Sampling error |
Chapter 1 One-Page Review
ALL individuals: Population
SOME individuals: Sample
Number describing sample: Statistic
Number describing population: Parameter
Category / label: Qualitative
Numerical measurement: Quantitative
Countable numerical values: Discrete
Any value on an interval: Continuous
Names only: Nominal
Ranked categories: Ordinal
Meaningful differences, zero ≠ none: Interval
True zero + meaningful ratios: Ratio
Observe only: Observational study → association
Manipulate + randomly assign: Designed experiment
Snapshot: Cross-sectional
Look backward: Case-control
Follow forward: Cohort
Some from every group: Stratified
Everyone from some groups: Cluster
Every kth: Systematic
Easiest people to reach: Convenience
Self-Check: Essential Skills
Identify population, sample, individual, statistic, and parameter.
Decide whether a variable is qualitative or quantitative.
Classify a quantitative variable as discrete or continuous.
Identify nominal, ordinal, interval, and ratio levels of measurement.
Distinguish an observational study from a designed experiment.
Identify explanatory and response variables.
Explain why observational studies show association rather than causation.
Distinguish cross-sectional, case-control, and cohort studies.
Recognize simple random, stratified, systematic, cluster, and convenience samples.
Identify sampling, nonresponse, and response bias, plus data-entry and sampling error.
Mini Practice Examples
A college randomly selects 50 students from each class year (freshman, sophomore, junior, senior). Sampling method: Stratified
A researcher surveys everyone in 8 randomly selected classrooms. Sampling method: Cluster
A survey asks people to rate service as poor, fair, good, or excellent. Level of measurement: Ordinal
A researcher records whether people who already exercise more tend to have lower resting heart rates, without assigning exercise plans. Study type: Observational study
A survey about campus dining is emailed to 500 randomly selected students, but students who dislike the dining hall are much more likely to respond. Type of bias: Nonresponse bias