뒤로Chapter 1: Introduction to Statistics – Study Guide
스터디 가이드 - 스마트 노트
자료에 맞춘 맞춤형 노트, 핵심 정의, 예시, 맥락을 확장해 제공합니다.
Introduction to Statistics
Statistical and Critical Thinking
Statistics is the science of collecting, analyzing, interpreting, and presenting data to draw meaningful conclusions. Understanding the foundational concepts and the process of statistical analysis is essential for making informed decisions based on data.
Data: Collections of observations, such as measurements, genders, or survey responses.
Statistics: The science of planning studies and experiments; obtaining, organizing, summarizing, presenting, analyzing, and interpreting data to draw conclusions.
Population: The complete collection of all measurements or data being considered, about which inferences are to be made.
Census: Data collection from every member of a population.
Sample: A subcollection of members selected from a population.
The Statistical Process: Prepare, Analyze, Conclude
Stage | Key Components & Questions |
|---|---|
Prepare |
|
Analyze |
|
Conclude |
|
Sampling Bias & Potential Pitfalls
Voluntary Response Sample: Respondents decide whether to participate (e.g., internet polls), often leading to bias.
Reported vs. Measured Data: Physical measurements are preferred over self-reported values to reduce bias.
Loaded Questions & Question Order: Wording or sequence can influence responses.
Nonresponse: When subjects refuse or are unavailable, increasing bias.
Misleading Percentages: Percentages can be misused, especially when changes over 100% are claimed inappropriately.
Types of Data
Parameters vs. Statistics
Parameter: A numerical measurement describing a characteristic of a population.
Statistic: A numerical measurement describing a characteristic of a sample.
Types of Data
Data Type | Definition | Example |
|---|---|---|
Quantitative (Numerical) | Numbers representing counts or measurements. | Weights, ages |
Categorical (Qualitative/Attribute) | Names, labels, or categories (not numerical counts). | Gender, shirt numbers |
Discrete Data: Quantitative data with a finite or countable number of values (e.g., number of coin tosses).
Continuous Data: Quantitative data with infinitely many possible values, not countable (e.g., distances).
Levels of Measurement
Level | Characteristics | Example |
|---|---|---|
Nominal | Categories, names, or labels only; cannot be ordered. | Yes/No/Undecided |
Ordinal | Can be ordered, but differences are meaningless. | Course grades (A, B, C, D, F) |
Interval | Ordered, meaningful differences, but no natural zero. | Years, temperatures (°F/°C) |
Ratio | Ordered, meaningful differences, and a natural zero; ratios are valid. | Time, distances, weights |
Big Data & Missing Data
Big Data: Extremely large and complex data sets requiring advanced computational methods.
Data Science: An interdisciplinary field combining statistics, computer science, and domain expertise.
Missing Completely at Random (MCAR): The probability of missing data is unrelated to any values in the data set.
Missing Not at Random (MNAR): The missingness is related to the value itself.
Handling Missing Data: Options include deleting cases or imputing (estimating) missing values.
Collecting Sample Data
Observational Studies vs. Experiments
Observational Study: Observes and measures characteristics without modifying subjects.
Experiment: Applies a treatment and observes its effects on subjects.
Key Experimental Concepts
Gold Standard: Randomization with treatment and placebo groups.
Placebo: An inactive treatment used as a control.
Replication: Repeating the experiment on enough subjects to observe effects.
Blinding: Subjects do not know if they receive the treatment or placebo.
Double-Blind: Neither subjects nor experimenters know group assignments.
Randomization: Assigning subjects to groups by chance.
Sampling Methods
Method | Description |
|---|---|
Simple Random Sample (SRS) | Every possible sample of size n has an equal chance of being chosen. |
Systematic Sampling | Select a starting point, then pick every k-th element. |
Convenience Sampling | Use data that are easy to obtain; high risk of bias. |
Stratified Sampling | Divide population into subgroups (strata) and sample from each. |
Cluster Sampling | Divide population into clusters, randomly select clusters, and sample all members in selected clusters. |
Multistage Sampling | Combine multiple sampling methods in stages. |
Types of Observational Studies
Cross-Sectional Study: Data collected at a single point in time.
Retrospective (Case-Control) Study: Data collected from past records or interviews.
Prospective (Longitudinal/Cohort) Study: Data collected in the future from groups sharing common factors.
Controlling Variable Effects & Experimental Design
Confounding: When the effect of one variable cannot be distinguished from another.
Completely Randomized Design: Subjects assigned to groups purely by chance.
Randomized Block Design: Subjects grouped by similar characteristics, then randomly assigned treatments within blocks.
Matched Pairs Design: Pairs of subjects matched closely, each receiving different treatments.
Rigorously Controlled Design: Groups are balanced on all key factors.
Errors in Statistics
Sampling Error (Random Sampling Error): The difference between a sample result and the true population result due to chance.
Nonsampling Error: Errors from human mistakes, such as data entry or biased survey wording.
Nonrandom Sampling Error: Errors from using non-random sampling methods (e.g., convenience sampling).
Example: Identifying Data Types
Example 1: The number of cars passing through an intersection in an hour is discrete quantitative data.
Example 2: The temperature in degrees Celsius is interval level data (no true zero).
Example 3: Survey responses of "Yes," "No," or "Undecided" are nominal data.