뒤로Chapter 1: Introduction to Data – Study Notes for Introductory Statistics
스터디 가이드 - 스마트 노트
자료에 맞춘 맞춤형 노트, 핵심 정의, 예시, 맥락을 확장해 제공합니다.
Introduction to Data
Section 1.1: What are Data?
Statistics is the foundational science for collecting, organizing, summarizing, and analyzing data to answer questions and draw meaningful conclusions. Understanding the nature and context of data is essential for accurate interpretation and analysis.
Statistics (Definition): The science of collecting, organizing, summarizing, and analyzing data to answer questions and/or draw conclusions.
Variation: Refers to differences or changes in an item or between items. Recognizing variation is crucial for understanding data patterns.
Data: Observations gathered to draw conclusions. Data must always be interpreted within its context to avoid misrepresentation.
Example: The numbers 90, 95, 80, 75, 92 could represent test scores, weights, speeds, or millions of people. Without context, their meaning is ambiguous.

Section 1.2: Classifying and Storing Data
Data can be sourced from entire populations or from samples. Understanding the difference is key to proper statistical inference and study design.
Population: The complete set of all data values of interest. Populations are often difficult to measure in their entirety.
Sample: A subset of the population, used to represent the population. Samples are easier to obtain and analyze.
Example: To determine the most common hair color among students at Kirkwood, surveying 100 students forms a sample, while all students at Kirkwood constitute the population.
Variable: A characteristic of people or things that can vary, such as hair color, height, or GPA.
Types of Variables:
Categorical: Describes a quality or class (e.g., letter grade, type of pet).
Numerical: Describes a quantity or measurement (e.g., height, GPA, hours worked).
Classification Example:
Height of a bridge – Numerical
GPA – Numerical
Letter grade in a class – Categorical
Hours worked each week – Numerical
Type of pet owned – Categorical
Model of a car – Categorical
Temperature outside – Numerical
Zip Code – Categorical
Section 1.3: Investigating Data
The data cycle is a systematic approach to answering questions using data. Each step is essential for drawing valid conclusions.
Ask Questions: Formulate detailed and specific questions to guide the investigation.
Consider Data: Identify and collect relevant data to answer the question.
Analyze Data: Use statistical methods, often starting with data visualization, to explore the data.
Interpret Data: Draw conclusions based on the analysis.
Example Questions: What is the typical age of students? Do older students have higher GPAs? Do students in mathematics study more?
Section 1.4: Organizing Categorical Data
Proper organization of categorical data is crucial for meaningful comparisons and interpretations.
Comparing Groups: Groups should be similar for valid comparisons. Percentages or rates are often more informative than raw counts.
Incomplete Data: Raw counts (e.g., number of injuries in sports) may not accurately reflect risk without considering exposure rates.
Injury Rate Formula:
Comparing injury rates, rather than just counts, can change the ranking of "most dangerous" sports.
Two-Way Tables: Used to organize and analyze the relationship between two categorical variables. They allow for calculation of joint, marginal, and conditional percentages.
Example: Survey responses about allowing a person to make a speech or teach, based on their beliefs, can be organized in a two-way table to analyze patterns in opinions.
Section 1.5: Collecting Data to Understand Causality
Understanding causality requires careful study design. Not all studies can establish cause-and-effect relationships.
Treatment Variable: The possible cause in a study.
Response Variable: The possible effect being measured.
Causality: Determining if the treatment variable causes a change in the response variable.
Types of Studies:
Anecdotal Evidence: Based on individual stories; not reliable for establishing causality.
Observational Study: Subjects are not assigned by the researcher; can show association but not causality.
Experiment: Researcher assigns subjects to groups; only experiments can establish causality.
Elements of a Study:
Treatment Group: Receives the treatment or characteristic of interest.
Comparison (Control) Group: Does not receive the treatment.
Placebo: Harmless substitute for the actual treatment, used to control for psychological effects.
Placebo Effect: Participants respond to the belief they are receiving treatment, even if they are not.
Gold Standards of Experiments:
Large Sample Size: Increases reliability of results.
Random Assignment: Subjects are randomly assigned to groups to minimize bias.
Blinding: Single-blind (participants unaware of group assignment) and double-blind (both participants and researchers unaware) designs reduce bias.
Example: Observing a correlation between ice cream sales and drownings does not imply causality. The study is observational, and a confounding variable (e.g., temperature) may explain the association.
Confounding Variable: An unaccounted variable that influences both the treatment and response variables, potentially leading to incorrect conclusions about causality.
Example Confounding Variable: In the ice cream and drowning example, temperature is a confounding variable because it affects both ice cream sales and swimming activity.