뒤로Chapter 1: Stats Starts Here – Foundations of Statistics and Data
스터디 가이드 - 스마트 노트
자료에 맞춘 맞춤형 노트, 핵심 정의, 예시, 맥락을 확장해 제공합니다.
Chapter 1: Stats Starts Here
What is Statistics?
Statistics is the science of collecting, organizing, analyzing, and interpreting data to make informed decisions. It helps us make sense of information by providing tools to understand and describe variability in data.
Data: Any collection of numbers, characters, images, or other items that provide information about something. Data can come from surveys, experiments, or observational studies.
Variation: Data values often differ from one observation to another. Understanding this variation is a key goal of statistics.
Application Example: Social media platforms like Facebook collect data on users to personalize advertisements and content.
Application Example: Studies on texting while driving use statistical methods to determine the impact on reaction times and safety.
Learning Outcomes
Interpret data and communicate results effectively.
Identify flaws in conclusions presented in articles or reports.
Become a more informed and critical consumer of information.
Data and Its Organization
Organizing Data
Raw data can be difficult to interpret without proper organization. Presenting data in tables or structured formats makes patterns and insights more apparent.

Example: The above table shows unorganized data, making it hard to extract meaningful information.

Example: The organized table above presents the same data in a more readable format, allowing for easier analysis and interpretation.
The Five "W"s and One "H"
To understand any dataset, consider the following questions:
Who: Who are the individuals or cases being described?
What: What variables are being measured?
When: When was the data collected?
Where: Where was the data collected?
Why: What is the purpose of the data collection?
How: How was the data collected (methodology)?
Types of Individuals in Data
Respondents: Individuals who answer surveys (e.g., customers at Amazon).
Subjects/Participants: People on whom experiments are conducted (e.g., patients in a clinical trial).
Experimental Units: Objects of study that are not people (e.g., rats in a maze).
Records: Rows in a database, each representing an individual case or transaction.

Example: The table above shows organized records, each row representing a unique case with multiple variables.
Sample and Population
Statistics often aims to make inferences about a population (the entire group of interest) using a sample (a subset of the population).
Population: The complete set of individuals or items of interest.
Sample: A representative subset of the population, used to draw conclusions about the whole.
Key Point: The sample must be representative to ensure valid inferences.
Think, Show, and Tell
Think: Identify the information you want to know.
Show: Display results clearly and accurately.
Tell: Interpret and communicate the findings from the data.
Variables
Categorical Variables
A categorical variable (also called nominal or qualitative) places an individual into one of several groups or categories.
Examples: Favorite color, country of birth, area code.
Drawback: Difficult to analyze with mathematical computations.
Quantitative Variables
A quantitative variable contains measured numerical values with units, representing amounts or degrees of something.
Examples: Age (in years), price (in dollars), temperature (in degrees Fahrenheit).
Categorical or Quantitative?
Some variables can be treated as either categorical or quantitative, depending on context.
Example: Age can be quantitative (measured in years) or categorical (child, teen, adult, senior).
Identifier Variables
An identifier variable uniquely identifies each individual in a dataset but does not describe the individual.
Examples: Login ID, customer number, transaction number, social security number.
Ordinal Variables
An ordinal variable reports order but not the exact difference between values. It has a meaningful sequence but no consistent unit of measurement.
Examples: Likert scale (Strongly Disagree, Disagree, Agree, Strongly Agree), Olympic medals (Gold, Silver, Bronze).
Ordinal variables can sometimes be treated as quantitative by assigning ranks (e.g., 1 = Strongly Disagree, 4 = Strongly Agree).
Models in Statistics
What is a Model?
A model is a simplified representation of reality, used to explain or predict phenomena. In statistics, models help us understand relationships between variables and make predictions.
Example: A model airplane in a wind tunnel is used to study flight dynamics.
Example: Kepler’s Laws model planetary motion.

Example: The graph above shows Tycho Brahe’s observations of Mars compared to the orbit calculated with modern methods, illustrating how models can fit real data.
What Can Go Wrong?
Do not label a variable as categorical or quantitative without considering the context.
Do not assume a variable is quantitative just because its values are numbers.
Always approach data with skepticism and critical thinking.
Chapter Review
Data are values (numerical or labels) with context.
The Five W’s and One H help clarify the context of data.
Identifying the cases (who), variables (what), and purpose (why) is essential for meaningful analysis.
Consider the source and motivation behind data collection.
Variables can be categorical or quantitative, and sometimes the same variable can be treated as either, depending on the research question.