Skip to main content
뒤로

Chapter 1: Introduction to Statistics – Study Guide

스터디 가이드 - 스마트 노트

자료에 맞춘 맞춤형 노트, 핵심 정의, 예시, 맥락을 확장해 제공합니다.

Introduction to Statistics

Statistical and Critical Thinking

Statistics is the science of collecting, analyzing, interpreting, and presenting data to draw meaningful conclusions. Understanding the foundational concepts and the process of statistical analysis is essential for making informed decisions based on data.

  • Data: Collections of observations, such as measurements, genders, or survey responses.

  • Statistics: The science of planning studies and experiments; obtaining, organizing, summarizing, presenting, analyzing, and interpreting data to draw conclusions.

  • Population: The complete collection of all measurements or data being considered, about which inferences are to be made.

  • Census: Data collection from every member of a population.

  • Sample: A subcollection of members selected from a population.

The Statistical Process: Prepare, Analyze, Conclude

Stage

Key Components & Questions

Prepare

  • Context: What do the data represent? What is the goal of the study?

  • Source of Data: Is there a special interest that might bias results?

  • Sampling Method: Was data collection unbiased or is there sampling bias?

Analyze

  • Graph the Data: Use visual representations.

  • Explore the Data: Check for outliers, central tendency, dispersion, distribution shape, missing data, or high refusal rates.

  • Apply Statistical Methods: Use statistical tools and technology to obtain results.

Conclude

  • Statistical Significance: Achieved if the likelihood of an event occurring by chance is 5% or less.

  • Practical Significance: Evaluate if a finding is meaningful in real-world terms.

Sampling Bias & Potential Pitfalls

  • Voluntary Response Sample: Respondents decide whether to participate (e.g., internet polls), often leading to bias.

  • Reported vs. Measured Data: Physical measurements are preferred over self-reported values to reduce bias.

  • Loaded Questions & Question Order: Wording or sequence can influence responses.

  • Nonresponse: When subjects refuse or are unavailable, increasing bias.

  • Misleading Percentages: Percentages can be misused, especially when changes over 100% are claimed inappropriately.

Types of Data

Parameters vs. Statistics

  • Parameter: A numerical measurement describing a characteristic of a population.

  • Statistic: A numerical measurement describing a characteristic of a sample.

Types of Data

Data Type

Definition

Example

Quantitative (Numerical)

Numbers representing counts or measurements.

Weights, ages

Categorical (Qualitative/Attribute)

Names, labels, or categories (not numerical counts).

Gender, shirt numbers

  • Discrete Data: Quantitative data with a finite or countable number of values (e.g., number of coin tosses).

  • Continuous Data: Quantitative data with infinitely many possible values, not countable (e.g., distances).

Levels of Measurement

Level

Characteristics

Example

Nominal

Categories, names, or labels only; cannot be ordered.

Yes/No/Undecided

Ordinal

Can be ordered, but differences are meaningless.

Course grades (A, B, C, D, F)

Interval

Ordered, meaningful differences, but no natural zero.

Years, temperatures (°F/°C)

Ratio

Ordered, meaningful differences, and a natural zero; ratios are valid.

Time, distances, weights

Big Data & Missing Data

  • Big Data: Extremely large and complex data sets requiring advanced computational methods.

  • Data Science: An interdisciplinary field combining statistics, computer science, and domain expertise.

  • Missing Completely at Random (MCAR): The probability of missing data is unrelated to any values in the data set.

  • Missing Not at Random (MNAR): The missingness is related to the value itself.

  • Handling Missing Data: Options include deleting cases or imputing (estimating) missing values.

Collecting Sample Data

Observational Studies vs. Experiments

  • Observational Study: Observes and measures characteristics without modifying subjects.

  • Experiment: Applies a treatment and observes its effects on subjects.

Key Experimental Concepts

  • Gold Standard: Randomization with treatment and placebo groups.

  • Placebo: An inactive treatment used as a control.

  • Replication: Repeating the experiment on enough subjects to observe effects.

  • Blinding: Subjects do not know if they receive the treatment or placebo.

  • Double-Blind: Neither subjects nor experimenters know group assignments.

  • Randomization: Assigning subjects to groups by chance.

Sampling Methods

Method

Description

Simple Random Sample (SRS)

Every possible sample of size n has an equal chance of being chosen.

Systematic Sampling

Select a starting point, then pick every k-th element.

Convenience Sampling

Use data that are easy to obtain; high risk of bias.

Stratified Sampling

Divide population into subgroups (strata) and sample from each.

Cluster Sampling

Divide population into clusters, randomly select clusters, and sample all members in selected clusters.

Multistage Sampling

Combine multiple sampling methods in stages.

Types of Observational Studies

  • Cross-Sectional Study: Data collected at a single point in time.

  • Retrospective (Case-Control) Study: Data collected from past records or interviews.

  • Prospective (Longitudinal/Cohort) Study: Data collected in the future from groups sharing common factors.

Controlling Variable Effects & Experimental Design

  • Confounding: When the effect of one variable cannot be distinguished from another.

  • Completely Randomized Design: Subjects assigned to groups purely by chance.

  • Randomized Block Design: Subjects grouped by similar characteristics, then randomly assigned treatments within blocks.

  • Matched Pairs Design: Pairs of subjects matched closely, each receiving different treatments.

  • Rigorously Controlled Design: Groups are balanced on all key factors.

Errors in Statistics

  • Sampling Error (Random Sampling Error): The difference between a sample result and the true population result due to chance.

  • Nonsampling Error: Errors from human mistakes, such as data entry or biased survey wording.

  • Nonrandom Sampling Error: Errors from using non-random sampling methods (e.g., convenience sampling).

Example: Identifying Data Types

  • Example 1: The number of cars passing through an intersection in an hour is discrete quantitative data.

  • Example 2: The temperature in degrees Celsius is interval level data (no true zero).

  • Example 3: Survey responses of "Yes," "No," or "Undecided" are nominal data.

Pearson Logo

스터디 프렙