뒤로Chapter 1: Introduction to Statistics – Key Concepts, Data Types, and Sampling Methods
스터디 가이드 - 스마트 노트
자료에 맞춘 맞춤형 노트, 핵심 정의, 예시, 맥락을 확장해 제공합니다.
Chapter 1: Introduction to Statistics
Section 1-1: Fundamental Concepts in Statistics
Statistics is the science of collecting, analyzing, presenting, and interpreting data to draw meaningful conclusions. Understanding the foundational vocabulary is essential for further study in statistics.
Data: Collections of observations, such as measurements, genders, or survey responses.
Statistics (the discipline): The science of planning studies and experiments; obtaining data; and organizing, summarizing, presenting, analyzing, and interpreting those data to draw conclusions.
Population: The complete collection of all measurements or data that are being considered. Typically, a population is the entire group about which inferences are to be made.
Census: The collection of data from every member of a population.
Sample: A subcollection of members selected from a population.
Voluntary Response Sample (Self-Selected Sample): A sample in which the respondents themselves decide whether to be included.
Statistically Significant: A result is statistically significant if the likelihood of an event occurring by chance is 5% or less.
Practically Significant: A result may be statistically significant but not large enough to be of practical importance in real-world terms.
Section 1-2: Types of Data and Levels of Measurement
Data can be classified by their nature and the level of measurement, which determines the types of statistical analyses that are appropriate.
Parameter: A numerical measurement describing some characteristic of a population.
Statistic: A numerical measurement describing some characteristic of a sample.
Quantitative Data: Consists of numbers representing counts or measurements (e.g., weights, ages).
Categorical Data (Qualitative): Consists of names or labels (e.g., gender, shirt numbers).
Discrete Data: Quantitative data where the number of possible values is finite or countable (e.g., number of coin tosses).
Continuous Data: Quantitative data with infinitely many possible values, not countable (e.g., lengths, time).
Levels of Measurement
Nominal: Data consists of names or labels with no inherent order (e.g., survey responses: yes, no, undecided).
Ordinal: Data can be arranged in order, but differences between values are not meaningful (e.g., course grades: A, B, C, D, F).
Interval: Data can be ordered, and meaningful differences can be found, but there is no natural zero point (e.g., years, temperature in Celsius or Fahrenheit).
Ratio: Data can be ordered, differences are meaningful, and there is a natural zero point; ratios are meaningful (e.g., time, money).
Section 1-3: Experimental Design and Sampling Methods
Proper experimental design and sampling methods are crucial for obtaining reliable and unbiased results in statistics.
Placebo: A harmless and ineffective treatment used for psychological benefit or as a control in experiments.
Experiment: A study in which a treatment is applied and its effects are observed.
Experimental Units (Subjects): The individuals on whom an experiment is conducted.
Observational Study: Observing and measuring specific characteristics without attempting to modify the subjects.
Replication: Repetition of an experiment on more than one individual to ensure reliability of results.
Blind Study: The subject does not know whether they are receiving the treatment or a placebo.
Double Blind Study: Both the subject and the experimenter do not know who receives the treatment or placebo.
Randomness: Assigning subjects to groups by random selection to reduce bias.
Sampling Methods
Simple Random Sample: Every possible sample of the same size has the same chance of being chosen.
Systematic Sampling: Select a starting point and then select every kth element in the population.
Convenience Sampling: Use data that are easy to obtain (e.g., surveying people in a study hall).
Stratified Sampling: Subdivide the population into subgroups (strata) with shared characteristics, then sample from each subgroup.
Cluster Sampling: Divide the population into clusters, randomly select some clusters, and include all members from those clusters.
Types of Studies
Cross-sectional Study: Data are collected at one point in time.
Retrospective (Case-Control) Study: Data are collected from past records or events.
Prospective (Longitudinal or Cohort) Study: Data are collected in the future from groups sharing common factors (cohorts).
Section 1-4: Sampling Bias
Sampling bias occurs when the method of collecting data causes the sample to differ from the population in a systematic way, leading to inaccurate results.
Sampling Bias: Systematic error due to a non-random sample of a population, causing some members to be less likely to be included than others.
Types of Sampling Bias: (Details not provided in the original notes; see additional info below.)
Additional info: Common types of sampling bias include selection bias, nonresponse bias, and response bias. Selection bias occurs when certain groups are systematically excluded from the sample. Nonresponse bias arises when individuals selected for the sample do not respond, and their nonresponse is related to the variable being studied. Response bias occurs when respondents give inaccurate answers, often due to question wording or social desirability.
Summary Table: Data Types and Levels of Measurement
Type | Description | Examples |
|---|---|---|
Quantitative (Discrete) | Countable numbers | Number of students, coin tosses |
Quantitative (Continuous) | Infinite possible values, measurable | Height, weight, time |
Categorical (Nominal) | Names or labels, no order | Gender, colors, survey responses |
Ordinal | Ordered categories, differences not meaningful | Grades (A, B, C), rankings |
Interval | Ordered, meaningful differences, no true zero | Temperature, years |
Ratio | Ordered, meaningful differences, true zero | Money, time, distance |
Key Formulas
Sample Mean:
Population Mean:
Additional info: These formulas are foundational for summarizing quantitative data and will be used extensively in later chapters.