Skip to main content
뒤로

Chapter 1: Data Collection – Foundations of Statistical Thinking

스터디 가이드 - 스마트 노트

자료에 맞춘 맞춤형 노트, 핵심 정의, 예시, 맥락을 확장해 제공합니다.

Section 1.1: Introduction to the Practice of Statistics

Definition and Scope of Statistics

Statistics is the science of collecting, organizing, summarizing, and analyzing information to draw conclusions or answer questions. It also involves providing a measure of confidence in any conclusions. The information used in statistics is called data, which are facts or propositions used to draw conclusions or make decisions. Data describe characteristics of individuals and are subject to variability—differences among individuals or within the same individual over time.

  • Key Point: Understanding and describing sources of variability is a central goal of statistics.

  • Example: Heights, sleep duration, and calorie intake are all data that vary among individuals.

The Process of Statistics

The process of statistics involves four main steps:

  1. Identify the research objective: Clearly state the question(s) and the population to be studied.

  2. Collect the data: Gather data from a sample, as studying the entire population is often impractical.

  3. Describe the data: Use descriptive statistics to summarize and organize the data.

  4. Perform inference: Apply inferential statistics to extend results from the sample to the population and report reliability.

Diagram showing population, sample, and individual

Additional info: The image illustrates the relationship between the population (entire group), sample (subset), and individual (single member).

Descriptive vs. Inferential Statistics

  • Descriptive statistics: Organize and summarize data using numerical summaries, tables, and graphs.

  • Inferential statistics: Use sample data to make generalizations about a population and measure reliability.

  • Parameter: Numerical summary of a population.

  • Statistic: Numerical summary based on a sample.

Section 1.1: Types of Variables and Data

Qualitative vs. Quantitative Variables

Variables are characteristics of individuals in a population. They can be classified as:

  • Qualitative (categorical) variables: Classify individuals based on attributes or characteristics (e.g., gender, education level).

  • Quantitative variables: Provide numerical measures (e.g., height, temperature, number of vending machines).

Diagram showing classification of variables: qualitative, quantitative, discrete, continuous

Additional info: The image visually distinguishes between qualitative variables and quantitative variables, which are further divided into discrete and continuous types.

Discrete vs. Continuous Variables

  • Discrete variable: Quantitative variable with a finite or countable number of possible values (e.g., number of children).

  • Continuous variable: Quantitative variable with an infinite number of possible values, not countable (e.g., height, gas mileage).

Levels of Measurement

  • Nominal: Values name, label, or categorize; no inherent order (e.g., gender).

  • Ordinal: Values can be ranked or ordered (e.g., income status, satisfaction ratings).

  • Interval: Differences between values have meaning; zero does not indicate absence (e.g., temperature in Celsius).

  • Ratio: Ratios of values have meaning; zero indicates absence (e.g., income, number of children).

Section 1.2: Observational Studies Versus Designed Experiments

Observational Studies vs. Experiments

  • Observational study: Measures the value of the response variable without influencing variables. Can show association, not causation.

  • Designed experiment: Researcher assigns treatments and controls variables to observe effects on the response variable. Can establish causation.

  • Confounding variable: Explanatory variable whose effect cannot be separated from another variable.

  • Lurking variable: Not considered in the study but affects the response variable.

Types of Observational Studies

  • Cross-sectional: Collects data at a specific point in time.

  • Case-control: Retrospective; compares individuals with and without a characteristic.

  • Cohort: Prospective; follows a group over time to observe outcomes.

Section 1.3 & 1.4: Sampling Methods

Simple Random Sampling

  • Every possible sample of size n from a population of size N has an equal chance of being selected.

  • Use random number tables or statistical software to select the sample.

Other Effective Sampling Methods

  • Stratified sampling: Divide population into homogeneous groups (strata), then randomly sample from each stratum.

  • Systematic sampling: Select every kth individual from a list, starting at a random point.

  • Cluster sampling: Divide population into groups (clusters), randomly select some clusters, and include all individuals from those clusters.

  • Convenience sampling: Select individuals easily obtained; not based on randomness and often leads to bias.

Diagram showing simple random, stratified, systematic, and cluster sampling

Additional info: The image provides visual examples of each sampling method, clarifying their differences.

Section 1.5: Bias in Sampling

Sources of Bias

  • Sampling bias: Sampling method favors one part of the population.

  • Nonresponse bias: Individuals selected do not respond, and their opinions differ from responders.

  • Response bias: Survey answers do not reflect true feelings due to interviewer error, misrepresented answers, or question wording.

  • Data-entry error: Mistakes in recording data can lead to inaccurate results.

  • Nonsampling error: Errors from bias or data-entry mistakes; can occur even in a census.

  • Sampling error: Error from using a sample to estimate population information.

Section 1.6: The Design of Experiments

Characteristics of an Experiment

  • Experiment: Controlled study to determine the effect of varying explanatory variables (factors) on a response variable.

  • Treatment: Any combination of values of the factors.

  • Experimental unit (subject): The individual to which a treatment is applied.

  • Control group: Baseline group for comparison; may receive a placebo.

  • Placebo effect: Improvement due to the belief in the treatment, not the treatment itself.

  • Blinding: Nondisclosure of treatment assignment; can be single-blind or double-blind.

Steps in Designing an Experiment

  1. Identify the problem to be solved (state the claim).

  2. Determine the factors that affect the response variable.

  3. Determine the number of experimental units.

  4. Determine the level of each factor (control or randomize).

  5. Conduct the experiment.

  6. Test the claim using inferential statistics.

Steps in designing an experiment Replication, data collection, and inferential statistics in experiments

Additional info: The images provide a step-by-step guide to experimental design, including replication and inferential statistics.

Experimental Designs

  • Completely randomized design: Experimental units are randomly assigned to treatments.

  • Matched-pairs design: Experimental units are paired based on similarity; each pair receives different treatments.

  • Randomized block design: Experimental units are divided into homogeneous blocks, then randomly assigned to treatments within each block.

Random Sampling vs. Random Assignment

  • Random sampling: Individuals are randomly selected from the population, allowing for population inference.

  • Random assignment: Individuals are randomly assigned to treatment groups, allowing for causal inference.

Pearson Logo

스터디 프렙