Skip to main content
Indietro

Chapter 1: Data Collection – Introductory Statistics Study Guide

Guida di studio - Note intelligenti

Appunti personalizzati basati sui tuoi materiali, ampliati con definizioni chiave, esempi e contesto.

Chapter 1: Data Collection

1.1 Introduction to the Practice of Statistics

Statistics is the science of collecting, organizing, summarizing, and analyzing information to draw conclusions or answer questions. It also provides a measure of confidence in any conclusions. The information used in statistics is called data, which describes characteristics of individuals and exhibits variability.

  • Statistics: The science of data analysis and interpretation.

  • Statistical Thinking: Understanding and accounting for variability in data.

  • Data: Facts or propositions used to draw conclusions or make decisions.

  • Variability: The tendency of data to differ among individuals or over time.

1.1.2 The Process of Statistics

The process of statistics involves several key steps: identifying the population, collecting data, describing the data, and performing inference.

  • Population: The entire group of individuals to be studied.

  • Sample: A subset of the population being studied.

  • Individual: A single member of the population.

Population, Sample, Individual diagram

  • Descriptive Statistics: Organizing and summarizing data using numerical summaries, tables, and graphs.

  • Inferential Statistics: Methods that use sample results to make generalizations about the population and measure reliability.

  • Parameter: Numerical summary of a population.

  • Statistic: Numerical summary based on a sample.

Example: If 84.9% of all students have a job (parameter) and 86.4% of a sample of 250 students have a job (statistic), the sample size is 250.

Steps in the Process of Statistics

  1. Identify the research objective.

  2. Collect the data needed to answer the question.

  3. Describe the data using descriptive statistics.

  4. Perform inference to extend sample results to the population and report reliability.

1.1.3 Qualitative vs. Quantitative Variables

Variables are characteristics of individuals within the population. They can be classified as qualitative or quantitative.

  • Qualitative (Categorical) Variables: Classify individuals based on attributes or characteristics.

  • Quantitative Variables: Provide numerical measures of individuals; values can be added or subtracted meaningfully.

Classification of variables diagram

Example: Nationality (qualitative), Number of children (quantitative), Household income (quantitative), Level of education (qualitative), Daily intake of whole grains (quantitative).

1.1.4 Discrete vs. Continuous Variables

Quantitative variables can be further classified as discrete or continuous.

  • Discrete Variable: Has a finite or countable number of possible values (e.g., number of children).

  • Continuous Variable: Has an infinite number of possible values and can be measured to any desired level of accuracy (e.g., household income, daily intake of whole grains).

Quantitative variables diagram

Example: Number of children (discrete), Household income (continuous), Daily intake of whole grains (continuous).

The type of variable dictates the methods used to analyze the data.

1.1.5 Levels of Measurement

Variables can be measured at different levels, which determine the types of statistical analyses that can be performed.

  • Nominal: Values name, label, or categorize; no ranking (e.g., hair color, zip code).

  • Ordinal: Values can be ranked; differences between values are not meaningful (e.g., letter grades).

  • Interval: Differences between values have meaning; zero does not indicate absence (e.g., temperature in Fahrenheit).

  • Ratio: Ratios of values have meaning; zero indicates absence (e.g., height, age).

Variable

Nominal

Ordinal

Interval

Ratio

Hair Color

Yes

No

No

No

Zip Code

Yes

No

No

No

Letter Grade

Yes

Yes

No

No

ACT Score

Yes

Yes

Yes

No

Height

Yes

Yes

Yes

Yes

Age

Yes

Yes

Yes

Yes

Temperature (F)

Yes

Yes

Yes

No

1.2 Observational Studies vs. Designed Experiments

Observational studies and designed experiments are two main approaches to data collection.

  • Observational Study: Measures the value of the response variable without influencing the outcome.

  • Designed Experiment: Researcher assigns individuals to groups, manipulates explanatory variables, and records response variables.

Example: Studying cell phone use and brain tumors by tracking individuals over time (observational study) vs. exposing rats to radio frequencies (designed experiment).

  • Confounding Variable: An explanatory variable whose effect cannot be distinguished from another.

  • Lurking Variable: Not considered in the study but affects the response variable.

Observational studies can only claim association, not causation.

Types of Observational Studies

  • Cross-sectional: Collects information at a specific point in time.

  • Case-control: Retrospective; compares individuals with and without certain characteristics.

  • Cohort: Prospective; follows a group over time.

  • Census: List of all individuals in a population with their characteristics.

1.3 Simple Random Sampling

Simple random sampling ensures every possible sample of size n from a population of size N has an equally likely chance of occurring.

  • Random Sampling: Using chance to select individuals for the sample.

  • Frame: List of all individuals in the population.

  • Random Number Table: Used to generate random samples.

Table of random numbers

Example: To select 5 members from 435, number them and use a random number generator to select the sample.

1.4 Other Effective Sampling Methods

There are several sampling methods besides simple random sampling:

  • Stratified Sample: Divide population into homogeneous strata and sample from each stratum.

  • Systematic Sample: Select every kth individual from the population after a random start.

  • Cluster Sample: Select all individuals within randomly chosen groups.

  • Convenience Sample: Individuals are easily obtained; results are often unreliable.

  • Multistage Sampling: Combines multiple sampling techniques.

Steps in systematic sampling Sampling methods diagram

1.5 Bias in Sampling

Bias occurs when the sample is not representative of the population. There are several sources of bias:

  • Sampling Bias: Technique favors one part of the population.

  • Nonresponse Bias: Individuals who do not respond differ from those who do.

  • Response Bias: Survey answers do not reflect true feelings due to interviewer error, misrepresented answers, wording, or order of questions.

  • Data-entry Error: Mistakes in recording data.

  • Nonsampling Error: Errors from bias or data-entry mistakes.

  • Sampling Error: Error from using a sample to estimate population information.

Caution sign for convenience sampling

1.6 The Design of Experiments

Experiments are controlled studies to determine the effect of explanatory variables on a response variable.

  • Treatment: Combination of values of factors.

  • Experimental Unit: Person, object, or item receiving treatment.

  • Control Group: Baseline for comparison.

  • Placebo: Inactive treatment used for comparison.

  • Blinding: Nondisclosure of treatment received.

  • Single-blind: Subject does not know treatment.

  • Double-blind: Neither subject nor researcher knows treatment.

Steps in Designing an Experiment

  1. Identify the problem and response variable.

  2. Determine factors affecting the response variable.

  3. Determine the number of experimental units.

  4. Determine the level of predictor variables (control and randomization).

  5. Conduct the experiment (replication and data collection).

  6. Test the claim using inferential statistics.

Completely Randomized Design

Each experimental unit is randomly assigned to a treatment.

Completely randomized design diagram

Example: Assigning 12 cars to three octane levels to compare miles per gallon.

Matched-Pairs Design

Experimental units are paired based on related characteristics, and each pair receives two levels of treatment.

  • Example: 75 children taste milk with and without Xylitol; order of tasting is randomized to avoid order effects.

  • Double-blind: Recommended to prevent bias.

Pearson Logo

Study Prep