Skip to main content
Back

Chapter 1: Data Collection – Applied Statistics Study Notes

Study Guide - Smart Notes

Tailored notes based on your materials, expanded with key definitions, examples, and context.

Chapter 1: Data Collection

Section 1.1: Introduction to the Practice of Statistics

Statistics is the science of collecting, organizing, summarizing, and analyzing information to draw conclusions or answer questions. It also involves providing a measure of confidence in any conclusions.

  • Statistics: The practice of collecting, organizing, summarizing, and analyzing information to draw conclusions or answer questions. It provides a measure of confidence in conclusions.

  • Data: Information collected for analysis. According to the American Heritage Dictionary, data is a collection of facts that can be analyzed or used to gain knowledge or make decisions.

Key Terms

  • Statistic: A numerical summary of a sample.

  • Parameter: A numerical summary of a population.

  • Descriptive Statistics: Use numerical summaries, data, and graphs to describe data.

  • Inferential Statistics: Attempt to extend results from samples to populations and measure their reliability.

Statistical Process

  1. Identify the question(s) to be answered and the population to be sampled.

  2. Collect data appropriately.

  3. Describe the data.

  4. Make an inference.

Example

Pew Research Center conducted a poll to determine if children are better off when both parents have a job versus only one. The process involved collecting data, summarizing results, and making inferences about the population.

Section 1.1 (continued): Types of Variables

Qualitative vs. Quantitative Variables

  • Qualitative (Categorical) Variables: Allow for classification of individuals based on some attribute or characteristic. Examples: Types of governments, presidential candidates, ZIP code, year.

  • Quantitative Variables: Provide numerical measures of individuals. Values can be added or subtracted and provide meaningful results. Examples: Time elapsed, number of players on a team, volume of a container.

Discrete vs. Continuous Variables

  • Discrete Variable: A quantitative variable with a finite or countable number of possible values. Examples: Money spent on shoes, number of customers at Walmart, populations of countries.

  • Continuous Variable: A quantitative variable with an infinite number of possible values between any two values. Examples: Distance traveled, weight, time, animal heights.

Discrete variables are counted, while continuous variables are measured.

Section 1.2: Observational Studies versus Designed Experiments

Observational Study vs. Experiment

  • Observational Study: Measures the value of the response variable without attempting to influence the value of either the response or explanatory variables. Keyword: "observes" Examples: Squirrel nut storage, migratory patterns of birds.

  • Designed Experiment: Researcher randomly assigns individuals to groups, intentionally manipulates the value of an explanatory variable, and records the response. Keyword: "designed" Examples: How a Nerf gun is fired affects bullet travel, vaccination tests.

Confounding and Lurking Variables

  • Confounding Variable: An explanatory variable whose effect cannot be distinguished from another explanatory variable in the study. Example: Wine consumption and heart disease rates may be confounded by other factors such as healthcare quality.

  • Lurking Variable: Not considered in a study but affects the value of the response variable. Example: Age, health status, or mobility in vaccine studies.

  • Important Note: Correlation does not imply causation.

Section 1.3: Simple Random Sampling

Sampling Methods

  • Random Sampling: Using chance to select individuals from a population. Definition: A frame is a list of all individuals in a population.

  • Simple Random Sample: Every possible sample of size n has an equally likely chance of occurring. Example: Assigning numbers to 5,000 people and randomly selecting for a survey.

  • Nonexample: Surveying the first 100 people who exit a store (not representative).

Section 1.4: Other Effective Sampling Methods

Stratified Sampling

  • Definition: Population is separated into nonoverlapping groups (strata), and a simple random sample is taken from each stratum. Example: Sampling students by grade level.

Systematic Sampling

  • Definition: Selecting every k-th individual from the population after a random start. Steps:

    1. Approximate population size.

    2. Determine sample size.

    3. Calculate and round down to nearest integer.

    4. Randomly select a starting point.

    5. Select every k-th individual.

    Example: Surveying every 5th customer at a restaurant.

Cluster Sampling

  • Definition: Selecting all individuals within a randomly selected group or collection. Example: Surveying every household in randomly selected city blocks.

Convenience and Multistage Sampling

  • Convenience Sample: Individuals are easily obtained and not based on randomness. Note: Often unreliable and not representative.

  • Multistage Sampling: Combines several sampling techniques, often used in large-scale surveys. Example: U.S. Census Bureau surveys.

Section 1.5: Bias in Sampling

Sources of Bias

  • Definition: If the sample is not representative of the population, it contains bias.

  • Sampling Bias: Technique favors one part of the population.

  • Nonresponse Bias: Selected individuals do not respond.

  • Response Bias: Answers do not reflect true feelings due to interviewer error, misrepresented answers, wording of questions, or data-entry error.

Types of Bias

Type of Bias

Description

Example

Sampling Bias

Technique favors one part of population

Convenience sampling

Nonresponse Bias

Selected individuals do not respond

Telephone surveys with low response rates

Response Bias

Answers do not reflect true feelings

Leading questions, interviewer error

Nonsampling Errors

  • Definition: Errors from undercoverage, nonresponse, response bias, or data-entry error. Sampling Error: Using a sample to estimate population information may give incomplete information.

  • Example: Measuring average height of basketball players but missing some data due to misreporting.

Section 1.6: The Design of Experiments

Characteristics of an Experiment

  • Experiment: Controlled study to determine the effect of varying explanatory variables (factors) on a response variable. Treatment: Any combination of factor values.

  • Elements:

    1. Experimental unit/subject

    2. Control group

    3. Placebo

    4. Blinding

  • Single-blind: Subject does not know treatment.

  • Double-blind: Neither subject nor researcher knows treatment.

Steps in Designing an Experiment

  1. Identify the problem.

  2. Determine factors influencing the response.

  3. Determine number of experimental units.

  4. Determine levels of each factor.

  5. Assign units to treatment groups.

  6. Conduct the experiment.

  7. Test the claim.

Completely Randomized Design

  • Definition: Each experimental unit is randomly assigned to a treatment.

  • Example: 150 participants randomly assigned to three groups for a drug test.

Matched-Pairs Design

  • Definition: Experimental units are paired, and each pair is assigned to different treatments. Example: Comparing two treatments on matched subjects (e.g., twins, husband and wife).

Important Equations

  • Sample Mean:

  • Sample Proportion:

Additional info: These notes expand on brief points with definitions, examples, and formulas for clarity and completeness.

Pearson Logo

Study Prep