BackChapter 1: Data Collection – Applied Statistics Study Notes
Study Guide - Smart Notes
Tailored notes based on your materials, expanded with key definitions, examples, and context.
Chapter 1: Data Collection
Section 1.1: Introduction to the Practice of Statistics
Statistics is the science of collecting, organizing, summarizing, and analyzing information to draw conclusions or answer questions. It also involves providing a measure of confidence in any conclusions.
Statistics: The practice of collecting, organizing, summarizing, and analyzing information to draw conclusions or answer questions. It provides a measure of confidence in conclusions.
Data: Information collected for analysis. According to the American Heritage Dictionary, data is a collection of facts that can be analyzed or used to gain knowledge or make decisions.
Key Terms
Statistic: A numerical summary of a sample.
Parameter: A numerical summary of a population.
Descriptive Statistics: Use numerical summaries, data, and graphs to describe data.
Inferential Statistics: Attempt to extend results from samples to populations and measure their reliability.
Statistical Process
Identify the question(s) to be answered and the population to be sampled.
Collect data appropriately.
Describe the data.
Make an inference.
Example
Pew Research Center conducted a poll to determine if children are better off when both parents have a job versus only one. The process involved collecting data, summarizing results, and making inferences about the population.
Section 1.1 (continued): Types of Variables
Qualitative vs. Quantitative Variables
Qualitative (Categorical) Variables: Allow for classification of individuals based on some attribute or characteristic. Examples: Types of governments, presidential candidates, ZIP code, year.
Quantitative Variables: Provide numerical measures of individuals. Values can be added or subtracted and provide meaningful results. Examples: Time elapsed, number of players on a team, volume of a container.
Discrete vs. Continuous Variables
Discrete Variable: A quantitative variable with a finite or countable number of possible values. Examples: Money spent on shoes, number of customers at Walmart, populations of countries.
Continuous Variable: A quantitative variable with an infinite number of possible values between any two values. Examples: Distance traveled, weight, time, animal heights.
Discrete variables are counted, while continuous variables are measured.
Section 1.2: Observational Studies versus Designed Experiments
Observational Study vs. Experiment
Observational Study: Measures the value of the response variable without attempting to influence the value of either the response or explanatory variables. Keyword: "observes" Examples: Squirrel nut storage, migratory patterns of birds.
Designed Experiment: Researcher randomly assigns individuals to groups, intentionally manipulates the value of an explanatory variable, and records the response. Keyword: "designed" Examples: How a Nerf gun is fired affects bullet travel, vaccination tests.
Confounding and Lurking Variables
Confounding Variable: An explanatory variable whose effect cannot be distinguished from another explanatory variable in the study. Example: Wine consumption and heart disease rates may be confounded by other factors such as healthcare quality.
Lurking Variable: Not considered in a study but affects the value of the response variable. Example: Age, health status, or mobility in vaccine studies.
Important Note: Correlation does not imply causation.
Section 1.3: Simple Random Sampling
Sampling Methods
Random Sampling: Using chance to select individuals from a population. Definition: A frame is a list of all individuals in a population.
Simple Random Sample: Every possible sample of size n has an equally likely chance of occurring. Example: Assigning numbers to 5,000 people and randomly selecting for a survey.
Nonexample: Surveying the first 100 people who exit a store (not representative).
Section 1.4: Other Effective Sampling Methods
Stratified Sampling
Definition: Population is separated into nonoverlapping groups (strata), and a simple random sample is taken from each stratum. Example: Sampling students by grade level.
Systematic Sampling
Definition: Selecting every k-th individual from the population after a random start. Steps:
Approximate population size.
Determine sample size.
Calculate and round down to nearest integer.
Randomly select a starting point.
Select every k-th individual.
Example: Surveying every 5th customer at a restaurant.
Cluster Sampling
Definition: Selecting all individuals within a randomly selected group or collection. Example: Surveying every household in randomly selected city blocks.
Convenience and Multistage Sampling
Convenience Sample: Individuals are easily obtained and not based on randomness. Note: Often unreliable and not representative.
Multistage Sampling: Combines several sampling techniques, often used in large-scale surveys. Example: U.S. Census Bureau surveys.
Section 1.5: Bias in Sampling
Sources of Bias
Definition: If the sample is not representative of the population, it contains bias.
Sampling Bias: Technique favors one part of the population.
Nonresponse Bias: Selected individuals do not respond.
Response Bias: Answers do not reflect true feelings due to interviewer error, misrepresented answers, wording of questions, or data-entry error.
Types of Bias
Type of Bias | Description | Example |
|---|---|---|
Sampling Bias | Technique favors one part of population | Convenience sampling |
Nonresponse Bias | Selected individuals do not respond | Telephone surveys with low response rates |
Response Bias | Answers do not reflect true feelings | Leading questions, interviewer error |
Nonsampling Errors
Definition: Errors from undercoverage, nonresponse, response bias, or data-entry error. Sampling Error: Using a sample to estimate population information may give incomplete information.
Example: Measuring average height of basketball players but missing some data due to misreporting.
Section 1.6: The Design of Experiments
Characteristics of an Experiment
Experiment: Controlled study to determine the effect of varying explanatory variables (factors) on a response variable. Treatment: Any combination of factor values.
Elements:
Experimental unit/subject
Control group
Placebo
Blinding
Single-blind: Subject does not know treatment.
Double-blind: Neither subject nor researcher knows treatment.
Steps in Designing an Experiment
Identify the problem.
Determine factors influencing the response.
Determine number of experimental units.
Determine levels of each factor.
Assign units to treatment groups.
Conduct the experiment.
Test the claim.
Completely Randomized Design
Definition: Each experimental unit is randomly assigned to a treatment.
Example: 150 participants randomly assigned to three groups for a drug test.
Matched-Pairs Design
Definition: Experimental units are paired, and each pair is assigned to different treatments. Example: Comparing two treatments on matched subjects (e.g., twins, husband and wife).
Important Equations
Sample Mean:
Sample Proportion:
Additional info: These notes expand on brief points with definitions, examples, and formulas for clarity and completeness.