BackIntroduction to Statistics: Populations, Parameters, Samples, and Inference
Study Guide - Smart Notes
Tailored notes based on your materials, expanded with key definitions, examples, and context.
Introduction to Statistics
Definitions and Key Concepts
Statistics is the science of collecting, analyzing, interpreting, and presenting data. The foundational concepts in statistics involve understanding populations, samples, parameters, statistics, and the process of making inferences about a population based on sample data.
Populations and Samples
Population
Population refers to the entire group of individuals or items that are the subject of a statistical study. The population is defined by the research question and includes all subjects of interest.
Example: All employees working at Capital One Financial Corporation.
Example: All properties sold by Century 21 Real Estate LLC.
Example: All students at a university.
Sample
A sample is a subset of the population selected for analysis. Samples are used because it is often impractical or impossible to collect data from every member of the population.
The sample should be representative of the population, meaning it reflects the important characteristics of the population.
Example: A group of 100 employees selected from all employees at Capital One Financial Corporation.
Example: 250 properties sold by Century 21 Real Estate LLC.
Example: 150 students selected from all students at a university.
Parameters and Statistics
Parameter
A parameter is a numerical characteristic of a population. Parameters are typically denoted by Greek letters and describe the entire population.
μ (mu): Population mean
η (eta): Population median
σ (sigma): Population standard deviation
σ^2 (sigma squared): Population variance
π (pi): Population proportion of successes
ρ (rho): Population correlation coefficient
Example: μ = mean age of all employees at Capital One Financial Corporation.
Example: π = proportion of all properties sold by Century 21 Real Estate LLC for $500,000 or more.
Example: ρ = correlation between years enrolled in a 401(k) plan and total value of the plan for all U.S. employees with 401(k) plans.
Statistic
A statistic is a numerical characteristic calculated from a sample. Statistics are typically denoted by English letters and are used to estimate population parameters.
n: Sample size (number of subjects in the sample)
\( \bar{X} \): Sample mean
M: Sample median
s: Sample standard deviation
S^2: Sample variance
\( \hat{p} \): Sample proportion of successes
r: Sample correlation coefficient
Example: \( \bar{X} = 38.4 \) years, the mean age of 100 sampled employees.
Example: \( \hat{p} = 0.10 \), the proportion of 250 sampled properties sold for $500,000 or more.
Hypotheses in Statistics
Statistical Hypothesis
A hypothesis (or statistical hypothesis) is a statement about a population parameter. Hypotheses are often conjectures or guesses about the value of a parameter, which are then tested using sample data.
Example: μ = 32.4 years (mean age of all employees at Capital One Financial Corporation)
Example: π = 0.12 (proportion of properties sold for $500,000 or more)
Example: μ = 70 inches (mean height of all students at a university)
Statistical Inference
Definition and Types
Statistical inference is the process of making statements about a population parameter based on statistics computed from sample data. There are two main types of statistical inference:
Estimation: Using sample data to estimate a population parameter, often by constructing a confidence interval.
Statistical Tests: Making a hypothesis about a parameter and using sample data to test the validity of the hypothesis.
Examples of Inference
Estimation Example: "Based on the sample of 100 employees, we estimate that the mean age of all employees is 38.4 years."
Hypothesis Test Example: "Based on the sample of 250 properties, we do not have sufficient evidence that the proportion of all property sales exceeding $500,000 is different from 0.12."
Summary Table: Parameters vs. Statistics
The following table summarizes the common parameters and their corresponding statistics:
Concept | Population (Parameter) | Sample (Statistic) |
|---|---|---|
Mean | μ (mu) | \( \bar{X} \) (X-bar) |
Median | η (eta) | M |
Standard Deviation | σ (sigma) | s |
Variance | σ² (sigma squared) | S² |
Proportion | π (pi) | \( \hat{p} \) (p-hat) |
Correlation | ρ (rho) | r |
Applications and Examples
Example 1: Estimating a Mean
Population: All Division I athletic competitions.
Parameter: μ = mean number of people attending all Division I athletic competitions.
Sample: n = 200 competitions, record attendance for each.
Statistic: \( \bar{X} \) = sample mean attendance.
Inference: Use \( \bar{X} \) to estimate μ or test a hypothesis such as μ = 1,725.
Example 2: Estimating a Proportion
Population: All people who have watched a specific film.
Parameter: π = proportion who enjoyed the film.
Sample: n = 500 viewers, ask if they enjoyed the film.
Statistic: \( \hat{p} \) = sample proportion who enjoyed the film.
Inference: Use \( \hat{p} \) to estimate π.
Key Formulas
Sample Mean:
Sample Proportion: where x = number of successes in the sample, n = sample size.
Sample Variance:
Sample Standard Deviation:
Conclusion
Understanding the distinction between populations and samples, parameters and statistics, and the process of statistical inference is fundamental to the study of statistics. These concepts form the basis for collecting data, analyzing results, and making informed decisions based on data.