Skip to main content
Indietro

Introductory Statistics Exam 1 Study Guide: Major Topics and Concepts

Guida di studio - Note intelligenti

Appunti personalizzati basati sui tuoi materiali, ampliati con definizioni chiave, esempi e contesto.

Statistical Thinking & Data Vocabulary

Understanding the Statistical Process

Statistics is the science of collecting, analyzing, and interpreting data. The process involves preparing data (considering context, source, and sampling), analyzing it, and drawing conclusions. Everything revolves around data, which consists of observations and measurements.

  • Population vs. Sample: A population parameter (e.g., , ) describes the entire population, while a sample statistic (e.g., , ) is computed from a subset.

  • Statistics: Refers both to quantities derived from data and the science of data analysis.

  • Data Types: Quantitative (numerical; discrete = countable, continuous = uncountable) vs. categorical (names and labels only).

  • EDA (Exploratory Data Analysis): The initial step to check for missing values, relationships, and data sanity.

Sampling Methods & Pitfalls in Drawing Conclusions

Sampling Techniques and Common Errors

Sampling is crucial for obtaining representative data. Poor sampling leads to biased results and unreliable conclusions.

  • Representative Sample: Should reflect the population's characteristics.

  • Voluntary Samples: Often biased (e.g., internet polls).

  • Convenience Samples: Use readily available subjects; may not be representative.

  • Random Sampling: Preferred for unbiased results.

  • Statistical Significance: Indicates results are unlikely due to chance.

  • Practical Significance: Results are important enough to warrant action.

  • Pitfalls: Correlation does not imply causation, loaded questions, question ordering, low response rates, null responses.

Frequency Distributions & Histograms

Organizing and Visualizing Data

Frequency distributions and histograms help summarize and visualize data, revealing patterns and shapes.

  • Frequency Distribution: Counts of data in categories or numerical classes (bins).

  • Class Terminology: Lower/upper class limits, class boundaries, class width, midpoint.

  • Derived Columns: Relative frequency (), cumulative frequency, relative cumulative frequency.

  • Shapes: Bell-shaped, uniform, right-skewed, left-skewed.

  • Other Displays: Stemplot, dotplot, time-series graph, Pareto chart.

Example Table: Frequency Distribution

Class Interval

Frequency

Relative Frequency

Cumulative Frequency

1-5

4

8%

4

6-10

10

20%

14

11-15

36

72%

50

Additional info: Example values inferred for illustration.

Scatterplots & Linear Correlation

Visualizing Relationships Between Variables

Scatterplots display paired data, helping to identify relationships. Correlation quantifies the strength and direction of linear relationships.

  • Scatterplot: Plots (x, y) pairs; independent variable on x-axis, dependent on y-axis.

  • Correlation: Measures linear relationship; does not imply causation.

  • Linear Correlation Coefficient (): ; sign indicates direction, magnitude indicates strength.

  • Interpretation: strong, moderate, weak, very weak/none.

Linear Regression

Predicting Values Using Regression Lines

Linear regression models the relationship between two variables, allowing predictions. The regression line minimizes the sum of squared residuals.

  • Regression Line: where is the y-intercept, is the slope.

  • Least Squares: Minimizes .

  • Slope and Intercept: , .

  • Coefficient of Determination (): Proportion of variation in y explained by x.

  • Prediction: Substitute x into .

  • Interpolation vs. Extrapolation: Interpolation (within data range) is safer.

Measures of Center

Describing the Central Tendency of Data

Measures of center summarize where most data values lie. Common measures include mean, median, mode, and midrange.

  • Mean: ; sensitive to outliers.

  • Median: Middle value; resistant to outliers.

  • Mode: Most frequent value; can be none, unimodal, bimodal, or multimodal.

  • Midrange: ; very sensitive to outliers.

  • Shape and Choice: Symmetric: mean ≈ median ≈ mode; right-skewed: mode < median < mean; left-skewed: mode > median > mean.

Measures of Variation: Range, IQR & Variance

Quantifying Data Spread

Variation measures describe how spread out data values are. Range, interquartile range (IQR), variance, and standard deviation are key measures.

  • Range: Max − Min; sensitive to outliers.

  • Interquartile Range (IQR): ; resistant to outliers.

  • Variance: Population: ; Sample: .

  • Bessel’s Correction: Using gives an unbiased estimate of population variance.

Standard Deviation & Its Calculation

Measuring Typical Distance from the Mean

Standard deviation is the square root of variance, indicating how much values typically differ from the mean.

  • Definition: ; always ≥ 0, same units as data.

  • Calculation Methods:

    • Manual Method 1: Mean → deviations → square → sum → divide by → square root.

    • Manual Method 2 (shortcut): .

  • Data Transformations: Adding/subtracting a constant does not change SD; multiplying by k scales SD by k.

  • Empirical Rule (Normal Data): 68% within 1 SD, 95% within 2 SD, 99.7% within 3 SD.

Coefficient of Variation, MAD & Choosing a Measure

Comparing Variability Across Datasets

Relative measures like coefficient of variation (CV) and mean absolute deviation (MAD) help compare variability across datasets with different units or scales.

  • Coefficient of Variation: ; lower CV = more consistent.

  • Mean Absolute Deviation (MAD): .

  • Selection Guide:

    • Symmetric, no outliers: variance or SD.

    • Skewed or outliers: IQR.

    • Different units/scales: CV.

    • Quick estimate: range.

  • Best Practices: Plot data first; report both center and variation; use resistant measures with outliers; do not compare SDs across different scales.

Z-Scores

Standardizing Data Values

Z-scores convert data values to a standard scale, indicating how many standard deviations a value is from the mean.

  • Sample Z-Score:

  • Population Z-Score:

  • Significance: marks significantly high or low values; is not significantly far from the mean.

Quartiles, Percentiles & Boxplots

Summarizing Data Distribution

Quartiles and percentiles divide data into equal parts, while boxplots visually summarize the distribution, highlighting center, spread, and outliers.

  • Quartile Location: (i = 1, 2, 3); round up if not integer.

  • Percentile of a Value:

  • Value at a Percentile:

  • Boxplot: Summarizes min, Q1, median, Q3, max; box spans Q1 to Q3 (IQR).

  • Modified Boxplot: Outliers are values beyond Q1 − 1.5·IQR or Q3 + 1.5·IQR.

Example Table: Quartiles and Percentiles

Statistic

Value

Q1

25

Median (Q2)

35

Q3

50

IQR

25

88th Percentile

60

Additional info: Example values inferred for illustration.

Example: For Space Mountain wait times (N = 50), 45 minutes is the 72nd percentile (36 values below).

Boxplot Outlier Example: Upper fence = Q3 + 1.5·IQR = 50 + 37.5 = 87.5; values 105 and 110 are outliers.

Pearson Logo

Study Prep