Skip to main content
Indietro

Introductory Statistics: Study Guide for Test 1 (Chapters 1–3)

Guida di studio - Note intelligenti

Appunti personalizzati basati sui tuoi materiali, ampliati con definizioni chiave, esempi e contesto.

Chapter 1: Introduction to Statistics and Collecting Data

The Process of Statistics

The process of statistics involves a series of steps used to collect, analyze, interpret, and present data. Understanding this process is fundamental to making informed decisions based on data.

  • Step 1: Identify the Research Objective – Clearly state what you want to learn from the data.

  • Step 2: Collect the Data – Gather data relevant to the research objective using appropriate methods.

  • Step 3: Describe the Data – Use tables, graphs, and summary statistics to organize and summarize the data.

  • Step 4: Make Inferences – Draw conclusions about the population based on the sample data.

Parameter vs. Statistic

It is important to distinguish between a parameter and a statistic when analyzing data.

  • Parameter: A numerical summary that describes a characteristic of a population.

  • Statistic: A numerical summary that describes a characteristic of a sample.

  • Example: The average height of all students in a university (parameter) vs. the average height of students in a sampled class (statistic).

Types of Variables

  • Qualitative (Categorical) Variable: Describes an attribute or category (e.g., eye color, type of car).

  • Quantitative Variable: Represents a measurable quantity (e.g., height, age).

Quantitative Variables: Discrete vs. Continuous

  • Discrete Variable: Takes on countable values (e.g., number of students).

  • Continuous Variable: Can take on any value within a range (e.g., weight, temperature).

  • Example: The number of books (discrete) vs. the time spent reading (continuous).

Simple Random Sampling

Simple random sampling is a method of selecting a sample from a population in such a way that every possible sample of the same size has an equal chance of being chosen.

  • Key Point: Reduces bias and ensures representativeness.

  • Example: Drawing names from a hat to select participants.

Chapter 2: Describing Data with Tables and Graphs

Histograms

A histogram is a graphical representation of the distribution of quantitative data, using bars to show the frequency of data within equal intervals (bins).

  • Key Point: Useful for visualizing the shape, center, and spread of data.

  • Example: A histogram showing the distribution of exam scores.

Frequency Distributions

A frequency distribution is a table that displays the number of occurrences (frequency) of each value or group of values in a dataset.

  • Relative Frequency: The proportion of observations within a category, calculated as frequency divided by total number of observations.

  • Cumulative Frequency: The sum of frequencies for all values up to and including a certain value.

Class Interval

Frequency

Relative Frequency

Cumulative Frequency

0–9

3

0.15

3

10–19

7

0.35

10

20–29

10

0.50

20

Additional info: Table entries are illustrative; actual data may vary.

Stem-and-Leaf Plot

A stem-and-leaf plot is a method of displaying quantitative data that shows the distribution while retaining the original data values.

  • Key Point: Each data value is split into a "stem" (all but the final digit) and a "leaf" (the final digit).

  • Example: For the numbers 23, 25, 27, the stem is 2 and the leaves are 3, 5, 7.

Chapter 3: Describing Data Numerically

Measures of Central Tendency

Measures of central tendency describe the center or typical value of a dataset.

  • Mean (Arithmetic Average): Sum of all data values divided by the number of values.

  • Median: The middle value when data are ordered from least to greatest.

  • Mode: The value(s) that occur most frequently in the dataset.

  • Example: For the data 2, 3, 3, 5, 7: Mean = 4, Median = 3, Mode = 3.

Measures of Spread (Dispersion)

  • Range: Difference between the maximum and minimum values.

  • Variance: Average of the squared differences from the mean.

  • Standard Deviation: Square root of the variance.

Empirical Rule

The Empirical Rule applies to bell-shaped (normal) distributions and describes the spread of data:

  • About 68% of data falls within 1 standard deviation of the mean.

  • About 95% within 2 standard deviations.

  • About 99.7% within 3 standard deviations.

Z-scores

A z-score indicates how many standard deviations a data value is from the mean.

  • Formula:

  • Interpretation: Positive z-scores are above the mean; negative are below.

Five-Number Summary and Box-and-Whisker Plot

The five-number summary provides a quick overview of the distribution of a dataset:

  • Minimum

  • First Quartile (Q1)

  • Median (Q2)

  • Third Quartile (Q3)

  • Maximum

A box-and-whisker plot (boxplot) visually displays the five-number summary, highlighting the spread and potential outliers.

Additional info: Students are encouraged to regularly review notes, lesson videos, homework, and quizzes to reinforce understanding and prepare for assessments.

Pearson Logo

Study Prep