BackExploring Data with Tables and Graphs: Frequency Distributions and Data Visualization
Study Guide - Smart Notes
Tailored notes based on your materials, expanded with key definitions, examples, and context.
Exploring Data with Tables and Graphs
Introduction
Organizing and summarizing data is a foundational skill in statistics. Frequency distributions and graphical representations help reveal patterns, trends, and differences within data sets, making large amounts of information more understandable and useful for analysis.
Frequency Distributions for Organizing and Summarizing Data
Definition and Purpose
Frequency Distribution (or Frequency Table): A table that shows how data are partitioned among several categories (or classes) by listing each category along with the number (frequency) of data values in each.
Purpose: To organize large data sets, making it easier to see the distribution and identify patterns such as clusters, gaps, or outliers.
Key Terms
Lower class limits: The smallest numbers that can belong to each class.
Upper class limits: The largest numbers that can belong to each class.
Class boundaries: Numbers used to separate classes, eliminating gaps between class limits.
Class midpoints: The value in the middle of each class, calculated as .
Class width: The difference between two consecutive lower class limits (or boundaries).
Procedure for Constructing a Frequency Distribution
Select the number of classes (usually between 5 and 20).
Calculate the class width:
Formula:
Round up to a convenient number.
Choose the first lower class limit (often the minimum value or a convenient value below it).
List the lower class limits by adding the class width successively.
Determine the upper class limits for each class.
Tally each data value into the appropriate class and count the frequencies.
Example: Commute Time in Los Angeles
Suppose we want to organize daily commute times using 7 classes.
Calculate class width, round up for convenience (e.g., from 12 to 15 minutes).
Start lower class limits at 0, then add 15 for each subsequent class: 0, 15, 30, 45, 60, 75, 90.
Upper class limits: 14, 29, 44, 59, 74, 89, 104.
Tally and count frequencies for each class.
Relative Frequency Distributions
Definition and Calculation
Relative Frequency Distribution: Each class frequency is replaced by a relative frequency (proportion) or percentage.
Calculation:
The sum of all relative frequencies (as percentages) should be close to 100% (allowing for rounding errors).
Comparisons Using Relative Frequency Distributions
Combining two or more relative frequency distributions in one table allows for easy comparison between data sets.
Example: Comparing commute times in New York, NY and Boise, ID shows that Boise has a much higher proportion of short commute times, reflecting differences in city size and population density.
Cumulative Frequency Distributions
Definition and Use
Cumulative Frequency Distribution: The frequency for each class is the sum of the frequencies for that class and all previous classes.
Useful for determining how many data values fall below a particular value.
Critical Thinking: Understanding Data Distributions
Normal Distributions
In a normal distribution, frequencies start low, increase to a maximum, then decrease, forming a symmetric pattern.
Frequencies before and after the maximum should be roughly mirror images.
Gaps in Data
The presence of gaps in a frequency distribution may indicate data from two or more different populations.
Example: A frequency distribution of penny weights shows a gap, suggesting two populations: pennies made before 1983 (mostly copper) and after 1983 (mostly zinc).
Tables
Example: Comparing Relative Frequency Distributions
The following table compares the relative frequency distributions of commute times in New York, NY and Boise, ID:
Commute Time (min) | New York, NY (%) | Boise, ID (%) |
|---|---|---|
0-14 | 10.2 | 45.3 |
15-29 | 18.7 | 30.5 |
30-44 | 35.1 | 15.2 |
45-59 | 20.0 | 6.0 |
60+ | 16.0 | 3.0 |
Additional info: Table values are illustrative; actual values may differ based on the original data.
Summary
Frequency distributions and their variants (relative, cumulative) are essential for summarizing and comparing data.
Key terms include class limits, boundaries, midpoints, and width.
Critical analysis of distributions can reveal underlying patterns, such as normality or the presence of multiple populations.
