Skip to main content
뒤로

Chapter 3: Numerical Descriptive Measures in Introductory Statistics

스터디 가이드 - 스마트 노트

자료에 맞춘 맞춤형 노트, 핵심 정의, 예시, 맥락을 확장해 제공합니다.

3.1 Measures of Central Tendency for Ungrouped Data

Definition and Importance

Measures of central tendency are statistical values that describe the center or typical value of a data set. The three main measures are the mean, median, and mode. These measures help summarize large data sets with a single representative value.

Section heading: Measures of Central Tendency for Ungrouped Data

Mean

The mean (or average) is calculated by dividing the sum of all values by the number of values in the data set. It is sensitive to every value, including outliers.

  • Population Mean:

  • Sample Mean:

  • Where is the sum of all values, is the population size, is the sample size, is the population mean, and is the sample mean.

Example: Table 3.1 shows cash donations by eight U.S. companies in 2010. The mean donation is calculated by summing all donations and dividing by 8.

Table 3.1: Cash Donations in 2010 by Eight U.S. Companies

Median

The median is the value of the middle term in a data set arranged in increasing order. If the number of values is odd, the median is the middle value; if even, it is the average of the two middle values. The median is less affected by outliers than the mean.

  • Steps to Find the Median:

    1. Rank the data in increasing order.

    2. Identify the middle term(s).

Example: Table 3.2 lists the number of homes foreclosed in seven states in 2010. The median is the fourth value when the data is ordered.

Table 3.2: Number of Homes Foreclosed in 2010Ordered data showing the median value

Example: Table 3.3 shows the total compensation of 12 highest-paid CEOs in 2010. The median is the average of the 6th and 7th values in the ordered list.

Table 3.3: Total Compensations of 12 Highest-Paid CEOs for the Year 2010Ordered data showing the median value for CEO compensation

Mode

The mode is the value that occurs most frequently in a data set. A data set may have no mode, one mode (unimodal), two modes (bimodal), or more than two modes (multimodal). The mode can be used for both quantitative and qualitative data.

  • Unimodal: One mode

  • Bimodal: Two modes

  • Multimodal: More than two modes

Relationships Among the Mean, Median, and Mode

The relationship among mean, median, and mode depends on the shape of the data distribution:

  • Symmetric Distribution: Mean = Median = Mode

  • Right-Skewed Distribution: Mean > Median > Mode

  • Left-Skewed Distribution: Mean < Median < Mode

Symmetric distribution: mean = median = modeRight-skewed distribution: mean > median > modeLeft-skewed distribution: mean < median < mode

3.2 Measures of Dispersion for Ungrouped Data

Section heading: Measures of Dispersion for Ungrouped Data

Range

The range is the difference between the largest and smallest values in a data set. It is a simple measure of dispersion but is highly sensitive to outliers.

  • Formula:

Example: Table 3.4 shows the total area of four western South-Central states. The range is calculated as the difference between the largest and smallest area values.

Table 3.4: Total Area of Four States

Variance and Standard Deviation

The variance and standard deviation measure how much the values in a data set deviate from the mean. The standard deviation is the square root of the variance and is the most commonly used measure of dispersion.

  • Population Variance:

  • Sample Variance:

  • Population Standard Deviation:

  • Sample Standard Deviation:

Example: Table shows baggage fee revenues for six airlines and the calculation of and for variance and standard deviation.

Baggage fee revenue data for airlinesSum and sum of squares for airline revenue data

Example: Table shows earnings for six employees and the calculation of and .

Sum and sum of squares for employee earnings data

3.3 Mean, Variance, and Standard Deviation for Grouped Data

Section heading: Mean, Variance, and Standard Deviation for Grouped Data

Mean for Grouped Data

For grouped data, the mean is calculated using the midpoints of the classes and their frequencies.

  • Population Mean:

  • Sample Mean:

  • Where is the class midpoint and is the class frequency.

Example: Table 3.8 shows the frequency distribution of daily commuting times for 25 employees.

Frequency distribution of daily commuting timesCalculation table for mean of grouped data

Example: Table 3.10 shows the frequency distribution of the number of orders received each day over 50 days.

Frequency distribution of number of ordersCalculation table for mean of grouped data (orders)

Variance and Standard Deviation for Grouped Data

  • Population Variance:

  • Sample Variance:

  • Short-cut formulas are also available for computational efficiency.

Example: Table shows the calculation for variance and standard deviation of daily commuting times.

Frequency and squared midpoints for commuting timesCalculation table for variance and standard deviation of grouped data

Example: Table shows the calculation for variance and standard deviation of number of orders received.

Frequency and squared midpoints for number of ordersCalculation table for variance and standard deviation of grouped data (orders)

Use of Standard Deviation

Chebyshev’s Theorem

Chebyshev’s theorem states that for any data set (regardless of shape), at least of the data values lie within standard deviations of the mean, where .

Chebyshev's theorem illustrationChebyshev's theorem for k=2Chebyshev's theorem for k=3

Example: For a mean of 187 and standard deviation of 22, at least 75% of values lie between 143 and 231.

Calculation for Chebyshev's theorem exampleChebyshev's theorem applied to systolic blood pressure

Empirical Rule

For bell-shaped (normal) distributions:

  • About 68% of values lie within 1 standard deviation of the mean

  • About 95% within 2 standard deviations

  • About 99.7% within 3 standard deviations

Empirical rule illustration

Example: For a mean age of 40 and standard deviation of 12, about 95% of people are between 16 and 64 years old.

Empirical rule applied to age distribution

Measures of Position

Quartiles and Interquartile Range (IQR)

Quartiles divide a ranked data set into four equal parts. The interquartile range (IQR) is the difference between the third and first quartiles and measures the spread of the middle 50% of the data.

  • Formula:

Quartiles illustration

Example: Table 3.3 (CEO compensation) is used to find quartiles and IQR.

Quartile calculation for CEO compensation

Example: Quartile calculation for employee ages.

Quartile calculation for employee ages

Percentiles and Percentile Rank

Percentiles divide a data set into 100 equal parts. The k-th percentile is the value below which k% of the data fall. The percentile rank of a value is the percentage of values in the data set that are less than that value.

Percentiles illustration

Box-and-Whisker Plot

A box-and-whisker plot visually displays the center, spread, and skewness of a data set using the median, quartiles, and extremes (excluding outliers).

Box-and-whisker plot constructionBox-and-whisker plot with outlier

Pearson Logo

스터디 프렙