Skip to main content
Back

Descriptive Statistics: Measures of Central Tendency, Variation, and Relative Position

Study Guide - Smart Notes

Tailored notes based on your materials, expanded with key definitions, examples, and context.

Descriptive Statistics

Introduction

Descriptive statistics are essential tools in business statistics, providing methods to summarize, organize, and interpret data. The main categories include measures of central tendency, measures of variability, and measures of relative position. These concepts help in understanding the distribution, spread, and ranking of data within a dataset.

Measures of Central Tendency

Definition and Overview

Measures of central tendency describe the center or typical value of a dataset. The three primary measures are the mean, median, and mode. Each measure provides different insights and is appropriate under different circumstances.

  • Mean (Arithmetic Average): The sum of all values divided by the number of observations.

  • Median: The middle value when data are arranged in order.

  • Mode: The value that appears most frequently in the dataset.

The Mean

  • Sample Mean Formula:

  • Population Mean Formula:

  • Weighted Mean Formula:

  • Example: Calculating the mean for a dataset: yields .

  • Weighted Mean Example: If exam, project, and homework scores are weighted 50%, 35%, and 15% respectively, the weighted mean is calculated as:

The Median

  • Definition: The value that divides the dataset into two equal halves.

  • Finding the Median: Arrange data in order. If is odd, the median is the middle value. If $n$ is even, it is the average of the two middle values.

  • Median Position Formula:

  • Example: For the dataset , the median is $86$ (the fifth value).

The Mode

  • Definition: The value with the highest frequency in the dataset.

  • Types: Unimodal (one mode), Bimodal (two modes), Multimodal (more than two modes), or No mode.

  • Example: In the dataset , the mode is $8$.

Choosing the Appropriate Measure

  • Mean: Best for symmetric distributions without outliers.

  • Median: Preferred when data are skewed or contain outliers.

  • Mode: Useful for categorical data.

Measures of Variability

Definition and Overview

Measures of variability describe the spread or dispersion of data. Common measures include the range, variance, standard deviation, and coefficient of variation.

  • Range: Difference between the highest and lowest values.

  • Variance (Sample):

  • Variance (Population):

  • Standard Deviation: Square root of variance. or

  • Coefficient of Variation (CV): Expresses standard deviation as a percentage of the mean.

    • Sample:

    • Population:

Example: Calculating Variance and Standard Deviation

  • Given data:

  • Sample variance:

  • Sample standard deviation:

Coefficient of Variation Example

  • Nike:

  • Google:

  • Interpretation: Lower CV indicates more consistency relative to the mean.

Using the Mean and Standard Deviation Together

Shapes of Frequency Distributions

  • Symmetric: Mean = Median

  • Left-skewed: Median < Mean

  • Right-skewed: Mean < Median

Quality Control Example

Histograms can be used to visualize the distribution of data and assess conformity to specifications. The mean and standard deviation together help determine the proportion of data within specification limits.

Histogram with mean and standard deviationHistogram with shifted meanHistogram with shifted meanHistogram with reduced standard deviation

The z-Score

Definition and Calculation

The z-score indicates how many standard deviations a value is from the mean. It is used to standardize values for comparison.

  • Population z-score:

  • Sample z-score:

  • Interpretation: A z-score of 0 means the value equals the mean; positive values are above the mean, negative values are below.

  • Outliers: Values with are considered extreme outliers.

Empirical Rule

  • For bell-shaped (normal) distributions:

  • ~68% of data within ±1 standard deviation

  • ~95% within ±2 standard deviations

  • ~99.7% within ±3 standard deviations

Chebyshev’s Theorem

  • For any distribution (not just normal):

  • At least of data falls within standard deviations of the mean, for .

  • At least 75% within ±2 standard deviations, 89% within ±3, 94% within ±4.

Measures of Relative Position

Percentiles and Quartiles

  • Percentiles: Divide data into 100 equal parts. The pth percentile is the value below which p% of the data fall.

  • Quartiles: Divide data into four equal parts:

    • Q1: 25th percentile

    • Q2: 50th percentile (median)

    • Q3: 75th percentile

  • Index for Percentile:

Box-and-Whisker Plots

A boxplot visually displays the five-number summary: minimum, Q1, median (Q2), Q3, and maximum. It also identifies outliers using the interquartile range (IQR).

  • IQR:

  • Outlier Limits:

    • Upper Limit:

    • Lower Limit:

Table of national park visitors with quartilesBoxplot of national park visitorsExcel boxplot of national park visitors

Using Excel for Descriptive Statistics

Descriptive Statistics Tool

Excel provides built-in functions and analysis tools for calculating descriptive statistics such as mean, median, mode, standard deviation, variance, percentiles, and quartiles.

Excel Data Analysis Toolpak Descriptive StatisticsExcel Descriptive Statistics dialogExcel Descriptive Statistics outputExcel Descriptive Statistics output with skewness and kurtosisExcel calculation of mean, median, mode

Summary

  • Descriptive statistics summarize data using measures of central tendency, variability, and relative position.

  • Choosing the appropriate measure depends on the data’s distribution and the presence of outliers.

  • Excel is a powerful tool for performing these calculations efficiently.

Pearson Logo

Study Prep