BackDescriptive Statistics: Measures of Central Tendency, Variation, and Relative Position
Study Guide - Smart Notes
Tailored notes based on your materials, expanded with key definitions, examples, and context.
Descriptive Statistics
Introduction
Descriptive statistics are essential tools in business statistics, providing methods to summarize, organize, and interpret data. The main measures include central tendency, variability, and relative position, each offering unique insights into data sets.
Measures of Central Tendency
The Mean (Arithmetic Mean)
The mean is the most common measure of central tendency, calculated by summing all values and dividing by the number of observations. It represents the average value in a data set.
Formula (Sample Mean):
Formula (Population Mean):
Example: For data set {87.2, 118.9, 76.2, 107.7, 61.5}, the mean is
Weighted Mean
The weighted mean assigns different weights to values, useful when some data points contribute more significantly than others.
Formula:
Example: If exam, project, and homework scores are weighted 50%, 35%, and 15% respectively, the weighted mean is calculated as
The Median
The median is the middle value when data are ordered. If the number of observations is even, it is the average of the two middle values.
Formula (Index Point):
Example: For sorted data {26, 28, 31, 39, 43, 45, 45, 50, 57, 62}, the median is the average of the 5th and 6th values:
The Mode
The mode is the value that appears most frequently in a data set. Data can be unimodal, bimodal, or have no mode.
Example: In {6, 7, 7, 8, 8, 8, 8, 8, 9, 9, 9, 10, 10, 11, 11, 11, 14, 14}, the mode is 8 (appears 5 times).
Choosing the Appropriate Measure
Mean: Easy to calculate, but sensitive to outliers.
Median: Not affected by outliers, better for skewed data.
Mode: Useful for categorical data, may not always exist or may be multiple.
Measures of Variation
Range
The range is the difference between the highest and lowest values in a data set.
Formula:
Limitation: Highly affected by outliers and does not consider data distribution shape.
Variance and Standard Deviation
Variance measures the average squared deviation from the mean, while standard deviation is its square root, representing spread in the same units as the data.
Sample Variance:
Population Variance:
Sample Standard Deviation:
Population Standard Deviation:


Coefficient of Variation (CV)
The coefficient of variation expresses the standard deviation as a percentage of the mean, allowing comparison of variability between data sets with different units or means.
Sample CV:
Population CV:
Using the Mean and Standard Deviation Together
Shapes of Frequency Distribution
Data distributions can be symmetric, left-skewed, or right-skewed. Skewness measures asymmetry, while kurtosis measures the peakedness of the distribution.
Symmetric: Mean = Median
Left-skewed: Mean < Median
Right-skewed: Mean > Median
Quality Control Example: Histograms
Histograms visually display the distribution of data and help identify the mean, standard deviation, and conformity to specifications.




The z-Score
Definition and Calculation
The z-score indicates how many standard deviations a value is from the mean, standardizing different data sets for comparison.
Population:
Sample:
Interpretation: A z-score < -3 or > +3 is considered an extreme outlier.
The Empirical Rule
For bell-shaped (normal) distributions:
Approximately 68% of values fall within ±1 standard deviation from the mean
Approximately 95% within ±2 standard deviations
Approximately 99.7% within ±3 standard deviations
Chebyshev’s Theorem
For any distribution (not just normal), at least % of values fall within z standard deviations from the mean, for z > 1.
At least 75% within ±2 standard deviations
At least 89% within ±3 standard deviations
At least 94% within ±4 standard deviations
Measures of Relative Position
Percentiles
Percentiles indicate the percentage of data values below a certain point. The pth percentile is the value below which p% of the data fall.
Index Point Formula:
If i is not a whole number, round up; if i is whole, average the ith and (i+1)th values.
Quartiles
Quartiles divide data into four equal parts:
Q1: 25th percentile
Q2: 50th percentile (median)
Q3: 75th percentile
Interquartile Range (IQR)
The IQR measures the spread of the middle 50% of data and is not influenced by outliers.
Formula:
Box-and-Whisker Plots
A boxplot visually displays the five-number summary (minimum, Q1, median, Q3, maximum) and identifies outliers.
Upper Limit:
Lower Limit:
Values outside these limits are considered outliers.



Excel/PHStat Applications
Descriptive Statistics in Excel
Excel provides tools for calculating mean, median, mode, standard deviation, variance, percentiles, quartiles, and creating boxplots.
Use Data Analysis → Descriptive Statistics for summary statistics.
Functions: =AVERAGE(), =MEDIAN(), =MODE.SNGL(), =STDEV.S(), =VAR.S(), =PERCENTILE.EXC(), =QUARTILE.EXC()





Summary
Measures of Central Tendency: Mean, Median, Mode
Measures of Variation: Range, Variance, Standard Deviation, Coefficient of Variation
Measures of Location: Percentile, Quartile, Box & Whisker Plot