IndietroMeasures of Variation in Descriptive Statistics
Guida di studio - Note intelligenti
Appunti personalizzati basati sui tuoi materiali, ampliati con definizioni chiave, esempi e contesto.
Measures of Variation
Introduction to Measures of Variation
Measures of variation describe how data values are spread out or dispersed within a data set. Understanding variation is essential for interpreting the reliability and consistency of data, especially when comparing different data sets with similar measures of central tendency.
Range
Range is the simplest measure of variation. It is calculated as the difference between the maximum and minimum values in a quantitative data set.
Formula:
Example: If the highest starting salary in Corporation A is $56,000 and the lowest is $46,000, then the range is $10,000.
Deviation, Variance, and Standard Deviation
While the range gives a basic idea of spread, variance and standard deviation provide more detailed information about how each data value differs from the mean.
Deviation: The difference between a data entry and the mean (population) or (sample).
Population Variance:
Population Standard Deviation:
Sample Variance:
Sample Standard Deviation:
Interpretation: Standard deviation measures the typical distance of data entries from the mean. A higher standard deviation indicates more spread out data.
Calculating Variance and Standard Deviation
List all data values and calculate the mean.
Find the deviation of each value from the mean.
Square each deviation and sum them.
Divide by (population) or (sample) for variance.
Take the square root for standard deviation.
Example: For a sample of recovery times: 8, 10, 4, 6, 7, 7, 9, 10, 7, 6, 5, 11, the sample variance is about 4.6 and the sample standard deviation is about 2.2 days.
Interpreting Standard Deviation: Empirical Rule (68–95–99.7 Rule)
The Empirical Rule applies to bell-shaped (normal) distributions:
About 68% of data lie within one standard deviation of the mean.
About 95% lie within two standard deviations.
About 99.7% lie within three standard deviations.
Example: If the mean height of women is 64.1 inches with a standard deviation of 2.6 inches, about 47.72% of women are between 58.9 inches and 64.1 inches tall (between the mean and two standard deviations below).
Chebychev’s Theorem
Chebychev’s Theorem applies to any data set, regardless of distribution shape. It states that the proportion of values within standard deviations of the mean is at least for .
For : At least 75% of data lie within two standard deviations.
For : At least about 88.9% of data lie within three standard deviations.
Example: If the mean age in Georgia is 41.7 years with a standard deviation of 20.85 years, at least 75% of the population is between 0 and 83.4 years old. An age of 90 is unusual as it is more than two standard deviations from the mean.
Standard Deviation for Grouped Data
When data are grouped into classes, estimate the mean and standard deviation using class midpoints and frequencies.
Sample Standard Deviation for Frequency Distribution:
Where is the frequency, is the class midpoint, and is the total number of entries.
Example: For the number of children in 50 households, the sample mean is about 1.8 children and the sample standard deviation is about 1.7 children.
Coefficient of Variation
The coefficient of variation (CV) expresses the standard deviation as a percentage of the mean, allowing comparison of variability between different data sets, even if units differ.
Population:
Sample:
Example: If the mean height of a basketball team is 74 inches with a standard deviation of 3.3 inches, . If the mean weight is 210 pounds with a standard deviation of 19.7 pounds, . The weights are more variable than the heights.