IndietroChapter 3
Guida di studio - Note intelligenti
Appunti personalizzati basati sui tuoi materiali, ampliati con definizioni chiave, esempi e contesto.
Measures of Central Tendency
Definition and Overview
Measures of central tendency are statistical values that describe the center point or typical value of a dataset. The three primary measures are the mean, median, and mode. These measures help summarize a large set of data with a single representative value.
Mean: The arithmetic average, calculated by summing all values and dividing by the number of observations.
Median: The middle value when data are arranged in order. If the number of observations is even, the median is the average of the two middle values.
Mode: The value that appears most frequently in the dataset. There can be more than one mode or no mode at all.
Calculating the Mean
Sample Mean Formula:
Population Mean Formula:
Weighted Mean Formula:
Example: Calculating the sample mean for the data set: 87.2, 118.9, 76.2, 107.7, 61.5 (in thousands of dollars):
Calculating the Weighted Mean
The weighted mean is used when different values in a dataset contribute unequally to the average.
Example: Exam (94, 50%), Project (92, 35%), Homework (100, 15%)
Median and Mode
Median: Arrange data in order and use the formula to find the position.
Mode: The value(s) with the highest frequency in the dataset.
Example: For the dataset 70, 74, 81, 83, 86, 88, 91, 92, 95, the median is the fifth value (86).
Choosing the Appropriate Measure
Mean: Best for symmetric distributions without outliers.
Median: Preferred when data are skewed or contain outliers.
Mode: Useful for categorical data or to identify the most frequent value.
Measures of Variation
Definition and Overview
Measures of variation describe the spread or dispersion of data values. Common measures include the range, variance, standard deviation, and coefficient of variation.
Range: Difference between the highest and lowest values.
Variance: Average of the squared differences from the mean.
Standard Deviation: Square root of the variance; measures average distance from the mean.
Coefficient of Variation (CV): Standard deviation expressed as a percentage of the mean, useful for comparing variability between datasets with different units or means.
Formulas
Sample Variance:
Population Variance:
Sample Standard Deviation:
Population Standard Deviation:
Coefficient of Variation (Sample):
Coefficient of Variation (Population):
Example: Calculating Variance and Standard Deviation
Given the data: 10, 10, 4, 8, 13, 6, 11
Sample variance:
Sample standard deviation:
Coefficient of Variation Example
Comparing Nike and Google stock prices:
Date | Nike ($) | Google ($) |
|---|---|---|
Mean | 49.77 | 719.62 |
Standard Deviation | 3.70 | 47.96 |
CV | 7.4% | 6.7% |
Google's stock price is more consistent because its CV is lower, despite a higher standard deviation.
Using the Mean and Standard Deviation Together
z-Score
The z-score indicates how many standard deviations a value is from the mean. It is used to standardize values and identify outliers.
Population z-score:
Sample z-score:
Example: For a value of 60, mean 50, and standard deviation 20:
The Empirical Rule
For bell-shaped (normal) distributions:
Approximately 68% of values fall within ±1 standard deviation from the mean.
Approximately 95% within ±2 standard deviations.
Approximately 99.7% within ±3 standard deviations.
Chebyshev’s Theorem
For any distribution (not necessarily normal), at least of values fall within z standard deviations from the mean, for .
At least 75% within ±2 standard deviations
At least 89% within ±3 standard deviations
At least 94% within ±4 standard deviations
Measures of Relative Position
Percentiles and Quartiles
Measures of relative position compare the position of one value in relation to others in the dataset.
Percentiles: Divide data into 100 equal parts. The pth percentile is the value below which p% of the data fall.
Quartiles: Divide data into four equal parts: Q1 (25th percentile), Q2 (50th percentile/median), Q3 (75th percentile).
Percentile Index Formula:
Box-and-Whisker Plot
A boxplot visually displays the five-number summary: minimum, Q1, median (Q2), Q3, and maximum. It also identifies outliers using the interquartile range (IQR).
IQR:
Upper Limit:
Lower Limit:
Values outside these limits are considered outliers.




Descriptive Statistics in Excel
Using Excel for Descriptive Statistics
Excel provides built-in functions and tools for calculating descriptive statistics:
Mean: =AVERAGE(data)
Median: =MEDIAN(data)
Mode: =MODE.SNGL(data) or =MODE.MULT(data)
Sample Variance: =VAR.S(data)
Population Variance: =VAR.P(data)
Sample Standard Deviation: =STDEV.S(data)
Population Standard Deviation: =STDEV.P(data)
Coefficient of Variation: =STDEV.S(data)/AVERAGE(data)*100
Percentiles: =PERCENTILE.EXC(array, k)
Quartiles: =QUARTILE.EXC(array, quart)





Interpreting Histograms and Boxplots
Histogram Analysis
Histograms display the distribution of data, showing frequency of values in intervals (bins). They help identify the shape (symmetric, skewed), modality (unimodal, bimodal), and outliers.
Boxplot Analysis
Boxplots are preferred for detecting outliers and comparing groups. The position of the median and the length of the whiskers indicate skewness and spread.
Practice Problems and Applications
Example: Music Downloads
Given hourly download data, you can construct a histogram, calculate mean, median, range, IQR, and standard deviation, and summarize the distribution's shape and spread.





Example: Food Store Sales
For a right-skewed distribution, the median is a better measure of central tendency than the mean, and the IQR is preferred over the standard deviation for measuring spread.



Example: Ozone Levels
Boxplots can be used to compare monthly ozone levels, identify months with the highest values, largest IQR, and smallest range, and observe annual patterns.






