Skip to main content
Indietro

Measures of Central Tendency in Descriptive Statistics

Guida di studio - Note intelligenti

Appunti personalizzati basati sui tuoi materiali, ampliati con definizioni chiave, esempi e contesto.

Descriptive Statistics

Measures of Central Tendency

Measures of central tendency are values that represent a typical or central entry of a data set. They are essential for summarizing and understanding the general trend of data. The most common measures are the mean, median, and mode.

  • Mean: The arithmetic average of all data entries.

  • Median: The middle value when the data set is ordered.

  • Mode: The data entry that occurs with the greatest frequency.

Mean (Average)

The mean is calculated by dividing the sum of all data entries by the number of entries. It is a reliable measure because it uses every value in the data set, but it can be greatly affected by outliers.

  • Population Mean:

  • Sample Mean:

Example: For the sample weights 274, 235, 223, 268, 290, 285, 235, the mean is calculated as follows:

  • Sum: 274 + 235 + 223 + 268 + 290 + 285 + 235 = 1810

  • Number of entries: 7

  • Mean: pounds

Median

The median is the value that lies in the middle of an ordered data set, dividing it into two equal parts. If the number of entries is odd, the median is the middle entry. If even, it is the mean of the two middle entries.

  • Odd number of entries: Median is the middle value.

  • Even number of entries: Median is the average of the two middle values.

Example: For the ordered weights 223, 235, 235, 268, 274, 285, 290 (7 entries), the median is 268 pounds (the fourth entry).

Example (Even): If 285 is removed, the ordered set is 223, 235, 235, 268, 274, 290. The median is pounds.

Mode

The mode is the data entry that occurs with the greatest frequency. A data set may have no mode, one mode (unimodal), or more than one mode (bimodal or multimodal).

  • If no entry is repeated, there is no mode.

  • If two entries occur with the same greatest frequency, both are modes (bimodal).

Example: In the data set 223, 235, 235, 268, 274, 285, 290, the mode is 235 (occurs twice).

Example (Categorical): If survey responses for political party are tabulated, the mode is the party with the highest frequency (e.g., Democrat).

Comparing the Mean, Median, and Mode

All three measures describe a typical entry of a data set, but each has advantages and disadvantages:

  • Mean: Uses all data values but is sensitive to outliers.

  • Median: Not affected by outliers; better for skewed distributions.

  • Mode: Useful for categorical data; may not represent a typical value in numerical data.

Example: In a class with ages mostly around 20, but one student aged 65, the mean is influenced by the outlier, while the median better represents the typical age.

Weighted Mean

The weighted mean is used when data entries have different weights. It is calculated as:

Example: Calculating a grade point average (GPA) where each grade has a different credit hour weight.

Mean of Grouped Data

When data is presented in a frequency distribution, the mean can be estimated using class midpoints and frequencies:

  • Where x is the class midpoint and f is the class frequency.

Example: For cell phone screen times grouped in intervals, the estimated mean is calculated using midpoints and frequencies, yielding an approximate mean (e.g., 287.7 minutes).

The Shape of Distributions

The shape of a data distribution affects the relationship between the mean, median, and mode:

  • Symmetric Distribution: The left and right halves are mirror images; mean, median, and mode are approximately equal.

  • Uniform Distribution: All values occur with equal frequency; symmetric and rectangular in shape.

  • Skewed Left (Negatively Skewed): The tail extends to the left; mean is less than the median.

  • Skewed Right (Positively Skewed): The tail extends to the right; mean is greater than the median.

Pearson Logo

Study Prep