Introductory Statistics Key Concepts
Termini in questo insieme (34)
Nominal level classifies data into categories without any order. Example: yes, no, undecided.
Ordinal level classifies data into categories with a meaningful order but no fixed interval. Example: course grades A, B, C, D, F.
Interval level has ordered categories with meaningful differences but no natural zero. Example: years.
Ratio level has ordered categories, meaningful differences, and a natural zero. Example: heights.
Every member of the population has an equal chance of being selected.
Select a starting point and then every kth element. Example: every 3rd car is chosen.
Divide the population into groups and randomly sample from each group.
Divide the population into sections, randomly select some sections, and include all members from those sections.
Cross-sectional: data collected at one point in time.
Retrospective: data collected from the past.
Prospective: data collected in the future.
The difference between two consecutive lower class limits in a frequency distribution. Calculated as (Max data value - Min data value) ÷ number of classes.
The value in the middle of a class interval, calculated as (Lower class limit + Upper class limit) ÷ 2.
Lower class limit: smallest number in a class.
Upper class limit: largest number in a class.
Numbers that separate classes without gaps, found by averaging adjacent class limits and adjusting by 0.5.
Organize data into equal intervals, count frequencies, and draw adjacent bars with heights representing frequencies and no gaps.
Data is normal if points lie close to a straight line without systematic deviations; otherwise, it is not normal.
Plot time on the x-axis and measured values on the y-axis, connecting points with lines to show trends over time.
Graphs with a nonzero vertical axis start above zero to exaggerate differences between groups.
The average of data values, found by summing all values and dividing by the number of values. Sensitive to outliers.
The middle value when data is ordered. Not affected by outliers.
The value that occurs most frequently in a data set.
The midpoint between the maximum and minimum data values, calculated as (Max + Min) ÷ 2. Sensitive to outliers.
Multiply each data point by its weight, sum these products, then divide by the sum of the weights.
Measures data spread around the mean. Calculated as the square root of the sum of squared deviations divided by n-1. Notation: s (sample), 𝞂 (population).
Most data lie within 2 standard deviations of the mean. Values beyond mean ± 2𝞂 are considered significantly low or high.
The number of standard deviations a data point is from the mean, calculated as (x - mean) ÷ standard deviation.
Calculated as the number of favorable outcomes divided by the total number of possible outcomes.
P(A and B) = P(A) × P(B|A). If events are independent, P(A and B) = P(A) × P(B).
P(A or B) = P(A) + P(B) - P(A and B) for events that are not mutually exclusive.
Independent: selection with replacement; dependent: selection without replacement.
A number of successes is significant if P(X or more) ≤ 0.05 (high) or P(X or fewer) ≤ 0.05 (low).
For large samples (n > 30), the sampling distribution of the sample mean is approximately normal, regardless of population distribution.
Calculate sample statistic, find critical value, compute margin of error, then add and subtract it from the sample statistic to get the interval.
Compare P-value to significance level α: if P ≤ α, reject null hypothesis; if P > α, fail to reject null hypothesis.
Type I: reject true null hypothesis.
Type II: fail to reject false null hypothesis.