Introductory Statistics Vocabulary and Concepts (Chapters 1-3)
Termini in questo insieme (28)
Statistics is the method of collecting, organizing, summarizing data, studying probability, and making inferences based on data analysis.
A population is the entire group of individuals or items that we want information about.
A sample is a subset of the population used to make inferences about the whole population.
A parameter is a numerical summary describing a characteristic of a population.
A statistic is a numerical summary describing a characteristic of a sample.
Four scales: nominal, ordinal, interval, and ratio, each with different properties and implications for analysis.
Nominal scale classifies data into distinct categories without any order (e.g., gender, colors).
Ordinal scale classifies data with a meaningful order but without consistent differences between ranks (e.g., rankings).
Interval scale has ordered categories with equal intervals but no true zero point (e.g., temperature in Celsius).
Ratio scale has all properties of interval scale and a true zero point, allowing for meaningful ratios (e.g., weight, height).
Measures that describe the center of a data set: mean, median, and mode.
The mean is the arithmetic average of data values, calculated by summing all values and dividing by the number of values.
The median is the middle value when data are ordered from smallest to largest.
The mode is the most frequently occurring value in a data set.
Measures that describe the spread or dispersion of data: range, variance, and standard deviation.
The range is the difference between the maximum and minimum data values.
Variance measures the average squared deviation of each data point from the mean.
The standard deviation is the square root of the variance, representing average distance from the mean.
The sample space is the set of all possible outcomes in a probability experiment.
Probability quantifies the likelihood of an event occurring, ranging from 0 (impossible) to 1 (certain).
For a discrete random variable, the mean is the expected value, and the variance measures spread around the mean.
The Central Limit Theorem states that the sampling distribution of the sample mean approaches a normal distribution as sample size increases.
A confidence interval estimates a population parameter with a range of values and a specified confidence level.
Hypothesis testing is a method to decide if sample data supports a specific claim about a population.
A Type I error occurs when a true null hypothesis is incorrectly rejected.
A Type II error occurs when a false null hypothesis is not rejected.
The p-value measures the probability of observing data as extreme as the sample, assuming the null hypothesis is true.
Sample distribution describes data from a sample; population distribution describes the entire population.