Introductory Statistics Key Concepts
Termini in questo insieme (26)
Statistics is the science of collecting, organizing, summarizing, and analyzing data to draw conclusions and provide a measure of confidence in those conclusions.
Population is the entire group studied. A sample is a subset of the population. An individual is a single member of the population.
A parameter is a numerical summary of a population, while a statistic is a numerical summary of a sample.
Descriptive statistics organize and summarize data. Inferential statistics use sample data to make generalizations about a population and measure reliability.
Qualitative variables classify individuals by attributes. Quantitative variables provide numerical measures that can be meaningfully added or subtracted.
Discrete variables have countable values (e.g., number of cars). Continuous variables have infinite possible values within an interval (e.g., distance traveled).
Nominal: categories without order.
Ordinal: categories with order.
Interval: ordered with meaningful differences, no true zero.
Ratio: interval with true zero and meaningful ratios.
An observational study observes without influencing variables. An experiment manipulates explanatory variables and records responses.
Confounding variables are explanatory variables whose effects cannot be separated. Lurking variables affect the response but are not considered in the study.
Cross-sectional: data at one point in time.
Case-control: retrospective, comparing groups.
Cohort: prospective, following a group over time.
A sample where every possible sample of size n has an equal chance of selection from the population.
Stratified: population divided into strata, random samples from each.
Systematic: select every kth individual.
Cluster: randomly select entire groups or clusters.
Sampling bias occurs when the sample does not represent the population, often due to undercoverage, nonresponse bias, or response bias.
Nonresponse bias arises when selected individuals do not respond.
Response bias occurs when answers do not reflect true feelings due to question wording, interviewer error, or misrepresentation.
An experiment studies effects of factors on a response variable. Key components include treatments, experimental units, control groups, placebos, and blinding.
Single-blind: subjects unaware of treatment.
Double-blind: neither subjects nor researchers know treatments.
1. Identify problem and response variable.
2. Determine factors.
3. Choose number of experimental units.
4. Set factor levels and randomize.
5. Conduct experiment with replication.
6. Test the claim.
Experimental units are randomly assigned to treatments without grouping.
Experimental units are paired based on similarity; each pair receives two treatments to compare effects.
Experimental units are divided into homogeneous blocks; units within each block are randomly assigned treatments.
Use frequency distributions, relative frequency distributions, bar graphs, Pareto charts, and pie charts to summarize categories.
Determine if data are discrete or continuous. Use frequency tables, relative frequencies, histograms, dot plots, and stem-and-leaf plots.
Histograms display quantitative data with touching bars; bar graphs display qualitative data with separated bars.
Uniform: frequencies evenly spread.
Bell-shaped: symmetric with peak in middle.
Skewed right: long tail on right.
Skewed left: long tail on left.
Starting vertical axis at nonzero, unequal bar widths, 3D effects, clutter, and truncating scales without indication can mislead interpretation.
Use clear titles and labels, avoid distortion, minimize white space, avoid clutter and 3D, use consistent design, and include scales and data sources.