Introductory Statistics: Probability and Distributions
Termini in questo insieme (23)
Probability is a numerical measure of how likely an event is to occur, ranging from 0 (impossible) to 1 (certain).
An experiment is a process with uncertain outcomes. The sample space (S) is all possible outcomes. An event is a subset of the sample space.
P(A) = Number of outcomes in A / Total number of outcomes in S.
For any event A, 0 ≤ P(A) ≤ 1. For the sample space S, P(S) = 1.
The intersection (A ∩ B) is the event where both A and B occur, containing outcomes common to both.
The union (A ∪ B) is the event where either A or B or both occur, containing all outcomes in A or B.
Two events are disjoint if they cannot occur simultaneously, so P(A ∩ B) = 0.
P(A ∪ B) = P(A) + P(B) - P(A ∩ B). For disjoint events, P(A ∪ B) = P(A) + P(B).
The complement (Ac) includes all outcomes in S not in A, with P(Ac) = 1 - P(A).
P(A|B) is the probability of event A occurring given that event B has occurred, calculated as P(A ∩ B) / P(B).
P(A ∩ B) = P(A) × P(B|A) = P(B) × P(A|B), the probability both A and B occur.
Events A and B are independent if P(A|B) = P(A), equivalently P(A ∩ B) = P(A) × P(B).
For independent events, P(A ∪ B) = P(A) + P(B) - P(A) × P(B).
A random variable assigns a numerical value to each outcome of a random experiment, denoted by a capital letter like X.
A probability distribution lists all possible values of a random variable and their associated probabilities.
All probabilities must be between 0 and 1, and the sum of all probabilities must equal 1.
The expected value is the long-run average outcome, calculated as the sum of each value times its probability.
For approximately normal data: ~68% within 1 SD, ~95% within 2 SDs, and ~99.7% within 3 SDs of the mean.
The normal distribution is symmetric, bell-shaped, defined by mean µ and standard deviation σ, and models many natural phenomena.
A normal distribution with mean 0 and standard deviation 1, denoted by Z.
Use StatCrunch's Normal Calculator with known mean and SD to find area under the curve for intervals.
Finding the value of a random variable corresponding to a given cumulative probability in a normal distribution.
P(A) = P(B) × P(A|B) + P(Bc) × P(A|Bc), combining probabilities over different pathways.