This lesson on Bayesian Inference — The Update Game is hands-on and example-driven. You will learn to algebraically derive Bayes' Theorem from the definition of conditional probability. This allows you to calculate an unknown conditional probability P(A|B) using its inverse P(B|A) and prior estimates when full population data is unavailable.
What You'll Be Able To Do
- Derive Bayes' Theorem using the symmetric property of joint probability.
- Calculate conditional probabilities using partial information and prior estimates.
- Distinguish between the prior probability P(A) and the posterior probability P(A|B).
- Apply the standard Bayesian formula to update initial probability estimates based on new evidence.
- Isolate the joint probability term P(A ∩ B) from conditional probability expressions.
Detailed Concept Walkthrough
1. Conditional Probability Review
Conditional probability defines the likelihood of an event occurring given that another event has already occurred, effectively shrinking the sample space. It is calculated by dividing the joint probability by the marginal probability of the known condition.
- Mechanism: P(A|B) is the fraction of times A occurs within the subset where B is true. The numerator P(A ∩ B) counts the joint occurrences, while the denominator P(B) normalizes this count to the new, smaller universe defined by B.
- The Scaling Effect: Changing the 'given that' condition (the denominator) fundamentally changes the resulting conditional probability, even if the numerator (the joint event) remains the same. Knowledge acts as a scaling factor on the probability space.
- Best Practice: Always ensure the condition P(B) is non-zero. If P(B)=0, the condition is impossible, and the conditional probability is undefined, indicating an impossible event has occurred.
Key Takeaway: Conditional probability is a ratio of the joint probability to the probability of the condition, representing a scaled view of the sample space.
2. Symmetric Joint Probability
The probability of two events A and B occurring together, P(A ∩ B), can be expressed in two equivalent ways, regardless of which event is treated as the condition first. This algebraic symmetry is the foundation of Bayes' Theorem.
- Algebraic Isolation: Starting with the definition P(A|B) = P(A ∩ B) / P(B), we isolate the joint term by multiplying both sides by P(B), yielding P(A ∩ B) = P(A|B) ⋅ P(B).
- Equivalency: Since the joint probability is symmetric (P(A ∩ B) = P(B ∩ A)), we can also write P(B ∩ A) = P(B|A) ⋅ P(A). Setting the two expressions for the joint probability equal establishes the core identity: P(A|B)P(B) = P(B|A)P(A).
- Under the Hood: This identity proves that the order in which you multiply the conditional probability by the marginal probability does not matter; the resulting joint probability is invariant, linking the forward and inverse conditional probabilities.
Key Takeaway: The joint probability P(A ∩ B) is symmetric and can be calculated using either P(A|B)P(B) or P(B|A)P(A).
3. The Bayesian Update Formula
Bayes' Theorem allows us to calculate a difficult-to-measure conditional probability (the posterior, P(A|B)) using its inverse (the likelihood, P(B|A)) and the prior probability (P(A)).
- Derivation: We take the symmetric joint probability equation P(A|B)P(B) = P(B|A)P(A) and solve for the desired term, P(A|B), by dividing both sides by P(B).
- Standard Form: The resulting formula is P(A|B) = [P(B|A) ⋅ P(A)] / P(B). This structure is universally described as Posterior = (Likelihood ⋅ Prior) / Evidence.
- Best Practice: When applying the theorem, clearly identify the prior P(A) (initial belief), the likelihood P(B|A) (how well the evidence fits the belief), and the marginal evidence P(B) (the overall probability of the evidence occurring).
P_D = 0.01 # Prior P(Disease)
P_T_given_D = 0.95 # Likelihood P(Test+|Disease)
P_T = 0.10 # Marginal Evidence P(Test+)
# Bayes' Theorem: P(D|T) = [P(T|D) * P(D)] / P(T)
P_D_given_T = (P_T_given_D * P_D) / P_T
print(f"Posterior P(D|T): {P_D_given_T:.4f}")
# The prior (0.01) is updated to the posterior (0.095)
Key Takeaway: Bayes' Theorem provides a formal, algebraic method for updating our initial belief (prior) in light of new data (evidence).
4. Necessity of Prior Beliefs
Bayes' Theorem is essential when full population data is unavailable, necessitating the use of estimated probabilities or 'priors' for the marginal probabilities, which are often derived from historical data.
- Incomplete Data: In many real-world applications, we know the likelihood P(B|A) (e.g., test accuracy) and the prior P(A) (e.g., base rate), but we lack the full contingency table needed to calculate the marginal evidence P(B) directly.
- The Prior's Role: The prior probability P(A) represents our initial belief about event A before observing evidence B. It is a necessary input that distinguishes Bayesian inference from classical frequentist methods.
- Calculating Evidence: If the marginal evidence P(B) is not directly given, it must be calculated using the Law of Total Probability, summing the joint probabilities over all possible conditions: P(B) = P(B|A)P(A) + P(B|¬A)P(¬A).
- Philosophical Nuance: The reliance on estimated priors introduces subjectivity, which is the core philosophical difference between Bayesian statistics (incorporating belief) and Frequentist statistics (relying solely on observed frequencies).
# Calculate Marginal Evidence P(B) using Law of Total Probability
P_A = 0.01 # Prior
P_not_A = 1 - P_A # 0.99
P_B_given_A = 0.95 # Likelihood
P_B_given_not_A = 0.05 # False positive rate
# P(B) = P(B|A)P(A) + P(B|not A)P(not A)
P_B = (P_B_given_A * P_A) + (P_B_given_not_A * P_not_A)
print(f"Marginal Evidence P(B): {P_B:.4f}")
Key Takeaway: When full data is missing, the prior P(A) and the calculated marginal evidence P(B) are necessary inputs to compute the posterior.
Topics Covered in Bayesian Inference — The Update Game
- Conditional Probability (0:22 - 2:11) — The definition of conditional probability is reviewed as a ratio of joint probability to the marginal probability of the condition.
- Scaling Effect (2:12 - 3:59) — The tutorial demonstrates how changing the 'given that' condition fundamentally changes the resulting conditional probability.
- Joint Equivalency Derivation (4:00 - 5:49) — The algebraic identity P(A|B)P(B) = P(B|A)P(A) is established by isolating the joint probability term.
- Bayes' Theorem Stated (5:50 - 7:49) — The symmetric equation is rearranged to derive the standard form of Bayes' Theorem, solving for the posterior P(A|B).
- Application with Priors (7:50 - 9:40) — The theorem is applied using hypothetical prior estimates, illustrating its utility when full data is unavailable.
- Notation and Philosophy (9:41 - 10:55) — The standard abbreviated notation is connected to the explicit algebraic derivation and the broader field of Bayesian statistics.
Statistics for Data Science Cheat Sheet
-
Conditional Probability P(A|B)— Probability of A given B has occurredP(A ∩ B) / P(B) -
Joint Probability Equivalency— Symmetric expression for the probability of A and BP(A|B) * P(B) = P(B|A) * P(A) -
Bayes' Theorem (Standard Form)— Calculates posterior probability from likelihood and priorP(A|B) = [P(B|A) * P(A)] / P(B) -
Prior Probability P(A)— Initial belief before observing any new evidence BP_A = 0.01 -
Law of Total Probability— Calculates marginal evidence P(B) over all conditionsP(B|A)P(A) + P(B|¬A)P(¬A)
Comparison Table
| Component | Standard Notation | Bayesian Terminology |
|---|---|---|
| P(A) | Marginal Probability | Prior |
| P(B | A) | Conditional Probability |
| P(A | B) | Conditional Probability |
| P(B) | Marginal Probability | Evidence |
Common Pitfalls
- Mistake: Confusing the likelihood P(B|A) with the posterior P(A|B). Avoid: Always check the order of events; the posterior is what you want to find.
- Mistake: Using only the prior P(A) as the final answer after observing evidence B. Avoid: The prior must be updated using the likelihood and marginal evidence P(B).
- Mistake: Assuming P(B) is simply P(B|A)P(A) when A and ¬A are possible. Avoid: Calculate P(B) using the Law of Total Probability over all possible conditions.
FAQs
- What is the difference between a prior and a posterior? The prior P(A) is your initial belief before seeing evidence B. The posterior P(A|B) is the updated belief after incorporating the evidence B.
- Why is the marginal probability P(B) so important in the denominator? P(B) acts as a normalizing constant, ensuring the posterior probability P(A|B) remains a valid probability between 0 and 1. It represents the overall probability of observing the evidence.
- Does the order of events matter in the joint probability P(A ∩ B)? No, the order does not matter; P(A ∩ B) is mathematically identical to P(B ∩ A). This symmetry is what allows the algebraic derivation of Bayes' Theorem.