Chapter 8 - Probability Reviews and Bayesian Networks

Updated 4 Oct 2026

8.1 Probabilities

  • A probability formally measures uncertainty by representing the degree of belief (0…1)
  • A random variable (ตัวแปรเชิงสุ่ม) is a variable in probability theory with:
    • A domain of possible values
    • An associated probability distribution

Examples of Random Variables

Toothache Example:

  • Variable: 'a person has toothache'
  • Domain: {true, false} → Boolean random variable
  • Probabilities:
    • p(Toothache=true)=p(toothache)=0.2p(\text{Toothache} = \text{true}) = p(\text{toothache}) = 0.2
    • p(Toothache=false)=p(¬toothache)=0.8p(\text{Toothache} = \text{false}) = p(\neg\text{toothache}) = 0.8
    • P(Toothache)=⟨0.2,0.8⟩\mathbf{P}(\text{Toothache}) = \langle 0.2, 0.8 \rangle
      • P\mathbf{P} เรียกว่า Probability distribution - collection of probability of all possible case
      • ∑\sum รวมกันทั้งหมดก็ต้องเป็น 1 ถูกมะ!!

เขียนได้สองแบบ ปกติกับ Short hand: Toothache=true≡toothache\text{Toothache} = \text{true}\equiv \text{toothache}

Weather Example:

  • Variable: 'a day's weather'
  • Domain: {sun, rain, cloud, snow}
  • Probabilities:
    • p(Weather=sun)=0.6p(\text{Weather} = \text{sun}) = 0.6
    • p(Weather=rain)=0.1p(\text{Weather} = \text{rain}) = 0.1
    • p(Weather=cloud)=0.29p(\text{Weather} = \text{cloud}) = 0.29
    • p(Weather=snow)=0.01p(\text{Weather} = \text{snow}) = 0.01
    • P(Weather)=⟨0.6,0.1,0.29,0.01⟩\mathbf{P}(\text{Weather}) = \langle 0.6, 0.1, 0.29, 0.01 \rangle
      • เช่นเดียวกัน รวมกัน 0.6+0.1+0.29+0.01=10.6+0.1+0.29+0.01=1

Basic Probability Rules

Two types or probability

Unconditional and Conditional Probability

  • Unconditional probability p(X)p(X): degree of belief without any other information
  • Conditional probability p(X∣Y)p(X | Y): likelihood of XX when ==YY is evidence== (YY is known!)
    • เวลาอ่าน ให้อ่านว่า “XX given YY”

Key Formulas

Conditional Probability:
p(X∣Y)=p(X∧Y)p(Y)=p(X∩Y)p(Y)=p(X,Y)p(Y)\boxed{p(X | Y) = \frac{p(X \land Y)}{p(Y)}=\frac{p(X \cap Y)}{p(Y)} = \frac{p(X, Y)}{p(Y)}}
(when p(Y)>0p(Y) > 0)

Product Rule:
p(X,Y)=p(X∣Y)p(Y)\boxed{p(X, Y) = p(X | Y)p(Y)}
p(X,Y)=p(Y∣X)p(X)\boxed{p(X, Y) = p(Y | X)p(X)}

Note: We use p(X,Y)p(X, Y) to denote p(X∩Y)p(X \cap Y)

For Three Variables:
p(X,Y,Z)=p(X∣Y,Z)p(Y∣Z)p(Z)\boxed{p(X, Y, Z) = p(X|Y,Z)p(Y|Z)p(Z)}
หรือ
p(X,Y,Z)=p(Z∣X,Y)p(Y∣X)p(X)\boxed{p(X, Y, Z) = p(Z|X,Y)p(Y|X)p(X)}

  • เช็คอีกทีว่าถูกจริงเหรอ

Chain Rule

For n variables:
p(Xn,Xn−1,...,X2,X1)=∏i=1np(Xi∣Xi−1,...,X1)\boxed{p(X_n, X_{n-1},...,X_2, X_1) = \prod_{i=1}^{n} p(X_i | X_{i-1},...,X_1)}

Expanded form:
=p(Xn∣Xn−1,...,X1)p(Xn−1∣Xn−2,...,X1)⋯p(X2∣X1)p(X1)= p(X_n | X_{n-1},...,X_1)p(X_{n-1} | X_{n-2},...,X_1) \cdots p(X_2 | X_1)p(X_1)

Boolean Variables Rules

For Boolean random variables AA and BB:

  • Negation: p(¬a)=1−p(a)p(\neg a) = 1 - p(a) or p(A=false)=1−p(A=true)p(A = \text{false}) = 1 - p(A = \text{true})
  • Inclusion-Exclusion Principle:
    p(A∨B)=p(A)+p(B)−p(A,B)\boxed{p(A \lor B) = p(A) + p(B) - p(A, B)}

Marginalization (Summing Out)

  • Given a full joint probability distribution of all combinations of values
  • The probability over a subset can be computed by summing out unwanted variables

Examples with Three Boolean Variables (A, B, C)

ถ้าเป็น 3 Boolean Variables อย่าง A, B, C → ก็มีได้ทั้งหมด 8 cases ถูกป่าว ทั้ง 8 อันรวมกัน เรียกว่า Full joint probability distrubution!

Marginalizing over C:
p(A,B)=p(A,B,c)+p(A,B,¬c)=∑C∈{true,false}p(A,B,C)p(A, B) = p(A, B, c) + p(A, B, \neg c) = \sum_{C \in \{\text{true}, \text{false}\}} p(A, B, C)

Marginalizing over A and C:
p(B)=p(a,B,c)+p(a,B,¬c)+p(¬a,B,c)+p(¬a,B,¬c)p(B) = p(a, B, c) + p(a, B, \neg c) + p(\neg a, B, c) + p(\neg a, B, \neg c)
=∑A∈{true,false}[∑C∈{true,false}p(A,B,C)]= \sum_{A \in \{\text{true}, \text{false}\}} \left[ \sum_{C \in \{\text{true}, \text{false}\}} p(A, B, C) \right]

Compact notation: ∑C∈{true,false}p(A,B,C)\sum_{C \in \{\text{true}, \text{false}\}} p(A, B, C) can be written as ∑Cp(A,B,C)\sum_C p(A, B, C)


Example 8.1: Full Joint Distribution

Given a full joint distribution for Toothache, Cavity, Catch:

toothache¬toothache
catch¬catchcatch¬catch
cavity0.1080.0120.0720.008
¬cavity0.0160.0640.1440.576

Find the following probability values:

  1. p(cavity)p(\text{cavity}) = [to be calculated]

  2. p(cavity∪toothache)p(\text{cavity} \cup \text{toothache}) = [to be calculated]

  3. p(catch∩¬cavity)p(\text{catch} \cap \neg\text{cavity}) = [to be calculated]

  4. p(cavity∣toothache)p(\text{cavity} | \text{toothache}) = [to be calculated]

  5. p(catch∣(toothache∩cavity))p(\text{catch} | (\text{toothache} \cap \text{cavity})) = [to be calculated]


8.2 Bayes' Rule

Bayes' rule flips a conditional probability:

p(Y∣X)=p(X∣Y)p(Y)p(X)\boxed{p(Y | X) = \frac{p(X | Y)p(Y)}{p(X)}}

Expanded form:
p(Y∣X)=p(X∣Y)p(Y)p(X,y)+p(X,¬y)=p(X∣Y)p(Y)p(X∣y)p(y)+p(X∣¬y)p(¬y)p(Y | X) = \frac{p(X | Y)p(Y)}{p(X, y) + p(X, \neg y)} = \frac{p(X | Y)p(Y)}{p(X | y)p(y) + p(X | \neg y)p(\neg y)}

Independence

Random variables XX and YY are independent if and only if:
p(X,Y)=p(X)p(Y)\boxed{p(X, Y) = p(X)p(Y)}

Example 8.2: Tuberculosis Test

Problem:

  • Reliability of skin test for active pulmonary tuberculosis (TB):
    • Of people with TB: 98% positive reaction, 2% negative reaction
    • Of people without TB: 99% negative reaction, 1% positive reaction
  • From a large population where 2 per 10,000 persons have TB
  • A person is selected at random, given a skin test → positive result
  • Question: What is the probability the person has active pulmonary TB?

Solution: [Space left for working]


8.3 Bayesian Networks

A Bayesian Network is a directed acyclic graph representing a probabilistic graphical model.

Definition

A Bayesian Network B=(G,Θ)B = (G, \Theta) where:

  • GG is a directed graph with no cycles
    • Each vertex XiX_i represents a random variable
    • Each edge from XiX_i to XjX_j represents statistical dependency (between two variables น้า)
      • XiX_i is a parent
      • XjX_j is a child
  • Θ\Theta represents the set of parameters specifying probability distributions for each variable

Example Structure

  • Parents of X: A, B
  • Children of X: C, D
  • Conditional probabilities: p(X∣A)p(X|A), p(X∣B)p(X|B), p(C∣X)p(C|X), p(D∣X)p(D|X)

8.4 Conditional Independence

XX and YY are independent, if the value of ZZ is known

Definition: Variables XX and YY are conditionally independent given ZZ if and only if:
p(X,Y∣Z)=p(X∣Z)p(Y∣Z)\boxed{p(X, Y | Z) = p(X | Z)p(Y | Z)}

Equivalent forms:

  • p(X∣Y,Z)=p(X∣Z)p(X | Y, Z) = p(X | Z)
  • p(Y∣X,Z)=p(Y∣Z)p(Y | X, Z) = p(Y | Z)

Joint Probability in Bayesian Networks

General chain rule:
p(X1,...,Xn)=∏i=1np(Xi∣X1,...,Xi−1)p(X_1, ..., X_n) = \prod_{i=1}^{n} p(X_i | X_1, ..., X_{i-1})
In Bayesian Networks:

  • Each XiX_i is conditionally independent of its ancestors given its parents
  • Therefore:
    p(X1,...,Xn)=∏i=1np(Xi∣parents of Xi)\boxed{p(X_1, ..., X_n) = \prod_{i=1}^{n} p(X_i | \text{parents of } X_i)}

Example 8.3: Joint Probability Distributions

Write the joint probability distribution for the following Bayesian networks:

ทำง่าย ๆ เลยล่ะ แค่เอาแม่ของมันมาเขียนไว้ข้างหลัง ของแต่ละตัว (ระวัง หลายตัวมีแม่หลายคนนะโกโก้)

1. Network with B, A, C, D, E

Answer: P(A,B,C,D,E)=p(A)p(B)p(C∣B)p(D∣A,C)p(E∣D)\mathbf{P}(A,B,C,D,E)=p(A)p(B)p(C|B)p(D|A,C)p(E|D)

  • IF! Without the Bayesian Network P(A,B,C,D,E)=p(A)p(B∣A)p(C∣A,B)p(D∣A,B,C)p(E∣A,B,C,D)\mathbf{P}(A,B,C,D,E)=p(A)p(B|A)p(C|A,B)p(D|A,B,C)p(E|A,B,C,D)

2. Network with A, B, C, D

Answer: P(A,B,C,D)=p(A)p(B∣A)p(C∣A)p(D∣B,C)\mathbf{P}(A,B,C,D)=p(A)p(B|A)p(C|A)p(D|B,C)


3. Chain Network

Answer: P(A,B,C,D)=p(A)p(B∣A)p(C∣B)p(D∣C)\mathbf{P}(A,B,C,D)=p(A)p(B|A)p(C|B)p(D|C)


8.5 Common Causes (Tail-to-Tail)

  • Vertex CC is tail-to-tail w.r.t. the path from AA to BB

Probability Relations

p(A,B,C)=p(A∣C)p(B∣C)p(C)p(A, B, C) = p(A|C)p(B|C)p(C)

p(A,B)=∑c∈Cp(A∣c)p(B∣c)p(c)≠p(A)p(B)p(A, B) = \sum_{c \in C} p(A|c)p(B|c)p(c) \neq p(A)p(B)

p(A,B∣C)=p(A∣C)p(B∣C)p(C)p(C)=p(A∣C)p(B∣C)p(A, B|C) = \frac{p(A|C)p(B|C)\cancel{p(C)}}{\cancel{p(C)}} = p(A|C)p(B|C)

Conclusion

  • AA and BB are NOT independent
  • AA and BB ARE conditionally independent given CC

8.6 Causal Chain (Head-to-Tail)

  • Vertex CC is head-to-tail w.r.t. the path from AA to BB

Probability Relations

p(A,B,C)=p(A)p(C∣A)p(B∣C)p(A, B, C) = p(A)p(C|A)p(B|C)

p(A,B)=∑c∈Cp(A)p(c∣A)p(B∣c)≠p(A)p(B)p(A, B) = \sum_{c \in C} p(A)p(c|A)p(B|c) \neq p(A)p(B)

p(A,B∣C)=p(A)p(C∣A)p(B∣C)p(C)=p(A∣C)p(B∣C)p(A, B|C) = \frac{p(A)p(C|A)p(B|C)}{p(C)} = p(A|C)p(B|C)

  • แอบใช้ Bayes’ rule นะ p(A)p(C∣A)p(C)=p(A∣C)\color{gray}\frac{p(A)p(C|A)}{p(C)}=p(A|C)

Conclusion

  • AA and BB are NOT independent
  • AA and BB ARE conditionally independent given CC

8.7 Common Effects (Head-to-Head)

  • Vertex CC is head-to-head w.r.t. the path from AA to BB

Probability Relations

p(A,B,C)=p(A)p(B)p(C∣A,B)p(A, B, C) = p(A)p(B)p(C|A, B)
p(A,B)=∑c∈Cp(A)p(B)p(c∣A,B)=p(A)p(B)∑c∈Cp(c∣A,B)=p(A)p(B)p(A, B) = \sum_{c \in C} p(A)p(B)p(c|A, B) = p(A)p(B) \sum_{c \in C} p(c|A, B) = p(A)p(B)
p(A,B∣C)=p(A)p(B)p(C∣A,B)p(C)≠p(A∣C)p(B∣C)p(A, B|C) = \frac{p(A)p(B)p(C|A, B)}{p(C)} \neq p(A|C)p(B|C)

Conclusion

  • AA and BB ARE independent
  • AA and BB are NOT conditionally independent given CC

Example 8.4: Burglar Alarm Network

Scenario:

  • You install a new burglar alarm
  • It reliably detects burglary but also responds to minor earthquakes
  • Two neighbors (John and Mary) promise to call police when they hear alarm
  • John: Always calls when hears alarm, but sometimes confuses alarm with phone ringing
  • Mary: Likes loud music, sometimes doesn't hear alarm
  • Goal: Estimate probability of burglary given evidence about who called

Bayesian Network Structure

Probability Tables

Prior Probabilities:

p(e)p(e)p(¬e)p(\neg e)
0.0020.998
p(b)p(b)p(¬b)p(\neg b)
0.0010.999

Alarm given Burglary and Earthquake:

BEp(a∣B,E)p(a \mid B,E)p(¬a∣B,E)p(\neg a \mid B,E)
be0.950.05
b¬e0.940.06
¬be0.290.71
¬b¬e0.0010.999

John Calls given Alarm:

Ap(j∣A)p(j \mid A)p(¬j∣A)p(\neg j \mid A)
a0.900.10
¬a0.050.95

Mary Calls given Alarm:

Ap(m∣A)p(m \mid A)p(¬m∣A)p(\neg m \mid A)
a0.700.30
¬a0.010.99

เรารู้ข้างบนแล้วสามารถคำนวณสิ่งนี้ได้ → p(B,E,A,J,M)=p(B)p(E)p(A∣B,E)p(J∣A)p(M∣A)p(B,E,A,J,M)=p(B)p(E)p(A|B,E)p(J|A)p(M|A)

Questions

1. What is the probability that the alarm has sounded but neither burglary nor earthquake has occurred, and both Mary and John call?

2. What is the probability that the alarm would sound, given that there was an earthquake?

3. What is the probability that the alarm will sound?


4. What is the probability that there was a burglary, given that John called?

[Space for solution]