8.1 Probabilities
- A probability formally measures uncertainty by representing the degree of belief (0…1)
- A random variable (ตัวแปรเชิงสุ่ม) is a variable in probability theory with:
- A domain of possible values
- An associated probability distribution
Examples of Random Variables
Toothache Example:
- Variable: 'a person has toothache'
- Domain: {
true, false} → Boolean random variable
- Probabilities:
- p(Toothache=true)=p(toothache)=0.2
- p(Toothache=false)=p(¬toothache)=0.8
- P(Toothache)=⟨0.2,0.8⟩
- P เรียกว่า Probability distribution - collection of probability of all possible case
- ∑ รวมกันทั้งหมดก็ต้องเป็น 1 ถูกมะ!!
เขียนได้สองแบบ ปกติกับ Short hand: Toothache=true≡toothache
Weather Example:
- Variable: 'a day's weather'
- Domain: {sun, rain, cloud, snow}
- Probabilities:
- p(Weather=sun)=0.6
- p(Weather=rain)=0.1
- p(Weather=cloud)=0.29
- p(Weather=snow)=0.01
- P(Weather)=⟨0.6,0.1,0.29,0.01⟩
- เช่นเดียวกัน รวมกัน 0.6+0.1+0.29+0.01=1
Basic Probability Rules
Two types or probability
Unconditional and Conditional Probability
- Unconditional probability p(X): degree of belief without any other information
- Conditional probability p(X∣Y): likelihood of X when ==Y is evidence== (Y is known!)
- เวลาอ่าน ให้อ่านว่า “X given Y”
Conditional Probability:
p(X∣Y)=p(Y)p(X∧Y)=p(Y)p(X∩Y)=p(Y)p(X,Y)
(when p(Y)>0)
Product Rule:
p(X,Y)=p(X∣Y)p(Y)
p(X,Y)=p(Y∣X)p(X)
Note: We use p(X,Y) to denote p(X∩Y)
For Three Variables:
p(X,Y,Z)=p(X∣Y,Z)p(Y∣Z)p(Z)
หรือ
p(X,Y,Z)=p(Z∣X,Y)p(Y∣X)p(X)
Chain Rule
For n variables:
p(Xn,Xn−1,...,X2,X1)=i=1∏np(Xi∣Xi−1,...,X1)
Expanded form:
=p(Xn∣Xn−1,...,X1)p(Xn−1∣Xn−2,...,X1)⋯p(X2∣X1)p(X1)
Boolean Variables Rules
For Boolean random variables A and B:
- Negation: p(¬a)=1−p(a) or p(A=false)=1−p(A=true)
- Inclusion-Exclusion Principle:
p(A∨B)=p(A)+p(B)−p(A,B)
Marginalization (Summing Out)
- Given a full joint probability distribution of all combinations of values
- The probability over a subset can be computed by summing out unwanted variables
Examples with Three Boolean Variables (A, B, C)
ถ้าเป็น 3 Boolean Variables อย่าง A, B, C → ก็มีได้ทั้งหมด 8 cases ถูกป่าว ทั้ง 8 อันรวมกัน เรียกว่า Full joint probability distrubution!
Marginalizing over C:
p(A,B)=p(A,B,c)+p(A,B,¬c)=∑C∈{true,false}p(A,B,C)
Marginalizing over A and C:
p(B)=p(a,B,c)+p(a,B,¬c)+p(¬a,B,c)+p(¬a,B,¬c)
=∑A∈{true,false}[∑C∈{true,false}p(A,B,C)]
Compact notation: ∑C∈{true,false}p(A,B,C) can be written as ∑Cp(A,B,C)
Example 8.1: Full Joint Distribution
Given a full joint distribution for Toothache, Cavity, Catch:
| toothache | | ¬toothache | |
|---|
| catch | ¬catch | catch | ¬catch |
| cavity | 0.108 | 0.012 | 0.072 | 0.008 |
| ¬cavity | 0.016 | 0.064 | 0.144 | 0.576 |
Find the following probability values:
-
p(cavity) = [to be calculated]
-
p(cavity∪toothache) = [to be calculated]
-
p(catch∩¬cavity) = [to be calculated]
-
p(cavity∣toothache) = [to be calculated]
-
p(catch∣(toothache∩cavity)) = [to be calculated]
8.2 Bayes' Rule
Bayes' rule flips a conditional probability:
p(Y∣X)=p(X)p(X∣Y)p(Y)
Expanded form:
p(Y∣X)=p(X,y)+p(X,¬y)p(X∣Y)p(Y)=p(X∣y)p(y)+p(X∣¬y)p(¬y)p(X∣Y)p(Y)
Independence
Random variables X and Y are independent if and only if:
p(X,Y)=p(X)p(Y)
Example 8.2: Tuberculosis Test
Problem:
- Reliability of skin test for active pulmonary tuberculosis (TB):
- Of people with TB: 98% positive reaction, 2% negative reaction
- Of people without TB: 99% negative reaction, 1% positive reaction
- From a large population where 2 per 10,000 persons have TB
- A person is selected at random, given a skin test → positive result
- Question: What is the probability the person has active pulmonary TB?
Solution: [Space left for working]
8.3 Bayesian Networks
A Bayesian Network is a directed acyclic graph representing a probabilistic graphical model.
Definition
A Bayesian Network B=(G,Θ) where:
- G is a directed graph with no cycles
- Each vertex Xi represents a random variable
- Each edge from Xi to Xj represents statistical dependency (between two variables น้า)
- Xi is a parent
- Xj is a child
- Θ represents the set of parameters specifying probability distributions for each variable
Example Structure

- Parents of X: A, B
- Children of X: C, D
- Conditional probabilities: p(X∣A), p(X∣B), p(C∣X), p(D∣X)
8.4 Conditional Independence
X and Y are independent, if the value of Z is known
Definition: Variables X and Y are conditionally independent given Z if and only if:
p(X,Y∣Z)=p(X∣Z)p(Y∣Z)
Equivalent forms:
- p(X∣Y,Z)=p(X∣Z)
- p(Y∣X,Z)=p(Y∣Z)
Joint Probability in Bayesian Networks
General chain rule:
p(X1,...,Xn)=∏i=1np(Xi∣X1,...,Xi−1)
In Bayesian Networks:
- Each Xi is conditionally independent of its ancestors given its parents
- Therefore:
p(X1,...,Xn)=i=1∏np(Xi∣parents of Xi)
Example 8.3: Joint Probability Distributions
Write the joint probability distribution for the following Bayesian networks:
ทำง่าย ๆ เลยล่ะ แค่เอาแม่ของมันมาเขียนไว้ข้างหลัง ของแต่ละตัว (ระวัง หลายตัวมีแม่หลายคนนะโกโก้)
1. Network with B, A, C, D, E

Answer: P(A,B,C,D,E)=p(A)p(B)p(C∣B)p(D∣A,C)p(E∣D)
- IF! Without the Bayesian Network P(A,B,C,D,E)=p(A)p(B∣A)p(C∣A,B)p(D∣A,B,C)p(E∣A,B,C,D)
2. Network with A, B, C, D

Answer: P(A,B,C,D)=p(A)p(B∣A)p(C∣A)p(D∣B,C)
3. Chain Network

Answer: P(A,B,C,D)=p(A)p(B∣A)p(C∣B)p(D∣C)
8.5 Common Causes (Tail-to-Tail)

- Vertex C is tail-to-tail w.r.t. the path from A to B
Probability Relations
p(A,B,C)=p(A∣C)p(B∣C)p(C)
p(A,B)=∑c∈Cp(A∣c)p(B∣c)p(c)=p(A)p(B)
p(A,B∣C)=p(C)p(A∣C)p(B∣C)p(C)=p(A∣C)p(B∣C)
Conclusion
- A and B are NOT independent
- A and B ARE conditionally independent given C
8.6 Causal Chain (Head-to-Tail)

- Vertex C is head-to-tail w.r.t. the path from A to B
Probability Relations
p(A,B,C)=p(A)p(C∣A)p(B∣C)
p(A,B)=∑c∈Cp(A)p(c∣A)p(B∣c)=p(A)p(B)
p(A,B∣C)=p(C)p(A)p(C∣A)p(B∣C)=p(A∣C)p(B∣C)
- แอบใช้ Bayes’ rule นะ p(C)p(A)p(C∣A)=p(A∣C)
Conclusion
- A and B are NOT independent
- A and B ARE conditionally independent given C
8.7 Common Effects (Head-to-Head)

- Vertex C is head-to-head w.r.t. the path from A to B
Probability Relations
p(A,B,C)=p(A)p(B)p(C∣A,B)
p(A,B)=∑c∈Cp(A)p(B)p(c∣A,B)=p(A)p(B)∑c∈Cp(c∣A,B)=p(A)p(B)
p(A,B∣C)=p(C)p(A)p(B)p(C∣A,B)=p(A∣C)p(B∣C)
Conclusion
- A and B ARE independent
- A and B are NOT conditionally independent given C
Example 8.4: Burglar Alarm Network
Scenario:
- You install a new burglar alarm
- It reliably detects burglary but also responds to minor earthquakes
- Two neighbors (John and Mary) promise to call police when they hear alarm
- John: Always calls when hears alarm, but sometimes confuses alarm with phone ringing
- Mary: Likes loud music, sometimes doesn't hear alarm
- Goal: Estimate probability of burglary given evidence about who called
Bayesian Network Structure

Probability Tables
Prior Probabilities:
| p(e) | p(¬e) |
|---|
| 0.002 | 0.998 |
| p(b) | p(¬b) |
|---|
| 0.001 | 0.999 |
Alarm given Burglary and Earthquake:
| B | E | p(a∣B,E) | p(¬a∣B,E) |
|---|
| b | e | 0.95 | 0.05 |
| b | ¬e | 0.94 | 0.06 |
| ¬b | e | 0.29 | 0.71 |
| ¬b | ¬e | 0.001 | 0.999 |
John Calls given Alarm:
| A | p(j∣A) | p(¬j∣A) |
|---|
| a | 0.90 | 0.10 |
| ¬a | 0.05 | 0.95 |
Mary Calls given Alarm:
| A | p(m∣A) | p(¬m∣A) |
|---|
| a | 0.70 | 0.30 |
| ¬a | 0.01 | 0.99 |
เรารู้ข้างบนแล้วสามารถคำนวณสิ่งนี้ได้ → p(B,E,A,J,M)=p(B)p(E)p(A∣B,E)p(J∣A)p(M∣A)
Questions
1. What is the probability that the alarm has sounded but neither burglary nor earthquake has occurred, and both Mary and John call?

2. What is the probability that the alarm would sound, given that there was an earthquake?

3. What is the probability that the alarm will sound?

4. What is the probability that there was a burglary, given that John called?
[Space for solution]