Notation
- Scalars are denoted by either small or capital letters
- e.g., , ,
- Vectors are denoted by small bold letters
- e.g., ,
- All vectors are column vectors
- Matrices and tensors are denoted by capital bold letters
- e.g., ,
- Sets are denoted by calligraphic letters
- e.g., ,
- Set members are enclosed by curly brackets
- e.g., , ,
- denotes the set of real numbers
- denotes the set of -dimensional vectors
- denotes the set of matrices of dimension
9.1 Machine Learning
Machine Learning (ML) studies how to make a computer system improve its performance using a collection of observed data.
ML focuses on:
- Building a model from a training set, and
- Using the model as a hypothesis about the world and a software to solve problems

Analogy: Think of ML like teaching a child to recognize animals. You show them many pictures of cats and dogs (training data), they learn patterns (the algorithm builds a model), and then they can identify new animals they haven't seen before (using the model).
9.2 Types of Machine Learning
1. Supervised Learning
Supervised Learning คือการ สอนให้โมเดล (model) เรียนรู้ ฟังก์ชัน (function) ที่สามารถแปลง input → output ได้ เหมือนเราพยายามหา ที่ “ซ่อนอยู่จริง ๆ” ในโลกนี้ แต่เราไม่รู้สมการของมัน จึงต้องให้โมเดล “เดา” จากตัวอย่าง (training data)
Definition: The task of learning a function that maps inputs to outputs from a given set of input-output pairs.
-
Given: A training set of input-output pairs
where each pair was generated by an unknown function
-
Goal: A supervised learning algorithm finds a function that accepts an input vector and returns a predicted output
-
Hypothesis: The function serves as an approximation of the true function . It is called a hypothesis.

Usage:
The hypothesis is used to predict an output from an unlabeled input .
Flow: Training pairs → Supervised Learning Algorithm → → Input → Output
Example: Handwritten digit recognition
- Training: Given images of handwritten digits with labels (0-9), learn function
- Testing: Use to predict the digit in a new image

Analogy: Like a teacher showing you labeled examples (this is a 5, this is a 0), and you learn to recognize new examples based on those labeled instances.
2. Unsupervised Learning
Definition: Learning from a set of unlabeled examples ().
Typical tasks include:
-
Learning a representation function that maps an unlabeled input into new representations:
-
Learning a generative model that captures the joint probability distribution of the data
- Examples of generative models:
- Bayesian Networks: Represent the probabilistic relationships among a set of variables
- Gaussian Mixture Models (GMMs): Assume that data is generated from a mixture of Gaussian distributions and can capture more complex data distributions than a single Gaussian
- Examples of generative models:
Example: Fingerprint representation
- Given a set of fingerprints without labels

- Learn a function that converts a fingerprint into a vector
- GMMs can learn parameters of a mixture of Gaussian distributions that model the data
- The learned parameters can be used to generate new samples that resemble the original data

- พอทำมาเป็น cluster, group ๆ แบบนี้สุดท้ายก็ได้ mean, covariance matrix มาเป็นข้อมูลใหม่ ๆ
เราเอา ที่ได้เนี่ย ไป compare 2 fingerprint ได้ง่าย (easier, more efficient) เช่นว่าแบบ The same มั้ยน้อออออ
Analogy: Like sorting a pile of mixed fruits without anyone telling you what categories to use. You might naturally group them by color, size, or shape - discovering patterns on your own.
3. Reinforcement Learning
Definition: Learns from interactions with an environment through rewards and penalties.
Process:
- At each time step , the agent receives:
- Current state
- Reward from the environment
- The agent then selects an action and sends it to the environment
- The environment responds by:
- Moving to a new state
- Producing a reward
- The agent uses the observed rewards to update its policy to maximize the expected cumulative future reward

Analogy: Like training a dog - it tries different actions, gets treats (rewards) for good behavior and nothing (or negative feedback) for bad behavior, gradually learning which actions lead to treats.
Example 9.1: Identifying ML Types
Identify whether the following descriptions refer to supervised, unsupervised, or reinforcement learning:
- Stock price prediction: You have historical data on stock prices, including features such as trading volume, past prices, and other market indicators. The goal is to build a model that predicts the stock price for the next day.
- Answer: Supervised learning (labeled data with features and target prices)
-
Gene expression dimensionality reduction: You have a high-dimensional dataset containing the expression levels of thousands of genes across various samples. You want to reduce the dimensionality of the data to reveal meaningful patterns that differentiate between different biological states, without any prior knowledge of these states.
- Answer: Unsupervised learning (no labels, finding patterns)
-
Satellite image object detection: You have a large set of satellite images, some of which are labeled with bounding boxes indicating where objects of interest (e.g., buildings, roads, trees) are located. You need to build a system that can detect and localize objects in new, unseen satellite images.
- Answer: Supervised learning (labeled with bounding boxes)
-
Self-driving car: A self-driving car must learn to navigate the streets by interacting with its environment. It receives input from sensors like cameras and LiDAR, and its goal is to safely drive from one location to another while avoiding obstacles and following traffic rules. The car continuously adjusts its actions based on rewards (e.g., reaching the destination safely) and penalties (e.g., collisions).
- Answer: Reinforcement learning (learns through interaction and rewards/penalties)
-
Customer segmentation: You have a collection of customer transaction records, but no labels are provided. Your goal is to group customers into different segments based on their purchasing behavior to better target marketing efforts.
- Answer: Unsupervised learning (clustering without labels)
-
Table tennis robot: A robot is trained to play table tennis by practicing against human players. It receives a positive reward when it successfully returns the ball and a negative reward when it misses. Over time, it learns to adjust its strategy to maximize successful returns.
- Answer: Reinforcement learning (learns through rewards/penalties)
9.3 Artificial Neurons
ต่อจากนี้เราจะเรียนเป็น Neural Network แล้วนะ — เริ่มจาก Artificial Neuron เลย เป็น Basic Components connection with AI (เราเอาอันนี้หลาย ๆ ตัวมา connect กันเพื่อ compute complicated task)
Definition: An artificial neuron is a basic computational unit of an artificial neural network. It is inspired by the structure of a biological neuron, but simplified for mathematical modeling.
Process:
- Takes a set of input values (มาจาก an input vector)
- Processes them with associated weights and a bias (another scalar value)
- Applies an activation function to produce an output

เรารับ Input เข้ามา เสร็จแล้วอาจจะมี Bias value มา แล้วเอามา Calculate เป็น Weighted sum (สีส้ม) หลังจากนั้นก็ส่งไปให้ (Activation function) → แล้ว Output ออกมาเป็น (a predicted output)
Basic Formulas
Given an input vector :
มันก็คือ Dot Product ของ Weight Vector กับ นั่นแหละ
where:
- = weight of input , representing importance of each input
- = bias, shifting the activation function
- = activation function
- = predicted output of the neuron
Analogy: Think of a neuron like a decision-maker at a company meeting. Each person (input) speaks, but some voices (weights) carry more influence. The decision-maker adds their own bias, processes all information, and makes a final decision (output).
9.3.1 Activation Functions
The activation function introduces non-linearity into the neuron's output.
Common activation functions:
-
Linear Function:
- Outputs the input directly
- Used in regression tasks

-
Sigmoid Function:
- Maps input to a real value between 0 and 1
- Useful for binary classification

-
Hyperbolic Tangent (tanh):
- Maps input to a real value between -1 and 1
- Zero-centered, often performs better than sigmoid
- Also binary classification

-
Rectified Linear Unit (ReLU):
- Outputs 0 for negative inputs
- Outputs the input itself for positive inputs
- Most popular in deep learning due to computational efficiency
- Used in hidden unit of neural network

Analogy: Activation functions are like different types of voting systems - linear is proportional voting, sigmoid is binary yes/no, tanh is approval/disapproval, and ReLU is "only count positive votes."
Example 9.2: Neuron Output Calculation
Calculate the output of a neuron with the following parameters:
- Input vector:
- Weights:
- Bias:
- Activation function: Sigmoid
Solution:
Example 9.3: ReLU Neuron
Calculate the output of a neuron with the following parameters:
- Input vector:
- Weights:
- Bias:
- Activation function: ReLU
Solution:
9.3.2 Artificial Neuron as a Binary Classifier
An artificial neuron can be used as a binary classifier by applying a threshold to its output.
Process:
- If output ≥ threshold → predict one class
- If output < threshold → predict the other class

- แนะนำว่า ควรเป็น Sigmoid or
Formulas:
- แค่เพิ่มมา 1 steps ข้างล่าง
มาจากอนาคต 9.4 Multilayer Perceptrons ถ้ามันอยู่บนเส้นพอดี ก็จะกลายเป็น Class 1 นะ เพราะมีเท่ากับอยู่
where:
- = threshold for classification (e.g., for sigmoid, for tanh — เราเลือกเลขที่อยู่ตรงกลาง)
- = predicted class label (0 or 1 for sigmoid, -1 or 1 for tanh)
Example 9.4: Binary Classification
Problem: A neuron with weights , bias , and sigmoid activation function is used as a binary classifier with threshold . Given input vector , determine the predicted output and predicted class label .
Solution:
อย่าลืมว่าเราเอา 0.5 เป็น Threshold
Decision Boundary
Definition: The decision boundary of a binary classifier is the set of points in the input space where the classifier changes its prediction from one class to another.
Decision Boundary of a Sigmoid Unit
For a neuron with sigmoid activation function and threshold , the decision boundary is defined by:
- ก็คือ
This is a linear equation in the input space, representing a hyperplane that separates the two classes.
Example 9.5: Decision Boundary Equation
Problem: A neuron with weights , bias , and sigmoid activation function is used as a binary classifier with threshold . Determine the equation of the decision boundary in the input space.
Solution:
The decision boundary is given by:
Or in slope-intercept form:

- ฝั่งสีแดง
- ฝั่งสีน้ำเงิน

9.3.3 Training an Artificial Neuron
Since the decision boundary of an artificial neuron is determined by its parameter values, training an artificial neuron involves:
- Adjusting its weights and bias
- ==To minimize the error between predicted outputs and actual target values==
- Using a training dataset
Flow: Training examples → Training Algorithm → Optimal parameters → Decision boundary
The function with parameters:

Analogy: Training is like tuning a musical instrument. You adjust the strings (weights and bias) bit by bit, testing the sound (predictions) against the correct notes (true labels) until it sounds right (minimizes error).
แต่สิ่งนี้ก็มี Limitation
9.4 Multilayer Perceptrons
Definition: A multilayer perceptron (MLP) is a type of artificial neural network that consists of multiple layers of neurons, including:
- An input layer
- One or more hidden layers
- An output layer
Why MLPs?
Problem: A single artificial neuron can only represent linear decision boundaries. It cannot model complex, non-linear relationships in data.
Example: The XOR function
This is a classic example of a problem that ==cannot be represented by a single neuron==.

ก็คือบางทีที่มีข้อมูลเป็นแบบนี้อะ มันไม่สามารถแบ่ง/classify แยกออกเป็นสองฝั่งได้ด้วยเส้นตรงเส้นเดียว!!
Analogy: A single neuron is like trying to separate apples from oranges with one straight cut - it works for simple cases. But for complex patterns (like separating mixed fruit salad), you need multiple cuts from different angles (multiple layers).
MLP Computation
Given an input vector , the output of the -th neuron in layer is computed as:
where:
- = number of neurons in the previous layer
- = weight connecting the -th neuron in layer to the -th neuron in layer
- = bias of the -th neuron in layer
- = output of the -th neuron in layer , with for the input layer
- = weighted sum input to the -th neuron in layer
- = output of the -th neuron in layer
Example 9.6: MLP for XOR Problem
Design a multilayer perceptron for the XOR problem.

- อ่านได้ว่า Weight from i-th neuron of layer to j-th neuron of layer
- Superscript represents layer
- Subscript represents from which unit to which unit
Architecture:
- Input layer (layer 0): ,
- Input vector:
- Hidden layer (layer 1): 2 neurons with sigmoid activation → ,
- Output layer (layer 2): 1 neuron with sigmoid activation → =
- All layers are fully connected (อันนี้แปลว่าแต่ละ Layer ไป Connect ตัวต่อไปแบบ every combination possible)
Computation (element-wise):
Matrix form:
Layer 1:
Layer 2:
Forward propagation:
Note: is a column vector, and is also a column vector.
Example 9.7: MLP Output Calculation
Calculate the output of a multilayer perceptron with the following parameters:
- Input vector:
- Weights and biases:
- Activation function: Sigmoid for all neurons
Solution:
Layer 1:
Layer 2:
หรือจะตอบได้อีกว่า
Visualization of the MLP for XOR
มาโชว์ว่ามันเวิร์คจริง ๆ แล้วนะ กับ Multilayer Perceptron เนี่ย
The hidden layer of the MLP for XOR transforms the input space into a new space where the classes are linearly separable.

Transformation table:
| 0 | 0 | 0.34 | 0.92 |
| 0 | 1 | 0.06 | 0.19 |
| 1 | 0 | 0.72 | 1.00 |
| 1 | 1 | 0.26 | 0.94 |

- สุดท้ายก็แบ่งโดยใช้ one line ได้เลยล่ะ
Example 9.8: MLP Architecture Analysis
From the following weights and biases of an MLP, answer the following questions:
All neurons use the sigmoid activation function.
Questions:
- How many input features does the MLP accept?
- Answer: 3 (from the number of columns in )
- How many neurons are there in the hidden layer?
- Answer: 2 (from the number of rows in )
- ข้อนี้มี 3 layer ไง Layer 0 (input), Layer 1 (hidden), Layer 2 (output)
- Write down the equations to compute the output of the MLP.
- Answer:
- Answer:
Number of Parameters in an MLP
The number of parameters (weights and biases) in an MLP can be calculated as:
where:
- = total number of layers (excluding the input layer)
- = number of neurons in the previous layer
- = number of neurons in the current layer
The formula accounts for:
- = weights connecting layer to layer (สองอันแรก)
- = biases for layer (อันท้าย ซึ่ง = )
Example 9.9: Counting Parameters
Calculate the total number of parameters in an MLP with the following architecture:
- Input layer: 3 neurons
- Hidden layer: 5 neurons
- Output layer: 2 neurons
Solution:
Layer 1 (Input → Hidden):
- Weights:
- Biases:
- Subtotal:
Layer 2 (Hidden → Output):
- Weights:
- Biases:
- Subtotal:
Total parameters:
Using the formula:
Summary
Key Concepts
-
Machine Learning Types:
- Supervised: Learn from labeled data
- Unsupervised: Find patterns in unlabeled data
- Reinforcement: Learn from rewards and penalties
-
Artificial Neurons:
- Compute weighted sum:
- Apply activation:
- Can classify with threshold
-
Activation Functions:
- Linear, Sigmoid, Tanh, ReLU
- Introduce non-linearity
-
Multilayer Perceptrons:
- Multiple layers enable non-linear decision boundaries
- Forward propagation through layers
- Parameters = weights + biases
Final Analogy: Building an MLP is like constructing a decision-making organization. The input layer receives information, hidden layers are departments that process it from different perspectives, and the output layer makes the final decision based on all processed information.