ก่อนหน้านี้ที่เราเรียนเป็น Supervised Learning หมดเลย
ก็คือเราเอา Pair of inputs
พอมาเป็น Unsupervised บ้าง เราให้แค่ No target output
14.1 Unsupervised Learning
- Unsupervised learning learns from a set of examples without labels/classes
- Tasks of unsupervised learning:
- Learn a function that maps an unlabeled input into new representations ()
- Learn a generative model such as a probability distribution to generate new examples
Think of unsupervised learning like exploring a new city without a map or guide - you discover patterns and structure on your own by observing the environment, rather than being told what everything is.
14.2 Autoencoder
A type of Neural Networks ของ Unsupervised Learning นั่นแหละ
Overview
- Autoencoder is an unsupervised learning technique that learns new representations of data
- Applications:
- Generate features of fingerprint images to make it easier to check if two fingerprints are from the same person
- Perform dimension reduction by transforming an input vector into a space with fewer dimensions
- ทำไงก็ remove irrelevant features
- construct new features by combining the existing ones (multiple features)
Imagine compressing a large file into a smaller one - the autoencoder learns to compress data into a compact representation while preserving the essential information.
Architecture
- Encoder creates a mapping that transforms input array into a new representation :
- Structure แบบนี้มัน Train ไม่ได้นะ (using Gradient Descent อะ) แล้วทำไงดี?
- เฉลย ต้องมี Decoder
- Since target outputs are not provided, we train this model by creating a decoder that accepts the encoded array and reconstructs the input array :
When and are approx. the same, represents
- The encoder and decoder are connected and trained to reconstruct the input values - this complete system is called an autoencoder
14.2.1 Fully-connected Autoencoder
- Architecture: Input → Latent Vector → Reconstructed Input
where:
- = input vector
- = latent vector (encoded representation)
- = reconstructed input
- = encoder weights and biases
- = decoder weights and biases
- = activation function
Loss Function
- Mean Squared Error (MSE) Loss:
-
Binary Cross Entropy: Can be used when for all
- Using sigmoid activation function
-
After training is complete, the decoder can be removed and only the encoder is used for feature extraction
Relationship to PCA
- When the activation function is set to the linear function, this autoencoder works almost identically to Principal Component Analysis (PCA) technique
- Important difference: The encoder extracted from an autoencoder does not guarantee orthogonal components
PCA is like organizing your closet by color and type in perfectly perpendicular sections, while an autoencoder organizes it in whatever way makes sense, which might not be perfectly aligned.
จะไม่ทำให้ดู ไม่ออกสอบแน่นอน
Techinique การทำ data analytics หาแกนหลักของข้อมูล ดูการตามกระจายของข้อมูล (คำนวณจากพวก covariance, ทิศไหนข้อมูลการกระจายมากที่สุด) → เราจึงลด Dimension ของมันได้ (ดูอะไรที่สำคัญที่สุด)
Example 14.1
Given:
- Input vector
- Fully-connected autoencoder with 2 hidden units and linear activation function
Encoder:
where ,
Decoder:
where ,
Task: Calculate the latent vector , reconstructed vector , and MSE loss between and .
Solution:
- Calculate latent vector :
(New representation of ) - Calculate reconstructed vector :
(Reconstruct of ) - Calculate MSE loss:
(แต่ ที่ reconstruct ก็ไม่ได้ดีขนาดนั้น ดังนั้น คำนวณหา Loss)
Multiple Hidden Layers
- We can introduce multiple hidden layers with nonlinear activation functions into an autoencoder
- This allows the autoencoder to learn more complex, non-linear representations

Example 14.2: MNIST Autoencoder
Given:
- MNIST dataset (handwritten digits)
- Encoder with 4 layers: 256, 128, 64, and 32 units
- Latent vector has 32 elements
Task: Fill in the decoder architecture
| Layer Type | Units | Activation Function |
|---|---|---|
| Input | 784 | - |
| Fully Connected Later (Encoder) | 256 | relu |
| Fully Connected Later (Encoder) | 128 | relu |
| Fully Connected Later (Encoder) | 64 | relu |
| Fully Connected Later (Encoder) | 32 | relu |
| Fully Connected Later (Encoder) | 64 | relu |
| Fully Connected Later (Encoder) | 128 | relu |
| Fully Connected Later (Encoder) | 256 | relu |
| Fully Connected Later (Encoder) | 784 | sigmoid |
![]() |
The decoder mirrors the encoder architecture but in reverse, like unfolding an origami creation back to a flat sheet.
14.2.2 Convolutional Autoencoder
- Convolutional autoencoder includes at least one convolutional layer
- Each layer in the decoding part needs to increase size of its input array to reconstruct the input
- Two main techniques for upsampling:
- Transposed Convolution
- Upsampling + Convolution
Transposed Convolution
Algorithm:
Given an input array with shape and a kernel with shape :
- Create intermediate arrays with shape , initialize all elements to zeros
- Multiply each element of the input array with the kernel
- Replace a portion of an intermediate array with the multiplication result
- Sum all intermediate arrays to create the output array
Think of transposed convolution as the opposite of regular convolution - instead of shrinking the image, you're expanding it by "painting" each input value across a larger canvas.
Example 14.3: Transposed Convolution
Given:
- Input array:
- Kernel:
Process:
For each element in input, multiply by kernel and place in output:
- (top-left)
- (top-right)
- (bottom-left)
- (bottom-right)
Sum all overlapping regions:
Example 14.4
Given:
- Input:
- Kernel:
Task: Compute the output of transposed convolution
[Left for student to complete]
Upsampling
- Upsampling conducts the reverse of a pooling operation
- When performing upsampling on a 2D array with size :
- Rows and columns are repeated by and respectively
Example: upsampling
Upsampling is like taking a pixelated image and making it bigger by duplicating each pixel - simple but effective for increasing size.
- Alternative approach: A combination of upsampling and convolutional layers can be used in place of transposed convolutional layer
14.2.3 Denoise Autoencoder
Training Process:
- Add noise to the input
- Train the autoencoder using the original input as the target
- This type of autoencoder is called denoise autoencoder
- Benefits:
- Makes the autoencoder more robust to noise
- Improves the performance of dimensionality reduction
Like learning to read messy handwriting - the autoencoder becomes better at extracting the core features by learning to ignore the noise.
\usepackage{tikz, amsmath, amssymb}
\usetikzlibrary{arrows.meta, positioning, shapes}
\begin{document}
\begin{tikzpicture}[node distance=2.5cm, >=Stealth, thick]
% Nodes
\node (clean) {Clean\\Image};
\node[right=of clean] (noisy) {Noisy\\Image};
\node[draw, rectangle, fill=cyan!30, right=of noisy, minimum height=1.5cm] (encoder) {Encoder};
\node[draw, rectangle, fill=green!30, right=of encoder, minimum height=1.5cm] (decoder) {Decoder};
\node[right=of decoder] (output) {Reconstructed\\Clean Image};
% Arrows
\draw[->] (clean) -- node[above] {Add noise} (noisy);
\draw[->] (noisy) -- (encoder);
\draw[->] (encoder) -- (decoder);
\draw[->] (decoder) -- (output);
\draw[->, dashed, red] (clean) to[bend left=45] node[above] {Target} (output);
\end{tikzpicture}
\end{document}[Image: Shows progression from clean digit → noisy digit → encoder → decoder → reconstructed clean digit]
14.3 Generative Adversarial Network (GAN)
Overview
- GAN (Generative Adversarial Network), proposed by Goodfellow et al. in 2014
- A generative model originally designed to generate new data
- Composed of two parts:
- Generator (): Learns to generate fake data
- Discriminator (): Learns to determine whether an example is real or fake
\usepackage{tikz, amsmath, amssymb}
\usetikzlibrary{arrows.meta, positioning, shapes}
\begin{document}
\begin{tikzpicture}[node distance=2.5cm, >=Stealth, thick]
% Nodes
\node[draw, rectangle, fill=gray!40] (noise) {Random\\Noise\\$z$};
\node[draw, rectangle, fill=blue!40, right=of noise, minimum height=1.5cm, minimum width=2cm] (generator) {Generator\\$G$};
\node[draw, rectangle, fill=red!30, below right=1cm and 1cm of generator] (fake) {Generated\\Example\\$G(z)$};
\node[draw, rectangle, fill=green!30, above right=1cm and 1cm of generator] (real) {Real\\Example\\$x$};
\node[draw, rectangle, fill=magenta!40, right=3cm of generator, minimum height=2cm, minimum width=2cm] (discriminator) {Discriminator\\$D$};
\node[right=of discriminator] (output) {$D(x)$ or\\$D(G(z))$};
% Arrows
\draw[->] (noise) -- (generator);
\draw[->] (generator) -- (fake);
\draw[->] (fake) -- (discriminator);
\draw[->] (real) -- (discriminator);
\draw[->] (discriminator) -- (output);
\end{tikzpicture}
\end{document}Think of GANs like an art forger (generator) trying to fool an art expert (discriminator). As the expert gets better at detecting fakes, the forger improves their technique, and vice versa.
Minimax Loss Function
Where:
- = real example
- = noise (random input)
- = probability that is real (estimated by discriminator)
- = example generated from noise
- = probability that generated example is real
- = expected value over all real examples
- = expected value over all random noises
Training Objective:
- Discriminator: Trained to maximize the loss
- Generator: Trained to minimize the loss
Training Algorithm
Algorithm 1: Minibatch stochastic gradient descent training of GANs
Hyperparameter: = number of discriminator steps per generator step (often )
For each training iteration:
Phase 1: Train Discriminator (repeat times)
-
Sample minibatch of noise samples from noise prior
-
Sample minibatch of examples from data distribution
-
Update discriminator by ascending its stochastic gradient:
Phase 2: Train Generator
-
Sample minibatch of noise samples from noise prior
-
Update generator by descending its stochastic gradient:
GAN Training Steps - Detailed
Step 1: Train the Discriminator
Freeze generator parameters
- Create a minibatch of real data
- Create a minibatch of noises and use generator to synthesize fake data
- Use gradient ascent to update discriminator parameters to maximize :
Equivalently, use gradient descent to minimize negative log likelihood:
Comparison to Binary Cross Entropy:
Key Insight: Training discriminator to maximize is equivalent to training with binary cross entropy where:
- Target = 1 for real data
- Target = 0 for generated/fake data
Step 2: Train the Generator
Freeze discriminator parameters
- Create a minibatch of noises and use generator to synthesize fake data
- Use gradient descent to update generator parameters to minimize :
Problem: At early training stage:
- Generator synthesizes data that doesn't look like real data
- Discriminator easily predicts generated data as fake with high confidence
- Causes vanishing gradient problem (error must propagate backward from discriminator to generator)
Solution: To avoid saturation, adjust training to:
Equivalently:
This is equivalent to using binary cross entropy with target = 1
Effect: Generator is trained to maximize the probability of generated data being predicted as real
Instead of trying to make fake data "less fake," we train the generator to make fake data "more real" - this provides stronger gradients early in training.
GAN Training Visualization
[Image: Four stages (a-d) showing the evolution of GAN training]
Training Progress:
(a) Initial state:
- (generator distribution, green) differs from (real data, black dots)
- (discriminator, blue dashed) is partially accurate
(b) Discriminator training:
- In inner loop, is trained to discriminate samples from data
- Converges to
(c) Generator update:
- After generator update, gradient of guides to flow toward regions more likely to be classified as data
(d) Convergence:
- After several training steps, if and have enough capacity:
- Reach point where
- Discriminator cannot differentiate:
Generated Samples
[Image: Examples of GAN-generated images]
Examples shown:
- a) MNIST digits (handwritten numbers)
- b) Face images
- c) Natural scenes
- d) Animals and objects
Rightmost column shows nearest training example (to verify not memorizing)
Vector Arithmetic in Latent Space
Concept: GANs learn meaningful latent representations that support vector arithmetic
Example:
[Image: Visual demonstration of vector arithmetic producing women with glasses]
The latent space learns abstract concepts like "glasses-ness" that can be added or subtracted from generated images, similar to how word embeddings work (king - man + woman = queen).
References
- Goodfellow et al. "Generative Adversarial Nets." Advances in Neural Information Processing Systems. 2014.
- Radford et al. "Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks." ICLR 2016.
Summary
Key Concepts
Autoencoders:
- Learn compressed representations through reconstruction
- Encoder:
- Decoder:
- Loss: MSE or Binary Cross Entropy
- Variations: Fully-connected, Convolutional, Denoise
GANs:
- Two-player game: Generator vs Discriminator
- Generator creates fake data
- Discriminator distinguishes real from fake
- Minimax objective:
- Training alternates between improving D and G
- Converges when
