Chapter 11 - Deep Neural Networks

Updated 4 Oct 2026

11.1 Shallow versus Deep Neural Networks

Definitions

  • Shallow Neural Network: A multilayer perceptron with 1 or 2 hidden layers
  • Deep Neural Network: A multilayer perceptron with more than 2 hidden layers

Why Deep Networks?

Universal Approximation Theorem: A feedforward network with a single hidden layer is sufficient to represent any function

However:

  • The single layer may be infeasibly large (ใหญ่ใกล้ ∞\infty เลยล่ะ)
  • May fail to learn and generalize correctly
  • Using deeper models can:
    • Reduce the number of units required
    • Reduce the amount of generalization error

Analogy: Think of building with LEGO blocks. You could build a complex structure using one massive layer of blocks, but it's often more efficient and stable to build it in multiple layers, with each layer adding specific features.


11.2 Vanishing Gradient Problem

What is it?

When training a deep neural network using gradient descent, the gradient may become very small as it is propagated backward to earlier layers.

Example: 4-Layer Network Gradient Calculation


For a network with 4 layers, the gradient is calculated using the chain rule:

∂L∂w1,1[3]=(∂L∂y^)(∂y^∂z1[3])(∂z1[3]∂w1,1[3])\frac{\partial L}{\partial w^{[3]}_{1,1}} = \left(\frac{\partial L}{\partial \hat{y}}\right) \left(\frac{\partial \hat{y}}{\partial z^{[3]}_1}\right) \left(\frac{\partial z^{[3]}_1}{\partial w^{[3]}_{1,1}}\right)

∂L∂w1,1[2]=(∂L∂y^)(∂y^∂z1[3])(∂z1[3]∂a1[2])(∂a1[2]∂z1[2])(∂z1[2]∂w1,1[2])\frac{\partial L}{\partial w^{[2]}_{1,1}} = \left(\frac{\partial L}{\partial \hat{y}}\right) \left(\frac{\partial \hat{y}}{\partial z^{[3]}_1}\right) \left(\frac{\partial z^{[3]}_1}{\partial a^{[2]}_1}\right) \left(\frac{\partial a^{[2]}_1}{\partial z^{[2]}_1}\right) \left(\frac{\partial z^{[2]}_1}{\partial w^{[2]}_{1,1}}\right)

∂L∂w1,1[1]=(∂L∂y^)(∂y^∂z1[3])(∂z1[3]∂a1[2])(∂a1[2]∂z1[2])(∂z1[2]∂a1[1])(∂a1[1]∂z1[1])(∂z1[1]∂w1,1[1])\frac{\partial L}{\partial w^{[1]}_{1,1}} = \left(\frac{\partial L}{\partial \hat{y}}\right) \left(\frac{\partial \hat{y}}{\partial z^{[3]}_1}\right) \left(\frac{\partial z^{[3]}_1}{\partial a^{[2]}_1}\right) \left(\frac{\partial a^{[2]}_1}{\partial z^{[2]}_1}\right) \left(\frac{\partial z^{[2]}_1}{\partial a^{[1]}_1}\right) \left(\frac{\partial a^{[1]}_1}{\partial z^{[1]}_1}\right) \left(\frac{\partial z^{[1]}_1}{\partial w^{[1]}_{1,1}}\right)

Key Problems:

  • Weights farther from the output layer have more terms in the product
  • If some terms are less than 1, the overall gradient becomes very small
    • จะต้องมาอัปเดตเนี่ย ถ้า Small Term มาคูณกันก็ยิ่งเล็กเลยล่ะ จนกว่ามันจะ converge เข้า เลขที่ถูกต้องก็ต้องใช้เวลานานไง เพราะสูตรในการ Update จำได้ใช่มั้ย w1,1[1]←w1,1[1]−η∂L∂w1,1[1]w_{1,1}^{[1]}\leftarrow w_{1,1}^{[1]}-\eta\frac{\partial L}{\partial w_{1,1}^{[1]}}
  • Weights in earlier layers are not properly updated

Analogy: Like playing telephone with whispers - by the time the message reaches the first person, it's barely audible. Similarly, the gradient signal weakens as it travels back through layers.


11.3 Mitigating Vanishing Gradient

จะทำแก้ปัญหานี้ยังไงล่ะ

Method 1: ReLU Activation Function

ลองเปลี่ยนใช้ ReLU แทนที่จะใช้ Sigmoid บ้างซิ

ReLU (Rectified Linear Unit):
φReLU(z)=max⁡(0,z)\boxed{\varphi_{\text{ReLU}}(z) = \max(0, z)}
φReLU′(z)={1z>00z≤0\boxed{\varphi'_{\text{ReLU}}(z) = \begin{cases} 1 & z > 0 \\ 0 & z \leq 0 \end{cases}}
Benefits:

  • Derivative is either 0 or 1
  • Prevents gradient from becoming too small
  • Commonly used in deep neural networks for hidden layers

Problem: "Dying ReLU"

  • Neurons can become inactive and only output zero

Solutions: Leaky ReLU and ELU

Leaky ReLU


φleaky(z)=max⁡(0,z)+α⋅min⁡(0,z)\boxed{\varphi_{\text{leaky}}(z) = \max(0, z) + \alpha \cdot \min(0, z)}
φleaky′(z)={1z>0αz≤0\boxed{\varphi'_{\text{leaky}}(z) = \begin{cases} 1 & z > 0 \\ \alpha & z \leq 0 \end{cases}}

  • Typically α=0.01\alpha = 0.01
  • Allows small, non-zero gradient when unit is not active

ELU (Exponential Linear Unit)


φelu(z)={zz>0α(ez−1)z≤0\boxed{\varphi_{\text{elu}}(z) = \begin{cases} z & z > 0 \\ \alpha(e^z - 1) & z \leq 0 \end{cases}}
φelu′(z)={1z>0α(ez)z≤0\boxed{\varphi'_{\text{elu}}(z) = \begin{cases} 1 & z > 0 \\ \alpha(e^z) & z \leq 0 \end{cases}}

Analogy: ReLU is like a gate that's either fully open (1) or closed (0). Leaky ReLU leaves the gate slightly open even when "closed", allowing a trickle of gradient through.


Method 2: He Initialization

Properly initialize the weight

Purpose: Maintain variance of activations and gradients throughout the network

Formula:

W∼N(0,2nin)\boxed{W \sim \mathcal{N}\left(0, \frac{2}{n_{in}}\right)}

Where ninn_{in} is the number of input units to the layer

Details:

  • Designed for ReLU activation functions
  • Weights initialized from normal distribution with:
    • Mean = 0
    • Variance = 2nin\frac{2}{n_{in}}

Analogy: Like tuning the volume at each stage of an audio system so the signal doesn't get too loud (explode) or too quiet (vanish) as it passes through.

  • Makes the gradient become stable, ละช่วยในเรื่อง Vanishing Gradient Problem ได้ดีเลยล่ะ

Method 3: Batch Normalization

Adjust the value in each mini-batch, ปรับค่าใหม่พยายามให้ Mean เป็น 0, Variance เป็น 1

Purpose: Normalize inputs to each layer for each mini-batch, ensuring mean of 0 and variance of 1

Formula:
For each minibatch:
μB=1B∑i=1Bxi(Mean)\mu_B = \frac{1}{B}\sum_{i=1}^{B} x_i\quad \quad \text{(Mean)}
σB2=1B∑i=1B(xi−μB)2(Variance)\sigma^2_B = \frac{1}{B}\sum_{i=1}^{B} (x_i - \mu_B)^2 \quad \quad \text{(Variance)}
x^i=xi−μBσB2+ϵ(Standardized value)\boxed{\hat{x}_i = \frac{x_i - \mu_B}{\sqrt{\sigma^2_B + \epsilon}}}\quad \quad \text{(Standardized value)}
yi=γx^i+β\boxed{y_i = \gamma\hat{x}_i + \beta}

Where:

  • γ\gamma and β\beta are learnable parameters → มันจะถูก updated by gradient descent
    • Allow the network to scale (γ\gamma) and shift (β\beta) the normalized output
  • ϵ\epsilon is a small constant for numerical stability

Benefits:

  • Stabilizes the learning process
  • Maintains stable gradients during training

Analogy: Like standardizing test scores across different exams - converting all scores to a common scale (mean 0, variance 1) so they can be fairly compared and processed.


Method 4: Skip Connections (ResNet)

Concept: Allow gradients to flow more directly through the network

Architecture:

x → Layer 1 → Layer 2 → F(x) 
↓                         ↓
└────────────────────────→ + → x + F(x)

Benefits:

  • Reduces likelihood of vanishing gradients in very deep networks
  • Creates "shortcuts" for gradient flow

Analogy: Like having express lanes on a highway - traffic (gradients) can bypass congested areas and reach their destination faster.


11.3 Input and Output Encoding

Boolean Attributes

  • False = 0
  • True = 1

Real-Valued Attributes

Should be normalized or standardized:

Normalization:
z=x−xmin⁡xmax⁡−xmin⁡\boxed{z = \frac{x - x_{\min}}{x_{\max} - x_{\min}}}

  • พอเรา Normalize ออกมาจะได้ค่า ∈[0,1]\in[0,1]
    Standardization:
    z=x−μσ\boxed{z = \frac{x - \mu}{\sigma}}
  • หลัง Standardize แล้ว N(0,1)\mathcal{N}(0,1) คือ mean = 0, variance (σ2\sigma^2) = 1

When to use:

  • Standardization is preferred when values follow a normal distribution
  • Helps smooth the optimization landscape
  • Speeds up convergence
  • If values change enormously among examples, map onto log scale

Analogy: Normalization is like converting temperatures to a 0-1 scale between the coldest and hottest days. Standardization is like expressing temperatures in terms of how many degrees away from average they are.

Categorical Attributes (More than 2 categories)

One-Hot Encoding:
For an attribute with dd possible values, use dd bits (vector with dd elements) where only one bit is 1.

Example: Weather attribute with 4 values: {sun, cloud, rain, snow}

sun=[1000],cloud=[0100],rain=[0010],snow=[0001]\text{sun} = \begin{bmatrix} 1 \\ 0 \\ 0 \\ 0 \end{bmatrix}, \quad \text{cloud} = \begin{bmatrix} 0 \\ 1 \\ 0 \\ 0 \end{bmatrix}, \quad \text{rain} = \begin{bmatrix} 0 \\ 0 \\ 1 \\ 0 \end{bmatrix}, \quad \text{snow} = \begin{bmatrix} 0 \\ 0 \\ 0 \\ 1 \end{bmatrix}

Analogy: Like having separate light switches for each option - only one switch is "on" at a time.

  • Put one in different position, to represent different attribute.

Ordinal Attributes

อันนี้คือมันก็เป็น categorical ถูกป้ะ แต่ว่ามันมี Order ของมันด้วย เช่น temperature∈\texttt{temperature}\in {low,mid,high}\{\texttt{low,mid,high}\}

Ordered Encoding:
For ordered categories where grp0<grp1<grp2<grp3\text{grp0} < \text{grp1} < \text{grp2} < \text{grp3}:

grp0=[000],grp1=[100],grp2=[110],grp3=[111]\text{grp0} = \begin{bmatrix} 0 \\ 0 \\ 0 \end{bmatrix}, \quad \text{grp1} = \begin{bmatrix} 1 \\ 0 \\ 0 \end{bmatrix}, \quad \text{grp2} = \begin{bmatrix} 1 \\ 1 \\ 0 \end{bmatrix}, \quad \text{grp3} = \begin{bmatrix} 1 \\ 1 \\ 1 \end{bmatrix}

Pattern: Cumulative binary representation preserves order

  • ถ้าเอาอย่าง temperature\texttt{temperature} ข้างบนก็จะเป็น [00],[10],[11]\begin{bmatrix} 0 \\ 0\end{bmatrix},\begin{bmatrix} 1 \\ 0\end{bmatrix},\begin{bmatrix} 1 \\ 1\end{bmatrix}

Analogy: Like filling a measuring cup - each level includes all the previous levels, preserving the ordering.

Example 11.2: Auto MPG Dataset

Objective: Predict fuel efficiency (MPG) from car attributes

Attributes:

AttributeTypeDescription
MPGContinuous[9.0, 46.6]y
cylindersMulti-valued discrete{3, 4, 5, 6, 8}x
displacementsContinuous[68.0, 455.0]x
horsepowerContinuous[46.0, 230.0]x
weightContinuous[1613.0, 5104.0]x
accelerationContinuous[8.0, 24.8]x
model yearMulti-valued discrete70 to 82x
originMulti-valued discrete1=US, 2=Europe, 3=Japanx
Encoding Recommendations:
  1. cylinders:
    • จะมองเป็น Ordinal Value ก็ได้ → 3=[0000],4=[1000],5=[1100],…3=\begin{bmatrix} 0\\0\\0\\0\end{bmatrix},4=\begin{bmatrix} 1\\0\\0\\0\end{bmatrix},5=\begin{bmatrix} 1\\1\\0\\0\end{bmatrix},\dotso
    • หรือจะมองเป็น Continuous → ให้ Normalize z=x−38−3z=\frac{x-3}{8-3}
  2. model_year' (transformed ordinal):
    model_year′={0if model_year<731if 73≤model_year<762if 76≤model_year<793if model_year≥79\text{model\_year}' = \begin{cases} 0 & \text{if model\_year} < 73 \\ 1 & \text{if } 73 \leq \text{model\_year} < 76 \\ 2 & \text{if } 76 \leq \text{model\_year} < 79 \\ 3 & \text{if model\_year} \geq 79 \end{cases}
    • Use ordered encoding (since it's ordinal)

0=\begin{bmatrix} 0\0\0\end{bmatrix},1=\begin{bmatrix} 1\0\0\end{bmatrix},2=\begin{bmatrix} 1\1\0\end{bmatrix},3=\begin{bmatrix} 1\1\1\end{bmatrix}

3. **displacements:** Standardization (continuous, use z-score normalization) - $\min = 68, \max=455$ $$ z=\frac{x-68}{455-68} $$ 4. **origin:** One-hot encoding (3 categories: US, Europe, Japan) $$\text{U.S.}=\begin{bmatrix} 1\\0\\0\end{bmatrix},\text{Europe}=\begin{bmatrix} 0\\1\\0\end{bmatrix},\text{Japan}=\begin{bmatrix} 0\\0\\1\end{bmatrix} $$ **Output Layer:** >Output = MPG = Continuous Value ($\min=9,\max=46.6$) → called **Regression Task ($y_i\in R$)** - **Number of units:** 1 (single continuous value - MPG) - **Activation function:** - Linear (no activation) — $\varphi(z)=z$ > [!question] ทำไมต้องเป็น 1 นะ เพราะเราอยากได้คำตอบเป็นเลขค่า ๆ เดียวหรอ > Answer? real number linear 9regression 0 or 1 ให้มช้ sigmoid -1 or 1 tanh (negative อ่าจจะใช้บางที) --- ## 11.4 Multi-Class Classification >ย้อนกลับไป Binary Classification is when $y_i\in\{0,1\}$ - ดังนั้นใช้แค่ 1 neuron + sigmoid ### Overview For training sets with more than 2 classes: $y_i \in \{0, 1, 2, \ldots, C-1\}$ where $C > 2$ **Architecture:** - Use ==one-hot encoding== for output - Construct MLP with $C$ units in output layer - Each output unit predicts probability for one class ### One-Hot Encoding for Targets When $y_i = k$: $$\boxed{\mathbf{y}_i = [0, \ldots, 0, \underbrace{1}_{\text{k-th position}}, 0, \ldots, 0]^{\top}}$$ $$\boxed{\mathbf{y}_i = \underbrace{[0, \ldots, 0, 1, 0, \ldots, 0]^{\top}}_{\text{k-th position}}}$$ > *Analogy: Like voting - each example "votes" for exactly one class by putting a 1 in that class's position and 0s everywhere else.* ### Example 11.3 When we have a training set with 3 classes, i.e. 0, 1, and 2, we design an MLP to have 3 output units ![[Pasted image 20251029140522.png|center|500]] ### 11.4.1 Softmax Activation Function **Purpose:** Generalization of sigmoid for multiple classes **Formula:** $$\boxed{\text{softmax}(\mathbf{z}) = [\sigma_0, \sigma_1, \ldots, \sigma_{C-1}]^{\top}}$$ Where each element: $$\boxed{\sigma_i = \frac{e^{z_i}}{\sum_{j=0}^{C-1} e^{z_j}} \quad \text{for } i \in \{0, 1, \ldots, C-1\}}$$ **Properties:** - Converts raw scores (logits) to probabilities - All outputs sum to 1 - Each output is between 0 and 1 #### Example 11.4: Calculate `softmax` Given $\mathbf{z} = [2, 3, 1]^{\top}$ where $e^1 = 2.718$, $e^2 = 7.389$, $e^3 = 20.086$:

\texttt{softmax}(\begin{bmatrix}2\3\1\end{bmatrix})

$$\sigma_0 = \frac{e^2}{e^2 + e^3 + e^1} = \frac{7.389}{30.193} = 0.245$$ $$\sigma_1 = \frac{e^3}{e^2 + e^3 + e^1} = \frac{20.086}{30.193} = 0.665$$ $$\sigma_2 = \frac{e^1}{e^2 + e^3 + e^1} = \frac{2.718}{30.193} = 0.090$$ $$\text{softmax}([2, 3, 1]^{\top}) = [0.2447, 0.6653, 0.0900]^{\top}$$ - สุดท้ายแล้ว Sum มันก็กลายเป็น 1 เลยล่ะ อย่างเจ๋งเลยล่ะ > *Analogy: Like converting exam scores to percentages - softmax converts raw scores into probabilities that show how confident the model is about each class. ### 11.4.2 Categorical Cross-Entropy Loss **Extension of binary cross-entropy to multiple classes:** From binary: $$L_{ce}(\mathbf{x}, y; \theta) = -[y \log(\hat{y}) + (1-y) \log(1-\hat{y})]$$ $$L_{ce}(\mathbf{x}, y; \theta) = -[\underbrace{y\log(\hat{y}}_\text{class 1}) + \underbrace{(1-y) \log(1-\hat{y})}_\text{class 0}]$$ To multi-class: $$\boxed{L_{cat}(\mathbf{x}, \mathbf{y}; \theta) = -\sum_{c=0}^{C-1} y_c \log(\hat{y}_c)}$$ $$= -[y_0 \log(\hat{y}_0) + y_1 \log(\hat{y}_1) + \cdots + y_{C-1} \log(\hat{y}_{C-1})]$$ **Key insight:** Since $\mathbf{y}$ is one-hot encoded, only one term in the sum is non-zero (the true class) ### 11.4.3 Argmax Function **Purpose:** Convert predicted probabilities back to class labels **Definition:** $$\boxed{\text{argmax}(\mathbf{v}) = \text{index of maximum value in } \mathbf{v}}$$ **Example:** $$\arg\max \begin{bmatrix} 0.15 \\ 0.80 \\ 0.05 \end{bmatrix} = 1$$ (Index starts at 0) #### Example 11.5: Combining Softmax and Argmax Given $\mathbf{z} = [2, 3, 1]^{\top}$: From previous calculation: $\text{softmax}(\mathbf{z}) = [0.245, 0.665, 0.090]^{\top}$ $$\arg\max(\text{softmax}([2, 3, 1]^{\top})) = \arg\max([0.245, 0.665, 0.090]^{\top}) = 1$$ **Predicted class: 1** > *Analogy: Like choosing the winner in an election - argmax picks the candidate (class) with the most votes (highest probability).* ### Example 11.6: MNIST Handwritten Digit Recognition ![[Pasted image 20251029143633.png|center|200]] **Dataset:** 70,000 handwritten digit images **Classes:** $\mathcal{Y} = \{0, 1, 2, \ldots, 9\}$ (10 classes) **Input Processing:** - Each image: 28 × 28 grayscale - Flattened to vector: 784 features ($28 \times 28 = 784$) - Pixel values: `[0.0, 1.0]` (normalized) **Network Architecture:** - **Input Layer:** 784 neurons (one per pixel) - **Hidden Layer 1:** 128 neurons (ReLU activation) - **Hidden Layer 2:** 64 neurons (ReLU activation) - **Output Layer:** 10 neurons (Softmax activation) **Example: Digit "3"** ![[Pasted image 20251029143747.png|center|500]] Construct a neural network to cope with the MNIST dataset. ![[Pasted image 20251029144027.png|center|500]] --- ### Example 11.7: Choosing Output Configuration **Task 1: Binary Disease Classification** - **Output neurons:** 1 - **Activation:** Sigmoid - **Loss function:** Binary Cross-Entropy **Task 2: Iris Flower Classification (3 classes)** - **Output neurons:** 3 - **Activation:** Softmax - **Loss function:** Categorical Cross-Entropy **Task 3: House Price Prediction (Regression)** - **Output neurons:** 1 - **Activation:** Linear (or ReLU for positive prices) - **Loss function:** Mean Squared Error (MSE) > *Analogy: The output layer is like choosing the right tool - a yes/no question needs a switch (sigmoid), a multiple-choice question needs multiple switches (softmax), and measuring a continuous value needs a gauge (linear).* --- ## 11.5 Underfitting and Overfitting ### Definitions **Underfitting:** - ==Model is not trained enough== to capture dataset characteristics - Poor performance on both training and validation sets - Model is too simple **Overfitting:** - Model fits training set ALL too well - Poor performance on unknown examples (validation/test sets) - Model memorizes training data instead of learning patterns ### Causes of Overfitting - Model is too complex for the data - Too many features with too few examples - Training for too many epochs ### Techniques to Reduce Overfitting 1. Using three sets (training, validation, test) 2. Early stopping 3. Regularization 4. Dropout 5. Batch normalization 6. Data augmentation > *Analogy: Underfitting is like studying too little for an exam - you don't understand the concepts. Overfitting is like memorizing exact exam questions - you do great on practice tests but fail when questions are slightly different.* ### 11.5.1 Three Sets of Data **Purpose:** Properly evaluate model performance **Three Sets:** 1. **Training Set** - Purpose: Build the model - Typically largest set 2. **Validation Set** - Purpose: Find best hyperparameters and architecture - Monitor for overfitting during training - **Determine when to stop training** 3. **Test Set** - Purpose: Final evaluation of model performance - Used only once at the end - Provides unbiased performance estimate ### Splitting Guidelines **Small datasets:** - Training: 70% - Validation: 15% - Test: 15% **Large datasets (millions of examples):** - Training: 95% - Validation: 2.5% - Test: 2.5% ### Early Stopping **Method:** Stop training when validation loss starts increasing **Indicators of Overfitting:** - Training loss continues decreasing - Validation loss starts increasing or plateaus at higher value - Growing gap between training and validation performance ![[Pasted image 20251029144744.png|center|500]] - อาจจะมีการ Early Stopping เพื่อให้มันไม่ Overfitting (Epoch 75 ก็พอแล้วจร้า) > *Analogy: Like practicing for a sports game - at some point, extra practice on the same drills stops helping your real game performance and might even hurt it (overtraining).* --- ## 11.6 Regularization ### Concept **Purpose:** Modify objective function to penalize model complexity **Key Idea:** Force the algorithm to build a less complex model by making many weights small (close to zero) ### L2 Regularization (Weight Decay) **Modified Loss Function:** $$\boxed{J(\theta) = \left[\frac{1}{N}\sum_{i=1}^{N} L(\mathbf{x}_i, y_i; \theta)\right] + \frac{\lambda}{2}\|\theta\|^2}$$ Where: - $\lambda$ = hyperparameter controlling regularization strength - $\|\theta\|^2 = \sum_j \theta_j^2$ (sum of squared weights) **Effect:** - Minimizes both prediction error AND weight magnitudes - Forces many weights toward zero - Simpler, less complex model ### Weight Update Rule with L2 Regularization $$\boxed{\theta \leftarrow \theta - \eta(\nabla_\theta L(\mathbf{x}_i, y_i; \theta) + \lambda\theta)}$$ **Equivalent form:** $$\boxed{\theta \leftarrow (1 - \eta\lambda)\theta - \eta\nabla_\theta L(\mathbf{x}_i, y_i; \theta)}$$ **Interpretation:** - Each weight is decayed by factor $(1 - \eta\lambda)$ before gradient update - Called "weight decay" **Effect on Training:** ![[Pasted image 20251029145608.png|center|500]] - ตรงท้าย ๆ สังเกตว่ามันจะ Flat กว่าด้านบนนะ ถ้าเรา Apply สิ่งนี้ > *Analogy: Like Marie Kondo-ing your model - regularization asks each weight "do you spark joy (contribute meaningfully)?" and shrinks those that don't, keeping the model tidy and focused.* --- ## 11.7 Dropout ### Concept **Purpose:** Prevent units from co-adapting too much **Method:** Randomly exclude units (and their connections) during training ### How It Works **During Training:** - Each unit retained with probability $p$ (e.g., $p = 0.7$) - Each unit dropped with probability $(1-p)$ (e.g., 0.3) - Different random subset each iteration **During Testing/Inference:** - All units are active - Weights are scaled by $p$ to compensate ### Where to Apply - Typically applied to fully connected layers - Less common in convolutional layers **Visualization:** ![[Pasted image 20251029145752.png|center|500]] **Effect on Training:** ![[Pasted image 20251029145954.png|center|500]] > *Analogy: Like practicing a team sport where random players sit out each practice - this prevents players from relying too much on specific teammates and forces everyone to be versatile. During the actual game (testing), everyone plays.* --- ## 11.8 Choosing Number of Hidden Layers and Neurons ### Strategy: Trial and Error with Validation Set **Process:** 1. **Start Small** - Begin with 1-2 hidden layers - Small number of neurons (32 or 64) 2. **Train and Evaluate** - Train the model - Evaluate on validation set 3. **Adjust and Retrain** **If Underfitting** (poor performance on both training and validation): - Increase number of neurons - Add more hidden layers - Train longer **If Overfitting** (good training, poor validation): - Reduce number of neurons - Remove layers - Apply regularization (L2, dropout) - Use batch normalization - Get more training data 4. **Find Optimal Architecture** - Continue iterating - Balance performance and complexity - Monitor validation performance ### Guidelines **Start with common architectures:** - 1 hidden layer: Good for simple problems - 2 hidden layers: Good for most problems - 3+ hidden layers: For complex problems (images, sequences, etc.) **Neuron counts:** - Common sizes: 32, 64, 128, 256, 512 - Usually decrease as you go deeper: 512 → 256 → 128 > *Analogy: Like Goldilocks finding the right porridge - you need to try different "temperatures" (architectures) until you find one that's "just right" - not too simple (cold), not too complex (hot).* --- ## Summary: Key Techniques for Deep Learning ### Fighting Vanishing Gradients 1. ✅ ReLU/Leaky ReLU/ELU activation functions 2. ✅ He initialization 3. ✅ Batch normalization 4. ✅ Skip connections (ResNet) ### Preventing Overfitting 1. ✅ Split data into training/validation/test sets 2. ✅ Early stopping 3. ✅ L2 regularization (weight decay) 4. ✅ Dropout 5. ✅ Batch normalization 6. ✅ Data augmentation ### Input/Output Encoding - **Boolean:** 0/1 - **Continuous:** Standardization or normalization - **Categorical:** One-hot encoding - **Ordinal:** Ordered encoding ### Output Layer Configuration - **Binary classification:** 1 neuron + Sigmoid + Binary Cross-Entropy - **Multi-class classification:** C neurons + Softmax + Categorical Cross-Entropy - **Regression:** 1 neuron + Linear/ReLU + MSE/MAE --- ## Important Formulas Summary ### Activation Functions - ReLU: $\varphi(z) = \max(0, z)$ - Leaky ReLU: $\varphi(z) = \max(0, z) + \alpha \cdot \min(0, z)$ - Softmax: $\sigma_i = \frac{e^{z_i}}{\sum_{j=0}^{C-1} e^{z_j}}$ ### Loss Functions - Categorical Cross-Entropy: $L_{cat} = -\sum_{c=0}^{C-1} y_c \log(\hat{y}_c)$ - L2 Regularization: $J(\theta) = L(\theta) + \frac{\lambda}{2}\|\theta\|^2$ ### Normalization - Standardization: $z = \frac{x - \mu}{\sigma}$ - Batch Normalization: $\hat{x}_i = \frac{x_i - \mu_B}{\sqrt{\sigma^2_B + \epsilon}}$ ### Initialization - He Initialization: $W \sim \mathcal{N}(0, \frac{2}{n_{in}})$