Skip to main content
Deep Neural Networks

Optimization and Regularization

📅 Published: 2026-07-15
🎓 Level: postgraduate
👥 Audience: Postgraduate students in Deep Neural Networks

Prerequisite Knowledge

This lecture builds on the following concepts from earlier lectures. If any feel unfamiliar, review the linked notes before proceeding.

Previously Covered in This Subject

  • Data Augmentation — covered in Session 1: Introduction and Overview of Deep Neural Networks
  • Adaptive Learning Rate Methods — covered in Session 1: Introduction and Overview of Deep Neural Networks
  • Backpropagation — covered in Session 1: Introduction and Overview of Deep Neural Networks
  • Layer Normalization — covered in Session 1: Introduction and Overview of Deep Neural Networks
  • Learning Rate Schedules — covered in Session 1: Introduction and Overview of Deep Neural Networks
  • Activation Functions — covered in Deep Neural Network Components and Perceptron
  • Weight Initialization — covered in Deep Neural Network Components and Perceptron
  • Learning Rate Schedules — covered in Deep Neural Network Components and Perceptron
  • Adaptive Learning Rate Methods — covered in Deep Neural Network Components and Perceptron
  • Data Augmentation — covered in Deep Neural Network Components and Perceptron
  • Layer Normalization — covered in Deep Neural Network Components and Perceptron
  • Optimization vs Regularization — covered in Deep Neural Network Components and Perceptron
  • Layer Normalization — covered in Perceptron Learning and Introduction to Regression
  • Early Stopping — covered in Perceptron Learning and Introduction to Regression
  • Data Augmentation — covered in Perceptron Learning and Introduction to Regression
  • Weight Initialization — covered in Perceptron Learning and Introduction to Regression
  • Xavier Initialization — covered in Perceptron Learning and Introduction to Regression
  • He Initialization — covered in Perceptron Learning and Introduction to Regression
  • Adaptive Learning Rate Methods — covered in Perceptron Learning and Introduction to Regression
  • Learning Rate Schedules — covered in Perceptron Learning and Introduction to Regression
  • Activation Functions — covered in Linear Neural Networks for Regression
  • Layer Normalization — covered in Linear Neural Networks for Regression
  • Weight Initialization — covered in Linear Neural Networks for Regression
  • Learning Rate Schedules — covered in Linear Neural Networks for Regression
  • Data Augmentation — covered in Linear Neural Networks for Regression
  • Gradient Clipping — covered in Linear Neural Networks for Regression
  • Adaptive Learning Rate Methods — covered in Linear Neural Networks for Regression
  • Momentum-Based Gradient Updates — covered in Linear Neural Networks for Regression
  • Gradient Descent — covered in Linear Neural Networks for Regression
  • Batch Normalization — covered in Linear Neural Networks for Regression
  • Activation Functions — covered in Gradient Descent Variants, Classification, and Evaluation
  • Gradient Clipping — covered in Gradient Descent Variants, Classification, and Evaluation
  • Weight Initialization — covered in Gradient Descent Variants, Classification, and Evaluation
  • Underfitting — covered in Gradient Descent Variants, Classification, and Evaluation
  • He Initialization — covered in Gradient Descent Variants, Classification, and Evaluation
  • Gradient Descent — covered in Gradient Descent Variants, Classification, and Evaluation
  • Xavier Initialization — covered in Gradient Descent Variants, Classification, and Evaluation
  • Batch Normalization — covered in Gradient Descent Variants, Classification, and Evaluation
  • Momentum-Based Gradient Updates — covered in Gradient Descent Variants, Classification, and Evaluation
  • Overfitting — covered in Gradient Descent Variants, Classification, and Evaluation
  • Batch Normalization — covered in Deep Feedforward Neural Networks
  • Underfitting — covered in Deep Feedforward Neural Networks
  • Data Augmentation — covered in Deep Feedforward Neural Networks
  • Backpropagation — covered in Deep Feedforward Neural Networks
  • Gradient Descent — covered in Deep Feedforward Neural Networks
  • Overfitting — covered in Deep Feedforward Neural Networks
  • Activation Functions — covered in Deep Feedforward Neural Networks
  • Momentum-Based Gradient Updates — covered in Deep Feedforward Neural Networks
  • Weight Initialization — covered in Deep Feedforward Neural Networks
  • Early Stopping — covered in Deep Feedforward Neural Networks
  • Gradient Clipping — covered in Deep Feedforward Neural Networks
  • Gradient Descent — covered in Deep Feed-Forward Neural Network Architecture Design and Training
  • Layer Normalization — covered in Deep Feed-Forward Neural Network Architecture Design and Training
  • Early Stopping — covered in Deep Feed-Forward Neural Network Architecture Design and Training
  • Activation Functions — covered in Deep Feed-Forward Neural Network Architecture Design and Training
  • Weight Initialization — covered in Deep Feed-Forward Neural Network Architecture Design and Training
  • Momentum-Based Gradient Updates — covered in Deep Feed-Forward Neural Network Architecture Design and Training
  • Gradient Clipping — covered in Deep Feed-Forward Neural Network Architecture Design and Training
  • Data Augmentation — covered in Deep Feed-Forward Neural Network Architecture Design and Training
  • Overfitting — covered in Deep Feed-Forward Neural Network Architecture Design and Training
  • Data Augmentation — covered in Exam Revision and Introduction to Convolutional Neural Networks
  • Xavier Initialization — covered in Exam Revision and Introduction to Convolutional Neural Networks
  • Overfitting — covered in Exam Revision and Introduction to Convolutional Neural Networks
  • Gradient Clipping — covered in Exam Revision and Introduction to Convolutional Neural Networks
  • Batch Normalization — covered in Exam Revision and Introduction to Convolutional Neural Networks
  • He Initialization — covered in Exam Revision and Introduction to Convolutional Neural Networks
  • Adaptive Learning Rate Methods — covered in Exam Revision and Introduction to Convolutional Neural Networks
  • Momentum-Based Gradient Updates — covered in Exam Revision and Introduction to Convolutional Neural Networks
  • Weight Initialization — covered in Exam Revision and Introduction to Convolutional Neural Networks
  • Gradient Descent — covered in Exam Revision and Introduction to Convolutional Neural Networks
  • Learning Rate Schedules — covered in Exam Revision and Introduction to Convolutional Neural Networks
  • Layer Normalization — covered in Exam Revision and Introduction to Convolutional Neural Networks
  • Adaptive Learning Rate Methods — covered in Convolutional Neural Networks
  • Activation Functions — covered in Convolutional Neural Networks
  • Normalization Techniques — covered in Convolutional Neural Networks
  • Learning Rate Schedules — covered in Convolutional Neural Networks
  • Layer Normalization — covered in Convolutional Neural Networks
  • Batch Normalization — covered in Convolutional Neural Networks
  • Learning Rate Schedules — covered in Recurrent Neural Networks — Foundations and Architecture
  • Data Augmentation — covered in Recurrent Neural Networks — Foundations and Architecture
  • Adaptive Learning Rate Methods — covered in Recurrent Neural Networks — Foundations and Architecture
  • Backpropagation — covered in Recurrent Neural Networks — Foundations and Architecture
  • Activation Functions — covered in Recurrent Neural Networks — Foundations and Architecture
  • Optimization vs Regularization — covered in Recurrent Neural Networks — Foundations and Architecture
  • Dropout — covered in Recurrent Neural Networks — Foundations and Architecture
  • Activation Functions — covered in Recurrent Neural Networks — Advanced Architectures
  • Gradient Clipping — covered in Recurrent Neural Networks — Advanced Architectures
  • Gradient Descent — covered in Recurrent Neural Networks — Advanced Architectures
  • Data Augmentation — covered in Recurrent Neural Networks — Advanced Architectures
  • Momentum-Based Gradient Updates — covered in Recurrent Neural Networks — Advanced Architectures
  • Layer Normalization — covered in Attention Mechanisms
  • Early Stopping — covered in Attention Mechanisms
  • Weight Initialization — covered in Attention Mechanisms
  • Batch Normalization — covered in Attention Mechanisms
  • Data Augmentation — covered in Attention Mechanisms
  • Adaptive Learning Rate Methods — covered in Attention Mechanisms
  • Optimization vs Regularization — covered in Attention Mechanisms
  • Normalization Techniques — covered in Attention Mechanisms
  • Learning Rate Schedules — covered in Attention Mechanisms
  • Layer Normalization — covered in Attention Mechanisms and Transformer Architecture
  • Normalization Techniques — covered in Attention Mechanisms and Transformer Architecture
  • Weight Initialization — covered in Attention Mechanisms and Transformer Architecture
  • Batch Normalization — covered in Attention Mechanisms and Transformer Architecture
  • Learning Rate Schedules — covered in Transformer Architectures and Optimization
  • Layer Normalization — covered in Transformer Architectures and Optimization
  • Normalization Techniques — covered in Transformer Architectures and Optimization
  • Momentum-Based Gradient Updates — covered in Transformer Architectures and Optimization
  • Gradient Descent — covered in Transformer Architectures and Optimization
  • Gradient Clipping — covered in Transformer Architectures and Optimization
  • Adaptive Learning Rate Methods — covered in Transformer Architectures and Optimization
  • Batch Normalization — covered in Transformer Architectures and Optimization
  • Optimization vs Regularization — covered in Transformer Architectures and Optimization

Optimization and Regularization

Training a deep network is a tug-of-war between two forces. One pushes the model to learn faster and more stably from the data. The other pulls it back so it does not memorize every wrinkle in the training set. This lecture shows you how both forces work, why they matter, and when to reach for each.

17.1 Optimization vs Regularization

Hook: You have spent hours tuning a model. The training accuracy climbs — 92%, 95%, 98%. You feel great. Then you run it on new data and the accuracy tanks to 64%. The model learned the training set perfectly but forgot how to generalize. Why did this happen, and how do you stop it?

17.1.1 Two Sides of the Training Coin

Think of training a deep network like coaching a student for an exam. Optimization is the drill sergeant who makes the student practice every problem in the textbook until they can solve it blindfolded. Regularization is the wise tutor who shuffles the problem wording and throws in curveballs. The real exam will not look exactly like the textbook. One builds raw skill. The other builds adaptability. You need both.

Optimization and regularization solve different problems during training. Optimization is about the learning process itself — you push the model to learn better, faster, and more stably. Regularization is about preventing the model from overfitting the training data, so it generalizes to unseen examples.

Optimization refines how the model learns. It answers: can we make gradient updates smarter? Can we give each parameter its own step size? Can we use history to smooth out noisy updates? Regularization constrains what the model learns. It answers: is the model memorizing noise instead of patterns? Can we force it to stay simple enough to generalize?

From a deep learning perspective: if your model is overfitting, you reach for regularization. If the model's capacity is not up to the mark and needs fine-tuning for the task, you reach for optimization. One refines the learning process; the other constrains it.

Dimension Optimization Regularization
Goal Improve learning speed and stability Prevent overfitting; improve generalization
Acts on Gradient updates, learning rates Model complexity, weight magnitudes
Signal Training loss not decreasing fast enough Validation loss diverging from training loss
Examples Momentum, Adam, learning rate schedules L1/L2 penalties, dropout, data augmentation
When to use Model underfits or converges too slowly Model overfits (low training error, high test error)
When to pick which: If both training and validation loss are high, optimize. If training loss is low but validation loss is climbing, regularize.
Scope: Optimization and regularization are not independent. A well-chosen optimizer (like AdamW) includes a regularization mechanism (weight decay). A regularization technique (like batch normalization) can speed up optimization. The boundary between the two can blur in practice. But the conceptual distinction still guides your debugging. Ask yourself: "is the model learning too poorly?" (optimization) or "is the model learning the training set too well?" (regularization).
Visual Intuition: Picture a plot with training epochs on the x-axis and loss on the y-axis. Two curves start together and descend. The training loss (blue) keeps dropping steadily. The validation loss (orange) drops at first, then flattens, then starts climbing back up — a U-shape. The gap between the two curves is the overfitting zone. Optimization techniques shift both curves down together. Regularization techniques close the gap between them.
Pitfalls: - Applying regularization when the model is actually underfitting makes both training and test performance worse. Always check if your model can even fit the training data first. - Using momentum or adaptive learning rates will not fix overfitting. They only change how the model descends the loss surface. They do not change what surface it is descending. - Skipping regularization because "I have a lot of data." Even large datasets can be overfit if the model is complex enough. Data quantity helps, but it is not a guarantee.

17.1.2 The Two Questions This Module Answers

The optimization module answers two core questions. First, how do you update the gradient component of the weight update equation? Not just using the current gradient. But also using historical momentum to smooth the trajectory. Second, how do you make the learning rate adaptive instead of fixed? So every parameter gets its own effective step size.

The regularization module answers: what do you do when the model overlearns the training data? How do you prevent memorization and ensure the model captures the underlying pattern rather than the noise?

Optimization makes the model learn faster and smarter. Regularization keeps it honest. The rest of this lecture walks through every major tool in both toolboxes — starting with momentum-based gradient updates.

Real-World Connection: Every production ML system balances these two forces. Google trains a translation model. The optimization team worries about convergence time on thousands of TPUs. The regularization team worries about slang, typos, and rare languages. Both concerns are equally critical to shipping a working product.

17.2 Momentum-Based Gradient Updates

Hook: Imagine you are hiking down a foggy mountain in the dark. At each step you feel the slope under your boots and step downhill. But the ground is rocky and uneven — one step says go left, the next says go right. You end up zigzagging, wasting energy. What if you could remember your last few steps and use that memory to smooth out the path?

17.2.1 The Standard Gradient Descent Update

Think of gradient descent like a blindfolded person trying to reach the bottom of a valley. At every step they feel which way the ground slopes and take a step in that direction. The size of the step is the learning rate. A small step means slow but safe progress. A large step means faster progress but risk of overshooting. This is the simplest possible learning rule — and its simplicity is also its weakness.

The basic weight update is:

Where:

- — the current parameter value (a weight or bias in your network) - — the learning rate, a scalar that controls step size. Typical values: 0.1, 0.01, 0.001. - — the gradient of the loss with respect to the parameter, evaluated at . Tells you the direction and magnitude of steepest ascent. You move opposite to it because you want to go downhill.

The update rule says: take the current weight, subtract a fraction of the gradient. That fraction is the learning rate. If the gradient is steep, the step is large. If the gradient is flat, the step is small.

17.2.2 Why Momentum Helps

When you use plain stochastic gradient descent, each parameter's gradient pulls the update in its own direction. One feature might pull gently; another might pull aggressively. The trajectory across the loss surface becomes noisy with oscillations. The problem is that every sampled instance or minibatch might have a different pattern. Because of this, successive weight adjustments see large swings while trying to converge.

The solution is to take help from the history of gradients. Instead of letting the current gradient alone decide the step, you accumulate a velocity term. This velocity carries forward information from previous gradients. It smooths out the oscillations and produces a straighter, faster trajectory toward the minimum.

17.2.3 SGD with Momentum

Momentum works like a heavy ball rolling down a hill. The ball has inertia — it keeps moving in roughly the same direction even when the ground bumps it sideways. A small bump does not reverse the ball's direction. Only a sustained slope change can redirect it. This is exactly what we want from our optimizer: resist noisy gradient fluctuations, follow the consistent downward trend.

The update becomes:

- — the momentum velocity at step . It is a running accumulation of past gradients. Same shape as . - — the momentum coefficient, a scalar in . Controls how much past velocity you carry forward. Common value: . How behaves: - : velocity is just the current gradient. You recover plain SGD. - : each step blends 90% of past velocity with 10% of the current gradient's influence (indirectly via the formula). Most of the direction comes from history. - : you add the entire accumulated gradient unchanged to the current gradient. Velocity never decays. Not used in practice — it would cause unbounded growth.

The physical intuition: imagine the gradient as an arrow pointing in some direction at each step. Without momentum, the update follows that arrow exactly, even if it oscillates wildly. With momentum, you take a scaled-down version of the previous arrow and add it to the current gradient arrow via vector addition. The resulting arrow is the actual direction the update moves. This compromise smooths out the zigzag.

Visual Intuition: On a contour plot of the loss surface (like a topographic map), plain SGD traces a jagged zigzag path. Each step corrects the overcorrection of the previous step. SGD with momentum traces a smoother, more direct curve. The velocity term acts like a low-pass filter on the gradient signal. It dampens high-frequency oscillations while preserving the low-frequency descent direction. The x-axis and y-axis represent two parameter dimensions. The contour lines show equal-loss levels, with the minimum at the center.
Worked Example — Momentum velocity calculation:

Start with , current gradient , and .

This new velocity replaces the plain gradient in the weight update. Without momentum, the update would use only the raw gradient 7.2. With momentum, the effective update signal is twice as strong — but smoothed by the contribution from past steps.

Sense-check: from the current gradient plus carried over from history. The past still dominates slightly since equals the fresh gradient — this shows that with , the velocity has a long memory.

17.2.4 Nesterov Accelerated Gradient (NAG)

Standard momentum is like driving a car while only looking in the rearview mirror. You base your steering on where you have been. Nesterov momentum is like glancing ahead through the windshield too. You first coast forward using your current velocity. Then you check the gradient at that future position. Only then do you decide your actual step. This look-ahead correction makes NAG more responsive to changes in the slope.

Standard momentum looks backward — it uses the previous velocity plus the current gradient. Nesterov's method also looks ahead. Instead of computing the gradient at the current position , you first take a look-ahead step using the previous velocity. Then you compute the gradient at that future position:

The critical difference: the gradient is evaluated at — the position you would land at if you just coasted on momentum. This is the "look-ahead" position. Then, as before, you combine this look-ahead gradient with the previous velocity.

Why it helps: Suppose momentum carries you toward a steep uphill. The look-ahead gradient at that future position will be strongly opposed to the current direction. NAG detects this early and slows down before overshooting. Standard momentum would charge ahead and only correct after the fact.
Pitfalls: - Setting too high (e.g., 0.999): the velocity barely decays. Past gradients from hundreds of steps ago still influence the current update. If the loss surface has changed, you are fighting stale information. - Setting too low (e.g., 0.5): momentum provides little smoothing. You are close to plain SGD and lose most of the benefit. - Using momentum on a problem that is already well-conditioned: if the loss surface is a perfect bowl, momentum can cause overshooting. It can also cause oscillation around the minimum. Plain SGD or adaptive methods may converge faster. - Forgetting that momentum adds a hyperparameter (): every hyperparameter you add increases the tuning burden. Default to unless you have a reason to change it.

17.2.5 Numerical Example: Momentum's Effect

Consider the same loss function, same initialization , and same learning rate used earlier for plain gradient descent. Without momentum, kept decreasing from the positive region toward zero. But it never crossed into negatives. With momentum, accelerates convergence and moves into the negative region. It overshoots the zero line, finds a better valley, and speeds up learning. The accumulated velocity carries it past local flat regions.

Trace — Momentum carrying a weight across zero:

At each step:

1. Take previous velocity 2. Multiply by (say 0.9) 3. Add the current gradient 4. Use the resulting to update the weight

The key insight: once velocity builds up in a consistent direction, it acts like a flywheel. Even when the current gradient at a particular point is small (a flat region), the accumulated velocity keeps the parameter moving. This is why can cross zero. The velocity carries it through, even though the local gradient at zero would not push it far enough on its own.

Scope — When momentum helps and when it does not: - Helps: Noisy gradients (small minibatches), ravines in the loss surface (steep in one direction, flat in another). Also problems where consistent descent direction exists. - Does not help (or hurts): Perfectly spherical loss surfaces (momentum causes overshoot). Problems where the gradient direction genuinely changes every step (momentum adds lag). Learning rate already well-tuned via an adaptive method like Adam (Adam already includes momentum internally).
Comparison — Standard Momentum vs Nesterov:
Property Standard Momentum Nesterov (NAG)
Gradient evaluated at Current position Look-ahead position
Responsiveness Reacts to past Anticipates future
Overshooting More likely near minima Less likely — corrects early
Convergence speed Faster than plain SGD Slightly faster than standard momentum
Extra computation None vs plain SGD One extra gradient evaluation (or clever reparameterization)
When to pick which: Use standard momentum as the default — it is simpler and implemented everywhere. Use NAG when you notice the optimizer oscillating near convergence or when training very deep networks where overshooting is costly.

Momentum turns the optimizer from a myopic step-taker into a boulder rolling downhill. It builds velocity in consistent directions and ignores noisy gradient fluctuations. The next section asks: what if we could also give each parameter its own learning rate?

Real-World Connection: Momentum is a standard ingredient in almost every production optimizer. The Adam optimizer — used to train GPT, BERT, and most modern transformers — includes momentum internally as its "first moment" estimate. When you use Adam, you are using momentum without having to tune from scratch. The default works well in most cases. Understanding momentum helps you debug when Adam misbehaves. If your model's loss oscillates, one culprit could be the momentum term reacting too slowly to gradient changes.

17.3 Adaptive Learning Rate Methods

Hook: You set the learning rate to 0.01. One weight updates beautifully — smooth, steady progress. Another weight barely budges. A third oscillates wildly, never settling. All three use the same learning rate. The problem is not the optimizer — it is that a single global step size cannot serve every parameter equally. What if every weight got its own learning rate, tuned automatically to its own gradient history?

17.3.1 Why a Fixed Learning Rate Fails

Imagine you are adjusting two dials on a radio. One dial (volume) needs tiny, precise turns — a big twist blows out the speakers. The other dial (tuning) needs broad sweeps to find the station — tiny nudges get you nowhere. A single "turn size" fails for both. Parameters in a network are the same: some have steep, volatile gradients and need small steps. Others have flat, consistent gradients and need larger steps to make progress.

In the same problem, one parameter's gradient flow changes gracefully, while the other parameter's gradient drops aggressively fast. A fixed learning rate is too fast for one parameter and too slow for the other. You need each parameter to have its own effective learning rate. Parameters with small, stable gradients can move faster. Parameters with large, volatile gradients need smaller, careful steps.

17.3.2 AdaGrad

AdaGrad adapts the learning rate per parameter using the accumulated squared gradient. The idea: parameters that have received large gradients get a smaller effective learning rate. Parameters that have received small gradients get a larger one.

The update is:

Where:

- — accumulated sum of squared gradients up to time : . Same shape as . - — tiny constant (e.g., ) to avoid division by zero. - — element-wise (Hadamard) multiplication. Each parameter gets its own scaling factor. - — global learning rate, divided per-parameter by . How it works: A parameter with consistently large gradients builds up a large . The denominator grows, shrinking the effective step. A parameter with small gradients keeps a small , so its effective step stays larger. Over time, every parameter's learning rate decays — but parameters that have already moved a lot decay faster.
Visual Intuition: Plot the effective learning rate on the y-axis against training iterations on the x-axis. For a frequently-updated parameter (steep gradient), the curve drops quickly and levels off at a small value. For a rarely-updated parameter (sparse feature), the curve stays high. The gap between the two curves is AdaGrad's adaptive advantage.
Worked Example — AdaGrad per-parameter scaling:

Two parameters, and . Global , .

After 3 steps, has seen gradients . Squared sum: .

After 3 steps, has seen gradients . Squared sum: .

Effective learning rate for : .

Effective learning rate for : .

Sense-check: learns about 26 times faster per step than because its gradients have been consistently small. This is exactly the desired behavior — needs bigger steps to catch up.
Scope — Where AdaGrad shines and where it fails: - Best for: Sparse features. In sentiment analysis with a bag-of-words model, most words never appear in a given review. Their gradient is zero most of the time. AdaGrad preserves a large learning rate for these rare features when they do appear. - Fails for: Dense data where all features are always active. The accumulated sum grows monotonically for every parameter. The learning rate decays to near-zero too fast, stopping learning prematurely. - Not for: Deep networks trained for many epochs. The aggressive decay kills training before convergence.
Pitfalls: - Using AdaGrad for dense problems: the monotonically growing denominator means the learning rate can only decrease. It never recovers. Training can stall entirely. - Forgetting that matters for very small gradients. If is too large compared to , the adaptation is weak. The denominator is dominated by . - Confusing with a moving average: is a cumulative sum, not an average. It has infinite memory — every gradient since step 1 contributes equally.

17.3.3 RMSProp

AdaGrad's problem is that it never forgets. Every gradient since the first step pulls down the learning rate. RMSProp fixes this by using a leaky memory — an exponentially weighted moving average of squared gradients. Old gradients fade away, so the denominator does not grow without bound. This makes RMSProp work for non-stationary problems where the gradient distribution shifts over time.

RMSProp fixes AdaGrad's aggressive decay by using an exponentially weighted moving average of squared gradients instead of a cumulative sum:

- — exponentially weighted moving average of squared gradients. Same shape as . - — decay rate, typically . Controls the memory length. How controls memory: - : past and present squared gradients get equal weight. Short memory. - : 90% from past, 10% from current. Long memory — about 10 steps of significant influence. - : Very long memory — about 100 steps. Close to AdaGrad-like behavior.

The key difference from AdaGrad: can go down as well as up. If recent gradients are small, shrinks and the learning rate recovers. AdaGrad's only grows.

RMSProp is best suited for non-stationary problems. These are situations where the underlying data pattern shifts over time. It also works well for sequential decision-making tasks — think RNNs and LSTMs.

Pitfall: RMSProp only adapts the learning rate. It does not include momentum. The gradient component is still the plain current gradient. So you lose the push that momentum provides. For problems where momentum matters (deep networks with noisy gradients), RMSProp alone may converge too slowly.

17.3.4 Adam

RMSProp gives you per-parameter learning rates. Momentum gives you smoothed gradient directions. Adam puts them together — like RMSProp and momentum had a child that inherited the best of both. Add a clever bias-correction trick for the first few steps. Now you have the most widely used optimizer in deep learning.

Adam combines RMSProp's adaptive learning rate with SGD's momentum. It maintains two moving averages:

- first moment estimate (mean of gradients). This is the momentum term. controls its decay (default: 0.9). - second moment estimate (uncentered variance of gradients). This is the adaptive scaling term. controls its decay (default: 0.999). - and are separate knobs — you can tune momentum memory and scaling memory independently.

The raw (uncorrected) update would be:

The cold-start problem: Both and start at zero. In the early iterations, this biases the estimates strongly toward zero. Imagine you ran a business and on day one made a profit of 5000 rupees. A naive moving average averages 0 and 5000 to get 2500. That is misleading — your actual performance is 5000, not 2500.

Adam corrects this with bias correction:

How bias correction works:

When is small (say ):

- , so . Dividing by 0.1 scales it up by 10x. This compensates for the fact that only captured 10% of the first gradient. - Same logic for .

When is large (say ):

- . . The correction factor is essentially 1. The raw moving averages are trustworthy.

The bias correction is a warm-up mechanism — it boosts early estimates when they are unreliable, then gracefully fades out.

With bias corrections, the final Adam update is:

Trace — Adam on the first two steps with : Step 1: gradient - - - ✓ (recovers full gradient) - - Update: Step 2: gradient - - - - - Update: Sense-check: The first step uses the full gradient (bias correction worked). The second step is smaller because the gradient reversed direction (the momentum term partially cancels). And the adaptive scaling grew (more variance observed). Both behaviors are correct.

For the numerical example with the same loss function and initialization , Adam produces smooth, steady updates. Both parameters progress gradually — neither too slow nor too fast. The bias corrections ensure unbiased estimates at the initial iterations.

Adam is the preferred choice in most deep learning architectures, including transformer models. A common strategy is to start training with Adam and, when fine-tuning is needed, switch to custom choices.

Student Q&A — Deduplicated:

> Q: So Adam combines momentum and adaptive learning — but does it also include the bias correction? > A: Yes. Adam includes bias correction for both the first moment (momentum) and the second moment (adaptive scaling). The first moment controls the gradient update; the second moment adapts the learning rate. Both get corrected using in the denominator during early iterations, which fades as training proceeds.

17.3.5 AdamW: Decoupled Weight Decay

Standard Adam mixes regularization into the gradient signal. The weight decay penalty gets absorbed into and , where the adaptive scaling distorts it. AdamW says: keep regularization separate. Let Adam handle the gradient. Apply weight decay as an independent shrink step afterward. This tiny architectural change turns out to matter a lot for generalization.

Standard Adam includes L2 regularization (weight decay), which adds to the loss. The gradient of this regularized loss is used in both the momentum and adaptive scaling calculations. But there is a conflict. The regularization term constrains the parameter values. Meanwhile, the adaptive mechanism uses those regularized gradients as if they were the true gradient. The original effect of the gradient is distorted.

AdamW fixes this by decoupling weight decay from the Adam update. After the normal Adam step computes the new parameters, you subtract the regularization component separately:

The key insight: The L2 regularization term is applied as an independent step, not baked into the gradient calculations. The adaptive scaling only affects the gradient-driven part of the update. The weight decay part uses the raw learning rate. Why this matters: In standard Adam, a parameter with a large adaptive scaling factor (small ) would have its regularization weakened. A parameter with small adaptive scaling (large ) would have its regularization amplified. The regularization strength becomes entangled with the gradient history — you lose control. AdamW restores independent control over both.

This decoupling leads to better generalization. AdamW is frequently used in GPT and other transformer architectures internally.

Comparison — Adam vs AdamW:
Property Adam AdamW
Weight decay Added to loss; gradient feeds into Subtracted after Adam update as separate step
Reg. strength Distorted by adaptive scaling Independent of gradient history
Generalization Good Better — decoupling prevents over-regularization of well-behaved parameters
When to use General default When weight decay is important (transformers, large models)

17.3.6 Symbol Registry — Adam and Variants

- — parameter value at iteration — scalar or vector - — learning rate (step size) — scalar in - — gradient of the loss with respect to — same shape as - — momentum velocity; first-moment estimate — same shape as - (RMS) — exponentially weighted squared gradient; second-moment estimate — same shape as - — momentum decay coefficient — scalar in - — RMS decay coefficient — scalar in - — small constant for numerical stability — scalar, e.g. - — accumulated sum of squared gradients (AdaGrad) — same shape as - — weight decay coefficient (L2 regularization strength) — scalar - — bias-corrected moment estimates (Adam)
Pitfalls across adaptive methods: - Using Adam with the default on problems with rapidly changing gradient statistics (like GAN training): the second moment adapts too slowly. Try or even 0.8. - Assuming Adam always beats SGD with momentum is a mistake. For some vision tasks (like ImageNet training), well-tuned SGD with momentum can generalize better than Adam. Adam converges faster but may not always find the best final solution. - Forgetting that Adam has three hyperparameters () plus : more knobs means more tuning. Start with defaults: . - Using AdamW without understanding that and interact is risky. A large with a small means the weight decay dominates the gradient update. The model shrinks weights faster than it learns.
Exam note: Know which optimizer for which scenario. Sparse data → AdaGrad. Non-stationary/RNN → RMSProp. General/transformers → Adam. Transformer training with weight decay → AdamW. Understand why Adam's bias correction exists (cold start from zero initialization) and how it fades with . The next section moves from adaptive step sizes to scheduled learning rates that change over time.
Real-World Connection: Adam (Kingma & Ba, 2014) is the default optimizer in PyTorch and TensorFlow. Nearly every transformer model — GPT, BERT, T5, LLaMA — uses either Adam or AdamW. The original Adam paper has over 150,000 citations, making it one of the most impactful papers in machine learning history. AdamW (Loshchilov & Hutter, 2017) is the standard for large language model pretraining. Its decoupled weight decay prevents the optimizer from silently under-regularizing parameters with large gradient histories.

17.4 Learning Rate Schedules

Hook: You pick a learning rate of 0.001. Training goes well for a while — the loss drops steadily. Then progress stalls. The optimizer is taking tiny steps around a minimum but cannot settle into it. If only the learning rate could start large (to explore) and shrink over time (to refine)... it can. That is what a learning rate schedule does.

17.4.1 Why Schedules Matter

Even with adaptive methods like Adam, the base learning rate is still a fixed hyperparameter chosen by the designer. There is selection bias in picking it. Learning rate schedules replace a fixed with one that changes over time according to a plan, reducing the hyperparameter sensitivity.

Think of a learning rate schedule like a photographer adjusting focus. At first, you twist the lens quickly to get near the right focal length (large learning rate — exploration). As the image sharpens, you slow down and make micro-adjustments (small learning rate — exploitation). If you twist too fast near the end, you overshoot and lose the shot. The schedule is your hand's memory of when to slow down.

A learning rate schedule is a function that maps the training step to a learning rate. The schedule replaces the fixed in whichever optimizer you are using — SGD, Adam, RMSProp, etc. The optimizer still does its internal adaptation. The schedule controls the baseline rate from which that adaptation starts.

17.4.2 Step Decay

A step decay schedule drops the learning rate by a fixed factor at predetermined epochs. For example, train with for 10,000 iterations, then drop to 0.01 for the next 40,000, then to 0.001 after a few million. The large initial rate encourages exploration — the optimizer takes big jumps across the loss surface. The smaller later rates encourage exploitation — fine-grained, careful steps near the minimum.

This is common in reinforcement learning, especially for AI agents learning to play complex games like chess or Go. Go is extremely complex to formulate programmatically; researchers often use step decay as a simple, predictable schedule.

Pitfall: The schedule is set by domain experts who know roughly when to shift from exploration to exploitation. However, choosing the exact epoch boundaries introduces its own bias. If you drop the rate too early, the model never escapes a poor region. If you drop it too late, you waste computation.

17.4.3 Exponential Decay

Exponential decay gives a smoother reduction:

- — the initial learning rate (e.g., 0.1) - — the decay rate, typically between 0.01 and 1 - — the epoch or iteration number

The learning rate decays gracefully as training progresses, rather than dropping abruptly. At each step, the rate is multiplied by a constant factor , which is less than 1.

Exponential decay is suited for simple feedforward neural networks or when the loss surface is noisy. The smooth decay avoids the sharp transitions of step decay.

Worked Example — Exponential decay experiment:

Synthetic linear regression: . True weights: (hidden from model).

Case 1 — Aggressive decay: - At : - At : - At : — near zero. Weights stop updating far from true values. Case 2 — Light decay: - At : - At : - At : — still learning. Weights approach true values. Sense-check: The interaction between learning rate and decay needs experimental tuning — there is no fixed formula. The aggressive decay kills learning too early. The light decay keeps the model improving but takes many iterations.
Scope: Exponential decay works well when you have a rough idea of how many epochs you need. The decay rate controls the half-life of the learning rate: halves every epochs.

17.4.4 Cosine Annealing

Cosine annealing decays the learning rate following a cosine curve: it starts high, drops rapidly, then slows down gracefully near zero. The formula uses the total number of epochs and two bounds — the maximum learning rate and the minimum :

- — starting (maximum) learning rate at - — target (minimum) learning rate at - — total number of epochs or iterations - — current epoch, ranging from to Verification against standard form: Some texts (Loshchilov & Hutter, 2016) write this as . The two forms are equivalent — just rename and . How it behaves: - : , so (full speed) - : , so (midpoint) - : , so (slowest)
Visual Intuition: Plot on the y-axis against training progress on the x-axis from 0 to 1. The curve looks like a downhill slope of a cosine wave. It starts flat at , drops with increasing steepness through the middle, then flattens out as it approaches . The shape is smooth everywhere with no sudden drops. This is unlike step decay (staircase shape) and exponential decay (fast initial drop then long tail).

Cosine annealing is frequently used in computer vision tasks. It is especially common in vision transformers. You use it when you have a fixed budget of epochs and a known range of learning rates.

17.4.5 Warm-Up Strategy

A warm-up is not a decay technique itself but a pre-decay trick. You start with a very small learning rate (near zero). Then you linearly increase it over the first few epochs — typically 5 to 10 epochs. The rate goes up until it reaches the intended initial learning rate. After warm-up, the actual learning rate schedule takes over.

Think of warm-up like letting a car engine idle for a minute on a cold morning before driving at highway speed. The oil needs to circulate. The pistons need to expand to their proper fit. In a neural network, the early gradients are the equivalent of a cold engine — noisy, unreliable, based on random weights. If you record those early, gibberish gradients into your momentum buffer, they pollute the optimizer's state for many steps afterward. Warming up gives the model time to produce sensible gradients before you start trusting them.

When training complex architectures from scratch — BERT, GPT, or large language models — the early gradients are essentially random. The model has not learned anything meaningful yet. Recording those gibberish gradients to drive momentum would pollute the optimizer state. By warming up, you let the model stabilize its initial weight trajectory before committing to serious training with momentum. After the warm-up epochs, the weights and gradients have settled into something sensible, and you begin the actual learning rate schedule.

Pitfalls: - Skipping warm-up when training transformers from scratch: the model may diverge in the first few hundred steps. The initial random gradients are too noisy for momentum-based optimizers. - Using too long a warm-up wastes compute. If warm-up is 50% of your total training budget, you waste half your compute on tiny learning rates. Aim for 1-5% of total steps. - Combining warm-up with step decay can backfire. If your first step drop happens during warm-up, the model never reaches the intended learning rate.
Comparison — Learning Rate Schedules:
Schedule Shape Best For Weakness
Step Decay Staircase RL, games (known phase transitions) Requires knowing when to drop
Exponential Decay Smooth, convex Feedforward nets, noisy surfaces Decay rate is sensitive
Cosine Annealing S-curve (cosine) Vision, fixed epoch budgets Needs known upfront
Warm-Up Ramp-up, then any schedule Large models from scratch Adds a hyperparameter (warm-up steps)

Learning rate schedules replace the guesswork of a fixed learning rate with a planned trajectory. Large steps early for exploration. Small steps later for refinement. Warm-up protects the optimizer from early random gradients. The next section shifts from improving the learning process (optimization) to constraining it (regularization).

Real-World Connection: Cosine annealing with warm-up is the default schedule for training Vision Transformers (ViT). It is widely used with the AdamW optimizer in the timm (PyTorch Image Models) library. Step decay dominated reinforcement learning for decades — DeepMind's AlphaGo used a manually tuned step decay schedule. Warm-up was popularized by the "Attention Is All You Need" paper (Vaswani et al., 2017) for training the original Transformer. It has been standard practice for large language model pretraining ever since.

17.5 Regularization: Core Concepts

Hook: Your model nailed every training example — 100% accuracy. You deploy it. The first real user types a query with a typo. The model panics and outputs gibberish. It did not learn the underlying pattern. It memorized the training set like a student who memorizes answers by page number. How do you stop a model from becoming a glorified flashcard?

17.5.1 Overfitting and Underfitting

Think of overfitting like a student who memorizes every slide in the exact order and position they appear. When the exam rewords a question or changes the numbers, the student is lost. Underfitting is like a student who only read the chapter titles. They know the broad topic. But they cannot answer any specific question. A well-trained model is like a student who understands the concepts well enough to solve problems they have never seen before.

Overfitting happens when a model memorizes the training data — it learns every point, every noise pattern, every idiosyncrasy. The model performs well on training data but fails on test data. It has not learned the underlying pattern — it has learned the dataset.

Underfitting is the opposite: the model is too simple even for the training data. It captures only a rough average pattern and performs poorly on both training and test sets.

A good model lives between these two: it should specialize for the specific task but generalize well within the task domain.

- Overfitting: Low training error, high test error. The model has high variance — it is too sensitive to the specific training examples. - Underfitting: High training error, high test error. The model has high bias — it is too simple to capture the true pattern. - Good fit: Low training error, low test error. The bias-variance tradeoff is balanced.
Visual Intuition: Picture a scatter plot. Black dots are training data points arranged roughly along a smooth curve but with some noise. An overfitted model (red line) snakes through every black dot perfectly — it wiggles to capture noise. An underfitted model (blue line) is a straight line through the middle — it misses the curve entirely. A well-fitted model (green line) traces the underlying curve without chasing every noisy fluctuation.

17.5.2 Explicit and Implicit Regularization

Explicit regularization directly adds a penalty term to the loss function: - L1 (Lasso): adds . Encourages sparse weights — many become exactly zero. - L2 (Ridge / weight decay): adds . Encourages small but non-zero weights. - Elastic Net: combines L1 and L2: . Implicit regularization emerges from the training setup itself without adding penalty terms: - More training data naturally regularizes — a network cannot memorize millions of examples. - The network architecture (number of layers, width, connectivity) constrains what can be learned. - The optimization algorithm (SGD noise, batch size) adds implicit regularization through stochasticity. - Early stopping implicitly regularizes by halting before overfitting sets in.

A network with 10 layers of 100 neurons each has far too much capacity for 100 instances with 3 features. It will memorize. But with millions of instances, the same network is forced to abstract and generalize. More data is a form of implicit regularization.

17.5.3 Deciding When to Apply Regularization

Track the training loss and validation loss over epochs. If the training loss keeps improving while the validation loss starts diverging (getting worse), you have overfitting. At that point, ask:

Decision flowchart for overfitting: 1. Working with a small dataset → try data augmentation to increase the effective dataset size. 2. Have enough data but the network is moderately complex → try batch normalization or dropout. 3. Network is very complex → consider advanced techniques (MixUp, CutMix, weight decay tuning). 4. Divergence is happening right now → try early stopping — halt training at the point where validation loss starts to separate.

If there is no overfitting (both curves track together), keep training. Adding unnecessary regularization to a model that is not overfitting only hurts performance.

Domain-specific considerations also matter:

- Computer vision: regularization must preserve spatial correlations — nearby pixels share patterns. DropBlock (not regular dropout) works for CNNs. - NLP: regularization must handle variable-length sequences and sequential dependencies. Attention dropout is preferred over neuron dropout for transformers. - Time series: must retain non-stationarity awareness. Regularization should not smooth away temporal shifts that are real signals. - Tabular data: features may interact in complex ways. You cannot assume independence — regularization that assumes uncorrelated features (like independent L1 penalties) may be suboptimal.
Student Q&A — Deduplicated:

> Q: The weights are learned by the model itself. In the Excel example, the true values 0.1 and 2 are hidden. How do we hyperparameter-tune then — don't we just keep trying different values? > A: Yes, hyperparameter tuning is inherently experimental — there is no rule book that says "3 layers, N nodes." It is an art, not a fixed framework. Anyone can randomly try values. But someone who understands what each hyperparameter actually does can tune intelligently. Know that momentum smooths oscillations. Know that RMSProp controls learning rate decay in non-stationary settings. Know that Adam is a good starting point. This knowledge lets you make informed choices rather than blind guesses, using minimal resources.

Pitfalls: - Adding regularization before checking if the model can even fit the training data. Always train without regularization first. If the model cannot overfit a small subset, it has an optimization problem, not a regularization problem. - Using the same regularization strength for all layers. Deeper layers learn more abstract features and may need different regularization than earlier layers. - Assuming more data always fixes overfitting. If your data is redundant (highly correlated examples), adding more similar data does not help — the model already memorized that pattern. - Ignoring the interaction between regularization and learning rate. A model with strong weight decay needs a larger learning rate to overcome the constant shrink toward zero.

Overfitting means your model memorized instead of learned. Regularization — explicit (L1, L2) or implicit (more data, architecture, early stopping) — is how you force generalization. The diagnostic signal is always the same: watch the gap between training loss and validation loss. When it widens, reach for regularization. The following sections give you the specific tools.

Real-World Connection: In production ML systems, overfitting is the most common cause of deployment failure. A fraud detection model trained on last year's transactions may perfectly flag known fraud patterns but completely miss new attack vectors. Teams combat this with continual retraining, data augmentation (synthetic fraud patterns), and implicit regularization through large, diverse training sets. The bias-variance tradeoff is not just a textbook concept. It is the difference between a model that works in the lab and one that works in the wild.

17.6 Gradient Clipping

Hook: You are training an RNN on long text sequences. Suddenly, at step 347, the loss jumps to NaN. Every weight in the network is now NaN — the model is dead. What happened? A gradient somewhere in your 200-layer unrolled network multiplied its way up to infinity. When a gradient explodes, it can irreversibly destroy weeks of training. Gradient clipping is the circuit breaker that stops this.

17.6.1 Vanishing and Exploding Gradients

When gradients are less than 1, the product shrinks toward zero during chain-multiplication across many layers. This is the vanishing gradient problem. When gradients are greater than 1, the product grows without bound — the exploding gradient problem.

Think of gradient flow through a deep network like a game of telephone. You whisper a number to the first person. Each person multiplies it by a factor before passing it on. After 100 people, if the factor is 0.9, the number shrinks to almost zero. No one at the end hears anything useful. That is vanishing. If the factor is 1.1, the number grows to over 13,000 — the last person hears a deafening roar (exploding).

The choice of activation function influences this:

- Sigmoid: outputs , derivative max is 0.25. Causes vanishing gradients. - Tanh: outputs , derivative max is 1.0. Less prone but still vanishes in deep nets. - ReLU: outputs 0 for negative inputs, passes positive inputs unchanged (derivative 1). Vanishing is mitigated on the positive side. But gradients can still explode with large positive weights.

Gradient clipping is one direct solution: if the gradient exceeds a threshold, clip it to that threshold.

17.6.2 Value-Based Clipping

Set a minimum and maximum threshold (typically in the range 0.1 to 5). For every individual gradient value in the gradient vector:

- If it lies between the thresholds, leave it unchanged. - If it is outside, clip it to the nearest threshold.

This clamps each element of the gradient independently. If one parameter's gradient is 20 and the threshold is 5, that element becomes 5. Other elements that are within bounds stay unchanged.

Value clipping is frequently applied in deep reinforcement learning. Environments like Atari games can produce huge reward spikes that create massive gradients for certain state-action pairs.

Pitfall: Value clipping can destroy the relative importance between gradients. If one parameter has gradient 5.0 and another has gradient 50.0, and the threshold is 5.0, both are clipped to 5.0. The optimizer now thinks these parameters are equally important, which is wrong.

17.6.3 Norm-Based Clipping

Norm-based clipping scales the entire gradient vector so that its L2 norm does not exceed a threshold . This preserves the relative proportions between gradient elements — it just caps the total magnitude.

If , leave unchanged.

- — the full gradient vector (all parameters concatenated) - — the L2 norm (Euclidean length) of the gradient vector - — the clipping threshold. Typical values: 1.0, 5.0, or 10.0. - — the scaling factor. Always positive since it is a ratio of two positive numbers.

The scaling factor multiplies every element of , so all signs are preserved. A negative gradient stays negative — it is just scaled down in magnitude.

Worked Example — Norm-based clipping:

Given and :

Step 1: Compute L2 norm: Step 2: Check: , so clipping is needed. Step 3: Compute scaling factor: Step 4: Scale every element: Step 5 — Verification: Compute the norm of the clipped vector: Sense-check: The norm is exactly . The relative proportions are preserved — 7 was the largest element and 3.75 is still the largest. The sign of is preserved as . ✓
Visual Intuition: In 2D, imagine the gradient vector as an arrow from the origin. A circle of radius is drawn around the origin. If the arrow tip lies inside the circle, do nothing. If the arrow tip lies outside, shrink it radially until it touches the circle. The direction of the arrow never changes — only its length. Norm-based clipping is like forcing every gradient update to fit inside a ball of radius .

This is conceptually similar to layer normalization — you scale the entire vector down rather than clamping individual elements. Norm-based clipping is most commonly observed in NLP, especially when training transformer models and LLMs for machine translation tasks.

Student Q&A — Deduplicated:

> Q: Gradients can be both positive and negative. If we clip based on the L2 norm, which always gives a positive value, do we lose the sign information? > A: No. The L2 norm is used only to compute the scaling factor . This scaling factor is always positive as a ratio of two positive numbers. It multiplies every element of the original gradient vector. All signs are preserved. A negative gradient stays negative; it is merely scaled down in magnitude.

Comparison — Value Clipping vs Norm-Based Clipping:
Property Value Clipping Norm-Based Clipping
What is clipped Each gradient element independently Entire gradient vector together
Relative proportions Can be destroyed Preserved
Sign preservation Yes Yes
Primary use case Deep RL (DQN, PPO) NLP, Transformers, LLMs
Hyperparameter Min/max per element Single threshold
Robustness to scale Sensitive — need per-layer tuning Robust — one works across layers
Pitfalls: - Setting too low: you effectively cap all updates, and the model learns nothing. The loss flatlines. - Setting too high: the clipping never triggers, and you get no protection against exploding gradients. - Using value clipping when you need norm clipping is a common mistake. In NLP, the gradient of the embedding layer can be orders of magnitude smaller than the output layer. Value clipping would treat them equally, potentially crushing the small gradients. - Forgetting that gradient clipping is a band-aid, not a cure. It prevents catastrophic failure. But exploding gradients are a symptom of poor initialization, bad learning rates, or problematic architecture. Address the root cause when possible.
Exam note: Value clipping clamps individual gradient elements and is common in RL. Norm-based clipping scales the whole gradient vector to fit within a ball of radius and is standard for training transformers and LLMs. Norm-based clipping preserves the relative proportions between gradients — it only caps the total magnitude.
Real-World Connection: Gradient clipping is a non-negotiable part of training any large language model. When OpenAI trained GPT-3 (175 billion parameters), norm-based clipping with was used throughout. Without it, the sheer scale of the model makes exploding gradients almost certain. This is especially true early in training when the weights are random and produce erratic outputs. The original Transformer paper (Vaswani et al., 2017) reported using norm-based clipping, and it has been standard practice in NLP ever since.

17.7 Weight Initialization

Hook: You have the perfect architecture, the best optimizer, and a great dataset. Training starts. The loss barely moves for the first 100 epochs. When it finally starts dropping, it converges to a mediocre solution. What went wrong? Your weights started in a terrible place. The optimizer spent most of its energy just crawling out of a bad neighborhood before it could even begin learning.

17.7.1 Why Initialization Matters

Starting weights at zero or at random positions can place you far from the global minimum. A poor initialization means the optimizer wastes many iterations just finding the right region of the search space. Smart initialization places the weights in a meaningful starting zone, so the learning converges faster and more stably.

Think of initialization like dropping a marble onto a bumpy landscape. If you drop it on a flat plain 10 kilometers from the deepest valley, it will take forever to roll there. If it ever does. If you drop it near the valley, it finds the bottom quickly. Initialization is choosing where to drop the marble. A good choice puts you in the right ballpark. A bad choice puts you in another country.

This is still an active area of research. There is no universally optimal initialization. But two practical schemes are widely used.

Why you cannot initialize all weights to zero. Every neuron in a layer starts with the same weight. They all compute the same output. They also all receive the same gradient. They stay identical forever — you effectively have one neuron replicated times. This is the symmetry problem. Random initialization breaks this symmetry.

17.7.2 Xavier / Glorot Initialization

For a given layer, let be the number of incoming connections (fan-in) and be the number of outgoing connections (fan-out). The key insight: if the variance of weights on incoming connections equals the variance on outgoing connections, the network learns stably.

Xavier initialization draws weights from a normal distribution with mean zero and variance:

- — fan-in: number of input units to this layer - — fan-out: number of output units from this layer - The factor 2 comes from averaging the fan-in and fan-out variances - The distribution: Why this variance? If you initialize too large, the activations explode as they propagate forward. If too small, the gradients vanish as they propagate backward. This variance balances the forward signal variance with the backward gradient variance — keeping both roughly constant across layers.
Worked Example — Xavier initialization for a hidden layer:

A fully connected layer with inputs and outputs.

Each weight is sampled from . About 68% of weights will fall in . About 95% in .

Sense-check: Larger layers (more fan-in/fan-out) get smaller initial weights. This prevents the weighted sum from growing with layer size. A layer with would get (much smaller). ✓

Xavier initialization works well with linear activations, sigmoid, and tanh activations. It assumes the activation function is about linear around zero. Sigmoid and tanh satisfy this — their derivatives are ~1 at zero.

17.7.3 He Initialization

For ReLU activations, He initialization uses:

Why only fan-in? ReLU kills negative values — it outputs zero for any negative input. On average, about half the incoming signals are squashed to zero, so the effective fan-in is roughly halved. To compensate, the variance uses instead of . - The factor 2 in the numerator compensates for ReLU zeroing out half the activations. - The outgoing fan-out is not considered because signals that land at zero (dead ReLUs) carry no information forward. The forward variance only depends on the "surviving" inputs.

The distribution: .

Sample from and use those samples to initialize the weights whenever ReLU activations are used. Both Xavier and He initialization are forms of implicit regularization — they stabilize the early training dynamics without adding explicit penalty terms.

Comparison — Xavier vs He:
Property Xavier (Glorot) He
Variance
Uses fan-out? Yes No
Best activations Linear, sigmoid, tanh ReLU, Leaky ReLU, PReLU
Why the difference Assumes symmetric activation ( linear near 0) Compensates for ReLU zeroing half the signal
Default in frameworks Used for tanh/sigmoid nets Default for ReLU nets (most modern CNNs)
Pitfalls: - Using Xavier with ReLU: the variance is too small (it accounts for fan-out and does not compensate for dying ReLUs). Gradients can vanish in deep networks. - Using He with tanh: the variance is too large. Activations can saturate early, killing gradients. - Forgetting that initialization interacts with batch normalization. If you use batch norm, the exact initialization matters less because normalization corrects the scale. But good initialization still speeds up convergence. - Applying the same initialization to biases. Biases are typically initialized to zero (or a small constant like 0.01) regardless of the weight initialization scheme. They do not cause symmetry problems because weights break symmetry.
Exam note: Xavier/Glorot initialization works for sigmoid and tanh activations. He initialization is designed for ReLU and its variants. The key formula difference: He uses only (not ) because ReLU kills half the incoming signals, effectively halving the fan-in. Good initialization is a form of implicit regularization — it puts the optimizer in the right neighborhood from step one.
Real-World Connection: He initialization (aka Kaiming initialization, after Kaiming He) was introduced alongside ResNet. It is the default in PyTorch and TensorFlow for convolutional layers with ReLU. Virtually every modern CNN — from ResNet-50 to EfficientNet — uses He initialization. The introduction of proper initialization schemes was one of the key breakthroughs that enabled training networks deeper than ~20 layers. This was nearly impossible with random uniform initialization before 2010.

17.8 Normalization Techniques

Hook: Layer 1 outputs values around 0.5. Layer 2 sees these as inputs and produces outputs around 3.2. Layer 3 gets 3.2 and outputs 17.8. By layer 10, the values have ballooned to 10,000 — gradients are either exploding or vanishing. Each layer is playing a different game with a different scale of numbers. What if you could force every layer to operate on the same playing field?

17.8.1 Covariate Shift

Consider an arbitrary layer in a deep network. The outputs of layer become the inputs to layer . The outputs of layer become the inputs to layer . If the distribution of these outputs is not centered around zero, every subsequent layer's inputs drift further away. Each layer compounds the shift. The variance grows with depth, making training unstable.

Covariate shift is the problem of distributions drifting across layers. The fix: force the output distribution of every layer to be centered around zero. Give it comparable variance. That way all layers operate on similarly-scaled values.

Think of a relay race where each runner hands off a baton. Imagine runner 1 hands off at chest height. Runner 2 grabs it at knee height. Runner 3 grabs it above their head. The fourth runner has no idea where to grab. Normalization is like forcing every handoff to happen at exactly waist height — same position, same orientation, every time. Each runner can focus on running faster instead of adjusting their grip.

17.8.2 Batch Normalization

Batch normalization applies Z-score normalization to the outputs of a layer before passing them to the next layer's activation:

- — the pre-activation output of a neuron (the raw weighted sum, before ReLU/sigmoid/etc.) - — mean of computed over the current mini-batch for that neuron - — variance of computed over the current mini-batch for that neuron - — tiny constant (e.g., ) to avoid division by zero

After normalization, two learnable parameters are applied:

- scale parameter (learnable). Lets the network undo the normalization if zero-mean unit-variance is suboptimal. - shift parameter (learnable). Lets the network shift the normalized distribution.

These let the network learn the optimal scale and shift for the normalized values. The network is not locked to zero-mean unit-variance. If the best representation for the next layer needs mean 2 and variance 5, and can learn that.

Training vs Inference:

During training, batch normalization uses the current batch's statistics ( and from the mini-batch). During inference, there is only one instance at a time — you do not have a batch. Instead, you maintain running statistics: after many training batches, you accumulate the mean and variance from all batches seen so far. For inference, you use these running averages instead of batch-specific statistics. This is why it is called batch normalization: the normalization happens across the batch dimension during training. Running statistics substitute during inference.

The correct placement of batch normalization is before the activation function. You normalize the layer's linear outputs, then apply the activation to the normalized values. This prevents the activation's squashing from masking distribution shifts.

Trace — Batch Normalization on a mini-batch of size 3 for one neuron:

Pre-activation values:

Step 1: Compute batch statistics: - - Step 2: Normalize (with ): - - - Step 3: Apply learned scale and shift (): - - - Sense-check: The normalized values are symmetric around 0 with comparable spread. The transformation then maps them to whatever scale the network needs. ✓

Batch normalization speeds up training and stabilizes learning. It also acts as a regularizer. The batch statistics add a small amount of noise since each mini-batch has slightly different and . That noise has a mild regularizing effect.

Student Q&A — Deduplicated:

> Q: If we apply a sigmoid activation function, the output is always between 0 and 1. Then why do we see a distribution shift between layers? > A: Batch normalization is applied before the activation, not after. The raw weighted sums from a layer can have any distribution — they are not constrained to . If you normalize those pre-activation values first and then apply the sigmoid, you avoid the distribution drift. The correct order is: layer output → batch normalization → activation function → next layer.

17.8.3 Layer Normalization

Layer normalization normalizes across the features within a single layer, rather than across the batch. For a layer with features:

Batch Norm vs Layer Norm — the critical difference: - Batch Norm normalizes per feature, across the batch dimension. For a batch of 32 images and 64 channels, it computes 64 separate means (one per channel). Each is averaged over the 32 batch items. The normalization axis is the batch. - Layer Norm normalizes per instance, across the feature dimension. For a single data point with 64 features, it computes one mean over all 64 features. The normalization axis is the features.
Property Batch Normalization Layer Normalization
Normalization axis Across batch, per feature Across features, per instance
Depends on batch size? Yes — needs large enough batch No — works with batch size 1
Depends on sequence length? No No — normalizes across features, not sequence
Inference behavior Uses running statistics Same computation as training
Primary use CNNs (ResNet, VGG) Transformers, RNNs, LSTMs
Why? CNNs have large batch sizes, spatial weight sharing Transformers have variable sequence lengths; batch norm fails with padding

Layer normalization is the standard choice in transformer architectures and RNNs. It works with variable sequence lengths and does not depend on batch size.

Pitfalls: - Using batch norm with small batch sizes (e.g., 2 or 4): the batch statistics are noisy and unreliable. The running mean/variance during inference may poorly match the training-time statistics. Switch to layer norm or group norm. - Placing batch norm after the activation: the activation squashes the distribution, hiding the shift. Always place normalization before activation. - Using batch norm in RNNs/transformers without accounting for sequence padding: padded positions contribute zeros that corrupt the batch statistics. Layer norm avoids this entirely. - Forgetting that batch norm adds learnable parameters (): these count toward your model size. In a large network, this overhead is negligible — but it exists. - Copying batch norm's running statistics incorrectly from PyTorch to a deployment framework (ONNX, TensorRT) causes silent failures. The running mean and variance must be frozen and exported. Missing this causes silent accuracy drops.
Comparison — Normalization Placement in Architectures:
Architecture Normalization Used Position
ResNet, VGG (CNNs) Batch Norm Conv → BatchNorm → ReLU
Transformer (BERT, GPT) Layer Norm Attention/FFN → LayerNorm (pre or post, depending on variant)
RNN, LSTM Layer Norm Hidden state → LayerNorm → activation
Vision Transformer (ViT) Layer Norm Patch embedding → LayerNorm → Attention
Exam note: Batch norm normalizes across the batch dimension per feature; layer norm normalizes across the feature dimension per instance. Batch norm needs running statistics during inference; layer norm computes the same way in training and inference. Batch norm dominates CNNs; layer norm dominates transformers. Both normalize before the activation function.
Real-World Connection: Batch normalization (Ioffe & Szegedy, 2015) was a breakthrough that made training networks deeper than ~10 layers practical. It is a required component in every modern CNN architecture. Layer normalization (Ba et al., 2016) became essential for transformers because sequence lengths vary and batch sizes are often small in NLP. GPT-3 uses layer normalization after every attention and feed-forward sublayer (172 layers of it). Without normalization, training networks at this depth is nearly impossible. The internal covariate shift would cause gradients to explode or vanish within the first few hundred steps.

17.9 Early Stopping

Hook: Training is going great. Loss drops every epoch. But something feels off. You check the validation loss and it has been climbing for the last 5 epochs while the training loss still drops. Your model is overfitting right now, this very epoch. Do you let it finish? Early stopping says: no. Save the model from 5 epochs ago and walk away.

17.9.1 The Patience Parameter

Think of early stopping like a chef tasting a simmering sauce. The flavor improves with time, but only up to a point. After that, the sauce reduces too much, burns, or becomes too salty. The chef does not wait until the smoke alarm goes off. They taste regularly and pull the pan off the heat when it peaks. The patience parameter counts how many tastings the chef waits. After the last improvement, they declare "this is as good as it gets."

Early stopping monitors the gap between training loss and validation loss. When the validation loss starts diverging from the training loss, overfitting has likely begun. The patience parameter tells the algorithm how many epochs to wait before making the stop decision.

- Patience — number of consecutive epochs the validation loss may worsen before stopping. Typical value: 3–10. - Best model — the checkpoint with the lowest validation loss seen so far. Not the final model.

If patience is 3 and the validation loss has been worsening for 3 consecutive epochs, training stops. The algorithm backtracks to the best model — the one from before the divergence began.

Why use patience rather than stopping instantly on the first divergence? A single epoch's divergence might be noise — a bad mini-batch with outliers, not a genuine overfitting trend. The next epoch might correct itself and the curves could re-converge. Patience gives the optimizer a chance to recover from temporary blips before making the irreversible stop decision.

17.9.2 Implementation

Early stopping algorithm (conceptual): 1. Initialize `best_val_loss = ∞`, `patience_counter = 0`, save initial model. 2. For each epoch:

a. Train one epoch. Compute training loss and validation loss. b. If validation loss < best_val_loss:

- Update `best_val_loss = current_val_loss` - Save current model as best model - Reset `patience_counter = 0`

c. Else (validation loss did not improve):

- `patience_counter += 1` - If `patience_counter > patience_threshold`: - Stop training - Restore best model - Return best model 3. If training completes all epochs without triggering, return best model.
Visual Intuition: Plot loss (y-axis) against epochs (x-axis). The training loss curve (blue) slopes downward continuously. The validation loss curve (orange) slopes downward at first, reaches a minimum around epoch 25, then curves upward. The gap between the curves widens from epoch 25 onward. Mark epoch 25 as the best model — stop here. Mark epoch 28 (patience = 3) as training halted. The region from epoch 25 to 28 is the patience window.
Pitfalls: - Setting patience too low (e.g., 1): you stop on noise. A single bad validation epoch from an unlucky mini-batch ends training prematurely. - Setting patience too high (e.g., 50): you waste compute. The model overfits badly before you stop, and the best checkpoint was epochs ago. - Using early stopping without saving checkpoints: you need the model snapshot from before the divergence. If you only save the final model, early stopping is useless. - Monitoring the wrong metric: validation loss is the standard. But in imbalanced classification, validation F1 or AUC may be more informative. Pick the metric that matters for deployment. - Confusing early stopping with learning rate decay: they serve different purposes. Early stopping halts training. Learning rate decay slows it down. You can use both — decay handles the long-term schedule; early stopping is the emergency brake.

Early stopping is a simple, effective regularizer. It costs almost nothing — just a patience counter and checkpoint saving. It works with any model, any optimizer, any task. It belongs in every training loop. The next section introduces a completely different approach to regularization: randomly disabling parts of the network during training.

Real-World Connection: Early stopping is so universally effective that it is a default component of almost every deep learning training framework. PyTorch Lightning, Keras, and Hugging Face Transformers all include built-in early stopping callbacks. In production settings where training runs can cost thousands of dollars, early stopping is not just a regularization technique. It is a cost-saving measure. It prevents wasting GPU hours on epochs that only make the model worse.

17.10 Dropout and Variants

Hook: What if, during training, you randomly disconnected half the neurons in your network on every single iteration? And it actually made the network better? That is dropout. By forcing the network to function with random subsets of neurons, you prevent it from developing fragile co-dependencies. Every neuron must pull its own weight.

17.10.1 Standard Dropout

Think of dropout like cross-training in sports. A basketball player who only practices free throws with the same routine, same lighting, and same crowd noise will crumble in a real game. Another player practices with random distractions — different ball weights, noisy speakers, one eye covered. That player builds strong skills that work in any condition. Dropout is cross-training for neurons. By randomly disabling teammates, each neuron learns to be useful on its own. It is not just part of a specific combination.

When you suspect a network is too complex and overlearning, you set a dropout rate (e.g., 0.5). At each training iteration, 50% of the neurons in the dropout layer are randomly deactivated. The neurons are not removed — their physical connections remain. The weights learned up to that point are preserved. But for that specific iteration, those neurons do not participate in the forward pass or receive gradient updates.

- Dropout rate — probability of dropping a neuron. Common values: 0.2 (input layers), 0.5 (hidden layers). - At training time: Each neuron is independently kept with probability . Dropped neurons output zero. - At inference time: All neurons are active. Their outputs are scaled down by to compensate for having more active neurons than during training. - Key effect: Prevents co-adaptation. Neurons cannot rely on specific partners always being present. Each neuron must learn features that are useful on their own.

At the next iteration, a different random subset of neurons is dropped. Over the course of training, every neuron gets dropped some of the time. And every neuron gets to learn some of the time. All neurons ultimately learn.

Student Q&A — Deduplicated:

> Q: Why not just design a network with fewer neurons in the first place, instead of using dropout? > A: If you build a smaller network from the start, you lose the representational capacity. A network with 4 nodes can express more patterns than a network with 2 nodes. Dropout lets you keep the large architecture's capacity while preventing overfitting. During training, subsets learn independently. During inference, the full capacity is available. You get the regularization benefit of a small network with the expressive power of a large one.

17.10.2 DropConnect

Dropout deactivates entire neurons. DropConnect is finer-grained: it randomly masks individual weight connections in the weight matrix between layers. At a dropout rate of 50%, half the individual weights are frozen for that iteration. They keep their previous values but do not receive gradient updates. The other half learn normally. In the next iteration, a different random set of 50% of weights are frozen.

This lets you regularize at the synapse level rather than the neuron level. DropConnect is useful for fully connected neural networks.

17.10.3 DropBlock

Dropout and DropConnect work for fully connected layers but fail for convolutional layers. In a CNN, nearby pixels carry spatial correlations. Dropping individual neurons in a feature map does little — adjacent neurons still pass nearly the same spatial information. The network barely notices a dropped neuron. The regularization effect is too weak.

DropBlock addresses this by masking contiguous blocks of neurons in the feature map. By dropping entire spatial regions at once, you force the network to learn from other parts of the image. It cannot rely on fine spatial detail in the dropped region. This works well for vision tasks with strong spatial correlation.

17.10.4 Attention Dropout

In transformer models, after computing the softmax attention weights, attention dropout randomly drops some of those attention weights. Suppose word attends to word and word via self-attention. If the attention weight from to is dropped, the model cannot rely on that token. It must use information from the remaining tokens for context at that iteration.

This prevents the model from over-relying on specific tokens. If a token is noisy or unavailable at inference time, the model can still extract meaning from the remaining tokens. Attention dropout is applied after the softmax in the attention mechanism.

Comparison — Dropout Variants:
Variant What is Dropped Architecture Reason
Standard Dropout Entire neurons Fully connected layers Prevent neuron co-adaptation
DropConnect Individual weights (synapses) Fully connected layers Finer-grained regularization
DropBlock Contiguous blocks CNNs Overcome spatial correlation
Attention Dropout Individual attention weights Transformers Prevent over-reliance on specific tokens
Pitfalls: - Using standard dropout in CNNs: the spatial redundancy makes it nearly useless. Use DropBlock or spatial dropout instead. - Setting dropout too high (e.g., 0.8): too few neurons are active, and the network cannot learn. The capacity is crippled. - Forgetting to scale activations at inference time causes a bug. During training, each neuron is active only of the time. So its expected output is scaled down. At inference, all neurons are active, so outputs must be multiplied by to match the training-time expectation. Most frameworks (PyTorch, TensorFlow) handle this automatically. - Using dropout with batch normalization: the noise from dropout and the noise from batch statistics can interact poorly. In practice, many architectures apply dropout before batch norm or in separate sub-networks. - Applying dropout during inference: dropout is a training-only technique. At inference, the full network is used (with scaling applied).
Exam note: Know dropout and its variants — DropConnect (weight-level, fully connected), DropBlock (block-level, CNNs), attention dropout (weight-level, transformers). Each exists because the standard neuron-level dropout fails for architectures with spatial or sequential structure. Dropout preserves the representational capacity of the full architecture while preventing overfitting through forced independence.
Real-World Connection: Dropout (Srivastava et al., 2014) was a key technique that enabled AlexNet to win the ImageNet competition in 2012. It helped kickstart the deep learning revolution. In modern architectures: ResNet uses no dropout (batch norm provides enough regularization). Transformers use attention dropout and residual dropout (dropping entire sub-layer outputs). EfficientNet uses DropBlock in later layers. The principle is the same everywhere — force the model to handle missing information.

17.11 Data Augmentation

Hook: You have 5,000 cat photos. That feels like a lot. But your network has 10 million parameters — 2,000 parameters per photo. It is going to memorize every whisker. You cannot collect more data. But you can make your existing data work harder by showing the network slightly altered versions of each image. A cat rotated 15 degrees is still a cat. A cat with adjusted brightness is still a cat. Suddenly, 5,000 photos become 50,000 training examples — all from the same disk.

17.11.1 Basic Image Augmentation

Think of data augmentation like a forger creating training samples for a detective. The forger takes an original document and produces slight variations — different handwriting angle, varied ink pressure, slightly different paper color. The detective must learn to recognize the underlying signature regardless of these surface changes. The forger is not creating new information — just teaching the detective which variations do not matter.

When the dataset is small but the network is complex, you can artificially increase the training data. For images, you apply transformations: rotate, scale, zoom, change brightness and contrast, add noise, or perform cutout. Each transformed version becomes a new training instance with the same label. One image might generate 8 augmented copies. None of these transformations change the class — a rotated cat is still a cat.

Another technique adds small Gaussian noise to images. This is similar to the idea behind denoising autoencoders. Train on noisy images so the model learns strong features that survive mild corruption.

Common image augmentations:

- Geometric: rotation (±15°), horizontal/vertical flip, scaling (zoom in/out), translation (shift), shearing - Photometric: brightness adjustment, contrast change, color jitter (hue/saturation), gamma correction - Noise-based: Gaussian noise, salt-and-pepper noise, Gaussian blur - Occlusion: random erasing / cutout (mask a random rectangular region with zeros or noise)

All preserve the class label — the network must learn invariance to these transformations.

17.11.2 MixUp

MixUp blends two training images and their labels. Given image (a cat) and image (a dog), and a mixing coefficient :

- — mixing coefficient, typically sampled from a Beta distribution: with - — pixel-wise weighted average of two images. A ghostly overlay. - — soft label: a weighted combination of the two one-hot labels. For , the label is 70% cat, 30% dog. Why it works: Real images often contain multiple objects or ambiguous boundaries. MixUp forces the model to produce calibrated, non-overconfident predictions. Instead of outputting "100% cat," the model learns to output "70% cat, 30% dog" on blended images, which improves generalization on clean images.

17.11.3 CutMix

CutMix is similar, but instead of blending pixels, it cuts a rectangular patch from one image and pastes it onto the other. The label is weighted by the area ratio of the patch.

- A random bounding box is sampled from image - The region in image is replaced with the corresponding patch from - The label mixing ratio equals the area ratio of the patch:

This is like occlusion — the cat might be partially hidden behind a cushion, with only its tail visible. The model must learn to recognize objects from partial views.

Comparison — MixUp vs CutMix:
Property MixUp CutMix
How images combine Pixel-wise weighted average (ghostly blend) Rectangular patch replacement (sharp boundaries)
Visual result Both images are faintly visible everywhere One region is one image, the rest is the other
Label mixing from Beta distribution = patch area ratio
Real-world analogy Looking through a semi-transparent overlay Object partially hidden behind another
Primary benefit Calibrated, non-overconfident predictions Robustness to occlusion and partial views

17.11.4 Progressive Augmentation

A practical training strategy: start with no data augmentation. Train until performance plateaus. Then introduce simple augmentations (flipping, rotation, scaling). If accuracy improves, continue. Then add harder augmentations (color jitter, cutout). If that helps further, try MixUp and CutMix. This progressive approach strengthens the model incrementally. Applying all augmentations at once can make training too difficult for the model early on. At that point, it has not yet learned the basic features.

Pitfalls: - Applying augmentations that change the class: rotating a "6" by 180° makes it a "9" — do not use rotation for digit recognition. Flipping text horizontally makes it unreadable. Always verify that your augmentations are label-preserving. - Augmenting the validation set: augmentation is for training only. The validation and test sets must use clean, unmodified images — otherwise you cannot measure true generalization. - Over-augmenting early in training: if the model has not learned basic features yet, heavy augmentation makes the task too hard. Use progressive augmentation. - Using MixUp or CutMix with too high (Beta parameter close to 0): this produces near-50-50 blends, making the task essentially a guessing game. to is typical. - Forgetting that augmentation changes the effective dataset size: if you generate 10 augmented versions per original image, an epoch now takes 10x longer. Budget your training time accordingly.
Exam note: Data augmentation artificially expands the training set using label-preserving transformations. Basic transforms (rotation, flip, noise) create variations of existing images. MixUp blends two images and labels with a mixing coefficient. CutMix pastes a patch from one image onto another. Progressive augmentation applies techniques in order of difficulty. Augmentation is one of the most cost-effective regularizers — it uses no extra parameters, no extra data collection, just extra computation.
Real-World Connection: Data augmentation is universal in production computer vision. ImageNet training pipelines apply random resized crops, horizontal flips, and color jitter as standard. Vision Transformers (ViT) use MixUp and CutMix during training to achieve state-of-the-art results. In medical imaging where labeled data is scarce and expensive, aggressive augmentation is often critical. Elastic deformations for tissue images can make the difference between a working model and a useless one. Self-driving car perception systems use augmentation to simulate different conditions. These include weather, lighting, and occlusions that would be dangerous or impossible to collect in real life.

Exam Guidance Summary

This lecture covers two major modules. Optimization techniques show how to make models learn better and faster. Regularization techniques show how to prevent memorization and improve generalization. Both are foundational for training deep networks.

Optimization: What to Know

- Conceptual comparisons: Be able to compare SGD, SGD with momentum, Nesterov (NAG), AdaGrad, RMSProp, Adam, and AdamW. For each, know what problem it solves (noisy gradients, fixed learning rate, sparse data, weight decay coupling). - Which optimizer for which scenario: Sparse data → AdaGrad. Non-stationary/sequential/RNN → RMSProp. General/transformers → Adam. Transformer training with weight decay → AdamW. - Adam bias correction: Know why it exists (cold start with biases early estimates toward zero), how it works (), and why it fades with . - Learning rate schedules: Step decay vs exponential decay vs cosine annealing vs warm-up. Know which is used where (RL → step decay; vision → cosine annealing; large models from scratch → warm-up).

Regularization: What to Know

- Gradient clipping: Value clipping (clamps per element, used in RL) vs norm-based clipping (scales whole vector, used in NLP/transformers). Understand why norm-based preserves proportions and how is the only hyperparameter. - Weight initialization: Xavier/Glorot (, for sigmoid/tanh) vs He (, for ReLU). Know why He uses only fan-in — ReLU kills half the signals. - Batch norm vs layer norm: What is normalized across in each (batch dimension vs feature dimension). Where each is applied (CNNs → batch norm; transformers → layer norm). How inference-time statistics work for batch norm (running mean/variance). Correct placement: before activation. - Dropout and variants: Standard dropout (neuron-level, fully connected). DropConnect (weight-level, fully connected). DropBlock (block-level, CNNs — overcomes spatial correlation). Attention dropout (weight-level, transformers — prevents token over-reliance). Know why different variants exist for different architectures. - Data augmentation: Basic image transforms, MixUp (pixel-wise blend with soft labels), CutMix (patch replacement with area-weighted labels). Rationale for progressive augmentation — strengthen the model incrementally.

Recurring Themes

- Optimization vs regularization: improve the learning process vs prevent overfitting. These are complementary, not competing, goals. - The diagnostic signal for overfitting is always the train/validation loss gap. When validation loss diverges from training loss, reach for regularization. - Hyperparameter tuning is experimental, not formulaic. Understanding what each technique does enables intelligent tuning rather than blind search.

Key Industry Applications

Optimization in Practice

- Adam is the default optimizer in PyTorch, TensorFlow, and JAX. It trains virtually every modern transformer model (GPT, BERT, T5, LLaMA). Its combination of momentum and adaptive learning rates provides strong convergence across diverse architectures with minimal tuning. - AdamW (decoupled weight decay) is used internally in GPT and other large language models. The decoupling prevents the adaptive scaling from distorting regularization strength, leading to better generalization at scale. - Norm-based gradient clipping is standard when training transformer models and LLMs for machine translation tasks. Without it, the chain multiplication of gradients through attention layers makes exploding gradients nearly certain. - Value-based gradient clipping is common in deep reinforcement learning algorithms (DQN, PPO). Atari games and robotics environments produce sparse, high-magnitude reward signals that create massive gradient spikes. - Step decay learning rate schedules are widely used in reinforcement learning for games like chess and Go. Domain experts set the drop epochs based on known phase transitions in the learning process. - Cosine annealing is popular in computer vision tasks, especially vision transformers (ViT). It provides smooth, predictable learning rate decay over a fixed epoch budget. - Warm-up strategies are essential when training large language models (BERT, GPT) from scratch. Starting with full learning rates on random weights produces unstable training; warm-up lets the optimizer accumulate reliable gradient statistics first.

Regularization in Practice

- Batch normalization is a standard component in CNN architectures (ResNet, VGG, EfficientNet). It enables training networks deeper than ~10 layers and speeds up convergence by 10-50x compared to unnormalized networks. - Layer normalization is the norm in transformers (BERT, GPT, T5). It handles variable sequence lengths and works with small batch sizes — critical for NLP workloads where sequence padding would corrupt batch statistics. - MixUp and CutMix are modern data augmentation techniques used in state-of-the-art image classification (training Vision Transformers, EfficientNet, ConvNeXt). They improve generalization and calibration without requiring more data. - Attention dropout is used in transformer training to prevent the model from over-relying on specific tokens. This is critical for machine translation, where the model must handle missing or noisy words at inference time. - AdaGrad excels with sparse features common in NLP tasks like sentiment analysis with bag-of-words features. In these tasks, most vocabulary terms are absent from any given document. - RMSProp is effective for RNN/LSTM training on sequential, non-stationary data such as time series forecasting and speech recognition. In these domains, gradient statistics shift over time. - Progressive augmentation strategies are used in production vision systems. They start simple (flip, rotation) and add complexity (color jitter, cutout, MixUp, CutMix) as the model improves. This approach is standard in self-driving car perception and medical imaging pipelines.

DNN Lecture 17 notes · Optimization and Regularization

Deep Neural Networks· postgraduate· 2026-07-15

Sections Breakdown

117.1 Optimization vs Regularization

Section covering 17.1 Optimization vs Regularization

217.2 Momentum-Based Gradient Updates

Section covering 17.2 Momentum-Based Gradient Updates

317.3 Adaptive Learning Rate Methods

Section covering 17.3 Adaptive Learning Rate Methods

417.4 Learning Rate Schedules

Section covering 17.4 Learning Rate Schedules

517.5 Regularization: Core Concepts

Section covering 17.5 Regularization: Core Concepts

617.6 Gradient Clipping

Section covering 17.6 Gradient Clipping

717.7 Weight Initialization

Section covering 17.7 Weight Initialization

817.8 Normalization Techniques

Section covering 17.8 Normalization Techniques

917.9 Early Stopping

Section covering 17.9 Early Stopping

1017.10 Dropout and Variants

Section covering 17.10 Dropout and Variants

1117.11 Data Augmentation

Section covering 17.11 Data Augmentation

12Exam Guidance Summary

Section covering Exam Guidance Summary

13Key Industry Applications

Section covering Key Industry Applications

Postgraduate students in Deep Neural Networks

Exam Revision Notes

Below is the distilled, exam-ready core of this lecture. Every entry is built from the full textbook notes above. Use this section for rapid review — but if something doesn't make sense, go back to the full explanation in the main content.

Momentum-Based Optimization

Must-know: Momentum accumulates a velocity of past gradients to smooth updates. Nesterov looks ahead for faster correction.

⚠️ Top pitfall: Setting β too high (0.999) — velocity decays too slowly and carries stale gradient information.

Self-check: What problem does momentum solve in SGD?

Connects to: Plain SGD (no momentum), Nesterov Accelerated Gradient, Adam (uses momentum internally)

Adaptive Learning Rate Methods

Must-know: AdaGrad accumulates squared gradients (aggressive decay). RMSProp uses a moving average. Adam combines momentum + adaptive scaling with bias correction.

⚠️ Top pitfall: Using Adam with default β₂=0.999 on rapidly changing gradient statistics (e.g., GANs) — the second moment adapts too slowly.

Self-check: Why does Adam need bias correction?

Connects to: AdaGrad, RMSProp, AdamW

Learning Rate Schedules

Must-know: Step decay drops the rate at fixed epochs. Exponential decay applies smooth decay. Cosine annealing follows a cosine curve. Warm-up increases from near zero before the main schedule.

⚠️ Top pitfall: Skipping warm-up when training transformers from scratch — early random gradients pollute the optimizer's momentum buffer.

Self-check: Which schedule is best for training Vision Transformers?

Connects to: Step Decay, Exponential Decay, Cosine Annealing, Warm-Up

Overfitting and Regularization

Must-know: Overfitting = low training error, high test error (high variance). Underfitting = high error on both (high bias). Explicit regularization (L1, L2) adds penalties; implicit regularization (more data, architecture, early stopping) emerges from training setup.

⚠️ Top pitfall: Adding regularization before confirming the model can even fit the training data — first check if underfitting or overfitting is the problem.

Self-check: What is the diagnostic signal for overfitting?

Connects to: L1 Regularization, L2 Regularization, Early Stopping, Dropout

Gradient Clipping

Must-know: Value clipping clamps each gradient element independently (used in RL). Norm-based clipping scales the whole gradient vector to a threshold τ, preserving relative proportions (standard in NLP/transformers).

⚠️ Top pitfall: Using value clipping for transformers — the embedding layer and output layer have vastly different gradient scales, and value clipping treats them equally.

Self-check: Why does norm-based clipping preserve relative proportions between gradients?

Connects to: Exploding Gradients, Value Clipping, Norm-Based Clipping

Weight Initialization

Must-know: Xavier/Glorot: σ² = 2/(n_in + n_out), for sigmoid/tanh. He: σ² = 2/n_in, for ReLU. He uses only fan-in because ReLU zeros out half the incoming signals.

⚠️ Top pitfall: Using Xavier with ReLU — the variance is too small, causing vanishing gradients in deep networks.

Self-check: Why does He initialization use only fan-in and not fan-out?

Connects to: Symmetry Problem, ReLU Activation, Batch Normalization

Normalization Techniques

Must-know: Batch norm normalizes across the batch dimension per feature (used in CNNs). Layer norm normalizes across the feature dimension per instance (used in transformers). Both use learnable scale (γ) and shift (β) parameters.

⚠️ Top pitfall: Using batch norm with very small batch sizes (e.g., 2-4) — the batch statistics become noisy and unreliable.

Self-check: Why does layer normalization work better than batch normalization for transformers?

Connects to: Covariate Shift, Batch Normalization, Layer Normalization

Early Stopping

Must-know: Monitors validation loss and halts training when it diverges from training loss for a set number of epochs (patience). Restores the best checkpoint before divergence.

⚠️ Top pitfall: Setting patience too low (e.g., 1) — you stop on noise instead of genuine overfitting.

Self-check: Why do we need a patience parameter instead of stopping on the first divergence?

Connects to: Overfitting, Validation Loss, Model Checkpointing

Dropout and Variants

Must-know: Standard dropout drops neurons (fully connected). DropConnect drops individual weights. DropBlock drops contiguous blocks (CNNs). Attention dropout drops attention weights (transformers).

⚠️ Top pitfall: Using standard dropout in CNNs — spatial redundancy makes neuron-level dropout ineffective. Use DropBlock instead.

Self-check: Why does DropBlock work better than standard dropout for convolutional layers?

Connects to: Standard Dropout, DropConnect, DropBlock, Attention Dropout

Data Augmentation

Must-know: Label-preserving transformations artificially expand the training set. MixUp blends two images and their labels. CutMix replaces a rectangular patch with another image. Progressive augmentation increases difficulty gradually.

⚠️ Top pitfall: Applying augmentations that change the class label — rotating a '6' by 180° makes it a '9'.

Self-check: What is the advantage of progressive augmentation over applying all augmentations at once?

Connects to: Basic Image Augmentation, MixUp, CutMix, Progressive Augmentation

Was this lecture useful?

Loading comments…
🤖

BitsNotes AI Assistant

Subject Notes Assistant

Configure AI Chat

Choose how to access the chatbot
Have your own API key?

Switch to "Bring Your Own Key" tab above for unlimited access with any OpenAI-compatible provider.

🔑 Enter API key above to fetch live models from provider, or enter model name manually.
OpenAI-Compatible API Support

Choose any provider preset (Gemini, DeepSeek, Kimi, GLM, MiniMax, Qwen, OpenAI, Groq, Ollama, etc.) or enter a custom endpoint URL.

Security & Privacy First

Your API key is sent directly from your browser to your specified provider. BitsNotes servers never store or see your key.