Optimization and Regularization
Prerequisite Knowledge
This lecture builds on the following concepts from earlier lectures. If any feel unfamiliar, review the linked notes before proceeding.
Previously Covered in This Subject
- Data Augmentation — covered in Session 1: Introduction and Overview of Deep Neural Networks
- Adaptive Learning Rate Methods — covered in Session 1: Introduction and Overview of Deep Neural Networks
- Backpropagation — covered in Session 1: Introduction and Overview of Deep Neural Networks
- Layer Normalization — covered in Session 1: Introduction and Overview of Deep Neural Networks
- Learning Rate Schedules — covered in Session 1: Introduction and Overview of Deep Neural Networks
- Activation Functions — covered in Deep Neural Network Components and Perceptron
- Weight Initialization — covered in Deep Neural Network Components and Perceptron
- Learning Rate Schedules — covered in Deep Neural Network Components and Perceptron
- Adaptive Learning Rate Methods — covered in Deep Neural Network Components and Perceptron
- Data Augmentation — covered in Deep Neural Network Components and Perceptron
- Layer Normalization — covered in Deep Neural Network Components and Perceptron
- Optimization vs Regularization — covered in Deep Neural Network Components and Perceptron
- Layer Normalization — covered in Perceptron Learning and Introduction to Regression
- Early Stopping — covered in Perceptron Learning and Introduction to Regression
- Data Augmentation — covered in Perceptron Learning and Introduction to Regression
- Weight Initialization — covered in Perceptron Learning and Introduction to Regression
- Xavier Initialization — covered in Perceptron Learning and Introduction to Regression
- He Initialization — covered in Perceptron Learning and Introduction to Regression
- Adaptive Learning Rate Methods — covered in Perceptron Learning and Introduction to Regression
- Learning Rate Schedules — covered in Perceptron Learning and Introduction to Regression
- Activation Functions — covered in Linear Neural Networks for Regression
- Layer Normalization — covered in Linear Neural Networks for Regression
- Weight Initialization — covered in Linear Neural Networks for Regression
- Learning Rate Schedules — covered in Linear Neural Networks for Regression
- Data Augmentation — covered in Linear Neural Networks for Regression
- Gradient Clipping — covered in Linear Neural Networks for Regression
- Adaptive Learning Rate Methods — covered in Linear Neural Networks for Regression
- Momentum-Based Gradient Updates — covered in Linear Neural Networks for Regression
- Gradient Descent — covered in Linear Neural Networks for Regression
- Batch Normalization — covered in Linear Neural Networks for Regression
- Activation Functions — covered in Gradient Descent Variants, Classification, and Evaluation
- Gradient Clipping — covered in Gradient Descent Variants, Classification, and Evaluation
- Weight Initialization — covered in Gradient Descent Variants, Classification, and Evaluation
- Underfitting — covered in Gradient Descent Variants, Classification, and Evaluation
- He Initialization — covered in Gradient Descent Variants, Classification, and Evaluation
- Gradient Descent — covered in Gradient Descent Variants, Classification, and Evaluation
- Xavier Initialization — covered in Gradient Descent Variants, Classification, and Evaluation
- Batch Normalization — covered in Gradient Descent Variants, Classification, and Evaluation
- Momentum-Based Gradient Updates — covered in Gradient Descent Variants, Classification, and Evaluation
- Overfitting — covered in Gradient Descent Variants, Classification, and Evaluation
- Batch Normalization — covered in Deep Feedforward Neural Networks
- Underfitting — covered in Deep Feedforward Neural Networks
- Data Augmentation — covered in Deep Feedforward Neural Networks
- Backpropagation — covered in Deep Feedforward Neural Networks
- Gradient Descent — covered in Deep Feedforward Neural Networks
- Overfitting — covered in Deep Feedforward Neural Networks
- Activation Functions — covered in Deep Feedforward Neural Networks
- Momentum-Based Gradient Updates — covered in Deep Feedforward Neural Networks
- Weight Initialization — covered in Deep Feedforward Neural Networks
- Early Stopping — covered in Deep Feedforward Neural Networks
- Gradient Clipping — covered in Deep Feedforward Neural Networks
- Gradient Descent — covered in Deep Feed-Forward Neural Network Architecture Design and Training
- Layer Normalization — covered in Deep Feed-Forward Neural Network Architecture Design and Training
- Early Stopping — covered in Deep Feed-Forward Neural Network Architecture Design and Training
- Activation Functions — covered in Deep Feed-Forward Neural Network Architecture Design and Training
- Weight Initialization — covered in Deep Feed-Forward Neural Network Architecture Design and Training
- Momentum-Based Gradient Updates — covered in Deep Feed-Forward Neural Network Architecture Design and Training
- Gradient Clipping — covered in Deep Feed-Forward Neural Network Architecture Design and Training
- Data Augmentation — covered in Deep Feed-Forward Neural Network Architecture Design and Training
- Overfitting — covered in Deep Feed-Forward Neural Network Architecture Design and Training
- Data Augmentation — covered in Exam Revision and Introduction to Convolutional Neural Networks
- Xavier Initialization — covered in Exam Revision and Introduction to Convolutional Neural Networks
- Overfitting — covered in Exam Revision and Introduction to Convolutional Neural Networks
- Gradient Clipping — covered in Exam Revision and Introduction to Convolutional Neural Networks
- Batch Normalization — covered in Exam Revision and Introduction to Convolutional Neural Networks
- He Initialization — covered in Exam Revision and Introduction to Convolutional Neural Networks
- Adaptive Learning Rate Methods — covered in Exam Revision and Introduction to Convolutional Neural Networks
- Momentum-Based Gradient Updates — covered in Exam Revision and Introduction to Convolutional Neural Networks
- Weight Initialization — covered in Exam Revision and Introduction to Convolutional Neural Networks
- Gradient Descent — covered in Exam Revision and Introduction to Convolutional Neural Networks
- Learning Rate Schedules — covered in Exam Revision and Introduction to Convolutional Neural Networks
- Layer Normalization — covered in Exam Revision and Introduction to Convolutional Neural Networks
- Adaptive Learning Rate Methods — covered in Convolutional Neural Networks
- Activation Functions — covered in Convolutional Neural Networks
- Normalization Techniques — covered in Convolutional Neural Networks
- Learning Rate Schedules — covered in Convolutional Neural Networks
- Layer Normalization — covered in Convolutional Neural Networks
- Batch Normalization — covered in Convolutional Neural Networks
- Learning Rate Schedules — covered in Recurrent Neural Networks — Foundations and Architecture
- Data Augmentation — covered in Recurrent Neural Networks — Foundations and Architecture
- Adaptive Learning Rate Methods — covered in Recurrent Neural Networks — Foundations and Architecture
- Backpropagation — covered in Recurrent Neural Networks — Foundations and Architecture
- Activation Functions — covered in Recurrent Neural Networks — Foundations and Architecture
- Optimization vs Regularization — covered in Recurrent Neural Networks — Foundations and Architecture
- Dropout — covered in Recurrent Neural Networks — Foundations and Architecture
- Activation Functions — covered in Recurrent Neural Networks — Advanced Architectures
- Gradient Clipping — covered in Recurrent Neural Networks — Advanced Architectures
- Gradient Descent — covered in Recurrent Neural Networks — Advanced Architectures
- Data Augmentation — covered in Recurrent Neural Networks — Advanced Architectures
- Momentum-Based Gradient Updates — covered in Recurrent Neural Networks — Advanced Architectures
- Layer Normalization — covered in Attention Mechanisms
- Early Stopping — covered in Attention Mechanisms
- Weight Initialization — covered in Attention Mechanisms
- Batch Normalization — covered in Attention Mechanisms
- Data Augmentation — covered in Attention Mechanisms
- Adaptive Learning Rate Methods — covered in Attention Mechanisms
- Optimization vs Regularization — covered in Attention Mechanisms
- Normalization Techniques — covered in Attention Mechanisms
- Learning Rate Schedules — covered in Attention Mechanisms
- Layer Normalization — covered in Attention Mechanisms and Transformer Architecture
- Normalization Techniques — covered in Attention Mechanisms and Transformer Architecture
- Weight Initialization — covered in Attention Mechanisms and Transformer Architecture
- Batch Normalization — covered in Attention Mechanisms and Transformer Architecture
- Learning Rate Schedules — covered in Transformer Architectures and Optimization
- Layer Normalization — covered in Transformer Architectures and Optimization
- Normalization Techniques — covered in Transformer Architectures and Optimization
- Momentum-Based Gradient Updates — covered in Transformer Architectures and Optimization
- Gradient Descent — covered in Transformer Architectures and Optimization
- Gradient Clipping — covered in Transformer Architectures and Optimization
- Adaptive Learning Rate Methods — covered in Transformer Architectures and Optimization
- Batch Normalization — covered in Transformer Architectures and Optimization
- Optimization vs Regularization — covered in Transformer Architectures and Optimization
Optimization and Regularization
Training a deep network is a tug-of-war between two forces. One pushes the model to learn faster and more stably from the data. The other pulls it back so it does not memorize every wrinkle in the training set. This lecture shows you how both forces work, why they matter, and when to reach for each.
17.1 Optimization vs Regularization
17.1.1 Two Sides of the Training Coin
Think of training a deep network like coaching a student for an exam. Optimization is the drill sergeant who makes the student practice every problem in the textbook until they can solve it blindfolded. Regularization is the wise tutor who shuffles the problem wording and throws in curveballs. The real exam will not look exactly like the textbook. One builds raw skill. The other builds adaptability. You need both.
Optimization and regularization solve different problems during training. Optimization is about the learning process itself — you push the model to learn better, faster, and more stably. Regularization is about preventing the model from overfitting the training data, so it generalizes to unseen examples.
From a deep learning perspective: if your model is overfitting, you reach for regularization. If the model's capacity is not up to the mark and needs fine-tuning for the task, you reach for optimization. One refines the learning process; the other constrains it.
| Dimension | Optimization | Regularization |
|---|---|---|
| Goal | Improve learning speed and stability | Prevent overfitting; improve generalization |
| Acts on | Gradient updates, learning rates | Model complexity, weight magnitudes |
| Signal | Training loss not decreasing fast enough | Validation loss diverging from training loss |
| Examples | Momentum, Adam, learning rate schedules | L1/L2 penalties, dropout, data augmentation |
| When to use | Model underfits or converges too slowly | Model overfits (low training error, high test error) |
17.1.2 The Two Questions This Module Answers
The optimization module answers two core questions. First, how do you update the gradient component of the weight update equation? Not just using the current gradient. But also using historical momentum to smooth the trajectory. Second, how do you make the learning rate adaptive instead of fixed? So every parameter gets its own effective step size.
The regularization module answers: what do you do when the model overlearns the training data? How do you prevent memorization and ensure the model captures the underlying pattern rather than the noise?
Optimization makes the model learn faster and smarter. Regularization keeps it honest. The rest of this lecture walks through every major tool in both toolboxes — starting with momentum-based gradient updates.
17.2 Momentum-Based Gradient Updates
17.2.1 The Standard Gradient Descent Update
Think of gradient descent like a blindfolded person trying to reach the bottom of a valley. At every step they feel which way the ground slopes and take a step in that direction. The size of the step is the learning rate. A small step means slow but safe progress. A large step means faster progress but risk of overshooting. This is the simplest possible learning rule — and its simplicity is also its weakness.
The basic weight update is:
Where:
The update rule says: take the current weight, subtract a fraction of the gradient. That fraction is the learning rate. If the gradient is steep, the step is large. If the gradient is flat, the step is small.
17.2.2 Why Momentum Helps
When you use plain stochastic gradient descent, each parameter's gradient pulls the update in its own direction. One feature might pull gently; another might pull aggressively. The trajectory across the loss surface becomes noisy with oscillations. The problem is that every sampled instance or minibatch might have a different pattern. Because of this, successive weight adjustments see large swings while trying to converge.
The solution is to take help from the history of gradients. Instead of letting the current gradient alone decide the step, you accumulate a velocity term. This velocity carries forward information from previous gradients. It smooths out the oscillations and produces a straighter, faster trajectory toward the minimum.
17.2.3 SGD with Momentum
Momentum works like a heavy ball rolling down a hill. The ball has inertia — it keeps moving in roughly the same direction even when the ground bumps it sideways. A small bump does not reverse the ball's direction. Only a sustained slope change can redirect it. This is exactly what we want from our optimizer: resist noisy gradient fluctuations, follow the consistent downward trend.
The update becomes:
The physical intuition: imagine the gradient as an arrow pointing in some direction at each step. Without momentum, the update follows that arrow exactly, even if it oscillates wildly. With momentum, you take a scaled-down version of the previous arrow and add it to the current gradient arrow via vector addition. The resulting arrow is the actual direction the update moves. This compromise smooths out the zigzag.
Visual Intuition: On a contour plot of the loss surface (like a topographic map), plain SGD traces a jagged zigzag path. Each step corrects the overcorrection of the previous step. SGD with momentum traces a smoother, more direct curve. The velocity term acts like a low-pass filter on the gradient signal. It dampens high-frequency oscillations while preserving the low-frequency descent direction. The x-axis and y-axis represent two parameter dimensions. The contour lines show equal-loss levels, with the minimum at the center.Start with , current gradient , and .
This new velocity replaces the plain gradient in the weight update. Without momentum, the update would use only the raw gradient 7.2. With momentum, the effective update signal is twice as strong — but smoothed by the contribution from past steps.
Sense-check: from the current gradient plus carried over from history. The past still dominates slightly since equals the fresh gradient — this shows that with , the velocity has a long memory.17.2.4 Nesterov Accelerated Gradient (NAG)
Standard momentum is like driving a car while only looking in the rearview mirror. You base your steering on where you have been. Nesterov momentum is like glancing ahead through the windshield too. You first coast forward using your current velocity. Then you check the gradient at that future position. Only then do you decide your actual step. This look-ahead correction makes NAG more responsive to changes in the slope.
Standard momentum looks backward — it uses the previous velocity plus the current gradient. Nesterov's method also looks ahead. Instead of computing the gradient at the current position , you first take a look-ahead step using the previous velocity. Then you compute the gradient at that future position:
The critical difference: the gradient is evaluated at — the position you would land at if you just coasted on momentum. This is the "look-ahead" position. Then, as before, you combine this look-ahead gradient with the previous velocity.
Why it helps: Suppose momentum carries you toward a steep uphill. The look-ahead gradient at that future position will be strongly opposed to the current direction. NAG detects this early and slows down before overshooting. Standard momentum would charge ahead and only correct after the fact.17.2.5 Numerical Example: Momentum's Effect
Consider the same loss function, same initialization , and same learning rate used earlier for plain gradient descent. Without momentum, kept decreasing from the positive region toward zero. But it never crossed into negatives. With momentum, accelerates convergence and moves into the negative region. It overshoots the zero line, finds a better valley, and speeds up learning. The accumulated velocity carries it past local flat regions.
At each step:
1. Take previous velocity 2. Multiply by (say 0.9) 3. Add the current gradient 4. Use the resulting to update the weightThe key insight: once velocity builds up in a consistent direction, it acts like a flywheel. Even when the current gradient at a particular point is small (a flat region), the accumulated velocity keeps the parameter moving. This is why can cross zero. The velocity carries it through, even though the local gradient at zero would not push it far enough on its own.
| Property | Standard Momentum | Nesterov (NAG) |
|---|---|---|
| Gradient evaluated at | Current position | Look-ahead position |
| Responsiveness | Reacts to past | Anticipates future |
| Overshooting | More likely near minima | Less likely — corrects early |
| Convergence speed | Faster than plain SGD | Slightly faster than standard momentum |
| Extra computation | None vs plain SGD | One extra gradient evaluation (or clever reparameterization) |
Momentum turns the optimizer from a myopic step-taker into a boulder rolling downhill. It builds velocity in consistent directions and ignores noisy gradient fluctuations. The next section asks: what if we could also give each parameter its own learning rate?
17.3 Adaptive Learning Rate Methods
17.3.1 Why a Fixed Learning Rate Fails
Imagine you are adjusting two dials on a radio. One dial (volume) needs tiny, precise turns — a big twist blows out the speakers. The other dial (tuning) needs broad sweeps to find the station — tiny nudges get you nowhere. A single "turn size" fails for both. Parameters in a network are the same: some have steep, volatile gradients and need small steps. Others have flat, consistent gradients and need larger steps to make progress.
In the same problem, one parameter's gradient flow changes gracefully, while the other parameter's gradient drops aggressively fast. A fixed learning rate is too fast for one parameter and too slow for the other. You need each parameter to have its own effective learning rate. Parameters with small, stable gradients can move faster. Parameters with large, volatile gradients need smaller, careful steps.
17.3.2 AdaGrad
AdaGrad adapts the learning rate per parameter using the accumulated squared gradient. The idea: parameters that have received large gradients get a smaller effective learning rate. Parameters that have received small gradients get a larger one.
The update is:
Where:
- — accumulated sum of squared gradients up to time : . Same shape as . - — tiny constant (e.g., ) to avoid division by zero. - — element-wise (Hadamard) multiplication. Each parameter gets its own scaling factor. - — global learning rate, divided per-parameter by . How it works: A parameter with consistently large gradients builds up a large . The denominator grows, shrinking the effective step. A parameter with small gradients keeps a small , so its effective step stays larger. Over time, every parameter's learning rate decays — but parameters that have already moved a lot decay faster.Two parameters, and . Global , .
After 3 steps, has seen gradients . Squared sum: .
After 3 steps, has seen gradients . Squared sum: .
Effective learning rate for : .
Effective learning rate for : .
Sense-check: learns about 26 times faster per step than because its gradients have been consistently small. This is exactly the desired behavior — needs bigger steps to catch up.17.3.3 RMSProp
AdaGrad's problem is that it never forgets. Every gradient since the first step pulls down the learning rate. RMSProp fixes this by using a leaky memory — an exponentially weighted moving average of squared gradients. Old gradients fade away, so the denominator does not grow without bound. This makes RMSProp work for non-stationary problems where the gradient distribution shifts over time.
RMSProp fixes AdaGrad's aggressive decay by using an exponentially weighted moving average of squared gradients instead of a cumulative sum:
The key difference from AdaGrad: can go down as well as up. If recent gradients are small, shrinks and the learning rate recovers. AdaGrad's only grows.
RMSProp is best suited for non-stationary problems. These are situations where the underlying data pattern shifts over time. It also works well for sequential decision-making tasks — think RNNs and LSTMs.
17.3.4 Adam
RMSProp gives you per-parameter learning rates. Momentum gives you smoothed gradient directions. Adam puts them together — like RMSProp and momentum had a child that inherited the best of both. Add a clever bias-correction trick for the first few steps. Now you have the most widely used optimizer in deep learning.
Adam combines RMSProp's adaptive learning rate with SGD's momentum. It maintains two moving averages:
The raw (uncorrected) update would be:
Adam corrects this with bias correction:
When is small (say ):
- , so . Dividing by 0.1 scales it up by 10x. This compensates for the fact that only captured 10% of the first gradient. - Same logic for .When is large (say ):
- . . The correction factor is essentially 1. The raw moving averages are trustworthy.The bias correction is a warm-up mechanism — it boosts early estimates when they are unreliable, then gracefully fades out.
With bias corrections, the final Adam update is:
For the numerical example with the same loss function and initialization , Adam produces smooth, steady updates. Both parameters progress gradually — neither too slow nor too fast. The bias corrections ensure unbiased estimates at the initial iterations.
Adam is the preferred choice in most deep learning architectures, including transformer models. A common strategy is to start training with Adam and, when fine-tuning is needed, switch to custom choices.
> Q: So Adam combines momentum and adaptive learning — but does it also include the bias correction? > A: Yes. Adam includes bias correction for both the first moment (momentum) and the second moment (adaptive scaling). The first moment controls the gradient update; the second moment adapts the learning rate. Both get corrected using in the denominator during early iterations, which fades as training proceeds.
17.3.5 AdamW: Decoupled Weight Decay
Standard Adam mixes regularization into the gradient signal. The weight decay penalty gets absorbed into and , where the adaptive scaling distorts it. AdamW says: keep regularization separate. Let Adam handle the gradient. Apply weight decay as an independent shrink step afterward. This tiny architectural change turns out to matter a lot for generalization.
Standard Adam includes L2 regularization (weight decay), which adds to the loss. The gradient of this regularized loss is used in both the momentum and adaptive scaling calculations. But there is a conflict. The regularization term constrains the parameter values. Meanwhile, the adaptive mechanism uses those regularized gradients as if they were the true gradient. The original effect of the gradient is distorted.
AdamW fixes this by decoupling weight decay from the Adam update. After the normal Adam step computes the new parameters, you subtract the regularization component separately:
This decoupling leads to better generalization. AdamW is frequently used in GPT and other transformer architectures internally.
Comparison — Adam vs AdamW:| Property | Adam | AdamW |
|---|---|---|
| Weight decay | Added to loss; gradient feeds into | Subtracted after Adam update as separate step |
| Reg. strength | Distorted by adaptive scaling | Independent of gradient history |
| Generalization | Good | Better — decoupling prevents over-regularization of well-behaved parameters |
| When to use | General default | When weight decay is important (transformers, large models) |
17.3.6 Symbol Registry — Adam and Variants
- — parameter value at iteration — scalar or vector - — learning rate (step size) — scalar in - — gradient of the loss with respect to — same shape as - — momentum velocity; first-moment estimate — same shape as - (RMS) — exponentially weighted squared gradient; second-moment estimate — same shape as - — momentum decay coefficient — scalar in - — RMS decay coefficient — scalar in - — small constant for numerical stability — scalar, e.g. - — accumulated sum of squared gradients (AdaGrad) — same shape as - — weight decay coefficient (L2 regularization strength) — scalar - — bias-corrected moment estimates (Adam)17.4 Learning Rate Schedules
17.4.1 Why Schedules Matter
Even with adaptive methods like Adam, the base learning rate is still a fixed hyperparameter chosen by the designer. There is selection bias in picking it. Learning rate schedules replace a fixed with one that changes over time according to a plan, reducing the hyperparameter sensitivity.
Think of a learning rate schedule like a photographer adjusting focus. At first, you twist the lens quickly to get near the right focal length (large learning rate — exploration). As the image sharpens, you slow down and make micro-adjustments (small learning rate — exploitation). If you twist too fast near the end, you overshoot and lose the shot. The schedule is your hand's memory of when to slow down.
A learning rate schedule is a function that maps the training step to a learning rate. The schedule replaces the fixed in whichever optimizer you are using — SGD, Adam, RMSProp, etc. The optimizer still does its internal adaptation. The schedule controls the baseline rate from which that adaptation starts.
17.4.2 Step Decay
A step decay schedule drops the learning rate by a fixed factor at predetermined epochs. For example, train with for 10,000 iterations, then drop to 0.01 for the next 40,000, then to 0.001 after a few million. The large initial rate encourages exploration — the optimizer takes big jumps across the loss surface. The smaller later rates encourage exploitation — fine-grained, careful steps near the minimum.
This is common in reinforcement learning, especially for AI agents learning to play complex games like chess or Go. Go is extremely complex to formulate programmatically; researchers often use step decay as a simple, predictable schedule.
17.4.3 Exponential Decay
Exponential decay gives a smoother reduction:
The learning rate decays gracefully as training progresses, rather than dropping abruptly. At each step, the rate is multiplied by a constant factor , which is less than 1.
Exponential decay is suited for simple feedforward neural networks or when the loss surface is noisy. The smooth decay avoids the sharp transitions of step decay.
Synthetic linear regression: . True weights: (hidden from model).
Case 1 — Aggressive decay: - At : - At : - At : — near zero. Weights stop updating far from true values. Case 2 — Light decay: - At : - At : - At : — still learning. Weights approach true values. Sense-check: The interaction between learning rate and decay needs experimental tuning — there is no fixed formula. The aggressive decay kills learning too early. The light decay keeps the model improving but takes many iterations.17.4.4 Cosine Annealing
Cosine annealing decays the learning rate following a cosine curve: it starts high, drops rapidly, then slows down gracefully near zero. The formula uses the total number of epochs and two bounds — the maximum learning rate and the minimum :
Cosine annealing is frequently used in computer vision tasks. It is especially common in vision transformers. You use it when you have a fixed budget of epochs and a known range of learning rates.
17.4.5 Warm-Up Strategy
A warm-up is not a decay technique itself but a pre-decay trick. You start with a very small learning rate (near zero). Then you linearly increase it over the first few epochs — typically 5 to 10 epochs. The rate goes up until it reaches the intended initial learning rate. After warm-up, the actual learning rate schedule takes over.
Think of warm-up like letting a car engine idle for a minute on a cold morning before driving at highway speed. The oil needs to circulate. The pistons need to expand to their proper fit. In a neural network, the early gradients are the equivalent of a cold engine — noisy, unreliable, based on random weights. If you record those early, gibberish gradients into your momentum buffer, they pollute the optimizer's state for many steps afterward. Warming up gives the model time to produce sensible gradients before you start trusting them.
When training complex architectures from scratch — BERT, GPT, or large language models — the early gradients are essentially random. The model has not learned anything meaningful yet. Recording those gibberish gradients to drive momentum would pollute the optimizer state. By warming up, you let the model stabilize its initial weight trajectory before committing to serious training with momentum. After the warm-up epochs, the weights and gradients have settled into something sensible, and you begin the actual learning rate schedule.
| Schedule | Shape | Best For | Weakness |
|---|---|---|---|
| Step Decay | Staircase | RL, games (known phase transitions) | Requires knowing when to drop |
| Exponential Decay | Smooth, convex | Feedforward nets, noisy surfaces | Decay rate is sensitive |
| Cosine Annealing | S-curve (cosine) | Vision, fixed epoch budgets | Needs known upfront |
| Warm-Up | Ramp-up, then any schedule | Large models from scratch | Adds a hyperparameter (warm-up steps) |
Learning rate schedules replace the guesswork of a fixed learning rate with a planned trajectory. Large steps early for exploration. Small steps later for refinement. Warm-up protects the optimizer from early random gradients. The next section shifts from improving the learning process (optimization) to constraining it (regularization).
17.5 Regularization: Core Concepts
17.5.1 Overfitting and Underfitting
Think of overfitting like a student who memorizes every slide in the exact order and position they appear. When the exam rewords a question or changes the numbers, the student is lost. Underfitting is like a student who only read the chapter titles. They know the broad topic. But they cannot answer any specific question. A well-trained model is like a student who understands the concepts well enough to solve problems they have never seen before.
Overfitting happens when a model memorizes the training data — it learns every point, every noise pattern, every idiosyncrasy. The model performs well on training data but fails on test data. It has not learned the underlying pattern — it has learned the dataset.
Underfitting is the opposite: the model is too simple even for the training data. It captures only a rough average pattern and performs poorly on both training and test sets.
A good model lives between these two: it should specialize for the specific task but generalize well within the task domain.
17.5.2 Explicit and Implicit Regularization
A network with 10 layers of 100 neurons each has far too much capacity for 100 instances with 3 features. It will memorize. But with millions of instances, the same network is forced to abstract and generalize. More data is a form of implicit regularization.
17.5.3 Deciding When to Apply Regularization
Track the training loss and validation loss over epochs. If the training loss keeps improving while the validation loss starts diverging (getting worse), you have overfitting. At that point, ask:
If there is no overfitting (both curves track together), keep training. Adding unnecessary regularization to a model that is not overfitting only hurts performance.
Domain-specific considerations also matter:
- Computer vision: regularization must preserve spatial correlations — nearby pixels share patterns. DropBlock (not regular dropout) works for CNNs. - NLP: regularization must handle variable-length sequences and sequential dependencies. Attention dropout is preferred over neuron dropout for transformers. - Time series: must retain non-stationarity awareness. Regularization should not smooth away temporal shifts that are real signals. - Tabular data: features may interact in complex ways. You cannot assume independence — regularization that assumes uncorrelated features (like independent L1 penalties) may be suboptimal.> Q: The weights are learned by the model itself. In the Excel example, the true values 0.1 and 2 are hidden. How do we hyperparameter-tune then — don't we just keep trying different values? > A: Yes, hyperparameter tuning is inherently experimental — there is no rule book that says "3 layers, N nodes." It is an art, not a fixed framework. Anyone can randomly try values. But someone who understands what each hyperparameter actually does can tune intelligently. Know that momentum smooths oscillations. Know that RMSProp controls learning rate decay in non-stationary settings. Know that Adam is a good starting point. This knowledge lets you make informed choices rather than blind guesses, using minimal resources.
Overfitting means your model memorized instead of learned. Regularization — explicit (L1, L2) or implicit (more data, architecture, early stopping) — is how you force generalization. The diagnostic signal is always the same: watch the gap between training loss and validation loss. When it widens, reach for regularization. The following sections give you the specific tools.
17.6 Gradient Clipping
17.6.1 Vanishing and Exploding Gradients
When gradients are less than 1, the product shrinks toward zero during chain-multiplication across many layers. This is the vanishing gradient problem. When gradients are greater than 1, the product grows without bound — the exploding gradient problem.
Think of gradient flow through a deep network like a game of telephone. You whisper a number to the first person. Each person multiplies it by a factor before passing it on. After 100 people, if the factor is 0.9, the number shrinks to almost zero. No one at the end hears anything useful. That is vanishing. If the factor is 1.1, the number grows to over 13,000 — the last person hears a deafening roar (exploding).
The choice of activation function influences this:
- Sigmoid: outputs , derivative max is 0.25. Causes vanishing gradients. - Tanh: outputs , derivative max is 1.0. Less prone but still vanishes in deep nets. - ReLU: outputs 0 for negative inputs, passes positive inputs unchanged (derivative 1). Vanishing is mitigated on the positive side. But gradients can still explode with large positive weights.Gradient clipping is one direct solution: if the gradient exceeds a threshold, clip it to that threshold.
17.6.2 Value-Based Clipping
Set a minimum and maximum threshold (typically in the range 0.1 to 5). For every individual gradient value in the gradient vector:
- If it lies between the thresholds, leave it unchanged. - If it is outside, clip it to the nearest threshold.This clamps each element of the gradient independently. If one parameter's gradient is 20 and the threshold is 5, that element becomes 5. Other elements that are within bounds stay unchanged.
Value clipping is frequently applied in deep reinforcement learning. Environments like Atari games can produce huge reward spikes that create massive gradients for certain state-action pairs.
17.6.3 Norm-Based Clipping
Norm-based clipping scales the entire gradient vector so that its L2 norm does not exceed a threshold . This preserves the relative proportions between gradient elements — it just caps the total magnitude.
If , leave unchanged.
The scaling factor multiplies every element of , so all signs are preserved. A negative gradient stays negative — it is just scaled down in magnitude.
Given and :
Step 1: Compute L2 norm: Step 2: Check: , so clipping is needed. Step 3: Compute scaling factor: Step 4: Scale every element: Step 5 — Verification: Compute the norm of the clipped vector: Sense-check: The norm is exactly . The relative proportions are preserved — 7 was the largest element and 3.75 is still the largest. The sign of is preserved as . ✓This is conceptually similar to layer normalization — you scale the entire vector down rather than clamping individual elements. Norm-based clipping is most commonly observed in NLP, especially when training transformer models and LLMs for machine translation tasks.
> Q: Gradients can be both positive and negative. If we clip based on the L2 norm, which always gives a positive value, do we lose the sign information? > A: No. The L2 norm is used only to compute the scaling factor . This scaling factor is always positive as a ratio of two positive numbers. It multiplies every element of the original gradient vector. All signs are preserved. A negative gradient stays negative; it is merely scaled down in magnitude.
| Property | Value Clipping | Norm-Based Clipping |
|---|---|---|
| What is clipped | Each gradient element independently | Entire gradient vector together |
| Relative proportions | Can be destroyed | Preserved |
| Sign preservation | Yes | Yes |
| Primary use case | Deep RL (DQN, PPO) | NLP, Transformers, LLMs |
| Hyperparameter | Min/max per element | Single threshold |
| Robustness to scale | Sensitive — need per-layer tuning | Robust — one works across layers |
17.7 Weight Initialization
17.7.1 Why Initialization Matters
Starting weights at zero or at random positions can place you far from the global minimum. A poor initialization means the optimizer wastes many iterations just finding the right region of the search space. Smart initialization places the weights in a meaningful starting zone, so the learning converges faster and more stably.
Think of initialization like dropping a marble onto a bumpy landscape. If you drop it on a flat plain 10 kilometers from the deepest valley, it will take forever to roll there. If it ever does. If you drop it near the valley, it finds the bottom quickly. Initialization is choosing where to drop the marble. A good choice puts you in the right ballpark. A bad choice puts you in another country.
This is still an active area of research. There is no universally optimal initialization. But two practical schemes are widely used.
17.7.2 Xavier / Glorot Initialization
For a given layer, let be the number of incoming connections (fan-in) and be the number of outgoing connections (fan-out). The key insight: if the variance of weights on incoming connections equals the variance on outgoing connections, the network learns stably.
Xavier initialization draws weights from a normal distribution with mean zero and variance:
- — fan-in: number of input units to this layer - — fan-out: number of output units from this layer - The factor 2 comes from averaging the fan-in and fan-out variances - The distribution: Why this variance? If you initialize too large, the activations explode as they propagate forward. If too small, the gradients vanish as they propagate backward. This variance balances the forward signal variance with the backward gradient variance — keeping both roughly constant across layers.A fully connected layer with inputs and outputs.
Each weight is sampled from . About 68% of weights will fall in . About 95% in .
Sense-check: Larger layers (more fan-in/fan-out) get smaller initial weights. This prevents the weighted sum from growing with layer size. A layer with would get (much smaller). ✓Xavier initialization works well with linear activations, sigmoid, and tanh activations. It assumes the activation function is about linear around zero. Sigmoid and tanh satisfy this — their derivatives are ~1 at zero.
17.7.3 He Initialization
For ReLU activations, He initialization uses:
The distribution: .
Sample from and use those samples to initialize the weights whenever ReLU activations are used. Both Xavier and He initialization are forms of implicit regularization — they stabilize the early training dynamics without adding explicit penalty terms.
Comparison — Xavier vs He:| Property | Xavier (Glorot) | He |
|---|---|---|
| Variance | ||
| Uses fan-out? | Yes | No |
| Best activations | Linear, sigmoid, tanh | ReLU, Leaky ReLU, PReLU |
| Why the difference | Assumes symmetric activation ( linear near 0) | Compensates for ReLU zeroing half the signal |
| Default in frameworks | Used for tanh/sigmoid nets | Default for ReLU nets (most modern CNNs) |
17.8 Normalization Techniques
17.8.1 Covariate Shift
Consider an arbitrary layer in a deep network. The outputs of layer become the inputs to layer . The outputs of layer become the inputs to layer . If the distribution of these outputs is not centered around zero, every subsequent layer's inputs drift further away. Each layer compounds the shift. The variance grows with depth, making training unstable.
Covariate shift is the problem of distributions drifting across layers. The fix: force the output distribution of every layer to be centered around zero. Give it comparable variance. That way all layers operate on similarly-scaled values.Think of a relay race where each runner hands off a baton. Imagine runner 1 hands off at chest height. Runner 2 grabs it at knee height. Runner 3 grabs it above their head. The fourth runner has no idea where to grab. Normalization is like forcing every handoff to happen at exactly waist height — same position, same orientation, every time. Each runner can focus on running faster instead of adjusting their grip.
17.8.2 Batch Normalization
Batch normalization applies Z-score normalization to the outputs of a layer before passing them to the next layer's activation:
After normalization, two learnable parameters are applied:
- — scale parameter (learnable). Lets the network undo the normalization if zero-mean unit-variance is suboptimal. - — shift parameter (learnable). Lets the network shift the normalized distribution.These let the network learn the optimal scale and shift for the normalized values. The network is not locked to zero-mean unit-variance. If the best representation for the next layer needs mean 2 and variance 5, and can learn that.
During training, batch normalization uses the current batch's statistics ( and from the mini-batch). During inference, there is only one instance at a time — you do not have a batch. Instead, you maintain running statistics: after many training batches, you accumulate the mean and variance from all batches seen so far. For inference, you use these running averages instead of batch-specific statistics. This is why it is called batch normalization: the normalization happens across the batch dimension during training. Running statistics substitute during inference.
The correct placement of batch normalization is before the activation function. You normalize the layer's linear outputs, then apply the activation to the normalized values. This prevents the activation's squashing from masking distribution shifts.
Pre-activation values:
Step 1: Compute batch statistics: - - Step 2: Normalize (with ): - - - Step 3: Apply learned scale and shift (): - - - Sense-check: The normalized values are symmetric around 0 with comparable spread. The transformation then maps them to whatever scale the network needs. ✓Batch normalization speeds up training and stabilizes learning. It also acts as a regularizer. The batch statistics add a small amount of noise since each mini-batch has slightly different and . That noise has a mild regularizing effect.
> Q: If we apply a sigmoid activation function, the output is always between 0 and 1. Then why do we see a distribution shift between layers? > A: Batch normalization is applied before the activation, not after. The raw weighted sums from a layer can have any distribution — they are not constrained to . If you normalize those pre-activation values first and then apply the sigmoid, you avoid the distribution drift. The correct order is: layer output → batch normalization → activation function → next layer.
17.8.3 Layer Normalization
Layer normalization normalizes across the features within a single layer, rather than across the batch. For a layer with features:
| Property | Batch Normalization | Layer Normalization |
|---|---|---|
| Normalization axis | Across batch, per feature | Across features, per instance |
| Depends on batch size? | Yes — needs large enough batch | No — works with batch size 1 |
| Depends on sequence length? | No | No — normalizes across features, not sequence |
| Inference behavior | Uses running statistics | Same computation as training |
| Primary use | CNNs (ResNet, VGG) | Transformers, RNNs, LSTMs |
| Why? | CNNs have large batch sizes, spatial weight sharing | Transformers have variable sequence lengths; batch norm fails with padding |
Layer normalization is the standard choice in transformer architectures and RNNs. It works with variable sequence lengths and does not depend on batch size.
| Architecture | Normalization Used | Position |
|---|---|---|
| ResNet, VGG (CNNs) | Batch Norm | Conv → BatchNorm → ReLU |
| Transformer (BERT, GPT) | Layer Norm | Attention/FFN → LayerNorm (pre or post, depending on variant) |
| RNN, LSTM | Layer Norm | Hidden state → LayerNorm → activation |
| Vision Transformer (ViT) | Layer Norm | Patch embedding → LayerNorm → Attention |
17.9 Early Stopping
17.9.1 The Patience Parameter
Think of early stopping like a chef tasting a simmering sauce. The flavor improves with time, but only up to a point. After that, the sauce reduces too much, burns, or becomes too salty. The chef does not wait until the smoke alarm goes off. They taste regularly and pull the pan off the heat when it peaks. The patience parameter counts how many tastings the chef waits. After the last improvement, they declare "this is as good as it gets."
Early stopping monitors the gap between training loss and validation loss. When the validation loss starts diverging from the training loss, overfitting has likely begun. The patience parameter tells the algorithm how many epochs to wait before making the stop decision.
If patience is 3 and the validation loss has been worsening for 3 consecutive epochs, training stops. The algorithm backtracks to the best model — the one from before the divergence began.
Why use patience rather than stopping instantly on the first divergence? A single epoch's divergence might be noise — a bad mini-batch with outliers, not a genuine overfitting trend. The next epoch might correct itself and the curves could re-converge. Patience gives the optimizer a chance to recover from temporary blips before making the irreversible stop decision.
17.9.2 Implementation
a. Train one epoch. Compute training loss and validation loss. b. If validation loss < best_val_loss:
- Update `best_val_loss = current_val_loss` - Save current model as best model - Reset `patience_counter = 0`c. Else (validation loss did not improve):
- `patience_counter += 1` - If `patience_counter > patience_threshold`: - Stop training - Restore best model - Return best model 3. If training completes all epochs without triggering, return best model.Early stopping is a simple, effective regularizer. It costs almost nothing — just a patience counter and checkpoint saving. It works with any model, any optimizer, any task. It belongs in every training loop. The next section introduces a completely different approach to regularization: randomly disabling parts of the network during training.
17.10 Dropout and Variants
17.10.1 Standard Dropout
Think of dropout like cross-training in sports. A basketball player who only practices free throws with the same routine, same lighting, and same crowd noise will crumble in a real game. Another player practices with random distractions — different ball weights, noisy speakers, one eye covered. That player builds strong skills that work in any condition. Dropout is cross-training for neurons. By randomly disabling teammates, each neuron learns to be useful on its own. It is not just part of a specific combination.
When you suspect a network is too complex and overlearning, you set a dropout rate (e.g., 0.5). At each training iteration, 50% of the neurons in the dropout layer are randomly deactivated. The neurons are not removed — their physical connections remain. The weights learned up to that point are preserved. But for that specific iteration, those neurons do not participate in the forward pass or receive gradient updates.
At the next iteration, a different random subset of neurons is dropped. Over the course of training, every neuron gets dropped some of the time. And every neuron gets to learn some of the time. All neurons ultimately learn.
> Q: Why not just design a network with fewer neurons in the first place, instead of using dropout? > A: If you build a smaller network from the start, you lose the representational capacity. A network with 4 nodes can express more patterns than a network with 2 nodes. Dropout lets you keep the large architecture's capacity while preventing overfitting. During training, subsets learn independently. During inference, the full capacity is available. You get the regularization benefit of a small network with the expressive power of a large one.
17.10.2 DropConnect
Dropout deactivates entire neurons. DropConnect is finer-grained: it randomly masks individual weight connections in the weight matrix between layers. At a dropout rate of 50%, half the individual weights are frozen for that iteration. They keep their previous values but do not receive gradient updates. The other half learn normally. In the next iteration, a different random set of 50% of weights are frozen.
This lets you regularize at the synapse level rather than the neuron level. DropConnect is useful for fully connected neural networks.
17.10.3 DropBlock
Dropout and DropConnect work for fully connected layers but fail for convolutional layers. In a CNN, nearby pixels carry spatial correlations. Dropping individual neurons in a feature map does little — adjacent neurons still pass nearly the same spatial information. The network barely notices a dropped neuron. The regularization effect is too weak.
DropBlock addresses this by masking contiguous blocks of neurons in the feature map. By dropping entire spatial regions at once, you force the network to learn from other parts of the image. It cannot rely on fine spatial detail in the dropped region. This works well for vision tasks with strong spatial correlation.
17.10.4 Attention Dropout
In transformer models, after computing the softmax attention weights, attention dropout randomly drops some of those attention weights. Suppose word attends to word and word via self-attention. If the attention weight from to is dropped, the model cannot rely on that token. It must use information from the remaining tokens for context at that iteration.
This prevents the model from over-relying on specific tokens. If a token is noisy or unavailable at inference time, the model can still extract meaning from the remaining tokens. Attention dropout is applied after the softmax in the attention mechanism.
Comparison — Dropout Variants:| Variant | What is Dropped | Architecture | Reason |
|---|---|---|---|
| Standard Dropout | Entire neurons | Fully connected layers | Prevent neuron co-adaptation |
| DropConnect | Individual weights (synapses) | Fully connected layers | Finer-grained regularization |
| DropBlock | Contiguous blocks | CNNs | Overcome spatial correlation |
| Attention Dropout | Individual attention weights | Transformers | Prevent over-reliance on specific tokens |
17.11 Data Augmentation
17.11.1 Basic Image Augmentation
Think of data augmentation like a forger creating training samples for a detective. The forger takes an original document and produces slight variations — different handwriting angle, varied ink pressure, slightly different paper color. The detective must learn to recognize the underlying signature regardless of these surface changes. The forger is not creating new information — just teaching the detective which variations do not matter.
When the dataset is small but the network is complex, you can artificially increase the training data. For images, you apply transformations: rotate, scale, zoom, change brightness and contrast, add noise, or perform cutout. Each transformed version becomes a new training instance with the same label. One image might generate 8 augmented copies. None of these transformations change the class — a rotated cat is still a cat.
Another technique adds small Gaussian noise to images. This is similar to the idea behind denoising autoencoders. Train on noisy images so the model learns strong features that survive mild corruption.
Common image augmentations:
- Geometric: rotation (±15°), horizontal/vertical flip, scaling (zoom in/out), translation (shift), shearing - Photometric: brightness adjustment, contrast change, color jitter (hue/saturation), gamma correction - Noise-based: Gaussian noise, salt-and-pepper noise, Gaussian blur - Occlusion: random erasing / cutout (mask a random rectangular region with zeros or noise)All preserve the class label — the network must learn invariance to these transformations.
17.11.2 MixUp
MixUp blends two training images and their labels. Given image (a cat) and image (a dog), and a mixing coefficient :
17.11.3 CutMix
CutMix is similar, but instead of blending pixels, it cuts a rectangular patch from one image and pastes it onto the other. The label is weighted by the area ratio of the patch.
This is like occlusion — the cat might be partially hidden behind a cushion, with only its tail visible. The model must learn to recognize objects from partial views.
| Property | MixUp | CutMix |
|---|---|---|
| How images combine | Pixel-wise weighted average (ghostly blend) | Rectangular patch replacement (sharp boundaries) |
| Visual result | Both images are faintly visible everywhere | One region is one image, the rest is the other |
| Label mixing | from Beta distribution | = patch area ratio |
| Real-world analogy | Looking through a semi-transparent overlay | Object partially hidden behind another |
| Primary benefit | Calibrated, non-overconfident predictions | Robustness to occlusion and partial views |
17.11.4 Progressive Augmentation
A practical training strategy: start with no data augmentation. Train until performance plateaus. Then introduce simple augmentations (flipping, rotation, scaling). If accuracy improves, continue. Then add harder augmentations (color jitter, cutout). If that helps further, try MixUp and CutMix. This progressive approach strengthens the model incrementally. Applying all augmentations at once can make training too difficult for the model early on. At that point, it has not yet learned the basic features.
Exam Guidance Summary
This lecture covers two major modules. Optimization techniques show how to make models learn better and faster. Regularization techniques show how to prevent memorization and improve generalization. Both are foundational for training deep networks.
Optimization: What to Know
- Conceptual comparisons: Be able to compare SGD, SGD with momentum, Nesterov (NAG), AdaGrad, RMSProp, Adam, and AdamW. For each, know what problem it solves (noisy gradients, fixed learning rate, sparse data, weight decay coupling). - Which optimizer for which scenario: Sparse data → AdaGrad. Non-stationary/sequential/RNN → RMSProp. General/transformers → Adam. Transformer training with weight decay → AdamW. - Adam bias correction: Know why it exists (cold start with biases early estimates toward zero), how it works (), and why it fades with . - Learning rate schedules: Step decay vs exponential decay vs cosine annealing vs warm-up. Know which is used where (RL → step decay; vision → cosine annealing; large models from scratch → warm-up).Regularization: What to Know
- Gradient clipping: Value clipping (clamps per element, used in RL) vs norm-based clipping (scales whole vector, used in NLP/transformers). Understand why norm-based preserves proportions and how is the only hyperparameter. - Weight initialization: Xavier/Glorot (, for sigmoid/tanh) vs He (, for ReLU). Know why He uses only fan-in — ReLU kills half the signals. - Batch norm vs layer norm: What is normalized across in each (batch dimension vs feature dimension). Where each is applied (CNNs → batch norm; transformers → layer norm). How inference-time statistics work for batch norm (running mean/variance). Correct placement: before activation. - Dropout and variants: Standard dropout (neuron-level, fully connected). DropConnect (weight-level, fully connected). DropBlock (block-level, CNNs — overcomes spatial correlation). Attention dropout (weight-level, transformers — prevents token over-reliance). Know why different variants exist for different architectures. - Data augmentation: Basic image transforms, MixUp (pixel-wise blend with soft labels), CutMix (patch replacement with area-weighted labels). Rationale for progressive augmentation — strengthen the model incrementally.Recurring Themes
- Optimization vs regularization: improve the learning process vs prevent overfitting. These are complementary, not competing, goals. - The diagnostic signal for overfitting is always the train/validation loss gap. When validation loss diverges from training loss, reach for regularization. - Hyperparameter tuning is experimental, not formulaic. Understanding what each technique does enables intelligent tuning rather than blind search.Key Industry Applications
Optimization in Practice
- Adam is the default optimizer in PyTorch, TensorFlow, and JAX. It trains virtually every modern transformer model (GPT, BERT, T5, LLaMA). Its combination of momentum and adaptive learning rates provides strong convergence across diverse architectures with minimal tuning. - AdamW (decoupled weight decay) is used internally in GPT and other large language models. The decoupling prevents the adaptive scaling from distorting regularization strength, leading to better generalization at scale. - Norm-based gradient clipping is standard when training transformer models and LLMs for machine translation tasks. Without it, the chain multiplication of gradients through attention layers makes exploding gradients nearly certain. - Value-based gradient clipping is common in deep reinforcement learning algorithms (DQN, PPO). Atari games and robotics environments produce sparse, high-magnitude reward signals that create massive gradient spikes. - Step decay learning rate schedules are widely used in reinforcement learning for games like chess and Go. Domain experts set the drop epochs based on known phase transitions in the learning process. - Cosine annealing is popular in computer vision tasks, especially vision transformers (ViT). It provides smooth, predictable learning rate decay over a fixed epoch budget. - Warm-up strategies are essential when training large language models (BERT, GPT) from scratch. Starting with full learning rates on random weights produces unstable training; warm-up lets the optimizer accumulate reliable gradient statistics first.Regularization in Practice
- Batch normalization is a standard component in CNN architectures (ResNet, VGG, EfficientNet). It enables training networks deeper than ~10 layers and speeds up convergence by 10-50x compared to unnormalized networks. - Layer normalization is the norm in transformers (BERT, GPT, T5). It handles variable sequence lengths and works with small batch sizes — critical for NLP workloads where sequence padding would corrupt batch statistics. - MixUp and CutMix are modern data augmentation techniques used in state-of-the-art image classification (training Vision Transformers, EfficientNet, ConvNeXt). They improve generalization and calibration without requiring more data. - Attention dropout is used in transformer training to prevent the model from over-relying on specific tokens. This is critical for machine translation, where the model must handle missing or noisy words at inference time. - AdaGrad excels with sparse features common in NLP tasks like sentiment analysis with bag-of-words features. In these tasks, most vocabulary terms are absent from any given document. - RMSProp is effective for RNN/LSTM training on sequential, non-stationary data such as time series forecasting and speech recognition. In these domains, gradient statistics shift over time. - Progressive augmentation strategies are used in production vision systems. They start simple (flip, rotation) and add complexity (color jitter, cutout, MixUp, CutMix) as the model improves. This approach is standard in self-driving car perception and medical imaging pipelines.DNN Lecture 17 notes · Optimization and Regularization
Sections Breakdown
Section covering 17.1 Optimization vs Regularization
Section covering 17.2 Momentum-Based Gradient Updates
Section covering 17.3 Adaptive Learning Rate Methods
Section covering 17.4 Learning Rate Schedules
Section covering 17.5 Regularization: Core Concepts
Section covering 17.6 Gradient Clipping
Section covering 17.7 Weight Initialization
Section covering 17.8 Normalization Techniques
Section covering 17.9 Early Stopping
Section covering 17.10 Dropout and Variants
Section covering 17.11 Data Augmentation
Section covering Exam Guidance Summary
Section covering Key Industry Applications
Exam Revision Notes
Below is the distilled, exam-ready core of this lecture. Every entry is built from the full textbook notes above. Use this section for rapid review — but if something doesn't make sense, go back to the full explanation in the main content.
Momentum-Based Optimization
Must-know: Momentum accumulates a velocity of past gradients to smooth updates. Nesterov looks ahead for faster correction.
⚠️ Top pitfall: Setting β too high (0.999) — velocity decays too slowly and carries stale gradient information.
Self-check: What problem does momentum solve in SGD?
Connects to: Plain SGD (no momentum), Nesterov Accelerated Gradient, Adam (uses momentum internally)
Adaptive Learning Rate Methods
Must-know: AdaGrad accumulates squared gradients (aggressive decay). RMSProp uses a moving average. Adam combines momentum + adaptive scaling with bias correction.
⚠️ Top pitfall: Using Adam with default β₂=0.999 on rapidly changing gradient statistics (e.g., GANs) — the second moment adapts too slowly.
Self-check: Why does Adam need bias correction?
Connects to: AdaGrad, RMSProp, AdamW
Learning Rate Schedules
Must-know: Step decay drops the rate at fixed epochs. Exponential decay applies smooth decay. Cosine annealing follows a cosine curve. Warm-up increases from near zero before the main schedule.
⚠️ Top pitfall: Skipping warm-up when training transformers from scratch — early random gradients pollute the optimizer's momentum buffer.
Self-check: Which schedule is best for training Vision Transformers?
Connects to: Step Decay, Exponential Decay, Cosine Annealing, Warm-Up
Overfitting and Regularization
Must-know: Overfitting = low training error, high test error (high variance). Underfitting = high error on both (high bias). Explicit regularization (L1, L2) adds penalties; implicit regularization (more data, architecture, early stopping) emerges from training setup.
⚠️ Top pitfall: Adding regularization before confirming the model can even fit the training data — first check if underfitting or overfitting is the problem.
Self-check: What is the diagnostic signal for overfitting?
Connects to: L1 Regularization, L2 Regularization, Early Stopping, Dropout
Gradient Clipping
Must-know: Value clipping clamps each gradient element independently (used in RL). Norm-based clipping scales the whole gradient vector to a threshold τ, preserving relative proportions (standard in NLP/transformers).
⚠️ Top pitfall: Using value clipping for transformers — the embedding layer and output layer have vastly different gradient scales, and value clipping treats them equally.
Self-check: Why does norm-based clipping preserve relative proportions between gradients?
Connects to: Exploding Gradients, Value Clipping, Norm-Based Clipping
Weight Initialization
Must-know: Xavier/Glorot: σ² = 2/(n_in + n_out), for sigmoid/tanh. He: σ² = 2/n_in, for ReLU. He uses only fan-in because ReLU zeros out half the incoming signals.
⚠️ Top pitfall: Using Xavier with ReLU — the variance is too small, causing vanishing gradients in deep networks.
Self-check: Why does He initialization use only fan-in and not fan-out?
Connects to: Symmetry Problem, ReLU Activation, Batch Normalization
Normalization Techniques
Must-know: Batch norm normalizes across the batch dimension per feature (used in CNNs). Layer norm normalizes across the feature dimension per instance (used in transformers). Both use learnable scale (γ) and shift (β) parameters.
⚠️ Top pitfall: Using batch norm with very small batch sizes (e.g., 2-4) — the batch statistics become noisy and unreliable.
Self-check: Why does layer normalization work better than batch normalization for transformers?
Connects to: Covariate Shift, Batch Normalization, Layer Normalization
Early Stopping
Must-know: Monitors validation loss and halts training when it diverges from training loss for a set number of epochs (patience). Restores the best checkpoint before divergence.
⚠️ Top pitfall: Setting patience too low (e.g., 1) — you stop on noise instead of genuine overfitting.
Self-check: Why do we need a patience parameter instead of stopping on the first divergence?
Connects to: Overfitting, Validation Loss, Model Checkpointing
Dropout and Variants
Must-know: Standard dropout drops neurons (fully connected). DropConnect drops individual weights. DropBlock drops contiguous blocks (CNNs). Attention dropout drops attention weights (transformers).
⚠️ Top pitfall: Using standard dropout in CNNs — spatial redundancy makes neuron-level dropout ineffective. Use DropBlock instead.
Self-check: Why does DropBlock work better than standard dropout for convolutional layers?
Connects to: Standard Dropout, DropConnect, DropBlock, Attention Dropout
Data Augmentation
Must-know: Label-preserving transformations artificially expand the training set. MixUp blends two images and their labels. CutMix replaces a rectangular patch with another image. Progressive augmentation increases difficulty gradually.
⚠️ Top pitfall: Applying augmentations that change the class label — rotating a '6' by 180° makes it a '9'.
Self-check: What is the advantage of progressive augmentation over applying all augmentations at once?
Connects to: Basic Image Augmentation, MixUp, CutMix, Progressive Augmentation
Was this lecture useful?
BitsNotes AI Assistant
Subject Notes AssistantConfigure AI Chat
Choose how to access the chatbotSigned in as
Powered by BitsNotes — 20 messages per day. No API key needed. Want unlimited access? Use "Bring Your Own Key" mode.
Sign in to use AI Chat
Get 20 free AI messages per day to ask questions about your lecture notes. Sign in with Google or GitHub — it takes 5 seconds.
Sign In to BitsNotesSwitch to "Bring Your Own Key" tab above for unlimited access with any OpenAI-compatible provider.