Taylor Series and Hessian-Based Optimization
Taylor Series and Hessian-Based Optimization
10.1 Rolle's Theorem
10.1.1 Definition and Intuition
The "Flat Spot" Guarantee. You start a hike at 1000m elevation and end at 1000m. Your trail can go up, then down. It can dip into a valley, then climb back. But one thing is certain: at some point along the way, you stood on perfectly flat ground. Your slope was zero. Rolle's theorem is the mathematical statement of that guarantee.
There are two fundamental theorems that build up Taylor series. The first is Rolle's theorem. It is the starting point.
The Thrown-Ball Analogy. Throw a ball straight up. It rises, reaches a highest point
And falls back. At that highest point, for one instant, the ball's rate of change is zero. It stopped going up. It has not yet started coming down. The slope at that instant is exactly zero. The function's path has a guaranteed "flat spot" somewhere between the start and the end.
Where the analogy breaks: Rolle's theorem only guarantees at least one flat spot —
The ball might have bounced, creating multiple. Also, the theorem covers minima too (a valley), not just maxima (the peak).
Rolle's theorem captures this idea for continuous functions. It has three conditions:
- The function is continuous on the closed interval . No breaks, no gaps, no jumps anywhere between and , endpoints included.
- The function is differentiable on the open interval . The curve must be smooth — no sharp corners, no sudden kinks — at every point between and .
- The function values at the endpoints are equal: .
Rolle's Theorem (Formal Statement). Let be a function where:
- is continuous on the closed interval
- is differentiable on the open interval
Then there exists at least one point such that:
The derivative at is zero. At that point, the tangent line is horizontal —
You have either a local maximum or a local minimum.
10.1.2 Why the Conditions Are Chosen That Way
Continuity on the closed interval —
The function must not break anywhere from start to finish. If there is a gap, you could jump past the critical point and never hit zero slope.
Differentiability on the open interval —
Differentiability means the curve is smooth. No sharp edges. But the condition uses the open interval, not the closed one. The endpoints and are excluded.
The reason is practical. You do not know whether the function has a sharp corner right at the starting point. Imagine a curve that is smooth between and
But at exactly it has a kink. The derivative at that edge is not well defined. To stay safe, you only require smoothness strictly between the endpoints.
Equal endpoint values —
Starting and ending at the same height guarantees that, during the journey, you either went up and came back down (a maximum) or went down and came back up (a minimum). Either way, the slope had to hit zero somewhere.
10.1.3 Worked Example
Verify Rolle's theorem for on .
Step 1 — Check conditions:
- is a polynomial, so it is continuous everywhere. True on .
- is a polynomial, so it is differentiable everywhere. True on .
- . All three conditions hold.
Step 2 — Find :
Set :
Step 3 — Verify is in : lies between 1 and 3.
Conclusion: Rolle's theorem is verified. The flat spot occurs at
Which is the vertex of the parabola —
A global minimum with value .
Sense-check: The parabola opens upward, starts at , dips to
And returns to . The lowest point must have a horizontal tangent. ✓
10.1.4 Assumptions & Scope
Scope: When Rolle's theorem applies and when it fails.
- Breaks on a gap: If has a jump discontinuity on , the theorem fails. A function can skip past a flat spot by teleporting through a gap.
- Breaks on a corner: If has a sharp kink (e.g., at ), the derivative does not exist at the corner. The curve changes direction instantly, but there is no point where the tangent is truly horizontal — the derivative simply does not exist at the kink.
- Breaks on a constant function: A constant function like on any interval satisfies Rolle's theorem — but every point has . The theorem says "at least one," which still holds.
- Breaks without equal endpoints: If , the theorem offers no guarantee at all. You could go from height 0 to height 10 without ever having zero slope.
10.1.5 Visual Intuition
Picture a smooth curve drawn on graph paper. The x-axis runs from to . The y-axis shows . The curve touches the same horizontal line at and . Now imagine sliding a horizontal ruler up and down until it barely kisses the curve from above or below. That point of contact is your —
Where the tangent is horizontal. If the curve goes above its starting level and comes back down, the ruler kisses the top (maximum). If it dips below and rises back, the ruler kisses the bottom (minimum). The one-sentence takeaway: equal heights at the ends force at least one horizontal tangent in between.
10.1.6 Pitfalls
- Forgetting the endpoints equality check. Students often jump straight to solving without verifying . If the endpoints are unequal, no is guaranteed — and you might waste time searching for one that does not exist.
- Confusing "continuous on " with "continuous on ". Rolle's requires continuity at the endpoints too. A removable hole at breaks the theorem even if the rest of the curve is perfect.
- Thinking Rolle's guarantees exactly one . The theorem only guarantees at least one. A wavy curve like on has and produces three points where .
- Assuming differentiability at the endpoints. The theorem explicitly avoids this by using the open interval . Do not try to compute or — they might not exist.
10.1.7 Student Questions and Answers
Q: Why do we require differentiability on an open interval but continuity on a closed interval? Could the starting point be a sharp edge?
A: Yes. At the starting point you do not know if the gradient changes abruptly. The curve might be smooth throughout but have a corner exactly at . Since you cannot guarantee the derivative exists at the edge, you only enforce differentiability inside.
10.1.8 Recap + Bridge
Rolle's theorem guarantees a horizontal tangent somewhere between equal-height endpoints. It is the building block —
The first two conditions of Rolle's (continuity on and differentiability on ) carry over exactly to the Mean Value Theorem. Only the third condition (equal endpoints) gets dropped in the next step.
Rolle's theorem is the seed from which the Taylor series grows. The full proof of Taylor's theorem cleverly constructs a special function that meets Rolle's three conditions, guaranteeing the existence of a point that defines the remainder term. This is why we start here.
Real-world connection: Civil engineers use the logic of Rolle's theorem when analyzing beam deflections. A beam supported at both ends with zero deflection at the supports must have at least one point of zero slope along its curve —
Which tells them where the beam bends the most. Any physical system that returns to its starting state (a pendulum swinging through one full period, a piston completing a cycle) must pass through at least one instant of zero rate of change.
---
10.2 Mean Value Theorem
10.2.1 Definition and Intuition
How can your instantaneous speed ever equal your average speed? You drive 200 km in 2 hours. Average speed: 100 km/h. You probably spent some time at 80 km/h, some at 120 km/h. But the MVT guarantees: at least once during the trip, your speedometer read exactly 100 km/h. The instantaneous matched the average.
The mean value theorem (MVT) inherits the first two conditions of Rolle's theorem and drops the third. It no longer requires . You can start at one height and end at another.
The tilted Rolle's. Rolle's theorem requires starting and ending at the same height —
Like a hiker returning to the same elevation. MVT is the tilted version. Now the trail starts at one elevation and ends at a different one. The guarantee shifts: instead of "somewhere the slope is zero," it becomes "somewhere the slope exactly matches the average slope of the whole trail." MVT tilts the horizontal tangent of Rolle's into a tangent parallel to the secant line.
The professor's highway driving analogy: you drive 200 km in 2 hours. Average = 100 km/h. The MVT says your speedometer must have hit exactly 100 km/h at some moment. Where the analogy breaks: the MVT assumes a smooth speed curve (no instantaneous jumps)
Which real driving has (braking, traffic lights)
But the core insight stands.
MVT says: if is continuous on and differentiable on , then there exists at least one point in where:
Mean Value Theorem (Formal Statement). Let be:
- continuous on the closed interval
- differentiable on the open interval
Then there exists at least one point such that:
The left side is the instantaneous rate of change at —
The slope of the tangent line at that one point. The right side is the average rate of change across the whole interval —
The slope of the secant line from to . MVT says these two slopes are equal at some interior point.
10.2.2 Comparison with Rolle's Theorem
Rolle's theorem is a special case of MVT. When , the right side becomes , and MVT reduces to Rolle's: .
| Aspect | Rolle's Theorem | Mean Value Theorem |
|---|---|---|
| Endpoint condition | required | Any allowed |
| Slope guaranteed | Zero (horizontal tangent) | Average slope (tangent ∥ secant) |
| Geometrically | Flat spot | Tangent parallel to secant line |
| Role in the story | Seed theorem — the starting point | Bridge to Taylor series |
| When you'd pick it | Need to prove existence of a critical point when endpoints match | Need to relate instantaneous change to overall average change |
10.2.3 Geometric Interpretation
Draw a smooth curve from to . The secant line connects the two endpoints —
That is the average rate of change. Now slide a tangent line along the curve. At some point , the tangent becomes exactly parallel to the secant. That point is where equals the average rate.
Real-world: functions like exponentials grow smoothly and always satisfy this. Pick any two points on . Somewhere between them, the instantaneous growth rate equals the average growth rate.
10.2.4 Worked Example
Find a point that satisfies the MVT for on .
Step 1 — Check conditions:
- is a polynomial: continuous everywhere on (including endpoints). ✓
- Differentiable everywhere on . ✓
Step 2 — Compute the average slope:
Step 3 — Find where :
Step 4 — Verify is in : lies between 1 and 3. ✓
Conclusion: At , the instantaneous slope equals the average slope over . The point is .
Sense-check: The curve steepens as grows. Near , the slope is (below average). Near , the slope is (above average). Somewhere in between, it must hit exactly 13. ✓
10.2.5 From MVT to Taylor Series — The Bridge
Rearrange the MVT equation:
You know the value at . You want the value at . MVT says: start with , then add the rate of change at some intermediate point , scaled by how far is from .
Now look at the first two terms of the Taylor series:
The only difference: Taylor uses the derivative at itself, not at some unknown in between. Taylor tweaks MVT by anchoring the slope at the starting point rather than searching for the magical point .
If is close to , this single-slope approximation works well. The slope at is a good proxy for the average slope of the short interval. But when is far from , the slope at no longer captures all the bending and turning that happens in between. You need more terms.
10.2.6 Assumptions & Scope
Scope: When MVT applies and when it fails.
- Both conditions are required. Losing continuity on or differentiability on kills the guarantee. The function on fails differentiability at and MVT does not apply there.
- The point is not the midpoint. Don't assume . It is wherever the tangent happens to match the secant — which could be near , near , or anywhere in between. The symmetic function on has (the midpoint), but on puts , closer to .
- MVT is an existence theorem, not a construction method. It tells you a exists but does not tell you how to find it directly. You still have to solve .
- Works for any real-valued function on a real interval. The theorem does not generalize cleanly to vector-valued functions — the vector MVT requires inequalities rather than equalities.
10.2.7 Visual Intuition
Plot your function on . Draw the secant line connecting to —
That is your reference slope. Now imagine a ruler that starts with the same slope as the secant and slides along the curve, always staying parallel to the secant. MVT guarantees there is at least one point where the ruler kisses the curve as a tangent line. That point of contact is . For on , the secant rises steeply from to . The tangent at runs parallel to it —
Touching the curve at a single point while matching the secant's slope. One-sentence takeaway: somewhere between and , the instantaneous slope matches the overall average slope.
10.2.8 Pitfalls
- Thinking is always the midpoint. It is not. The location depends on how the curve bends. For highly asymmetric functions, can be arbitrarily close to one endpoint.
- Applying MVT to vector-valued functions directly. The standard MVT applies to functions . For curves in , you need a different form involving norms and inequalities.
- Forgetting to check continuity at the endpoints. MVT requires continuity on the closed interval. A removable discontinuity at — even a single missing point — breaks the guarantee.
- Confusing MVT with the Intermediate Value Theorem. IVT guarantees a function takes every value between and . MVT guarantees a derivative value equals the average slope. These are different theorems with different conclusions.
- **Assuming the tangent is unique.** MVT guarantees at least one . A function that oscillates may have many parallel tangents.
10.2.9 Recap + Bridge
MVT guarantees that the instantaneous slope equals the average slope somewhere inside any smooth interval. Rearranging MVT gives —
A formula that directly prefigures the Taylor series. The Taylor series simply swaps the unknown midpoint slope for the known starting-point slope , then adds higher-order corrections for curvature.
The connection from MVT to Taylor is the conceptual bridge of this entire lecture. MVT says "there exists a giving the exact correction." Taylor says "instead of hunting for , I'll use as a proxy and add more terms to fix the error." This is the shift from an existence theorem (guaranteeing some point works) to a constructive approximation (building a polynomial that gets arbitrarily close).
Real-world connection: The Mean Value Theorem underpins error bounds in numerical methods. When a physics simulation approximates a continuous trajectory with discrete time steps, MVT provides the theoretical guarantee that the real velocity at some moment matches the computed average. Speed cameras on highways rely on the same logic: measure your time between two points, compute average speed
And —
By MVT —
You must have hit that speed at some instant, making the ticket mathematically sound.
---
10.3 Taylor Series — The Full Intuition
10.3.1 Position, Slope, Curvature, and Beyond
How can you clone a function using only information from a single point? Imagine you are a master forger trying to replicate a complex signature. You cannot see the whole signature at once. But you can analyze it deeply at its starting point —
The exact position, the angle of the pen, how the curve is bending, how the bend itself is changing. The Taylor series is what lets you build a polynomial imitation that gets closer and closer to the real thing with every extra detail you match.
A Taylor series approximation builds up in layers. Each new term adds a geometric feature to the polynomial clone:
- Degree 0 — Position: . Where the function sits. The forger places a dot at the starting point.
- Degree 1 — Slope: . How fast the function is changing at , times the step you take. The forger draws a straight line matching the starting angle.
- Degree 2 — Curvature: . How the slope itself is bending. The forger bends the line into a parabola matching the initial curve.
- Degree 3+ — Higher-order curvatures: finer and finer bends. The forger matches how the curvature changes, then how that change changes, and so on.
Each added term gives the approximation a new way to bend. First you place the point. Then you tilt a straight line. Then you bend that line into a parabola. Then you add more nuanced curves on top. The Taylor series is the ultimate forger —
Using all the derivative information at a single point to build a perfect polynomial clone in that neighborhood.
Where the master forger analogy breaks. A real forger eventually runs out of steady-hand precision. The Taylor series can keep adding terms forever (the infinite Taylor series) to get an arbitrarily close match, as long as the function is infinitely differentiable. Also, the clone is only perfect near the starting point —
Walk too far away and the approximation diverges.
10.3.2 The Full 1D Formula
Taylor Series (1D). For a function that is times differentiable at a point , the Taylor polynomial of degree about is:
Expanded:
Each piece:
- : the -th derivative evaluated at . The 0-th derivative is the function itself.
- : factorial. , , , , , and so on.
- : the step from , raised to the -th power. The farther you step, the larger this term.
- : the coefficient for the -th degree term — the pure -th derivative value, corrected for the factorial accumulation.
When , the series is called a Maclaurin series —
A special name for a Taylor series centered at zero. The formula simplifies because for all terms.
10.3.3 Where the Factorials Come From
Taylor series approximates a non-polynomial function (like , , ) with a polynomial. A non-polynomial function has a Taylor expansion as a polynomial series.
Set up a generic polynomial:
You want the coefficients so that matches and all its derivatives at .
Evaluate at :
- , so .
- . At : , so .
- . At : , so .
- . At : , so .
The factorial appears because repeated differentiation pulls down the exponent as a multiplicative factor. Each time you differentiate , you multiply by , then , then
And so on. The accumulated product is exactly . Dividing by cancels that buildup, leaving the pure derivative value as the coefficient.
10.3.4 Why Higher-Order Terms Matter
The first derivative tells you how the function changes right at . That is a local property. If is close to , the local slope is enough —
The function has not had room to bend much.
If is far from , the local slope at does not know about all the curving that happened along the way. The second derivative captures how the slope itself changes —
It adds curvature. The third derivative captures how the curvature changes. Each higher order gives the approximation more room to follow the real function's bends.
10.3.5 Worked Examples from the Companion Document
Example 1: Taylor series of about .
The function is special because its derivative is always itself. All derivatives at equal 1.
Plug into the Taylor formula:
Sense-check: At , the approximation gives , which matches . At , the true value . The degree-3 approximation gives — excellent agreement.
Example 2: Taylor series of about .
Derivatives of cosine cycle every four steps and alternate between 0 and ±1:
Only the even-order terms survive (odd-order coefficients are zero):
Sense-check: For small angles, is the standard small-angle approximation used in physics and engineering. At rad, the true value is . The approximation —
Accurate to five decimal places from just two terms.
10.3.6 Assumptions & Scope
Scope: What Taylor series needs to work.
- Infinite differentiability for the infinite series. For the finite Taylor polynomial (degree ), you need to be times differentiable at . For the infinite series, must be infinitely differentiable — and even that is not enough. Some infinitely differentiable functions have a Taylor series that converges to the wrong value.
- Convergence radius. Even when the Taylor series converges, it only matches the function within a certain radius around . For about , the series only works for , even though the function is smooth at .
- The remainder is real. A degree- Taylor polynomial has an error term — the remainder . This remainder is what the full Taylor's theorem (the one with Rolle's theorem in its proof) quantifies. For practical work, you either ignore the remainder (staying close to ) or bound it using the Lagrange or integral form of the remainder.
- Not all functions equal their Taylor series. The classic counterexample is for with . All derivatives at 0 are zero, so the Taylor series is identically zero — but the function is positive for . The series converges to the wrong value.
10.3.7 Visual Intuition
Plot on the x-axis from to . At , plot the degree-0 approximation: a horizontal line . Then overlay the degree-1 approximation: the tangent line . Then degree 2: —
A parabola that hugs the exponential upward. Then degree 3:
Which bends upward more aggressively for positive . As you increase the degree, the polynomial curves more and more like the exponential, staying close to it over a wider band around . Landmarks: all approximations pass through —
That is the position match. The tangent line (degree 1) touches with the right slope. The parabola (degree 2) has the right bend. One-sentence takeaway: each added term lets the polynomial curve in one more way, buying you accuracy over a larger neighborhood.
10.3.8 Pitfalls
- Forgetting the factorial. Writing instead of is a common slip. The denominator is , not . For , the coefficient uses , not 3.
- Mixing up the center. If the expansion is around , all terms use , , , not , , . Every must have the right .
- Thinking more terms always means better approximation. More terms give better accuracy near , but far from , a high-degree polynomial can oscillate wildly (Runge's phenomenon). The Taylor series is a local tool, not a global one.
- Assuming the Taylor series converges to the function. As noted in scope, some functions have Taylor series that converge to something else entirely. Always know your function's convergence radius.
- **Confusing the Taylor polynomial with the Taylor series.** The polynomial truncates at degree . The series is the infinite sum. In practice you always truncate, so the remainder matters.
10.3.9 Student Questions and Answers
Q: I understand the 2nd derivative adds curvature. What is the intuition for the 3rd and higher derivatives?
A: In one variable, it is hard to visualize beyond the 2nd. The 1st derivative is the rate of change. The 2nd is the rate of change of the rate of change —
How the slope bends. The 3rd is how the curvature itself changes. Beyond that, you are adding finer and finer corrections.
Q: Is this like distance, velocity, acceleration, jerk?
A: Yes
That is a good parallel. Position is the 0th derivative. Velocity is the 1st. Acceleration is the 2nd. Jerk is the 3rd. Each tells a different story about the motion.
Q: In multiple dimensions, do higher-order terms make more sense?
A: Yes. In a multivariable function like a 3D surface, the function changes in the -direction, the -direction
And along the -plane. The 3rd-order terms capture mixed rates of change —
How the curvature in changes as you move in
And so on. The cross-derivative terms multiply once you go multivariable.
10.3.10 Pedagogical Insight
When you fit a Taylor polynomial, think of adding one geometric feature at a time. The 0th term places a horizontal line at the right height. The 1st term tilts it. The 2nd term bends it into a parabola. The 3rd term adds an S-shape. Each term refines the shape. This is the geometric meaning of Taylor series.
10.3.11 Recap + Bridge
The Taylor series clones a function at a point by matching its position, slope, curvature
And all higher-order bends using only derivative information from that single point. Each added derivative buys one more degree of shape-matching freedom. The factorial denominators undo the chain-rule accumulation of repeated differentiation.
The Taylor series is the central theorem that motivates everything else in this lecture. Rolle's and MVT built the logical foundation. Taylor provides the constructive tool. The next section extends this tool from one variable to two —
Opening the door to the gradient and the Hessian.
Real-world connection: Taylor series approximate solutions to differential equations that have no closed form. Weather models, orbital mechanics simulations
And electronic circuit analysis all use truncated Taylor expansions to turn impossible nonlinear equations into solvable polynomial approximations. In quantitative finance, the Black-Scholes option pricing formula is often approximated by a second-order Taylor expansion (delta-gamma approximation) for real-time risk calculations. Any time you see code that computes on a computer, it is almost certainly using a truncated Taylor (or Chebyshev) polynomial —
A degree-5 or degree-7 approximation that delivers double-precision accuracy for all inputs.
---
10.4 Two-Variable Taylor Series and the Gradient
10.4.1 The Hilly Landscape
What if your function is not a curve
But a whole surface? So far we approximated single-variable functions —
A road winding through hills. Now we step onto a full 2D landscape where elevation depends on two coordinates. How do you approximate the height around any point? The answer is the same idea —
Match the position, the slope
And the curvature —
But now "slope" means two numbers (east-west and north-south) and "curvature" means three.
A function defines a surface over the -plane. The Taylor expansion in two variables builds a polynomial surface —
A tangent plane at degree 1, a paraboloid at degree 2 —
That hugs the real surface near the anchor point.
10.4.2 Symbol Registry
- — the function being expanded — real-valued function on
- — the anchor point where you know the function's value — constants in
- — step size in the -direction — scalar
- — step size in the -direction — scalar
- — partial derivative with respect to — scalar-valued function
- — partial derivative with respect to — scalar-valued function
- — the gradient vector: — vector in
- — second partial derivative in — scalar
- — second partial derivative in — scalar
- — mixed second partial derivative — scalar
10.4.3 The First-Order Expansion
The tangent plane. When you stand on a hillside, the ground immediately under your feet is approximately flat. That flat approximation —
The tangent plane —
Is exactly the first-order Taylor expansion. You know the height at your feet. To estimate the height one step east and one step north, you add: (east slope) × (east step) + (north slope) × (north step). That is all the first-order expansion does.
For a function of two variables, you want to approximate . You already know the function's value at . You take steps in and steps in .
The first-order Taylor polynomial in two variables:
Verbally: the value at the new point equals the value at the old point, plus the slope in times the -step, plus the slope in times the -step.
10.4.4 Gradient Form
The two first-order terms can be written as a dot product:
The gradient is the vector of all partial derivatives. The step vector is how far you move. Taking their dot product gives the total first-order change.
This is like saying: if you walk meters east, your elevation changes by (slope east) . If you walk meters north, your elevation changes by (slope north) . Add both contributions.
The gradient is the vector pointing in the direction of steepest ascent. The first-order term is the directional derivative —
It tells you how much the function changes when you step in direction . When points along the gradient, the dot product is maximal. When is perpendicular to the gradient, you walk along a contour line and the first-order change is zero.
10.4.5 The Second-Order Expansion
The degree-2 term adds curvature information:
You need three pieces:
- : how the slope in changes as you move in — the pure- curvature.
- : how the slope in changes as you move in — the pure- curvature.
- : how the slope in changes as you move in — the cross-curvature, or how the two directions interact.
The factor 2 appears before the term because the mixed derivative shows up twice in the expansion —
Once from and once from . Since (Clairaut's theorem), these two contributions are equal
So they combine to .
10.4.6 The Full Two-Variable Taylor Formula (to Degree 3)
The third-order coefficients are the binomial expansion of . This pattern repeats for all degrees: the coefficients are binomial coefficients. At degree , the coefficient of is .
For the degree- term, the general pattern is:
Where is the binomial coefficient.
10.4.7 At a Critical Point — What the Taylor Expansion Reveals
At a critical point where the gradient is zero (), the first-order term vanishes. The Taylor expansion simplifies to:
The left side is the change in height when you move away from the critical point. The right side is a quadratic form in . The sign of this quadratic form —
Positive, negative
Or mixed —
Determines whether the critical point is a minimum, maximum
Or saddle. This is why the Hessian matrix
Which organizes these second derivatives, is the tool for classifying extrema.
10.4.8 Assumptions & Scope
Scope: What the two-variable Taylor expansion needs.
- Existence of partial derivatives. The expansion to degree requires all partial derivatives up to order to exist at . For most smooth functions (, polynomials, etc.), this holds everywhere.
- Local tool. Like the 1D case, the 2D Taylor polynomial is accurate near and diverges for large steps. The error grows with the distance from the anchor point.
- At a critical point, the first-order term vanishes. This is by design — you solve to find candidate extrema, then use the second-order term (Hessian) to classify them. If the Hessian test is inconclusive, higher-order terms may be needed.
- The number of terms grows combinatorially. At degree in variables, you have distinct partial derivatives to compute. For : terms. For : terms. This combinatorial explosion is why high-dimensional Taylor expansions are rarely used past degree 2.
10.4.9 Visual Intuition
Picture a 3D surface . At the anchor point , the height is —
A single dot in 3D space. The first-order expansion adds a tangent plane —
The flat surface that just kisses the real surface at , matching its slope in both the and directions. The second-order expansion bends that plane into a paraboloid —
It can curve upward (bowl), downward (dome)
Or twist into a saddle shape, depending on the signs of the Hessian entries. Landmarks: at a minimum, the tangent plane is horizontal (gradient zero) and the paraboloid opens upward. At a maximum, the plane is horizontal and the paraboloid opens downward. At a saddle, the plane is horizontal but the paraboloid curves up in one direction and down in another. One-sentence takeaway: the 2D Taylor expansion builds a polynomial surface that matches the real surface in height, tilt
And bend.
10.4.10 Pitfalls
- Forgetting the factor of 2 before . The middle term of the degree-2 expansion is , not . You get the factor of 2 because the mixed derivative contributes from both and .
- Writing instead of when the anchor is not zero. If the expansion is around , every term uses and . The formula template stays the same; only the numbers and variable expressions change.
- Thinking the gradient is a scalar. The gradient is a vector of partial derivatives. The first-order change is the dot product of the gradient with the step — a scalar.
- Confusing with . is a single second-derivative — first differentiate in , then in . It is not the product of the two first derivatives.
- Mixing up the factorial. At degree 2, the denominator is , not 1. At degree 3, it is , not 3.
10.4.11 Student Questions and Answers
Q: In a 1D Taylor series, you do not worry about direction. In 2D, how does the 2nd derivative tell you which direction to go?
A: The mixed term captures how the function changes when both and move together. The gradient tells you the steepest direction of change. The Hessian tells you how that steepest direction itself curves. You combine the gradient (direction) and Hessian (curvature) to make a better step.
Q: If I have a single-variable function like , is that still a Taylor series in one variable?
A: Yes. It has only one variable, . The exponent being just makes the coefficients work out differently —
You can expand it using the chain rule. It is still a 1-variable Taylor expansion. A function becomes multivariable only when you have two or more distinct variables, like or .
Q: In the question format, how do we know whether to write or ?
A: If the expansion is around , every term uses and . So the first-order term becomes , the second-order terms become
And so on. If the expansion is around , then and directly —
No subtraction needed.
10.4.12 Gradient Descent Connection
The gradient always points in the direction where the function increases fastest. In machine learning, you want the minimum. You take the negative gradient —
The opposite direction —
And step downhill. Each step: compute the gradient, move opposite to it, repeat. This iterative process is gradient descent. It is covered fully in later topics
But the first-order Taylor term is its mathematical foundation.
10.4.13 Recap + Bridge
The two-variable Taylor series extends the 1D idea: match position , then slope in each direction , then three curvature terms , with binomial coefficients governing the higher-order terms. The gradient points uphill
And its dot product with the step vector gives the first-order change.
The second-order term —
The quadratic form in —
Organizes the curvature information that the next two sections will formalize as the Hessian matrix. Once you have the Hessian, you can classify critical points and take informed steps in optimization.
Real-world connection: Gradient descent on a loss function is the workhorse of every neural network training loop. The first-order Taylor term is exactly what tells you which direction to update each weight. Weather models use 2D Taylor expansions to interpolate pressure and temperature readings from scattered weather stations onto a uniform grid for forecasting. In computational fluid dynamics, the Navier-Stokes equations are solved by discretizing space into a grid and using truncated Taylor expansions (finite difference methods) to approximate derivatives —
The very same idea, applied to partial differential equations over entire flow fields.
---
10.5 The Hessian Matrix
10.5.1 From Taylor Expansion to Matrix Form
The Hessian is the organizer. In the 1D world, the second derivative at a critical point is a single number: positive means minimum, negative means maximum. In 2D, you have three second derivatives () and you need to know the curvature in every direction —
Not just the pure and axes. The Hessian matrix packs those three numbers into a symmetric matrix
And the quadratic form tells you the curvature along any direction .
The second-order term of the two-variable Taylor expansion is:
This rearranges into a matrix-vector product. Define the step vector and the matrix as:
The Hessian Matrix. For a function with continuous second partial derivatives:
The second-order Taylor term then equals:
is the Hessian matrix — the matrix of all second partial derivatives. It is always symmetric because (Clairaut's theorem).
Why the quadratic form works. Multiply it out:
The off-diagonal appears twice (once from and once from ), automatically producing the factor of 2 in the original formula. The quadratic form is the way to express the curvature of a multivariable function in any direction.
10.5.2 Symbol Registry
- — the Hessian matrix — symmetric (or ) matrix of second partial derivatives
- — the step vector — vector in
- — second partial derivative in — scalar (rate of change of the slope in )
- — second partial derivative in — scalar
- — mixed second partial derivative — scalar (how slope in changes with )
- — quadratic form — scalar (curvature in the direction )
- positive definite — for all non-zero — signals a minimum
- negative definite — for all non-zero — signals a maximum
10.5.3 What the Hessian Represents
The Hessian captures how the gradient itself changes. Each entry has a geometric meaning:
- : how the slope in the -direction changes when you move along .
- : how the slope in the -direction changes when you move along .
- : how the slope in changes when you move along — the cross-interaction.
When you vary both and , the function's slope changes in all these ways. The Hessian bundles them into one matrix.
10.5.4 Connection to Definiteness
You have seen the form before —
In the context of positive definite matrices. A matrix is positive definite if for every non-zero . It is negative definite if the product is always negative.
The Hessian being positive definite means the function is curving upward in every direction —
You are at a local minimum. Negative definite means the function is curving downward in every direction —
A local maximum. Mixed signs (some directions up, some down) give a saddle point.
10.5.5 Derivation: How the Hessian Connects to the Taylor Second-Order Term
Starting from the 2D Taylor degree-2 term:
Rewrite as a matrix-vector product:
This is the standard form. At a critical point where , the full Taylor expansion simplifies to:
The sign of for every possible direction determines whether is a minimum (always positive), maximum (always negative), or saddle (mixed).
10.5.6 The Hessian in Variables
For a function , the Hessian generalizes to an symmetric matrix:
. The Hessian has distinct entries (due to symmetry). For
That is 5050 entries —
Already computationally expensive. For (a modest neural network), a dense Hessian would be entries
Which is why second-order optimization methods require approximations like L-BFGS rather than computing the full Hessian.
10.5.7 Alternative Notation
Some texts write the Hessian with Leibniz notation:
Both notations mean the same thing. The subscript notation is shorter. The Hessian is also sometimes written as —
The "Laplacian of gradient" notation —
Emphasizing it is the gradient of the gradient.
10.5.8 Hessian is Always Symmetric
For functions with continuous second derivatives, . The order of differentiation does not matter. This is Clairaut's theorem (also called Schwarz's theorem). It makes the Hessian symmetric. Symmetry is important —
It guarantees real eigenvalues, orthogonal eigenvectors
And lets you classify critical points cleanly.
10.5.9 Worked Example — Build a Hessian
Compute the Hessian of at the point .
Step 1 — First derivatives (gradient):
Step 2 — Second derivatives:
Step 3 — Evaluate at :
Step 4 — Hessian:
Sense-check: The matrix is symmetric ✓. —
Indicating this point would be a saddle if it were a critical point. (Check: the gradient at is
So it is not a critical point;
The Hessian still exists and describes the local curvature.)
10.5.10 Assumptions & Scope
Scope: When the Hessian is defined and what it can do.
- Continuous second derivatives required. The Hessian needs to be twice differentiable. For functions with discontinuities in second derivatives (e.g., ReLU at zero, where for and for ), the Hessian is undefined at the kink.
- Symmetric only when Clairaut's theorem holds. For pathological functions where mixed partials differ, the Hessian is not symmetric. These functions rarely appear in applied ML problems.
- **The Hessian is a local descriptor.** Like the gradient, the Hessian at a point describes behavior in an infinitesimal neighborhood. A function can have positive-definite Hessian at one point and negative-definite at another.
- Computational cost scales as . Computing the full Hessian for variables requires second derivative evaluations. For large , this is prohibitive. In deep learning, you use first-order methods (SGD, Adam) that never touch the Hessian.
10.5.11 Visual Intuition
At a critical point, imagine standing on a surface. The ground under you is flat (gradient zero). To figure out whether you are in a bowl, on a dome
Or on a saddle, you need to know what happens when you take a tiny step in any direction. The Hessian describes this: gives the curvature along direction . If , then —
Always positive, a circular bowl. If , then —
Positive along the x-axis, negative along the y-axis, a saddle. The eigenvalues of capture the principal curvatures;
The eigenvectors capture the principal curvature directions. One-sentence takeaway: the Hessian converts "how does the surface bend" into a matrix whose definiteness classifies the type of critical point.
10.5.12 Pitfalls
- Confusing the Hessian with the Jacobian. The Hessian is for scalar-valued functions () — it is a square matrix of second derivatives. The Jacobian is for vector-valued functions () — it is an matrix of first derivatives.
- Hessian entries might not be constants. In the worked example for Section 10.8, the Hessian entries were constant (). But for most functions (like ), the Hessian entries depend on and must be evaluated at the specific critical point.
- Forgetting the Hessian exists even at non-critical points. You can compute anywhere. At a critical point, it classifies the extremum. Elsewhere, it describes local curvature for steps — useful in Newton's method.
- Assuming definiteness implies global behavior. A positive-definite Hessian at a critical point only guarantees a local minimum. The global minimum could be elsewhere.
10.5.13 Recap + Bridge
The Hessian matrix organizes all second partial derivatives into a symmetric matrix, converting the second-order Taylor term into the quadratic form . The definiteness of (all-positive, all-negative
Or mixed) reveals whether a critical point is a minimum, maximum
Or saddle.
The Hessian is the judge at the critical point. Once you have found where the gradient vanishes, you form , compute its determinant and
And deliver the verdict. The next section formalizes the classification rules.
Real-world connection: In Newton's method for optimization, the update step uses the inverse Hessian: . The Hessian tells you not just which direction to go (gradient)
But how far —
By accounting for curvature. This is what lets Newton's method converge quadratically near the optimum, compared to gradient descent's linear convergence. In structural engineering, the Hessian of the potential energy function determines the stability of a structure —
A positive-definite Hessian means the structure is stable under load;
A non-positive-definite Hessian signals buckling or collapse. In Gaussian process regression, the Hessian of the log-likelihood (the Fisher information matrix) quantifies the uncertainty in hyperparameter estimates.
---
10.6 Classifying Critical Points Using the Hessian
10.6.1 The Classification Rules
You found flat ground —
Now what? You stand at a point where the gradient is zero. The ground is flat under your feet. But are you at the bottom of a bowl, the top of a dome
Or on a mountain pass that goes up in one direction and down in another? The Hessian ends the suspense. It tells you which kind of flat ground you are standing on.
Given a function and a critical point where , form the Hessian matrix at that point. Compute its determinant:
The Second Derivative Test (2D). Let at a critical point where :
Then classify:
| Condition | Result |
|---|---|
| and | Local minimum |
| and | Local maximum |
| Saddle point | |
| Test inconclusive |
10.6.2 Symbol Registry for Classification
- — determinant of the Hessian — scalar
- — second partial derivative in at the critical point — scalar
- — second partial derivative in at the critical point — scalar
- — mixed partial derivative at the critical point — scalar
10.6.3 Intuition for Each Case
The professor's landscape metaphor. At a critical point, the tangent plane is perfectly horizontal. The Hessian tells you what the landscape looks like immediately around you.
- Minimum (, ): The bowl. Any step you take leads uphill. The surface is concave-up in all directions. Think of the bottom of a cereal bowl.
- Maximum (, ): The dome. Any step you take leads downhill. The surface is concave-down in all directions. Think of the top of a hill.
- Saddle point (): The Pringles chip. In one direction you go up; in the perpendicular direction you go down. The surface looks like a horse saddle or a mountain pass.
- Inconclusive (): The test cannot tell. The landscape might be completely flat (a plateau), or have a trough, or a higher-order shape. You need to look at directional derivatives or the next-order Taylor terms.
10.6.4 Intuitive Explanation of the Determinant Rule
Look at the diagonal entries and . If both are positive, the function is curving upward along both pure directions. But that alone is not enough —
The cross-term could overpower them and flip the sign when you move diagonally. The determinant test accounts for that cross-interaction.
When and , the positive curvature along is strong enough that no mixed direction can make it go negative. You are at a minimum. Conversely, when and , no mixed direction can make it go positive —
A maximum.
When , the cross-term dominates. Some directions curve up, others down. Saddle point.
10.6.5 One-Variable Analogy
In single-variable calculus, you found critical points by setting . You classified them with :
- → minimum.
- → maximum.
The Hessian test is the multivariable version of this. The 2nd derivative becomes a matrix. "Positive" becomes "positive definite." The determinant test is a shortcut for checking definiteness of a symmetric matrix.
10.6.6 Why Suffices When
For a symmetric matrix, means the two diagonal entries have the same sign. If , then must also be positive (otherwise could not be positive —
A negative would make the product negative and the determinant would only be saved if were even more negative
Which is impossible since squares are non-negative).
Proof sketch: If , then
So . This means and share the same sign. Checking just is enough —
You get 's sign for free.
For , you check all leading principal minors, not just the determinant and one diagonal entry.
10.6.7 Worked Example — Saddle Point Detection
Find and classify the critical point of .
Step 1 — Gradient:
Critical point: .
Step 2 — Hessian entries:
Step 3 — Determinant:
Step 4 — Classify: → saddle point at .
Sense-check: —
Along the x-axis, the function curves upward (minimum at 0). —
Along the y-axis, the function curves downward (maximum at 0). A classic saddle. The function value at the critical point is ;
Nearby it is positive in some directions and negative in others.
10.6.8 Assumptions & Scope
Scope: Limitations of the second derivative test.
- Critical point required. The test only classifies points where . If the gradient is non-zero, you are on a slope — there is no extremum there in the first place.
- is a dead end. When the determinant is zero, the Hessian test gives no answer. The point could be anything — a minimum, maximum, saddle, or something more exotic. You need to examine the function directly (graph it, check directional derivatives along suspected trouble directions, or compute higher-order Taylor terms).
- Local, not global. A "local minimum" via Hessian tells you nothing about other critical points elsewhere. The function could dip lower at some faraway point.
- The generalization uses leading principal minors, not a single determinant. For a function of variables, the definiteness test checks the signs of , , …, . All positive → positive definite (minimum). Signs alternate starting negative → negative definite (maximum).
10.6.9 Visual Intuition
Plot a 3D surface with x and y axes on the horizontal plane and z = f(x,y) on the vertical axis. For a minimum, the surface near the critical point looks like an upright bowl —
Contour lines are concentric ellipses. For a maximum, the surface looks like an inverted bowl —
Contour lines are concentric ellipses but the z-values decrease outward. For a saddle point, the surface twists —
Along one axis it rises;
Along the perpendicular axis it falls. Contour lines form hyperbolas crossing at the saddle. The determinant is the Gaussian curvature (up to a factor) of the surface at that point. Positive curvature = local extremum. Negative curvature = saddle. One-sentence takeaway: the sign of distinguishes a bowl/dome from a saddle;
The sign of distinguishes a bowl from a dome when .
10.6.10 Pitfalls
- Forgetting to verify first. The Hessian test only makes sense at critical points. Applying it elsewhere gives curvature information but does not classify extrema.
- Checking instead of is fine when . Both have the same sign. But the canonical rule says . Stick to it for consistency.
- **Assuming and means global minimum.** It only means local. There could be a deeper valley elsewhere.
- Treating as equivalent to a saddle. means the test is inconclusive. A saddle requires . These are different categories.
- Forgetting the Hessian entries must be evaluated at the critical point. If , you must plug in the -coordinate of the critical point before checking its sign. The sign of changes with .
10.6.11 Comparison: 1D vs 2D Classification
| Aspect | 1D () | 2D () |
|---|---|---|
| Critical point condition | ||
| Curvature measure | Single number | Hessian matrix (3 numbers) |
| Minimum test | and | |
| Maximum test | and | |
| Saddle/inflection | (inconclusive/possible inflection) | (definite saddle) |
| Possible outcomes | Min, max, inflection | Min, max, saddle, inconclusive |
| Key difference | No saddle possible in 1D | Saddle points are a new category in 2D+ |
10.6.12 Student Questions and Answers
Q: Why do the minima/maxima conditions only check ? What if the function is decreasing in ?
A: If and , the positive definiteness guarantees that must also be positive. The determinant condition bundles the cross-check. A positive definite symmetric matrix has both diagonal entries positive. Similarly, negative definite means both diagonals are negative.
Q: Is there a way to tell if a critical point is a global minimum or maximum just from the equation?
A: No. There is no mathematical method to guarantee a global extremum in general. You can only verify locally. If a function has only one critical point and the Hessian says it is a minimum, then it is the global minimum. But with multiple critical points, each must be evaluated separately. This is exactly the problem that makes machine learning hard —
With hundreds of variables, you cannot solve for all critical points directly. Iterative methods like gradient descent are used instead
And they can get stuck in local minima.
Q: From an exam perspective, what exactly do we do?
A: Find where . Solve the system. Compute the Hessian entries at each critical point. Compute . Also note . Classify using the table. If there are multiple critical points, evaluate each one separately.
10.6.13 The Saddle Point Problem in Deep Learning
In high-dimensional optimization (neural networks), saddle points are far more common than local minima. The Hessian becomes huge. Most directions are flat or mixed. The saddle point —
Where some directions curve up and others down —
Is the norm. This is why pure second-order methods are rarely used directly. The Hessian is too expensive to compute and too mixed to trust.
The probability that a random critical point in dimensions is a local minimum decays exponentially with . The ratio of saddle points to minima grows combinatorially. A 1000-dimensional loss surface is overwhelmingly saddle —
Which is why gradient descent with momentum (which can "roll through" shallow saddles) works so well in practice
And why Newton-type methods (which get attracted to saddles) require careful damping.
10.6.14 Recap + Bridge
At a critical point, compute . If , the sign of decides min vs max. If , it is a saddle. If , the test cannot decide. This is the multivariable upgrade of the 1D second derivative test
And it will be applied in the next worked example.
The Hessian is the judge. You have the crime scene (the critical point where the gradient vanished). Now you need to deliver the verdict. The next two sections walk through full worked examples —
One for Taylor expansion, one for extrema classification.
Real-world connection: Hessian-based saddle point detection is central to understanding why deep learning works at all. The loss surfaces of neural networks are dominated by saddle points, not poor local minima —
Meaning gradient descent rarely gets permanently trapped in bad spots. It mostly slows down near saddle points (where the gradient is small but not zero in all directions). Techniques like second-order optimization (Newton, L-BFGS), natural gradient descent
And trust-region methods explicitly use the Hessian or its approximations to account for curvature when deciding update steps. In computational chemistry, Hessian eigenvalue analysis determines whether a molecular geometry is a stable equilibrium (all positive eigenvalues) or a transition state (exactly one negative eigenvalue —
A saddle point on the potential energy surface).
---
10.7 Worked Example — Two-Variable Taylor Series
10.7.1 Problem Statement
This is the canonical Taylor exam problem. You are handed a two-variable function, told an anchor point and a degree
And asked to crank out the expansion. The process is mechanical —
List derivatives, evaluate, plug in —
But it rewards systematic work. Every skipped derivative is a lost mark.
Expand as a Taylor polynomial around up to degree 3.
The anchor point is , so and .
10.7.2 Step 1 — List All Required Derivatives
The general degree-3 formula from Section 10.4:
You need: , , , , , , , , , . That is 10 partial derivatives in total (counting itself as the 0-th derivative). All evaluated at .
10.7.3 Step 2 — Compute Each Derivative and Evaluate at (0, 0)
Degree 0 — The function itself:
Degree 1 — First partial derivatives:
- : differentiate with respect to . The factor is constant with respect to . The derivative of is .
- : differentiate with respect to . The factor is constant with respect to . The derivative of is .
Degree 2 — Second partial derivatives:
- : differentiate with respect to again. Same as .
- : differentiate with respect to . The derivative of is .
- : start from . Differentiate with respect to . is constant. The derivative of is .
Degree 3 — Third partial derivatives:
- : differentiate with respect to . Same result.
- : differentiate with respect to . The derivative of is .
- : start from . Differentiate with respect to . Derivative of is .
- : start from . Differentiate with respect to . Derivative of is .
10.7.4 Step 3 — Assemble
Plug all computed values into the formula:
Degree 0 term:
Degree 1 term:
Degree 2 term:
Degree 3 term:
10.7.5 Final Answer
The approximation is valid near and includes terms up to degree 3.
10.7.6 Numerical Sense-Check
Verify the approximation at a small point, say , :
True value:
Approximation:
Error: — accurate to four decimal places. The degree-3 polynomial is practically exact for small near the origin. ✓
10.7.7 Assumptions & Scope
Scope: Limitations of this expansion.
- Local only. The approximation degrades as or grow. At , , the true value is . The degree-3 approximation gives — still decent (~2% error). At , , the error blows up because the exponential factor dominates and the truncated polynomial cannot keep up.
- The function must be smooth. is infinitely differentiable everywhere, so no problems. Functions like or have derivative issues that limit or prevent Taylor expansion at certain points.
- Higher-degree terms are not zero. The degree-3 expansion stopped at 3 — it ignored degree-4 terms like evaluated at , which are zero for this particular function at the origin. But for other anchor points, or for functions like , the higher terms matter.
10.7.8 Visual Intuition
The surface near looks like a gently twisted sheet. At the origin, the function value is zero —
You are at sea level. Moving in the pure -direction ( is near-zero for tiny ), the surface is nearly flat because . Moving in the pure -direction, the surface rises linearly with slope 1 (the term ). Moving in the -diagonal, the surface rises with an extra push from the mixed term . The cubic term bends the surface slightly downward in the -direction as you move farther out. One-sentence takeaway: the polynomial is the degree-3 polynomial that best matches near the origin in position, slope
And all curvatures.
10.7.9 Pitfalls
- Missing the factor of 2 before . In the degree-2 term, it is , not . The factor of 2 comes from the binomial coefficient .
- Mixing up with . is a second derivative — differentiate once in , then once in . It is not the product of the two first derivatives.
- Forgetting to evaluate at the anchor point. is the general formula. You must plug in to get the coefficient. The fact that means that specific term vanishes — not that you skip computing it.
- Omitting the factorial denominators. Degree 2 uses . Degree 3 uses . These are not optional — they undo the chain of factors that repeated differentiation accumulates.
- Writing instead of . At , and , so the shortcut works. At any other anchor, you must write and everywhere.
10.7.10 Student Questions and Answers
Q: Why is computed by taking and differentiating with respect to ?
A: The notation means: first differentiate with respect to , then differentiate the result with respect to . You compute first (getting ), then treat that as a new function and take its -derivative. Since is constant with respect to , only gets differentiated, giving .
Q: Does the order matter? Is the same as ?
A: For smooth functions (continuous second derivatives), yes. . You can take the derivatives in either order and get the same result. This is Clairaut's theorem. In this example, computing (first , then ): , then —
Identical to .
Q: What if the anchor point is not ? Say it is ?
A: Then every term uses instead of
And instead of . Compute all derivatives at instead of . For example, , ,
And so on. The structure of the formula stays the same;
Only the numbers change.
10.7.11 Exam Guidance
Exam note: The question format is predictable —
You get a function like ,
Or . You are told an anchor point and a degree (usually 2 or 3). Compute all partial derivatives to that degree, evaluate at the point
And plug into the formula. Show each derivative step. Do not skip intermediate derivatives. The marking scheme rewards systematic work —
Even if you make an arithmetic error, a clear table of derivatives earns partial credit. Practice at least three of these before the exam.
10.7.12 Recap + Bridge
The degree-3 Taylor expansion of about is . The process is mechanical but requires care —
List all derivatives, compute each one, evaluate at the anchor, then assemble term by term. The Bernoulli coefficients at each degree are binomial coefficients divided by the factorial of the degree.
Now you have seen the full Taylor expansion pipeline. The next section flips the lens —
Instead of building a Taylor expansion, we use the gradient and Hessian to find and classify extrema.
Real-world connection: Every physics engine in a game or simulation uses truncated Taylor expansions for numerical integration. The Verlet integration algorithm updates positions using a second-order Taylor expansion of the equations of motion —
Matching position and velocity exactly, with error in the jerk term. When you see cloth ripple or water splash in a video game, each particle's trajectory is computed with exactly this kind of local Taylor approximation.
---
10.8 Worked Example — Extrema Using the Hessian
10.8.1 Problem Statement
This is the canonical extrema exam problem. You are given a two-variable function, told to find and classify its critical points. The four-step recipe —
Gradient zero, Hessian entries, determinant, classify —
Works every time. This is the problem that tests whether you understand what the Hessian is for.
Find and classify the extrema of:
10.8.2 Step 1 — Find Critical Points
Set the gradient to zero:
Solve the system of two linear equations:
From the first equation: . Substitute into the second:
Then .
The critical point is .
10.8.3 Step 2 — Compute the Hessian Entries
All three second derivatives are constants —
They do not depend on or . This means the curvature is the same everywhere on the surface.
The Hessian at :
10.8.4 Step 3 — Classify
Compute the determinant:
- (positive).
- (negative).
Apply the classification table from Section 10.6:
and → local maximum.
Conclusion: the critical point is a local maximum.
10.8.5 Step 4 — Compute the Maximum Value
Plug into :
The maximum value is .
10.8.6 Step 5 — Verify by Intuition
Both and are negative . The function curves downward in both the and directions. The determinant is positive, confirming the curvature is consistently downward in every direction —
The quadratic form is always negative for non-zero because the cross-term is never large enough to overcome (since by the AM-GM inequality). You are at the top of a hill.
10.8.7 Second Worked Example — Multiple Critical Points
Find and classify all critical points of . (From the companion document.)
Step 1 — Gradient:
Substitute into :
So or .
- If , then . Critical point: .
- If , then . Critical point: .
Step 2 — Hessian entries:
Step 3 — Classify each point:
At :
→ Saddle point at .
At :
and → Local minimum at .
Values: . .
Sense-check: The function is cubic in both variables. Near , in the direction , —
Curving downward. In the direction , —
Curving upward. Mixed curvature confirms the saddle. Near , the function dips to a minimum value of 0.
10.8.8 Assumptions & Scope
Scope: What makes Hessian classification easy or hard.
- Constant Hessian entries are rare. For this function, are constants. For most functions (like ), the Hessian entries vary with and . Always evaluate them at the critical point.
- Linear gradient equations are also rare. The system , was two linear equations — easy to solve. For nonlinear gradients (like ), you may get zero, one, or many critical points.
- A single critical point with constant negative-definite Hessian means global maximum. The function is a concave quadratic. No other critical point can exist, because the gradient is a linear system with a unique solution. The Hessian negativity means the function is concave everywhere, so the unique stationary point is the global maximizer.
- With multiple critical points, no global guarantee without comparing values. You must evaluate at each candidate and at the boundaries (if the domain is bounded) to find the global extremum.
10.8.9 Visual Intuition
The function is a concave quadratic —
An inverted paraboloid. Completing the square shows it is
Which after rotation amounts to a downward-opening bowl centered at with maximum value 8. The contour lines are concentric ellipses, getting smaller and higher as they approach the center. The gradient arrows all point toward the peak. The Hessian has eigenvalues and —
Both negative, confirming the surface bends downward in all directions, more steeply along one axis (eigenvalue ) than the other (). One-sentence takeaway: this function is a textbook concave quadratic —
One peak, no surprises, perfectly classified by the Hessian.
10.8.10 Pitfalls
- Solving the gradient equations sloppily. For linear systems, double-check your algebra. A sign error in vs flips the entire classification.
- Forgetting to compute at the critical point. The question asks you to find the extrema — that means reporting both the point AND the function value there. The Hessian only classifies the type; it does not give the value.
- Assuming constant Hessian implies constant classification everywhere. For this function, the Hessian is constant, so one classification suffices. For non-constant Hessians (like with ), the sign depends on where you are. Classify each critical point independently.
- Stopping at the determinant without checking . narrows it to min or max, but you need 's sign to finish. For symmetric matrices, means both diagonal entries have the same sign, so either one works — but the convention is to state .
10.8.11 Student Questions and Answers
Q: Can we say this is a global maximum?
A: Yes, in this case. The Hessian entries are constants —
The curvature is the same everywhere. There is only one critical point. Since the Hessian is negative definite everywhere
That single critical point is the global maximum. The function value there is 8
And for all .
Q: What if we get multiple critical points?
A: Evaluate the Hessian at each one separately. Some might be minima, some maxima, some saddle points. You cannot declare a global extremum without checking all of them and comparing their function values. Also check the function's behavior as —
For unbounded domains, the global extremum might not exist at all.
Q: From an exam perspective, what exactly do we do?
A: Four steps: (1) Set and solve for critical point(s). (2) Compute and evaluate at each critical point. (3) Compute . (4) Classify using the table: → min;
→ max;
→ saddle;
→ inconclusive. Also report the function value at each extremum.
10.8.12 Exam Guidance
Exam note: The extrema question follows this exact template. Expect one function to classify —
Possibly with multiple critical points (like ). Practice the four-step recipe until it is automatic. Common variants: the function might be or . For the second example in this section, you got two critical points —
One saddle at , one minimum at . The professor expects you to classify all critical points found. Missing one is a lost mark.
10.8.13 Recap + Bridge
has a global maximum at with value 8. The classification pipeline —
Gradient zero → Hessian entries → determinant → classify —
Is the same for every two-variable function. The only variation is whether the Hessian entries are constants (easy) or functions of (requires plugging in the critical point).
Now you have seen both the Taylor expansion pipeline (Section 10.7) and the Hessian classification pipeline (this section). The next two sections close the lecture with two more directly applicable topics —
Linearization (the simplest Taylor approximation) and backpropagation (the chain rule on graphs).
Real-world connection: This exact problem type —
Find and classify the extrema of a multivariate function —
Is the central task in maximum likelihood estimation. When you fit a Gaussian distribution to data, you maximize the log-likelihood function. You set its gradient to zero, compute the Hessian of the log-likelihood (which gives you the Fisher information matrix)
And verify the point is a maximum. In portfolio optimization (Markowitz mean-variance), you minimize portfolio variance subject to return constraints —
The objective is a quadratic function whose Hessian is the covariance matrix of asset returns. Classifying the critical point confirms you have the global minimum-variance portfolio.
---
10.9 Linearization
10.9.1 Definition
What is the simplest possible Taylor approximation? Take the full Taylor series and stop after degree 1. That's linearization —
Just the position and the slope, no curvature. It turns a curved function into a straight line (or flat plane) that touches it at one point. It's the math equivalent of saying "for small changes, everything looks linear."
Linearization is the first-order Taylor approximation. It approximates near a point using only the position and slope:
1D Linearization:
2D Linearization:
That is it. No second derivatives. No curvature. No factorials. Just the function's value at the anchor point plus the gradient dotted with the step.
In vector form, for a function :
10.9.2 Symbol Registry
- — the linearization (linear approximation) of at — real-valued function of
- — the anchor point where you know the function's value — constant in
- — the function value at — scalar
- — the first derivative (slope) at — scalar
- — the step size from the anchor — scalar
10.9.3 When Linearization Works
Linearization is accurate when is close to . The slope at is a good proxy for the average slope over a tiny interval. As the step size grows, the error from ignored curvature grows quadratically with the step size —
That quadratic error is exactly the second-order Taylor term you dropped:
This is why linearization is a "local" tool. Near , the error is tiny. Far from , the curvature terms you ignored dominate and the approximation becomes useless.
10.9.4 Worked Example 1 — 1D
Find the linearization of at . (From the lecture.)
Step 1 — Value of the function at :
Step 2 — First derivative:
Evaluate at :
Step 3 — Assemble:
Simplify if needed:
Numerical sense-check: . At ,
Which matches . At (a small nudge of 0.1), the linearization predicts . The true value —
The linearization is excellent at close range. ✓
10.9.5 Worked Example 2 — 2D
Find the linearization of at . (From the companion document.)
Step 1 — Value of the function:
Step 2 — Gradient:
Evaluate at :
Step 3 — Assemble:
Conclusion: Near the origin, . The function behaves like a plane tilted in the -direction only —
The first-order change in is zero because the cosine is flat at .
Sense-check: At , the true value is . The linear approximation gives . Error: — incredibly accurate at this small step. ✓
10.9.6 Assumptions & Scope
Scope: When linearization is good enough and when it is not.
- Step size matters. If is tiny, linearization is excellent. If the step is large, the error from dropped curvature and higher-order terms dominates.
- The function must be differentiable at . If does not exist (sharp corner, cusp), there is no tangent line to build the linearization around.
- Linearization replaces the function with its tangent. For a concave function like , the linearization always overestimates. For a convex function like , it always underestimates. Knowing the sign of tells you which direction the error goes.
- Linearization is the foundation of gradient descent. Each step of gradient descent moves along the direction of steepest descent as predicted by linearization. The first-order Taylor approximation justifies taking a small step opposite the gradient.
- Linearization cannot capture curvature-dependent behavior. If the function bends sharply near (large ), linearization breaks quickly. The second-order Taylor term would be needed for even moderate step sizes.
10.9.7 Visual Intuition
Plot over . At , the function has value and is sloping downward (since ). The linearization is the tangent line at that point —
A straight line passing through with slope . Near , the tangent line hugs the curve closely. Move out to or
And the curve pulls away from the straight line —
The linearization underestimates the function at both ends because is convex (its second derivative is positive). One-sentence takeaway: linearization replaces a curve with its tangent line —
Accurate near the touch point, diverging as you move away.
10.9.8 Pitfalls
- Forgetting the chain rule in the derivative. For , the derivative is , not alone. The chain rule brings the factor .
- Dropping the sign of the derivative. At , — it is negative. The tangent line slopes downward. Forgetting the negative sign gives an upward slope, which is completely wrong.
- Using instead of . The formula is . If , then . Writing instead is a common slip.
- Confusing linearization with interpolation. Linearization matches the function value and derivative at one point. Interpolation matches function values at multiple points. These serve different purposes.
10.9.9 Exam Guidance
Exam note: A linearization question is a Taylor series question cut off at degree 1. Compute and
Or and the gradient for 2D. Plug into . That is the whole answer. A common variant gives a 2D function and asks for the linearization at a specific point —
Compute both partial derivatives and assemble. Do not overthink it;
No factorials, no second derivatives.
10.9.10 Recap + Bridge
Linearization is the degree-1 Taylor approximation. It captures position and slope, ignoring curvature and all higher-order bends. It is accurate near the anchor point and forms the mathematical basis for gradient descent —
At each step, you use the linear approximation to decide where to go.
The last concept in this lecture, backpropagation, uses exactly the machinery built so far —
Partial derivatives, the chain rule
And the gradient —
To compute gradients through networks of functions efficiently.
Real-world connection: Linearization is the first step in virtually every nonlinear control system. When a drone stabilizes itself mid-flight, the flight controller linearizes the nonlinear aerodynamics around the current orientation and solves a linear quadratic regulator (LQR) problem 100 times per second. Self-driving cars linearize their motion models around the current state estimate in Extended Kalman Filters. In circuit design, transistors are nonlinear devices
But small-signal analysis linearizes them around a DC operating point, enabling frequency response analysis with simple linear techniques. Any time you hear "small-signal model" or "linear regime," you are hearing about linearization in the wild.
---
10.10 Backpropagation and Computational Graphs
10.10.1 What Backpropagation Is
You know the chain rule. So you already know backpropagation. If you can compute by saying "derivative of sin is cos, times derivative of the inside which is ," you have done backpropagation. Backprop is nothing more than the chain rule applied to a network of operations —
Doing it backwards, one layer at a time
So you never repeat work.
Backpropagation is the chain rule applied to a network of connected computations. It is the algorithm that trains neural networks. But the core idea is simple.
Consider a function defined through intermediate steps. Say flows into and . Both and flow into .
Forward pass: compute all intermediate values, working left to right. Backward pass: compute how the output changes with respect to the input , working right to left.
10.10.2 Why Backpropagation Exists
Purpose. Manual differentiation of deeply nested functions —
Like neural networks with millions of parameters —
Is impossible. Backpropagation automates gradient computation by breaking a complex function into elementary operations whose derivatives are known, then applying the chain rule systematically in reverse. The result: the gradient of the output with respect to every input and every intermediate parameter, computed in one backward pass.
10.10.3 Inputs & Outputs
Inputs:
- A computational graph (DAG) of elementary operations (+, ×, sin, exp, matrix multiply, etc.)
- An input value for each leaf node
- A final output node whose gradient we want to propagate
Outputs:
- The gradient for every leaf node in the graph — efficiently, in a single backward pass
The forward pass computes all intermediate values. The backward pass computes all gradients, reusing the forward values where needed.
10.10.4 The Chain Rule on a Graph
To get , trace every path from to and multiply derivatives along each path, then sum:
If a node feeds into multiple downstream nodes , its total influence on the output is the sum over all paths:
This is the multivariate chain rule. Each edge stores a local derivative. The backward pass accumulates them along all paths.
10.10.5 Steps — The Forward/Backward Algorithm
Forward Pass (left to right):
- Start with the input node(s), with their actual numeric values.
- At each node, apply its operation using the values from its parents, producing a new value.
- Store each intermediate result — these will be needed during the backward pass.
- Continue until you reach the output node.
Backward Pass (right to left):
- Initialize the gradient of the output with respect to itself: .
- At each node producing output that feeds into downstream nodes:
- Compute each local derivative for every input .
- Multiply the downstream gradient by each local derivative to get .
- If feeds into multiple nodes, sum all incoming gradient contributions.
- Continue until all leaf nodes have accumulated gradients.
The notation (read "u-bar") is the total gradient of the final output with respect to node .
10.10.6 Trace — Worked Example from the Companion Document
Compute at for:
Step 1 — Decompose into elementary operations:
Step 2 — Forward pass (compute all values with ):
Step 3 — Backward pass (compute gradients right to left):
Start:
Final answer:
Sense-check: The function decreases as increases at —
The negative gradient confirms this. If we nudge upward by a tiny amount , should drop by roughly . The magnitude is reasonable for a function whose value is ~8.5 and whose dominant term is with dominating the square root.
10.10.7 Complexity & Cost
The forward pass requires one evaluation per node. The backward pass requires one gradient computation per edge. The total cost of computing the gradient with respect to all inputs via backpropagation is roughly 2× the cost of a single forward pass —
One forward compute, one backward compute. This is what makes it feasible to train networks with millions of parameters. Without backprop, computing each parameter's gradient independently would cost forward passes for parameters. With backprop: one forward + one backward, regardless of .
The memory cost scales with the graph size —
You must store all intermediate forward-pass values because the backward pass needs them to compute local derivatives. For deep learning, this is the leading memory bottleneck. Techniques like activation checkpointing trade compute for memory by recomputing some forward values during the backward pass.
10.10.8 When to Use / Alternatives
- Use backpropagation when: you need gradients through a differentiable computational graph — neural network training, variational inference, neural ODEs, differentiable physics simulators.
- Use numerical differentiation (finite differences) when: you only need to verify a handful of gradients, or the function is not analytically differentiable. cost makes it impractical for training.
- Use symbolic differentiation when: the function is small and you want an exact analytical expression. This blows up exponentially on large graphs (expression swell).
- Use automatic differentiation (backprop is the reverse-mode variant) when: you have many inputs and few outputs — the standard case in ML (millions of parameters → one scalar loss). Forward-mode AD is better when you have few inputs and many outputs.
10.10.9 Symbol Registry
- — nodes in a computational graph — can be scalars, vectors, or tensors
- — the total derivative of output with respect to input — scalar
- — "u-bar": shorthand for , the gradient of final output w.r.t. node
- — local derivative of with respect to intermediate node — scalar
- — local derivative of with respect to input — scalar
10.10.10 How It Connects to Neural Networks
A neural network is one giant computational graph. Each layer computes:
Where is the input from the previous layer, and are the parameters (weights and biases), and is an activation function (sigmoid, tanh, ReLU).
The goal: find parameters minimizing a loss, e.g., squared error:
Backpropagation computes and for every layer in one backward pass. Update rule (gradient descent):
Backpropagation is not a new mathematical idea. It is the chain rule, applied systematically to a densely connected graph. Once you understand partial derivatives and the chain rule, backpropagation follows naturally.
10.10.11 Assumptions & Scope
Scope: When backpropagation works and when it doesn't.
- The graph must be differentiable. Every operation in the graph must have a well-defined local derivative. Operations like "if x > 0 then y = 1 else y = 0" break differentiability. ReLU is differentiable except at , where the subgradient is used.
- Directed acyclic graph (DAG). Backprop works on feedforward graphs. For recurrent networks with cycles, you unroll the graph in time (backpropagation through time, BPTT), creating a DAG.
- Vanishing/exploding gradients. For very deep networks, the chain of multiplied derivatives can shrink to zero (vanishing) or blow up to infinity (exploding). This is a practical problem, not a mathematical failure — it is addressed by residual connections, careful initialization, gradient clipping, and normalization layers.
- Memory scales with graph depth. Storing all intermediate activations for the backward pass is expensive. For a 100-layer network with batch size 64 and activation size , the activation memory alone can exceed GPU RAM.
10.10.12 Visual Learning Resource
Interactive visualizations help make backpropagation intuitive. MLU Explain (Machine Learning University Explain) has an interactive neural network demo that steps through the forward and backward passes. Watching each weight update as the error flows backward makes the chain of derivatives concrete.
10.10.13 Student Questions and Answers
Q: How does this relate to the gradient and Taylor series we covered?
A: The gradient tells you the direction of steepest change for a function. Backpropagation is how you actually compute that gradient for a neural network —
A function defined by many composed operations. You use the chain rule (backprop) to get the gradient
And then you use the gradient to update weights (gradient descent). The first-order Taylor approximation is the theoretical foundation: you move weights opposite the gradient to reduce the loss. The full chain is: Taylor → gradient → gradient descent
And backprop → efficiently computing that gradient for huge functions.
10.10.14 Pitfalls
- Forgetting to sum gradients when a node has multiple children. If node flows into and , then . Missing one path gives the wrong gradient.
- Accumulating gradients instead of resetting between batches. In frameworks like PyTorch, calling
.backward()accumulates gradients in.grad. Forgetting to zero them (optimizer.zero_grad()) causes the gradient to be the sum over all previous batches — a common debugging nightmare. - Confusing the backward pass with backpropagation-through-time (BPTT). Standard backprop works on DAGs. Recurrent networks need BPTT, which unrolls the recurrence into a DAG over time steps.
- Assuming backprop is always reverse-mode AD. Reverse-mode AD (what backprop is) is optimal for scalar outputs with many inputs. Forward-mode AD is optimal for few inputs with many outputs. The choice matters for applications beyond neural network training.
10.10.15 Recap + Bridge
Backpropagation is the chain rule applied backwards through a computational graph, computing the gradient of a scalar output with respect to every input in a single right-to-left pass. It is not magic —
It is the systematic application of the chain rule, reusing intermediate gradients to avoid redundant computation. The forward pass computes values;
The backward pass computes derivatives by multiplying local derivatives along each edge and summing at every node.
Backpropagation ties together every concept in this lecture. Rolle's theorem and MVT built the logical foundation for approximation. Taylor series gave the polynomial-building tool. The gradient tells you which direction to move. The Hessian tells you about curvature. And backpropagation is how you actually compute the gradient for real-world functions that are too complex to differentiate by hand.
Real-world connection: Every deep learning framework —
TensorFlow, PyTorch, JAX —
Is built on automatic differentiation (backprop). When you write loss.backward() in PyTorch, you are running exactly the backward pass algorithm described here, scaled to millions of parameters. The same technique powers differentiable physics engines (Brax, DiffTaichi) that optimize robot controllers by backpropagating through simulated physics. In computational finance, the "Greeks" (sensitivities of option prices to underlying parameters) are computed via AAD (adjoint algorithmic differentiation) —
The same reverse-mode AD that backs neural network training. Google's PageRank can be understood as a fixed-point computation whose sensitivity to link structure is computed via the same chain-rule accumulation.
---
Exam Guidance Summary
Exam note: The exam has 5 to 6 problems total. Lectures 6 and 7 carry the most weight —
Expect at least two or three problems from those sections. Lecture 8 (linearization, backpropagation) is lighter.
Question distribution (approximate):
| Problem | Topic | Weight |
|---|---|---|
| 1 | REF / RREF — guaranteed | Standard |
| 2 | Linear combination, independence, or basis — one question from these related concepts | Standard |
| 3 | Vector spaces — proving a set is a vector space or subspace | Standard |
| 4 | Decomposition — SVD, eigendecomposition, or diagonalization | Standard |
| 5 | Taylor series — possibly combined with gradient concepts | Standard |
| 6 | Hessian matrix — classification of critical points | Standard |
Topic priority for exam preparation:
- Lectures 6 and 7 carry the most weight. Expect at least two or three problems from these sections.
- Lecture 8 (linearization, backpropagation) is lighter. Give it less time if you are short.
- Practice Taylor series and Hessian problems. These are reliably tested.
For Taylor series problems (2D):
Exam note —
Taylor recipe: (1) List all required partial derivatives up to the stated degree. (2) Compute each derivative and evaluate at the given anchor point. (3) Plug into the two-variable Taylor formula, respecting the binomial coefficients and factorial denominators. (4) Show every derivative step —
Skipping intermediate derivatives costs marks. The anchor point determines whether you write or . If the anchor is , and . Otherwise, every term uses and .
For Hessian / extrema problems (2D):
Exam note —
Extrema recipe: (1) Set and solve the system to find all critical points. (2) Compute —
These may be functions of
So evaluate them at each critical point. (3) Compute . (4) Classify using the table: → min;
→ max;
→ saddle;
→ inconclusive. (5) Report the function value at each extremum —
Not just the classification. If there are multiple critical points, classify all of them.
General exam advice:
- Companion documents are the best study resource. They are concise — a few pages each. Recordings take much longer to review. Use the documents.
- Practice from the exam preparation problem set uploaded to the course platform.
- Work through the full worked examples in Sections 10.7 and 10.8 until the four-step recipes become automatic — these are the two most reliably tested problems from this lecture.
---
Key Industry Applications
The concepts in this lecture are not just exam material — they are the mathematical backbone of modern machine learning systems.
Gradient descent is the workhorse of all neural network training. Every time you call model.fit() in Keras or optimizer.step() in PyTorch, the framework computes a gradient via backpropagation and updates weights by stepping opposite to it. The first-order Taylor approximation is the mathematical guarantee that a small step opposite the gradient reduces the loss. Stochastic gradient descent (SGD) and its variants (Adam, RMSProp, Adagrad) are used in every production ML system from recommendation engines to large language models.
Hessian-based optimization powers second-order methods like Newton's method and quasi-Newton approximations (L-BFGS). These are standard in logistic regression, support vector machines
And any convex optimization problem where the Hessian is tractable. L-BFGS is the default optimizer in scikit-learn's LogisticRegression and is widely used in maximum entropy models, conditional random fields
And classical structure-from-motion in computer vision. The Hessian (or its approximation) informs both the direction and the step size, enabling quadratic convergence near the optimum —
Dramatically faster than gradient descent's linear convergence.
Backpropagation and computational graphs are the engine behind all deep learning frameworks —
TensorFlow, PyTorch, JAX, Flux.jl. These libraries build a computational graph from your code, then apply reverse-mode automatic differentiation (backprop) to compute gradients. The same technique powers differentiable renderers (for inverse graphics), differentiable physics simulators (Brax, DiffTaichi —
Used in robotics for optimizing controllers through simulated dynamics)
And neural ODEs. Financial institutions use adjoint algorithmic differentiation (AAD) —
The same reverse-mode AD —
To compute the "Greeks" (market risk sensitivities) for complex derivative portfolios.
Saddle points are a central research topic in deep learning theory. In high dimensions, the loss surface is overwhelmingly dominated by saddle points, not local minima. The probability that a random critical point is a local minimum decays exponentially with dimension. This is why gradient descent with momentum —
Which can "roll through" shallow saddles —
Works so well empirically
And why pure Newton methods fail in deep learning without significant damping or trust-region modifications. Understanding saddle points is the starting point for grasping why training large networks is possible at all.
Linearization is the foundational technique in control theory and state estimation. Extended Kalman Filters (used in every GPS receiver, drone autopilot
And self-driving car perception stack) linearize nonlinear motion and observation models around the current state estimate at each time step. Small-signal analysis in circuit design linearizes transistor behavior around a DC operating point. Model predictive control (MPC) for autonomous vehicles repeatedly linearizes the vehicle dynamics to solve a quadratic program in real time. Every "linear regime" or "small-signal model" in engineering is linearization in the wild —
The degree-1 Taylor approximation, applied at scale.
---
MFML Lecture 10 notes · Taylor Series and Hessian-Based Optimization
Sections Breakdown
The flat-spot guarantee and why its three conditions are chosen that way.
Average slope equals an instantaneous slope; the bridge to Taylor series.
Position, slope, curvature, factorials, and higher-order terms in one variable.
Extending Taylor to surfaces, the gradient, and the second-order expansion.
Collecting second derivatives and what the symmetric matrix represents.
The determinant test that separates minima, maxima, and saddles.
A full degree-3 expansion of e^x sin y about the origin.
Finding and classifying critical points step by step.
The degree-1 approximation and where it stays accurate.
Reverse-mode differentiation and the chain rule on a graph.
Question distribution and the two reliable four-step recipes.
Gradient descent, Hessian methods, backpropagation, saddles, and linearization in practice.
Exam Revision Notes
Below is the distilled, exam-ready core of this lecture. Every entry is built from the full textbook notes above. Use this section for rapid review — but if something doesn't make sense, go back to the full explanation in the main content.
Rolle's Theorem
Must-know: A function that starts and ends at the same height must flatten somewhere between. This is the stepping stone to the Mean Value Theorem.
⚠️ Top pitfall: Forgetting that the endpoint equality is required. Without it the theorem does not apply.
Self-check: Give one example where yet a zero slope still occurs in .
Connects to: Mean Value Theorem, Taylor series foundations.
Mean Value Theorem
Must-know: Somewhere between the endpoints, the instantaneous slope equals the average slope across the interval. This bridges Rolle's theorem to Taylor series.
⚠️ Top pitfall: Reading the conclusion as if it held at every point. The slope matches only at one point .
Self-check: For on , find the value of the theorem promises.
Connects to: Rolle's theorem, Taylor remainder term.
One-Variable Taylor Series
Must-know: Near a point, a smooth function is built from its value, slope, curvature, and higher derivatives. Each term carries a factorial in the denominator.
⚠️ Top pitfall: Dropping the factorial denominators. They undo the chain of factors that repeated differentiation produces.
Self-check: Write the degree-2 Taylor polynomial of about and state where it stays accurate.
Connects to: Two-variable Taylor series, linearization.
Two-Variable Taylor Series and the Gradient
Must-know: In two variables the first-order change is the dot product of the gradient with the step. The degree-2 term adds three curvature pieces, with a factor of 2 on the mixed term.
⚠️ Top pitfall: Writing instead of . The mixed derivative contributes twice by Clairaut's theorem.
Self-check: Explain why the mixed term carries the coefficient 2 rather than 1.
Connects to: Hessian matrix, gradient descent, worked Taylor example.
Two-Variable Taylor Worked Example
Must-know: The exam recipe is mechanical. List every derivative to the stated degree, evaluate at the anchor, then plug into the formula. Show each derivative step.
⚠️ Top pitfall: Forgetting to evaluate derivatives at the anchor. A zero coefficient means that term vanishes, not that you skip it.
Self-check: At the origin, which single term gives the leading linear rise in the -direction?
Connects to: Two-variable Taylor series, Hessian classification.
The Hessian Matrix
Must-know: The Hessian collects all second partial derivatives into a symmetric matrix. It organises the curvature that the Taylor series captures at a critical point.
⚠️ Top pitfall: Thinking the Hessian is not symmetric. For smooth functions , so the off-diagonal entries match.
Self-check: Why must the Hessian be symmetric when second derivatives are continuous?
Connects to: Critical point classification, Newton's method.
Classifying Critical Points with the Hessian
Must-know: Set the gradient to zero to find candidates, then use the Hessian determinant . The sign of and of decides the type.
⚠️ Top pitfall: Treating as a saddle. It is inconclusive, so use higher-order terms or test nearby points.
Self-check: For , classify the critical point at the origin.
Connects to: Hessian matrix, two-variable Taylor series, extrema worked example.
Linearization
Must-know: Linearization is the degree-1 Taylor polynomial. It replaces a curved function with its tangent line near a point and underpins many engineering models.
⚠️ Top pitfall: Using the linearization far from the anchor. It is a local tool and the error grows with distance.
Self-check: Give the linearization of about and state where it stays accurate.
Connects to: One-variable Taylor series, Extended Kalman Filters.
Backpropagation and Computational Graphs
Must-know: A neural network is a graph of operations. Backpropagation applies the chain rule backwards to compute each weight's gradient in one pass.
⚠️ Top pitfall: Believing backprop invents new math. It is the chain rule applied to a graph, nothing more.
Self-check: Why does reverse-mode automatic differentiation need only one backward pass for all parameters?
Connects to: Gradient descent, first-order Taylor term.
Was this lecture useful?
BitsNotes AI Assistant
Subject Notes AssistantConfigure AI Chat
Choose how to access the chatbotSigned in as
Powered by BitsNotes — 20 messages per day. No API key needed. Want unlimited access? Use "Bring Your Own Key" mode.
Sign in to use AI Chat
Get 20 free AI messages per day to ask questions about your lecture notes. Sign in with Google or GitHub — it takes 5 seconds.
Sign In to BitsNotesSwitch to "Bring Your Own Key" tab above for unlimited access with any OpenAI-compatible provider.