Part of Quantitative Finance
Covers derivatives, gradients, and optimization techniques essential for quantitative finance. Topics include Greeks, portfolio optimization.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Differential Calculus and Optimization Basics
Quantitative finance often asks two practical questions: how does a modeled value respond when an input changes, and which feasible choice best meets a stated objective? A trader asking how much an option's model value changes when the stock moves by one dollar is asking a calculus question. A portfolio manager choosing an allocation to maximize a specified risk-adjusted objective is posing an optimization problem. Differential calculus describes local rates of change, while optimization compares feasible choices under stated constraints.
Sensitivity analysis and optimization often complement one another. Derivatives describe local changes, while first- and second-order conditions use those changes to characterize candidate optima when the required smoothness and constraint conditions hold. The gradient gives local directional information that a gradient-based method can use to search for improvement. Derivative-free and discrete methods handle problems where that information is unavailable or insufficient.
This chapter builds the calculus foundation you will use throughout this book. We begin with derivatives as measures of sensitivity, then extend to functions of multiple variables using partial derivatives and gradients. From there, we develop the machinery for finding optimal solutions, both unconstrained and constrained, ending with the Lagrange multiplier technique used in modern portfolio theory. Finally, we examine convexity. When minimizing a convex objective over a convex feasible set, every local minimum is global, although a numerical candidate still requires feasibility, convergence, and optimality checks.
Derivatives and Rates of Change
The derivative of a function measures its instantaneous rate of change. If represents a quantity that depends on , then the derivative tells us how rapidly changes as changes. In finance, this sensitivity analysis is essential. We constantly need to understand how outputs respond to changes in inputs.
To understand derivatives, consider what it means to measure change. If you know the value of your portfolio today and knew its value yesterday, you can compute how much it changed over that day. But that average change over a day may obscure important dynamics that occurred during trading hours. The derivative captures the limit of the average rate of change as the time interval tends to zero, when that limit exists.
The formal definition of the derivative captures this limiting process precisely. We start by computing the average rate of change over an interval of size , then we ask what happens to this average as becomes vanishingly small. When this limit exists and is well-defined, we have successfully captured the instantaneous rate of change.
The derivative of a function at a point is defined as the limit of the difference quotient:
where:
- : the function value at point
- : a small increment approaching zero
- : the change in function value over the interval
When this limit exists, we say is differentiable at . The derivative represents the slope of the tangent line to the function at that point.
The geometric interpretation of the derivative illuminates its meaning. If you graph the function and zoom in on any differentiable point, the curve begins to look more and more like a straight line. The derivative gives the slope of this line, the tangent line that best approximates the function at that point. This slope tells us the direction and steepness of the function's change. A positive derivative means the function is increasing; a negative derivative means it is decreasing. A zero derivative means the tangent is horizontal, so there is no first-order change at that point. It may signal a maximum or minimum, but by itself it does not determine whether the function is increasing or decreasing nearby.


Financial Interpretation of Derivatives
In finance, calculus derivatives provide local sensitivity measures. Consider a simple example: the profit function of a market maker.
Financial derivatives such as options and futures depend on an underlying asset, rate, index, or other reference. Calculus derivatives are one way to measure how a differentiable pricing model responds to a small change in that reference or another model input. For example, the sensitivity of an option pricing function to the underlying stock price is its partial derivative with respect to that price. Discrete scenarios, finite differences, and lattice methods can measure related changes without requiring an analytic derivative.
Suppose a market maker's daily profit depends on trading volume :
where:
- : daily profit in dollars
- : trading volume in units
- : revenue from bid-ask spread ($0.001 per unit)
- : fixed daily operating costs
- : market impact costs that grow quadratically with volume
This profit function captures three fundamental aspects of market making. The first term represents the revenue the market maker earns from the bid-ask spread: every time a unit is traded, the market maker captures a small spread of $0.001. If this were the only factor, profit would grow linearly forever. The constant term of 500 represents the fixed costs that must be paid regardless of volume: technology, salaries, rent, and regulatory fees. These costs must be covered before any profit is realized.
The quadratic term is an illustrative market-impact model rather than a universal empirical scaling law. Within this example, total impact cost grows with , so marginal impact cost grows linearly with volume and eventually offsets the fixed per-unit spread revenue. Actual impact depends on liquidity, execution horizon, order type, and market conditions; the negative quadratic coefficient is a modeling assumption chosen to create this trade-off.
The first term represents revenue (a small spread earned per unit volume), the constant represents fixed costs, and the quadratic term captures market impact costs that grow with volume. The derivative:
where:
- : the marginal profit at volume level
- : the marginal revenue per unit (the bid-ask spread)
- : the marginal market impact cost, which increases linearly with volume
This is the marginal profit, meaning the additional profit from one more unit of trading volume. When , increasing volume increases profit. When , the market impact costs outweigh the spread revenue.
Notice how the derivative transforms our understanding of the profit function. The original function tells us total profit at any volume level, but the derivative tells us whether we should increase or decrease our volume from the current level. This marginal perspective is precisely what an economist or trader needs for decision-making. We don't ask "what is our profit?" but rather "should we trade more or less?" The derivative answers this second, more actionable question.
# Define profit function and its derivative
def profit(V):
return 0.001 * V - 500 - 0.00000001 * V**2
def marginal_profit(V):
return 0.001 - 0.00000002 * V
# Find optimal volume where marginal profit = 0
optimal_volume = 0.001 / 0.00000002Optimal trading volume: 50,000 units Maximum profit at optimal volume: $-475.00 Marginal profit at optimal volume: 0.000000
At 50,000 units, the marginal profit is zero, indicating we have found the profit-maximizing volume. Beyond this point, each additional unit of volume reduces total profit due to market impact.
For this differentiable, strictly concave quadratic and its interior optimum, zero marginal profit means marginal revenue from one more trade (the bid-ask spread of $0.001) equals marginal market-impact cost. Below 50,000 units the derivative is positive; above it the derivative is negative, so this stationary point is the unique maximum. Marginal benefit equals marginal cost is a first-order condition for an interior differentiable choice, not a universal definition of an optimum: boundary, constrained, or nonsmooth problems require the corresponding optimality conditions.


Option Sensitivities: The Greeks
A major application of derivatives in finance is the calculation of option sensitivities, known as "the Greeks." These measure how an option's model price responds locally to changes in specified inputs while the other inputs are held fixed.
The Greeks earned their collective name because most of them are denoted by Greek letters. The five introduced here summarize several important local sensitivities under a specified pricing model and held-fixed inputs. They are useful for model-based pricing and hedging, but they neither guarantee exact dollar changes nor capture every exposure, such as jumps, volatility-surface dynamics, cross-Greeks, higher-order terms, transaction costs, or model error.
Consider a European call option with price that depends on the underlying stock price , time to expiration , volatility , and risk-free rate . Each Greek is a partial derivative of with respect to one of these inputs:
- Delta (): measures sensitivity to stock price changes
- Gamma (): measures the rate of change of delta
- Theta (): measures sensitivity to time decay
- Vega (): measures sensitivity to volatility changes
- Rho (): measures sensitivity to interest rate changes
where:
- : the call option price
- : the current stock price
- : time to expiration
- : the model volatility input; when inferred from an option price, it is the implied volatility for that specified contract, strike, maturity, and model
- : risk-free interest rate
Delta is widely used in option risk management. If an option quoted per share has a delta of 0.5, then a small $1 increase in the stock price changes the model price by approximately $0.50 when other inputs are fixed. Traders can offset an option portfolio's delta with stock or other options. A delta-neutral portfolio has net first-order underlying-price sensitivity near zero, so small stock moves have little first-order effect under the model; the hedge must be rebalanced as delta and other inputs change.
Gamma adds depth to the delta picture by measuring how delta itself changes. An option with high gamma has a delta that changes rapidly as the stock price moves. This is particularly important near expiration for near-the-money options. For a dynamically delta-hedged long-gamma position, sufficiently large realized moves can produce rebalancing gains, but those gains are not automatic and must be weighed against theta, transaction costs, jumps, and other exposures. A short-gamma position faces the opposite convexity and can require frequent hedge adjustments.
These Greeks provide a useful but incomplete local sensitivity summary under the chosen model. Delta and Gamma describe first- and second-order underlying-price effects; Theta describes passage-of-time sensitivity; Vega describes model-volatility sensitivity; and Rho describes interest-rate sensitivity. Hedging them addresses those modeled local exposures, not jumps, cross-effects, volatility-surface changes, liquidity costs, or model risk.
We will derive these explicitly when we cover the Black-Scholes model in a later chapter. For now, understand that each Greek answers a practical question: "If this input changes by a small amount, how much does my option position change in value?"


Multivariable Calculus
Financial models rarely depend on a single variable. A portfolio's return depends on the returns of all constituent assets. An option's value depends on price, time, volatility, and interest rates simultaneously. This requires extending calculus to functions of multiple variables.
The transition from single-variable to multivariable calculus is conceptually straightforward and practically important. In single-variable calculus, we had one direction of change. In multivariable calculus, we have infinitely many directions we could move. Understanding how the function changes in each of these directions requires new mathematical machinery.
Partial Derivatives
When a function depends on multiple variables, the partial derivative measures the rate of change with respect to one variable while holding all others constant.
The key conceptual insight is that partial derivatives treat other variables as if they were constants. If you want to know how portfolio return changes when you adjust the weight in one particular asset, you imagine all other weights frozen in place and observe how the return responds to that single change. This "one variable at a time" approach simplifies the analysis but also limits what we can learn. The true behavior of the function involves simultaneous changes in multiple variables, which we will address with gradients and directional derivatives.
For a function , the partial derivative with respect to is:
where:
- : the variable with respect to which we differentiate
- : a small increment approaching zero
- : all other variables, held constant during differentiation
We treat all variables except as constants and differentiate normally.
The notation uses the curly "partial" symbol ∂ rather than the regular "d" used in single-variable calculus. This notational distinction is a reminder that we are holding other variables fixed. The computation itself follows the same rules as ordinary differentiation: we simply treat the other variables as constants and differentiate with respect to the variable of interest.
Consider a simple two-asset portfolio where total return depends on the returns and of each asset and their weights and :
The partial derivatives tell us how sensitive total return is to each input:
The first equation says that the sensitivity of portfolio return to the first weight equals the return of that asset. The second says that sensitivity to the first asset's return equals its weight in the portfolio. Both are intuitive, but partial derivatives make this reasoning precise.
These results have immediate practical interpretations. Holding fixed, increasing by 0.01 changes portfolio return by but also increases net exposure and therefore requires additional financing. In a fully invested two-asset portfolio, , so the relevant reallocation derivative is : shifting 0.01 from asset 2 to asset 1 changes return by . Separately, if asset 1's return changes by 0.01 while its weight is 0.40, portfolio return changes by 0.004, or 0.4 percentage points. The same local linearization idea also appears in factor-based risk decompositions.
The Gradient Vector
Collecting all partial derivatives of a function into a single vector gives us the gradient, a fundamental object in optimization.
The gradient represents the multivariable generalization of the derivative. While a single-variable derivative tells us the rate of change in the only direction available, the gradient in multiple dimensions tells us how to combine changes in each direction to achieve the maximum rate of increase. This makes the gradient more than a collection of sensitivities; it is a directional guide pointing toward improvement.
The gradient of a function is the vector of all partial derivatives:
where:
- : the gradient operator (nabla) applied to , producing a vector
- : the partial derivative of with respect to the -th variable
- : the dimension of the input space
When is nonzero, it points in the direction of steepest ascent among Euclidean unit directions.
The gradient's direction has a precise interpretation when it is nonzero: among Euclidean unit directions, maximizes the directional derivative. Its magnitude is that maximum first-order rate of increase. This property makes the gradient the workhorse of optimization algorithms, as we will see shortly.
To understand this result, consider the directional derivative for a Euclidean unit vector . The Cauchy-Schwarz inequality gives , with equality in the normalized gradient direction when the gradient is nonzero; the opposite direction gives steepest descent. If the gradient is zero, every first-order directional derivative is zero, so no unique steepest direction exists.
The magnitude of the gradient, computed as , tells us how rapidly the function is changing in the steepest direction. A large gradient magnitude indicates a region where the function is changing rapidly; a small magnitude indicates a relatively flat region. When the gradient magnitude is exactly zero, we have reached a critical point where the function has no preferred direction of change, which typically indicates a local minimum, maximum, or saddle point.
import numpy as np
def f(x, y):
"""A quadratic function of two variables"""
return x**2 + 2 * y**2 - 2 * x * y + 4 * x - 6 * y
def gradient_f(x, y):
"""Gradient of f computed analytically"""
df_dx = 2 * x - 2 * y + 4
df_dy = 4 * y - 2 * x - 6
return np.array([df_dx, df_dy])
# Evaluate gradient at a specific point
point = np.array([1.0, 2.0])
grad = gradient_f(point[0], point[1])Function value at (1, 2): -3.00 Gradient at (1, 2): [2.00, 0.00] Gradient magnitude: 2.00
The gradient at (1, 2) is [2, 0], meaning the function increases most rapidly in the positive direction at that point, with no change in the direction. This tells us that from (1, 2), moving in the positive direction increases the function value, while moving in the positive direction has no first-order effect.
The fact that the -component of the gradient is zero at this point reveals that we are on a "ridge" or "valley" in the -direction. The function is neither increasing nor decreasing as we vary while holding fixed at 1. This doesn't mean we're at the optimum; it just means we're at a stationary point with respect to at this particular -value. To find the true minimum, we need to find a point where both components of the gradient are zero simultaneously.

Numerical Differentiation
While analytical derivatives are preferred when available, we often need to compute derivatives numerically, especially for complex models or when working with black-box functions.
In many practical situations, we have access to a function only through evaluations. We can plug in values and observe outputs, but we don't have a closed-form expression that allows symbolic differentiation. This occurs frequently when dealing with simulation-based models, legacy software systems, or complex financial instruments where the pricing function involves numerous nested calculations. Numerical differentiation provides a way to approximate derivatives using only function evaluations.
The simplest approach uses the finite difference approximation:
However, the central difference formula provides better accuracy:
where:
- : the derivative we are approximating
- : a small step size (often around to for central differences in double precision, depending on scaling)
- and : function evaluations at points symmetric around
The central difference achieves accuracy compared to for the forward difference, because the symmetric evaluation cancels first-order error terms. In practice, choosing involves a trade-off: too large introduces truncation error, too small introduces floating-point rounding error.
The error analysis behind these approximations reveals why the central difference is superior. In the Taylor expansions of and , the first-order terms have opposite signs, so subtraction produces . The even-order terms have the same signs and cancel. In particular,
when the required derivatives exist near . Thus the central quotient has truncation error, whereas the forward difference generally has truncation error.
The choice of step size represents a fundamental tension in numerical computation. A smaller reduces truncation error because we're better approximating the limit as . But computers store numbers with finite precision, and when becomes too small, the difference involves subtracting nearly equal numbers, which amplifies rounding errors. The optimal balances these effects and often falls around to for central differences in double-precision floating-point numbers, depending on scaling.
def numerical_gradient(f, x, h=1e-6):
"""
Compute gradient numerically using central differences.
Parameters:
f: function that takes array x and returns scalar
x: point at which to evaluate gradient (numpy array)
h: step size for finite differences
Returns:
gradient vector (numpy array)
"""
n = len(x)
grad = np.zeros(n)
for i in range(n):
x_plus = x.copy()
x_minus = x.copy()
x_plus[i] += h
x_minus[i] -= h
grad[i] = (f(x_plus) - f(x_minus)) / (2 * h)
return grad
# Define gradient_f for use in this block
def gradient_f(x, y):
"""Gradient of f computed analytically"""
df_dx = 2 * x - 2 * y + 4
df_dy = 4 * y - 2 * x - 6
return np.array([df_dx, df_dy])
# Test numerical vs analytical gradient
def f_array(x):
"""Wrapper function that takes array input for numerical gradient."""
return x[0] ** 2 + 2 * x[1] ** 2 - 2 * x[0] * x[1] + 4 * x[0] - 6 * x[1]
point = np.array([1.0, 2.0])
numerical_grad = numerical_gradient(f_array, point)
analytical_grad = gradient_f(point[0], point[1])Analytical gradient: [2. 0.] Numerical gradient: [2. 0.] Difference: [2.79555934e-10 0.00000000e+00]
The numerical gradient matches the analytical result to within floating-point precision, validating our implementation. This technique becomes invaluable when analytical derivatives are unavailable or too complex to derive by hand.
Unconstrained Optimization
Optimization is the process of finding the best solution among all feasible alternatives. In unconstrained optimization, we seek to minimize or maximize a function without restrictions on the input variables.
The goal of optimization is to find the point or points where a function achieves its best value. "Best" might mean smallest (for costs and risks) or largest (for returns and profits). In unconstrained optimization, we are free to choose any values for the input variables; there are no boundaries or restrictions to respect. While this may seem simpler than constrained optimization, the principles we develop here form the foundation for handling constraints as well.
First-Order Conditions
The first step in finding an optimum is identifying critical points, where the gradient equals zero.
The first-order condition follows from the gradient's directional interpretation. At a point where the function achieves a maximum or minimum, you cannot improve by moving in any direction. If the gradient were nonzero, it would point toward a direction of ascent, meaning you could increase the function by moving that way. At a minimum, there's no direction of descent available; at a maximum, there's no direction of ascent. The only way both can fail to exist is if the gradient is the zero vector.
If has a local minimum or maximum at an interior point , and is differentiable at , then:
Points satisfying this condition are called critical points or stationary points.
It is essential to understand that the first-order condition is necessary but not sufficient for an optimum. A zero gradient tells us we have found a critical point, but not all critical points are minima or maxima. Some are saddle points, where the function increases in certain directions and decreases in others. Think of a mountain pass: it is a minimum when you traverse the pass but a maximum when you go perpendicular to it. Both directions through this critical point have zero derivative, yet the point is neither a global maximum nor a global minimum.
Setting the gradient to zero gives us a system of equations. For our earlier quadratic function:
Solving this system:
from scipy.optimize import fsolve
def gradient_system(vars):
x, y = vars
return [2 * x - 2 * y + 4, 4 * y - 2 * x - 6]
critical_point = fsolve(gradient_system, [0, 0])Critical point: x = -1.0000, y = 1.0000 Function value at critical point: -5.0000 Gradient at critical point: [0. 0.]
The critical point is at (-1, 1), where the gradient is indeed zero.
Second-Order Conditions and the Hessian
A zero gradient tells us we have found a critical point, but not whether it is a minimum, maximum, or saddle point. To classify critical points, we need second-order information encoded in the Hessian matrix.
The Hessian matrix extends the concept of the second derivative to multiple dimensions. Just as the second derivative of a single-variable function tells us about the function's curvature (whether it curves upward or downward), the Hessian tells us about curvature in all directions simultaneously. Because there are infinitely many directions in multidimensional space, this information is naturally organized into a matrix rather than a single number.
The Hessian matrix of a function is the matrix of second partial derivatives:
For functions with continuous second derivatives, the Hessian is symmetric.
The diagonal entries of the Hessian capture the pure second derivatives: how the rate of change with respect to one variable itself changes as that variable varies. The off-diagonal entries capture the mixed partial derivatives: how the rate of change with respect to one variable changes as another variable varies. The symmetry of the Hessian, known as the equality of mixed partials or Clairaut's theorem, holds for functions with continuous second derivatives and reflects the fundamental fact that the order of differentiation doesn't matter for such functions.
At a critical point of a twice-differentiable function, a definite or indefinite Hessian gives the nondegenerate second-order classification below. If any eigenvalue is zero and the remaining eigenvalues do not already have mixed signs, the quadratic test is inconclusive and higher-order or other analysis is required.
- All eigenvalues positive: strict local minimum
- All eigenvalues negative: strict local maximum
- Mixed signs: saddle point
The connection between eigenvalues and the nature of critical points comes from a deep result in linear algebra. The eigenvalues of the Hessian represent the curvatures of the function along the principal directions (the eigenvectors). If all these curvatures are positive, the function curves upward in every direction, creating a bowl shape characteristic of a minimum. If all are negative, it curves downward in every direction, creating an inverted bowl characteristic of a maximum. Mixed signs create the saddle shape, curving up in some directions and down in others.
For our quadratic function, the Hessian is constant:
import numpy as np
# Compute Hessian analytically (constant for quadratic)
# For f(x,y) = x² + 2y² - 2xy + 4x - 6y:
# ∂²f/∂x² = 2, ∂²f/∂y² = 4, ∂²f/∂x∂y = -2
H = np.array([[2, -2], [-2, 4]])
# Find eigenvalues to classify the critical point
eigenvalues = np.linalg.eigvals(H)Hessian matrix: [[ 2 -2] [-2 4]] Eigenvalues: [0.76393202 5.23606798] All eigenvalues positive: True Classification: Local minimum
Both eigenvalues are positive, confirming that the critical point (-1, 1) is a local minimum. Since the function is quadratic with a positive definite Hessian, this minimum is also global.
The fact that the Hessian is constant for this function is a special property of quadratic functions. The second derivatives of a quadratic function are constants because taking two derivatives of a quadratic term yields a constant. This constancy means the curvature is the same everywhere, which simplifies both analysis and computation. For more general functions, the Hessian varies from point to point, and we must evaluate it specifically at the critical point to determine the point's nature.




Gradient Descent
Gradient descent is a foundational first-order optimization algorithm. Since the gradient points uphill, we find a minimum by repeatedly stepping in the opposite direction.
Gradient descent is conceptually straightforward. At any point, we compute the gradient to determine the direction of steepest ascent, then move in the opposite direction to descend toward lower values. By repeating this process, we trace out a path that, under appropriate conditions, leads to a minimum. This iterative approach is particularly valuable for high-dimensional problems where analytical solutions are impossible or impractical.
The update rule is:
where:
- : the current iterate at step
- : the next iterate after the update
- : the learning rate (step size), controlling how far we move in each iteration
- : the gradient evaluated at the current point, indicating the direction of steepest ascent
The negative sign makes the update descend: we subtract the gradient because we want to move downhill, not uphill. The learning rate scales the step size, determining how far we move in the descent direction. This single parameter has an outsized influence on the algorithm's behavior.
The choice of controls whether the algorithm converges. If it is too large, the algorithm may overshoot and diverge. If it is too small, convergence becomes painfully slow.
The learning rate embodies a fundamental trade-off in optimization. A large learning rate allows rapid progress when far from the minimum, but risks overshooting when close. Imagine trying to land on a narrow valley floor: large steps might cause you to bounce from one side to the other, never settling down. A small learning rate ensures stable progress but may require many iterations to reach the minimum, particularly in flat regions where the gradient is small. Advanced variants of gradient descent, such as momentum methods and adaptive learning rates, address these issues by adjusting the step size dynamically.
import numpy as np
def gradient_descent(f, grad_f, x0, learning_rate=0.1, max_iter=200, tol=1e-6):
"""
Minimize f using gradient descent.
Parameters:
f: objective function
grad_f: gradient function
x0: starting point
learning_rate: step size
max_iter: maximum iterations
tol: convergence tolerance on gradient norm
Returns:
x_history: list of iterates
f_history: list of function values
"""
x = x0.copy()
x_history = [x.copy()]
f_history = [f(x)]
for i in range(max_iter):
grad = grad_f(x)
# Check convergence
if np.linalg.norm(grad) < tol:
break
# Update step
x = x - learning_rate * grad
x_history.append(x.copy())
f_history.append(f(x))
return np.array(x_history), np.array(f_history)
# Run gradient descent on our quadratic function
def f_vec(x):
return x[0] ** 2 + 2 * x[1] ** 2 - 2 * x[0] * x[1] + 4 * x[0] - 6 * x[1]
def grad_f_vec(x):
return np.array([2 * x[0] - 2 * x[1] + 4, 4 * x[1] - 2 * x[0] - 6])
x0 = np.array([3.0, 0.0])
x_history, f_history = gradient_descent(
f_vec, grad_f_vec, x0, learning_rate=0.15
)Starting point: [3. 0.] Final point: [-0.999999, 1.000001] True minimum: [-1, 1] Iterations run before ||gradient|| < 1e-6: 120 Final function value: -5.000000


For this smooth strongly convex quadratic, the Hessian's largest eigenvalue is about 5.236, so the fixed step lies within the stable interval . With the 200-iteration cap, the run satisfies the gradient-norm tolerance before stopping and exhibits linear convergence. Convexity alone does not make an arbitrary step size converge.



Constrained Optimization and Lagrange Multipliers
Many financial optimization problems involve constraints. An unlevered, fully invested mandate may require portfolio weights to sum to 100%, while a leveraged or margin-financed portfolio can have gross exposure above 100% of capital. A risk mandate may cap modeled volatility, and a trading mandate may impose position limits. These rules lead to constrained optimization problems.
Constraints reflect the rules of a financial decision. Regulations or mandates may prohibit specified positions, risk limits may restrict concentration, and a budget constraint limits admissible spending or leverage. Constrained optimization finds the best objective value within the resulting feasible set. Sensitivity analysis can separately ask how that optimum would change if a constraint were relaxed slightly.
The Lagrangian Method
The method of Lagrange multipliers handles equality constraints by converting a constrained problem into an unconstrained one. Consider minimizing subject to the constraint .
Lagrange's approach augments the objective with a free multiplier rather than a fixed penalty. Differentiating the Lagrangian with respect to recovers , while differentiation with respect to encodes the required gradient alignment. Solving these stationarity equations produces constrained candidates, not guaranteed optima; each candidate still needs feasibility, constraint-qualification, and second-order or global optimality checks.
The Lagrangian combines the objective function and constraints:
where:
- : the Lagrangian function
- : the decision variables we are optimizing
- : the Lagrange multiplier, measuring the sensitivity of the optimum to the constraint
- : the objective function we want to minimize or maximize
- : the equality constraint that must be satisfied
At a constrained optimum, the gradient of the Lagrangian with respect to all variables (including ) equals zero.
The key insight is geometric. At a constrained optimum, the gradient of the objective function must be parallel to the gradient of the constraint function. If they were not parallel, you could move along the constraint surface in a direction that decreases the objective.
To understand this geometric insight more deeply, imagine standing on a hill (the objective function) while constrained to walk along a path (the constraint surface). At the optimal point along this path, the steepest uphill direction on the original surface points directly toward or away from the path itself; there is no component along the path. If there were such a component, you could walk along the path in that direction and climb higher on the hill, contradicting the optimality of your current position. This perpendicularity between the gradient and the constraint surface is precisely what the Lagrange conditions capture.
Mathematically, this means:
Combined with the constraint , these equations form a system we can solve for both the optimal point and the multiplier .
The system consists of equations in the components of and the multiplier . A nonsingular Jacobian at a solution can make that solution locally isolated, but matching the number of equations and unknowns does not guarantee existence or a unique global solution for a nonlinear system. Each candidate still requires constraint-qualification, feasibility, and second-order or global optimality checks.

Worked Example: Portfolio Allocation
Consider allocating between two assets to minimize portfolio variance, subject to achieving a target expected return.
This example studies the mean-variance trade-off within a specified asset universe and constraint set. For each feasible target expected return, the model selects the portfolio with the lowest variance. Along the upper efficient branch beyond the global-minimum-variance point, increasing the target return raises that minimum attainable variance. Outside this model and opportunity set, a portfolio can be dominated by another portfolio with both higher expected return and lower risk.
Let and be the weights in assets 1 and 2, with expected returns and , variances and , and correlation . The portfolio variance is:
where:
- : portfolio variance
- : portfolio weights for assets 1 and 2
- : variances of assets 1 and 2
- : correlation between the two assets
- : covariance contribution from the interaction between assets
This variance formula reveals the mathematics of diversification. The first two terms represent the variance contributions from each asset individually, scaled by the square of their weights. The third term captures the interaction between assets through their covariance. When correlation is less than 1, this interaction term is smaller than it would be for perfectly correlated assets, allowing the portfolio variance to be less than the weighted average of individual variances. This reduction is the diversification benefit.
We want to minimize subject to:
- Target return:
- Full investment:
The Lagrangian is:
where:
- : Lagrange multiplier for the return constraint, representing the marginal variance cost of achieving additional return
- : multiplier for the net-budget constraint; its sign depends on the Lagrangian convention, and its magnitude is a local objective sensitivity per unit change in the constraint's right-hand side
The two constraints serve different purposes. The return equality fixes expected return. The equation fixes net exposure and implies no net cash position, but it does not rule out leveraged long-short portfolios: for example, weights sum to one and have gross exposure three. Nonnegative weights or an explicit gross-exposure constraint are needed to exclude leverage. With normalized weights, has variance units per unit change in the budget right-hand side; it is not directly a value per dollar unless the problem is formulated in dollar holdings and currency units.
from scipy.optimize import minimize
# Asset parameters
mu = np.array([0.10, 0.05]) # Expected returns
sigma = np.array([0.20, 0.10]) # Standard deviations
rho = 0.3 # Correlation
target_return = 0.08
# Covariance matrix
cov_matrix = np.array(
[
[sigma[0] ** 2, rho * sigma[0] * sigma[1]],
[rho * sigma[0] * sigma[1], sigma[1] ** 2],
]
)
def portfolio_variance(w):
"""Portfolio variance given weights w"""
return w @ cov_matrix @ w
def portfolio_return(w):
"""Portfolio expected return given weights w"""
return w @ mu
# Constraints
constraints = [
{
"type": "eq",
"fun": lambda w: portfolio_return(w) - target_return,
}, # Target return
{"type": "eq", "fun": lambda w: np.sum(w) - 1}, # Full investment
]
# Initial guess
w0 = np.array([0.5, 0.5])
# Optimize
result = minimize(
portfolio_variance, w0, method="SLSQP", constraints=constraints
)
optimal_weights = result.xOptimal Portfolio Allocation ------------------------------ Weight in Asset 1 (μ=10%, σ=20%): 0.6000 Weight in Asset 2 (μ=5%, σ=10%): 0.4000 Portfolio Statistics ------------------------------ Expected return: 0.0800 (8.00%) Portfolio variance: 0.018880 Portfolio std dev: 0.1374 (13.74%)
The full-investment and 8% target-return equalities determine a 60% allocation to the higher-return, higher-risk asset and a 40% allocation to the lower-return, lower-risk asset before variance enters the calculation. With two assets, this is the only feasible portfolio at that target, so it is also the minimum-variance portfolio for that target.
Interpreting Lagrange Multipliers
Under regularity conditions, a Lagrange multiplier has a local envelope interpretation: after fixing the constraint's right-hand-side parameterization and the Lagrangian sign convention, the corresponding signed derivative gives the first-order change in the optimal objective value.
For a binding constraint, the correctly signed envelope derivative describes only a local, first-order response around the current optimum. A large magnitude means the optimized objective is locally sensitive to that specified right-hand-side change; whether this is a benefit or a cost depends on whether the problem minimizes or maximizes, how relaxation is parameterized, and the multiplier sign convention. A small magnitude likewise indicates local insensitivity, not that the constraint is globally unimportant.
For our minimization problem and the convention , the return multiplier equals the local derivative of minimum variance with respect to decimal target return. Its sign would reverse under the opposite Lagrangian convention.
Suppose the return-constraint multiplier is 0.02 variance units per unit of decimal expected return. Raising the target from 8% to 8.1% changes decimal return by 0.001, so the first-order variance change is approximately . The corresponding standard-deviation change is nonlinear and must be computed from the square root rather than assigned the same units.
# Examine how variance changes with target return
# Trace the upper frontier by solving minimum variance at each target.
gmv_return = (
mu
@ np.linalg.solve(cov_matrix, np.ones(len(mu)))
/ (np.ones(len(mu)) @ np.linalg.solve(cov_matrix, np.ones(len(mu))))
)
target_returns = np.linspace(gmv_return, 0.10, 50)
min_variances = []
for target in target_returns:
constraints = [
{"type": "eq", "fun": lambda w, t=target: portfolio_return(w) - t},
{"type": "eq", "fun": lambda w: np.sum(w) - 1},
]
result = minimize(
portfolio_variance, w0, method="SLSQP", constraints=constraints
)
min_variances.append(result.fun)
min_variances = np.array(min_variances)
min_std = np.sqrt(min_variances)
The plotted curve is the upper, mean-variance-efficient branch. It begins at the global-minimum-variance portfolio, whose expected return is about 5.53%; targets below that point lie on a dominated lower branch and are omitted. In this two-asset fully invested example, the budget and exact return equalities determine one portfolio for each target. The return-constraint multiplier measures the local change in minimum variance as the target increases.


Convexity
Convexity is a key structural property in optimization. When a function is convex, any local minimum is automatically a global minimum, dramatically simplifying the search for optimal solutions.
Convexity provides essential guarantees in optimization. A non-convex objective can have many valleys, and gradient-based methods can stop in any of them with no guarantee of finding the deepest one. In convex optimization, there is a single global valley (sometimes with a flat bottom), and any local minimum you reach is globally optimal. Under standard finite-dimensional representations and regularity assumptions, this structure often makes global optimization algorithmically tractable. Actual computational cost still depends on the problem representation, dimension, oracle access, conditioning, constraints, and requested accuracy.
Definition and Intuition
The formal definition of convexity captures a simple geometric idea. A function is convex if the line segment connecting any two points on its graph lies above the graph itself. Alternatively, if you interpolate between two input points and evaluate the function at the interpolated point, you get a value no larger than if you had interpolated between the function values at the original points.
A function is convex if for any two points and and any :
where:
- : any two points in the domain of
- : a convex combination weight (interpolation parameter)
- : a point on the line segment between and
- : the corresponding point on the chord connecting and
Geometrically, this means the line segment connecting any two points on the graph lies above the graph itself.
The definition says that interpolating between any two points in the domain and evaluating the function gives a value at most as large as interpolating between the function values at those points. Intuitively, a convex function "curves upward" everywhere.
Consider a bowl placed on a table. If you pick any two points on the rim of the bowl and stretch a string between them, the string remains above or on the surface of the bowl everywhere along its length. This is the defining characteristic of convexity. A non-convex surface might dip below such a string, creating regions where the surface rises above the interpolated line.
The parameter can be thought of as a "mixing" proportion. When , we are entirely at point ; when , we are entirely at point ; and intermediate values of give us intermediate points along the segment. The convexity inequality states that the function value at any such intermediate point is at most the correspondingly weighted average of the function values at the endpoints.


Testing for Convexity
For twice-differentiable functions on an open convex domain, convexity can be checked using the Hessian matrix.
A function on such a domain is convex if and only if its Hessian is positive semidefinite at every point. Checking eigenvalues at only one point establishes local curvature there, not global convexity.
The function is strictly convex if is positive definite everywhere (all eigenvalues ).
The Hessian captures how the gradient changes as we move through the space. Positive semidefiniteness everywhere means the function curves upward or remains flat in every direction, ruling out suboptimal local minima. It is not by itself an algorithmic convergence guarantee: for a differentiable convex function with an -Lipschitz gradient and an attained minimum, fixed-step gradient descent also needs an appropriate step such as .
To understand why the Hessian condition characterizes convexity, recall that the Hessian appears in the second-order Taylor expansion of the function. Around any point , the function behaves approximately as a quadratic: . The quadratic term determines the local curvature. For a positive semidefinite , this term is always non-negative, meaning the function curves upward or remains flat in every direction. This local property, holding everywhere, implies global convexity.
For our quadratic portfolio variance function:
where is the covariance matrix. The Hessian of this function is:
Since covariance matrices are positive semidefinite by construction (variances of portfolios cannot be negative), the portfolio variance function is convex. This guarantees that the minimum-variance portfolio we found is the global minimum.
This result is mathematically elegant and practically significant. When we solve for the minimum-variance portfolio, we know we have found the truly optimal solution rather than a merely local optimum. There are no hidden better solutions lurking elsewhere in the feasible region. This guarantee of global optimality provides confidence in the solution and justifies the widespread use of mean-variance optimization in portfolio management.
import numpy as np
# Define covariance matrix (from earlier portfolio example)
sigma = np.array([0.20, 0.10]) # Standard deviations
rho = 0.3 # Correlation
cov_matrix = np.array(
[
[sigma[0] ** 2, rho * sigma[0] * sigma[1]],
[rho * sigma[0] * sigma[1], sigma[1] ** 2],
]
)
# Verify convexity of portfolio variance
# Hessian is 2 * covariance matrix
H_portfolio = 2 * cov_matrix
eigenvalues_portfolio = np.linalg.eigvals(H_portfolio)Covariance matrix: [[0.04 0.006] [0.006 0.01 ]] Hessian of variance function (2Σ): [[0.08 0.012] [0.012 0.02 ]] Eigenvalues of Hessian: [0.08231099 0.01768901] All eigenvalues ≥ 0: True Conclusion: Portfolio variance is convex
Why Convexity Matters in Finance
Convexity has significant practical implications in quantitative finance:
Portfolio optimization can be convex. With a positive-semidefinite covariance matrix and a convex feasible set such as linear equalities and inequalities, mean-variance minimization is a convex program. Every local minimum is then global, but a solver's returned point still requires success, feasibility, and numerical-tolerance checks.
Convexity removes suboptimal local-minimum traps from this formulation, but it does not guarantee that an arbitrary numerical run finds or accurately certifies a solution. Conditioning, tolerances, infeasibility, and implementation details still matter. Nonconvexity can introduce multiple basins and make global certification harder; its computational cost depends on the problem's structure rather than following a universal exponential search rule.
Convex risk measures reward mixing relative to an average benchmark. For a convex risk functional, the risk of a weighted combination cannot exceed the same weighted average of the component risks. Equality can occur, and the combination can still be riskier than one component; convexity does not promise a strict reduction in every comparison. Value-at-Risk is not convex in general.
If is a convex risk functional and we combine two portfolios and with weights and , then . The inequality supplies an upper bound relative to the weighted average. Whether the mixture is strictly safer than either component depends on the positions, dependence structure, weights, and risk functional.
Non-convexity creates computational challenges. When objective functions are non-convex, optimization becomes much harder. Multiple local minima may exist. Gradient-based methods can get stuck. Problems involving integer constraints (such as selecting a discrete number of assets) or certain complex risk measures lose the convexity guarantee, requiring more sophisticated solution techniques.
Non-convex problems arise frequently in practice despite the convenience of convex formulations. Transaction costs with fixed components (minimum fees regardless of trade size) introduce non-convexity. Cardinality constraints (hold at most 30 stocks) are inherently non-convex. Some risk measures like Value-at-Risk under certain distributions violate convexity. Depending on the structure and required guarantee, practitioners may use local methods, multi-start searches, relaxations, exact discrete algorithms, or global heuristics; heuristic methods do not by themselves certify a global optimum.


Practical Implementation with SciPy
SciPy's optimize module implements methods for many smooth unconstrained and constrained examples. Suitability for a financial application depends on the objective, constraints, derivatives, scaling, conditioning, and the guarantees required; always inspect solver status and residuals.
import numpy as np
from scipy.optimize import minimize
# Example 1: Unconstrained optimization
def rosenbrock(x):
"""The Rosenbrock function - a classic test function.
Known for its curved valley that makes optimization challenging.
Global minimum at x = [1, 1, ..., 1] with f(x*) = 0.
"""
return sum(100.0 * (x[1:] - x[:-1] ** 2.0) ** 2.0 + (1 - x[:-1]) ** 2.0)
def rosenbrock_gradient(x):
"""Gradient for the two-dimensional Rosenbrock example."""
return np.array(
[
-400 * x[0] * (x[1] - x[0] ** 2) - 2 * (1 - x[0]),
200 * (x[1] - x[0] ** 2),
]
)
# Different optimization methods
x0 = np.array([-1.0, -1.0])
results = {}
for method in ["Nelder-Mead", "BFGS", "CG"]:
result = minimize(rosenbrock, x0, method=method)
results[method] = {
"x": result.x,
"fun": result.fun,
"nfev": result.nfev,
"success": result.success,
"message": str(result.message),
"gradient_norm": np.linalg.norm(rosenbrock_gradient(result.x)),
}Optimization Results on Rosenbrock Function ================================================== True minimum: [1, 1], f(x*) = 0 Nelder-Mead: Solution: [0.999999, 0.999995] Function value: 5.31e-10 Function evaluations: 125 Solver success: True (Optimization terminated successfully.) Gradient norm: 1.03e-03 BFGS: Solution: [0.999996, 0.999991] Function value: 2.00e-11 Function evaluations: 120 Solver success: True (Optimization terminated successfully.) Gradient norm: 5.92e-06 CG: Solution: [0.999997, 0.999995] Function value: 7.46e-12 Function evaluations: 210 Solver success: False (Desired error not necessarily achieved due to precision loss.) Gradient norm: 4.19e-05
Different optimization algorithms have different strengths. BFGS is a quasi-Newton method that uses first derivatives and builds an inverse-Hessian approximation. Nelder-Mead is derivative-free. For smooth problems with trustworthy derivatives, derivative-based methods often need fewer function evaluations, but relative performance depends on scaling, dimension, noise, initialization, and stopping tolerances.
Handling Bounds and Constraints
Real financial problems often involve bounds (no short selling: ) and multiple constraints (budget, sector limits, etc.).
from scipy.optimize import Bounds
# Three-asset portfolio optimization with constraints
n_assets = 3
mu_3 = np.array([0.12, 0.08, 0.05])
sigma_3 = np.array([0.22, 0.15, 0.08])
corr_matrix = np.array([[1.0, 0.4, 0.2], [0.4, 1.0, 0.3], [0.2, 0.3, 1.0]])
cov_3 = np.outer(sigma_3, sigma_3) * corr_matrix
def port_var_3(w):
return w @ cov_3 @ w
def port_ret_3(w):
return w @ mu_3
# Constraints
target_ret = 0.09
constraints = [
{"type": "eq", "fun": lambda w: np.sum(w) - 1}, # Fully invested
{
"type": "eq",
"fun": lambda w: port_ret_3(w) - target_ret,
}, # Target return
]
# Bounds: no short selling (all weights >= 0)
bounds = Bounds(lb=np.zeros(n_assets), ub=np.ones(n_assets))
# Optimize
w0 = np.ones(n_assets) / n_assets
result = minimize(
port_var_3,
w0,
method="SLSQP",
bounds=bounds,
constraints=constraints,
options={"ftol": 1e-12, "maxiter": 500},
)
w_optimal = result.x
# Check solver status, feasibility, bounds, and first-order stationarity.
equality_residuals = np.array(
[constraint["fun"](w_optimal) for constraint in constraints]
)
lower_bound_violation = np.maximum(bounds.lb - w_optimal, 0).max()
upper_bound_violation = np.maximum(w_optimal - bounds.ub, 0).max()
constraint_jacobian = np.vstack([np.ones(n_assets), mu_3])
objective_gradient = 2 * cov_3 @ w_optimal
lagrange_multipliers = np.linalg.lstsq(
constraint_jacobian.T, -objective_gradient, rcond=None
)[0]
stationarity_residual = np.linalg.norm(
objective_gradient + constraint_jacobian.T @ lagrange_multipliers
)
checks_passed = (
result.success
and np.max(np.abs(equality_residuals)) < 1e-8
and max(lower_bound_violation, upper_bound_violation) < 1e-10
and stationarity_residual < 1e-8
)
if not checks_passed:
raise RuntimeError(
f"Portfolio optimization checks failed: {result.message}"
)Three-Asset Portfolio Optimization (No Short Sales) ================================================== Solver success: True (Optimization terminated successfully) Maximum equality residual: 4.94e-13 Maximum bound violation: 0.00e+00 First-order stationarity residual: 5.05e-11 Asset Parameters: Asset Return Std Dev 1 12.0% 22.0% 2 8.0% 15.0% 3 5.0% 8.0% Optimal Weights for 9% Target Return: Asset 1: 43.86% Asset 2: 31.00% Asset 3: 25.14% Portfolio Statistics: Expected Return: 9.00% Standard Deviation: 12.96% Sharpe Ratio (rf=2%): 0.540
The successful solver result allocates across all three assets. The reported equality, bound, and first-order residuals verify this interior solution before we describe it as the minimum-variance portfolio for the 9% target under the stated constraints.
Key Parameters
The key parameters for portfolio optimization are:
- μ (mu): Expected returns vector. These estimates affect the return constraint or objective, but the direction of an asset's optimal-weight response also depends on covariance estimates, the other assets, and the active constraints.
- Σ (Sigma): Covariance matrix capturing asset variances and correlations. Holding weights and marginal volatilities fixed, a lower pairwise correlation reduces that pair's covariance contribution to portfolio variance.
- Target return: The required portfolio return, which determines the position on the efficient frontier.
- Bounds: Constraints on individual weights (e.g., no short selling requires w ≥ 0).
- Learning rate (α): For gradient descent, controls step size. A sufficiently large rate can cause divergence; a small rate slows convergence.
Limitations and Practical Considerations
While calculus and optimization are useful tools for quantitative finance, practitioners should know their limitations.
Model sensitivity. Optimization results can be highly sensitive to input parameters. In portfolio optimization, small changes in expected returns or covariances can produce dramatically different optimal allocations. This phenomenon (called estimation error amplification) means that optimizers may act as "error maximizers" when inputs are estimated with uncertainty. Techniques like shrinkage estimation, robust optimization, and resampling can help mitigate this sensitivity.
Local vs. global optima. A convex objective over a convex feasible set has no suboptimal local minima, while nonconvex problems may have multiple basins. Gradient descent can then depend on initialization. Depending on problem structure and the guarantee required, multi-start local search, relaxations, exact methods, simulated annealing, or genetic algorithms may be useful, but heuristics generally do not certify a global optimum.
Numerical precision. Finite precision arithmetic can cause issues in optimization, especially near singular or ill-conditioned matrices. Covariance matrices estimated from returns can become nearly singular when the number of assets approaches the number of observations. Regularization techniques and careful numerical implementations help maintain stability.
Constraints in practice. Real trading constraints are often more complex than simple equality or inequality constraints. Transaction costs create path-dependence. Integer constraints (minimum lot sizes) make problems combinatorially hard. Regulatory constraints may involve complex interactions between positions. These practical complexities often require specialized algorithms beyond standard calculus-based optimization.
Despite these limitations, the framework in this chapter supports portfolio management, derivatives pricing, and risk management. Derivatives provide local sensitivities, gradients provide local search directions, and convexity makes every local minimum global only when the objective and feasible set meet the stated convexity conditions; numerical solutions still require validation.
Summary
This chapter established the calculus and optimization foundations for quantitative finance:
Derivatives as sensitivities: The derivative measures how a function changes in response to its inputs. In finance, derivatives quantify sensitivities. Marginal profit, option Greeks, and portfolio risk exposures all rely on this fundamental concept.
Multivariable calculus: Partial derivatives extend this to functions of many variables, measuring sensitivity to each input individually. The gradient vector collects these sensitivities and points toward steepest ascent.
Unconstrained optimization: Interior differentiable critical points occur where the gradient vanishes. A definite or indefinite Hessian classifies nondegenerate critical points, while zero-eigenvalue cases are inconclusive and need higher-order or other analysis. Gradient descent provides an iterative algorithm for finding minima.
Constrained optimization: Lagrange multipliers incorporate constraints into a stationarity system. Under regularity conditions, their signed values give local first-order objective sensitivity to the specified constraint right-hand sides, with units and signs determined by the formulation.
Convexity: Every local minimum of a convex objective over a convex feasible set is global. Portfolio variance is convex when its covariance matrix is positive semidefinite, so mean-variance minimization is a convex program when its feasible set is also convex. For a twice-differentiable function on an open convex domain, the Hessian must be positive semidefinite throughout the domain; eigenvalues at one point classify only local curvature.
These tools appear throughout quantitative finance. The next chapter develops integral calculus and differential equations, building from local rates of change to accumulated quantities and continuous-time dynamics. Later chapters reuse optimization in portfolio construction and parameter estimation.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about differential calculus and optimization in quantitative finance.
Differential Calculus and Optimization
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Quantitative Finance. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Quantitative FinanceStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
1 comment
Hello, I greatly appreciate this chapter of the book, what other books do you recommend that covers calculus/optimization like this page?
Hi Yanj! Glad you find it useful. You could look at some chapters in this book, as they cover many topics at a foundational level. https://mbrenndoerfer.com/books/machine-learning-from-scratch