Probability Distributions in Finance: Normal, Lognormal

Michael BrenndoerferOctober 21, 202565 min read

Part of Quantitative Finance

Covers probability distributions essential for quantitative finance: normal, lognormal, binomial, Poisson, and fat-tailed distributions with Python examples.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Common Probability Distributions in Finance

Probability distributions are a core modeling tool in quantitative finance. Many portfolio-risk estimates, option-pricing models, and default-likelihood models require assumptions about how uncertain outcomes are distributed. A misspecified distribution can materially change estimated tail probabilities, option values, and risk limits.

The choice becomes especially visible in the tails. Later in this chapter, the same simulated heavy-tailed sample produces a lower 95% VaR but higher expected shortfall and higher 99% risk measures than a fitted normal model. That mixed result is the practical reason to inspect a distribution's assumptions rather than relying on a blanket rule about normal models.

This chapter covers several common models in quantitative finance: the normal distribution as a baseline for returns and mean-variance analysis, the lognormal distribution for positive quantities such as modeled asset prices, discrete distributions like the binomial and Poisson for counts and event models, and heavier-tailed alternatives. Whether any of them fits a particular dataset or decision well must be checked against the relevant horizon and evidence. For each distribution, we examine both the mathematical properties and the financial intuition behind its use.

The Normal Distribution

The normal distribution, also called the Gaussian distribution, is a common baseline in financial modeling. One mathematical reason is the classical Central Limit Theorem: standardized sums of independent, identically distributed random variables with finite variance converge in distribution to a normal law. Thinking of a return as the aggregate effect of many small shocks can motivate a normal approximation, but market shocks need not be independent or identically distributed, so this is a modeling heuristic rather than a theorem about returns.

Normal Distribution

A continuous probability distribution characterized by its symmetric, bell-shaped curve. It is fully specified by two parameters: the mean μ\mu (center) and variance σ2\sigma^2 (spread). Also known as the Gaussian distribution.

Mathematical Properties

To understand what it means for a random variable to follow a normal distribution, we must first grasp the concept of a probability density function. Unlike discrete random variables that take on specific values with certain probabilities, a continuous random variable like one following a normal distribution can take any value along the real line. Probability over a region is the integral of the density over that region; density at one point is not itself a probability. A higher density at a point means that sufficiently small, equal-width neighborhoods around that point have greater probability than neighborhoods around a lower-density point.

A random variable XX follows a normal distribution with mean μ\mu and variance σ2\sigma^2, written X∼N(μ,σ2)X \sim N(\mu, \sigma^2), if its probability density function is:

f(x)=1σ2πexp⁡(−(x−μ)22σ2),−∞<x<∞f(x) = \frac{1}{\sigma\sqrt{2\pi}} \exp\left(-\frac{(x-\mu)^2}{2\sigma^2}\right), \quad -\infty < x < \infty

Each component serves a specific purpose. Let's examine each component:

  • xx: the value at which we evaluate the density
  • μ\mu: the mean (center of the distribution)
  • σ\sigma: the standard deviation (controls the spread)
  • 1σ2π\frac{1}{\sigma\sqrt{2\pi}}: a normalization constant ensuring the density integrates to 1
  • (x−μ)2/(2σ2)(x-\mu)^2/(2\sigma^2): one half of the squared standardized distance from xx to the mean

The heart of the normal distribution lies in the exponential term. The quantity (x−μ)2(x-\mu)^2 measures the squared distance between the point xx and the center of the distribution μ\mu. Dividing by 2σ22\sigma^2 scales this distance relative to the spread of the distribution. A value two units away from the mean is much more unusual when σ=0.5\sigma = 0.5 than when σ=10\sigma = 10, because the standardized distance is larger. The negative sign in the exponent ensures that the function decreases as we move away from the mean, and the exponential function turns this squared distance into a probability density that decays smoothly and symmetrically.

The exponential of a negative quadratic creates the characteristic bell shape: values near the mean have high probability density, while values far from the mean have exponentially decreasing density. For a standard normal variable, the two-sided probability beyond four standard deviations is about 6.3×10−56.3 \times 10^{-5}, and beyond five it is about 5.7×10−75.7 \times 10^{-7}. Whether those benchmark probabilities adequately describe a financial return series is an empirical question tied to the asset, horizon, and sample. The quadratic term in the exponent is the key: because squaring grows faster than linear growth, the tails of the normal distribution thin out quickly.

The normalization constant 1σ2π\frac{1}{\sigma\sqrt{2\pi}} deserves special attention. This factor ensures that when we integrate the density function over all possible values from negative infinity to positive infinity, we obtain exactly 1. This is the total probability. The 2π\sqrt{2\pi} arises from the Gaussian integral, one of the most famous results in calculus, while the σ\sigma in the denominator accounts for how wider distributions (larger σ\sigma) must have lower peak heights to maintain the same total area under the curve.

Key parameters:

  • Mean (μ\mu): The center of the distribution and the expected value of XX. In finance, this represents the expected return.
  • Variance (σ2\sigma^2): Measures the spread of the distribution. The square root, σ\sigma, is the standard deviation, which in finance we call volatility.
  • Skewness: Zero for the normal distribution, meaning it is perfectly symmetric around the mean.
  • Kurtosis: Equal to 3 for the normal distribution (or excess kurtosis of 0). Kurtosis is the fourth standardized moment; by itself, it does not uniquely determine how probability is divided among the center, shoulders, and tails.

The symmetry of the normal distribution is both its elegance and its limitation. Symmetry is a modeling assumption: it assigns equal probabilities to equal-sized deviations above and below the mean. Some financial return series exhibit asymmetry, with the direction and magnitude depending on the asset, sampling horizon, and market regime. This possibility is one reason to treat the normal distribution as a baseline rather than a universal model of returns.

The standard normal distribution is the special case where μ=0\mu = 0 and σ2=1\sigma^2 = 1. Any normal variable can be standardized:

Z=X−μσ∼N(0,1)Z = \frac{X - \mu}{\sigma} \sim N(0, 1)

This standardization process turns any normal random variable into the canonical standard normal form. The operation X−μσ\frac{X - \mu}{\sigma} first centers the distribution by subtracting the mean (so the new center is at zero), then scales it by dividing by the standard deviation (so the new spread equals one). This step preserves the essential shape of the distribution while placing it in a standard coordinate system.

This standardization is central because it allows us to use a single table of probabilities for any normal distribution. The cumulative distribution function of the standard normal, denoted Φ(z)\Phi(z), gives the probability that a standard normal variable is less than zz. Before computers, statisticians relied on printed tables of Φ(z)\Phi(z) values, and the standardization formula allowed them to answer questions about any normal distribution using this single reference. Even today, the standard normal is the universal reference point: when we speak of a "3-sigma event," we mean an outcome more than three standard deviations from the mean, regardless of what the actual mean and standard deviation are.

The 68-95-99.7 Rule

One of the most useful properties of the normal distribution is how probability mass is concentrated around the mean:

  • Approximately 68% of observations fall within one standard deviation of the mean
  • Approximately 95% fall within two standard deviations
  • Approximately 99.7% fall within three standard deviations

These percentages emerge directly from integrating the normal density function, but their approximate values (68, 95, and 99.7) are worth committing to memory because they provide instant intuition for assessing how unusual any particular observation is. If someone tells you that an event was a "two-sigma move," you immediately know that such events occur roughly 5% of the time, or about one trading day in twenty.

This rule provides quick mental estimates for risk. If daily stock returns have a standard deviation of 1%, then under normality assumptions, the probability of an absolute move larger than 3% is about 0.27%. That is about once every 370 trading days, or roughly once every 1.5 trading years. This calculation reveals both the power and the danger of the normal assumption: it gives us precise quantitative predictions, but those predictions can be spectacularly wrong if the true distribution differs from normal.

Out[2]:
Visualization
Bell curve for the standard normal distribution with shaded standard-deviation probability regions.
The standard normal distribution with shaded standard-deviation bands. The central band contains about 68.3% of the probability mass, the two bands from one to two standard deviations add about 27.2% (95.4% cumulative), and the bands from two to three standard deviations add about 4.3% (99.7% cumulative).
Out[3]:
Visualization
Line plot of normal probability density functions with different means and fixed standard deviation.
Normal distributions with different means (σ fixed). Changing the mean shifts the distribution horizontally.
Line plot of normal probability density functions with fixed mean and different standard deviations.
Normal distributions with different standard deviations (μ fixed). Changing the standard deviation affects the spread and peak height.

Applications in Finance

The normal distribution appears throughout quantitative finance:

Portfolio theory

Portfolio theory often summarizes portfolios by expected return and variance. Normality is sufficient for those two moments to determine the full distribution of portfolio returns, but the mean-variance framework itself does not require normally distributed returns. When candidate portfolio returns are normal, expected utility for a fixed utility function can be evaluated from the mean and variance because those parameters determine the distribution.

Value at Risk (VaR)

VaR calculations often assume normal returns. Under that model, the signed VaR loss quantile at confidence level α\alpha is:

VaRα=zασ−μ\text{VaR}_{\alpha} = z_{\alpha} \sigma - \mu

This formula translates the abstract concept of a tail quantile into a concrete loss estimate. Under the assumption that returns are normally distributed with mean μ\mu and standard deviation σ\sigma, we seek the threshold such that returns fall below this level only (1−α)(1-\alpha) percent of the time. The standard normal quantile zαz_\alpha tells us how many standard deviations below the mean this threshold lies.

Each component of this formula plays a specific role:

  • VaRα\text{VaR}_\alpha: the signed loss quantile at confidence level α\alpha; under a continuous model, losses exceed this threshold with probability 1−α1-\alpha
  • μ\mu: the expected return over the time horizon
  • σ\sigma: the volatility (standard deviation of returns) over the time horizon
  • zαz_\alpha: the α\alpha-quantile of the standard normal distribution (e.g., z0.95≈1.645z_{0.95} \approx 1.645 for 95% confidence)

The formula works because under normality, the (1−α)(1-\alpha)-percentile of returns is exactly μ−zασ\mu - z_\alpha \sigma. For example, the 95% VaR uses z0.95≈1.645z_{0.95} \approx 1.645, meaning we expect losses to exceed this threshold only 5% of the time. Taking the negative of this return quantile yields the signed loss quantile zασ−μz_\alpha \sigma - \mu. It is positive when zασ>μz_\alpha\sigma>\mu, zero at equality, and negative when the modeled lower-tail return quantile is still a gain. A reporting policy that disallows negative VaR can use max⁡(0,zασ−μ)\max(0, z_\alpha\sigma-\mu), but that floor is a reporting convention rather than the raw quantile.

Out[4]:
Visualization
Normal distribution with VaR threshold and shaded tail region showing Expected Shortfall.
Value at Risk (VaR) and Expected Shortfall (ES) visualized on a return distribution. With tail probability p=0.05, the plotted line is the lower-return threshold with probability p of a worse return; for these parameters, its negative is a positive 95% VaR loss threshold. ES is the average loss in that tail.

Option pricing

The Black-Scholes model assumes that log-returns are normally distributed, as we will explore in a later chapter.

In[5]:
Code
import numpy as np
from scipy import stats


def calculate_var_normal(mu, sigma, confidence_level=0.95, horizon=1):
    """
    Calculate Value at Risk assuming normal returns.

    Parameters:
        mu: Expected return (daily)
        sigma: Volatility (daily standard deviation)
        confidence_level: Confidence level (e.g., 0.95 for 95% VaR)
        horizon: Time horizon in days

    Returns:
        Signed VaR loss quantile; it is positive only when the modeled
        lower-tail return quantile is negative
    """
    # Scale parameters to the horizon
    mu_scaled = mu * horizon
    sigma_scaled = sigma * np.sqrt(horizon)

    # Find the quantile
    z = stats.norm.ppf(1 - confidence_level)

    # VaR is the negative of the return quantile; the result may be negative
    var = -(mu_scaled + z * sigma_scaled)
    return var


# Example: Daily return of 0.05% with 1.5% volatility
daily_return = 0.0005
daily_vol = 0.015

var_95_1day = calculate_var_normal(daily_return, daily_vol, 0.95, 1)
var_99_1day = calculate_var_normal(daily_return, daily_vol, 0.99, 1)
var_95_10day = calculate_var_normal(daily_return, daily_vol, 0.95, 10)
Out[6]:
Console
Portfolio with daily return 0.05% and volatility 1.50%
1-day 95% VaR: 2.42%
1-day 99% VaR: 3.44%
10-day 95% VaR: 7.30%

These VaR figures are loss thresholds, not maximum or expected losses. The 1-day 95% VaR of approximately 2.4% means that, under the fitted normal model, losses exceed 2.4% with 5% probability. The 10-day VaR is higher here because the volatility term grows with n\sqrt{n} and outweighs the mean term, which grows with nn. Under the iid normal model, VaRn=zασn−nμ\text{VaR}_n=z_\alpha\sigma\sqrt{n}-n\mu, so VaR need not increase with the horizon for every parameter choice. The square-root scaling of the volatility term follows from variance addition: if daily variances are σ2\sigma^2, then the variance over nn days is nσ2n\sigma^2, and the standard deviation is nσ\sqrt{n}\sigma.

The Lognormal Distribution

Returns may be approximately normal, but a nonnegative stock-price level cannot follow an exact nondegenerate normal law because that law assigns some probability below zero. A local normal approximation may still be useful. A stock price can rise 50% but cannot fall more than 100%. This asymmetry is captured by the lognormal distribution, which constrains values to be positive. The lognormal distribution arises when multiplicative returns produce an accumulated log return that is normally distributed: exponentiating that normal log return gives a positive, lognormally distributed price. Multiplicative compounding by itself does not determine the distribution.

Lognormal Distribution

A probability distribution of a random variable whose logarithm is normally distributed. If ln⁡(X)∼N(μ,σ2)\ln(X) \sim N(\mu, \sigma^2), then XX follows a lognormal distribution. Asset prices are often modeled as lognormal.

From Returns to Prices

The connection between normal returns and lognormal prices is one of the most important relationships in quantitative finance. If log-returns are normally distributed, then prices are lognormally distributed. Consider a stock with initial price S0S_0 and continuously compounded return rr over some period:

ST=S0erS_T = S_0 e^r

This equation captures a key relationship in how prices evolve. The exponential function ere^r converts the additive world of log-returns into the multiplicative world of price changes. When r=0.10r = 0.10, the stock has grown by a factor of e0.10≈1.105e^{0.10} \approx 1.105, or about 10.5%. The slight difference between the log-return and the percentage return becomes more pronounced for larger moves.

The components of this price equation are:

  • STS_T: stock price at time TT
  • S0S_0: initial stock price
  • rr: continuously compounded return over the period
  • ere^r: the growth factor, converting log-return to price ratio

If r∼N(μ,σ2)r \sim N(\mu, \sigma^2), then STS_T follows a lognormal distribution. This relationship is fundamental: the exponential function turns the symmetric, unbounded normal distribution into a right-skewed distribution bounded below by zero, which matches the behavior of asset prices that can rise without limit but cannot fall below zero. The step preserves the elegant properties of the normal distribution while mapping them onto the economically sensible domain of positive prices.

The probability density function of the lognormal distribution takes a form that reveals its connection to the normal:

f(x)=1xσ2πexp⁡(−(ln⁡x−μ)22σ2),x>0f(x) = \frac{1}{x\sigma\sqrt{2\pi}} \exp\left(-\frac{(\ln x - \mu)^2}{2\sigma^2}\right), \quad x > 0

Compare this to the normal density to see the main differences. The argument of the exponential now involves ln⁡x\ln x rather than xx, which reflects that we are measuring distance in the log-space. The presence of 1/x1/x in front of the exponential arises from the chain rule when transforming from the normal to the lognormal through the change of variables.

Each element of this formula serves a distinct purpose:

  • xx: the value at which we evaluate the density (must be positive)
  • μ\mu: the mean of ln⁡(X)\ln(X), not the mean of XX itself
  • σ\sigma: the standard deviation of ln⁡(X)\ln(X), not of XX itself
  • 1/x1/x: this factor (compared to the normal PDF) accounts for the change of variables from ln⁡(x)\ln(x) to xx

The key insight is that μ\mu and σ\sigma are parameters of the underlying normal distribution of log-values, not the lognormal variable itself. This distinction often causes confusion: when we say a stock has lognormal volatility of 20%, we mean the standard deviation of its log-returns, not of the returns themselves. The actual mean and variance of a lognormal variable are:

E[X]=eμ+σ2/2E[X] = e^{\mu + \sigma^2/2} Var(X)=(eσ2−1)e2μ+σ2\text{Var}(X) = (e^{\sigma^2} - 1)e^{2\mu + \sigma^2}

These formulas reveal how the parameters of the underlying normal distribution translate into the moments of the lognormal:

  • E[X]E[X]: the expected value (mean) of the lognormal variable XX
  • Var(X)\text{Var}(X): the variance of XX
  • μ\mu: the mean of ln⁡(X)\ln(X), the underlying normal distribution
  • σ2\sigma^2: the variance of ln⁡(X)\ln(X)
  • eμ+σ2/2e^{\mu + \sigma^2/2}: the mean is greater than eμe^\mu when σ>0\sigma > 0, and equal to it in the degenerate case σ=0\sigma = 0

The σ2/2\sigma^2/2 term in the mean formula is a convexity adjustment: because the exponential function is convex, Jensen's inequality tells us that E[eY]>eE[Y]E[e^Y] > e^{E[Y]} for any non-constant random variable YY. Specifically, the curvature of the exponential means that high values of YY contribute disproportionately to the mean of eYe^Y. This is why the expected value of a lognormal variable includes a variance correction rather than equaling eμe^\mu, since higher variance increases this convexity effect. Holding the mean log-return μ\mu fixed, greater σ\sigma therefore raises E[X]E[X]. In the GBM notation below, however, μ\mu denotes arithmetic drift: holding that drift fixed, the log-mean falls by σ2T/2\sigma^2T/2 as variance rises, and E[ST]=S0eμTE[S_T]=S_0e^{\mu T} does not depend on σ\sigma.

Geometric Brownian Motion

A standard baseline model for stock price dynamics is geometric Brownian motion (GBM), which produces lognormally distributed prices. This model has been central to quantitative finance since Samuelson's work in the 1960s and forms the foundation for the Black-Scholes option pricing formula. Understanding GBM requires familiarity with stochastic calculus, but the core idea is intuitive. Prices evolve continuously with both a deterministic trend and random fluctuations proportional to the current price.

Under GBM, the stock price evolves according to:

dS=μS dt+σS dWdS = \mu S \, dt + \sigma S \, dW

This stochastic differential equation describes infinitesimal changes in the stock price. Each term has a clear interpretation:

  • dSdS: infinitesimal change in stock price
  • SS: current stock price
  • μ\mu: drift rate (expected return per unit time)
  • σ\sigma: diffusion volatility (annualized when time is measured in years); over an interval Δt\Delta t, the log-return standard deviation is σΔt\sigma\sqrt{\Delta t}
  • dtdt: infinitesimal time increment
  • dWdW: increment of a Wiener process (Brownian motion), with dW∼N(0,dt)dW \sim N(0, dt)

The first term μS dt\mu S \, dt represents the deterministic trend, while the second term σS dW\sigma S \, dW captures random fluctuations proportional to the current price. The key feature is that both terms are proportional to SS: a $100 stock and a $10 stock with the same parameters will have the same percentage return distribution, but the absolute dollar changes will be ten times larger for the more expensive stock. This proportionality is what makes GBM appropriate for modeling asset prices, which tend to grow (or shrink) in percentage terms.

Solving this stochastic differential equation yields:

ST=S0exp⁡[(μ−σ22)T+σWT]S_T = S_0 \exp\left[\left(\mu - \frac{\sigma^2}{2}\right)T + \sigma W_T\right]

This solution tells us exactly where the stock price will be at any future time, given the realized path of the Brownian motion. The formula has a clear structure:

  • STS_T: stock price at time TT
  • S0S_0: initial stock price at time 0
  • μ\mu: drift rate (annualized expected return)
  • σ\sigma: volatility (annualized standard deviation)
  • TT: time horizon in years
  • WTW_T: value of the Wiener process at time TT, with WT∼N(0,T)W_T \sim N(0, T)
  • σ2/2\sigma^2/2: Itô correction term arising from the quadratic variation of Brownian motion

The Itô correction σ2/2\sigma^2/2 appears because we apply Itô's lemma to ln⁡(S)\ln(S): the second derivative term 12∂2ln⁡S∂S2(σS)2=−σ22\frac{1}{2}\frac{\partial^2 \ln S}{\partial S^2}(\sigma S)^2 = -\frac{\sigma^2}{2} contributes to the drift. This correction is one of the key differences between ordinary calculus and stochastic calculus. In ordinary calculus, if we know dS/SdS/S, we would conclude that d(ln⁡S)=dS/Sd(\ln S) = dS/S. However, when SS follows a diffusion process, the additional −σ2/2-\sigma^2/2 term arises from the non-zero quadratic variation of Brownian motion. For σ>0\sigma>0, the expected GBM log-growth rate μ−σ2/2\mu-\sigma^2/2 is below its arithmetic drift μ\mu; at σ=0\sigma=0, the two are equal.

This model forms the foundation of the Black-Scholes option pricing formula.

Out[7]:
Visualization
Line plot of a normal probability density over daily returns with a vertical line at zero.
Normal distribution of returns, centered near zero and symmetric.
Line plot of a lognormal probability density over stock prices with a vertical line at the initial price.
Lognormal distribution of prices, bounded below by zero and positively skewed.

Simulating Price Paths

Monte Carlo simulation of stock prices under GBM is simple. The approach below draws exact GBM transition increments at each point on the chosen time grid and cumulatively sums them. Unlike an Euler approximation, this produces the exact GBM distribution at the grid points for constant parameters; it does not simulate between-grid extrema.

In[8]:
Code
import numpy as np


def simulate_gbm_paths(S0, mu, sigma, T, n_steps, n_paths):
    """
    Simulate geometric Brownian motion paths.

    Parameters:
        S0: Initial stock price
        mu: Annualized drift (expected return)
        sigma: Annualized volatility
        T: Time horizon in years
        n_steps: Number of time steps
        n_paths: Number of simulation paths

    Returns:
        Array of shape (n_paths, n_steps + 1) containing price paths
    """
    dt = T / n_steps

    # Generate random increments
    np.random.seed(42)
    Z = np.random.standard_normal((n_paths, n_steps))

    # Calculate log returns for each step
    log_returns = (mu - 0.5 * sigma**2) * dt + sigma * np.sqrt(dt) * Z

    # Cumulative sum to get log prices relative to initial
    log_prices = np.zeros((n_paths, n_steps + 1))
    log_prices[:, 0] = np.log(S0)
    log_prices[:, 1:] = np.log(S0) + np.cumsum(log_returns, axis=1)

    # Convert to prices
    prices = np.exp(log_prices)
    return prices


# Simulate 1 year of daily prices
S0 = 100
mu = 0.10  # 10% expected annual return
sigma = 0.25  # 25% annual volatility
T = 1.0
n_steps = 252  # Trading days
n_paths = 1000

paths = simulate_gbm_paths(S0, mu, sigma, T, n_steps, n_paths)
final_prices = paths[:, -1]
Out[9]:
Console
Initial price: $100.00
Expected final price (theoretical): $110.52
Mean simulated final price: $110.44
Median simulated final price: $108.19
5th percentile: $71.15
95th percentile: $160.84

Notice that the mean simulated price is close to the theoretical expected value of S0eμTS_0 e^{\mu T}, but the median is lower. This is typical of a non-degenerate lognormal distribution: its mean exceeds its median, and its mode is lower still. These quantities answer different questions. The median describes the middle terminal outcome across simulated paths, the mean describes ensemble expected terminal wealth, and long-run performance of repeatedly compounded wealth is governed by average log growth rather than by the arithmetic mean terminal price alone.

Out[10]:
Visualization
Spaghetti plot of simulated GBM price paths with an overlaid expected path.
Fifty sample paths from a Monte Carlo simulation of 1,000 stock-price paths under geometric Brownian motion. The fan-like shape illustrates how uncertainty increases with time.
Histogram of simulated final stock prices with vertical lines marking the mean and median.
Distribution of final simulated prices under geometric Brownian motion after one year, illustrating the right-skewed outcomes of a lognormal model.

Discrete Distributions: Binomial and Poisson

Not all financial phenomena are continuous. Events like defaults, dividend announcements, or large jumps in prices are discrete occurrences. The binomial and Poisson distributions model these situations and give a framework for counting how many times something happens rather than measuring continuous quantities. These distributions are particularly important in credit risk, where we count defaults in a portfolio, and in derivatives pricing, where discrete lattice models approximate continuous price processes.

The Binomial Distribution

The binomial distribution models the number of successes in a fixed number of independent trials, each with the same probability of success. The classic example is flipping a coin multiple times and counting heads, but the same mathematics applies to questions such as: How many stocks in a portfolio will outperform? How many corporate bonds will default? How many days this month will have positive returns?

Binomial Distribution

The probability distribution of the number of successes kk in nn independent Bernoulli trials, where each trial has success probability pp. Used in discrete option pricing models and credit risk modeling.

If X∼Binomial(n,p)X \sim \text{Binomial}(n, p), the probability mass function is:

P(X=k)=(nk)pk(1−p)n−k,k=0,1,…,nP(X = k) = \binom{n}{k} p^k (1-p)^{n-k}, \quad k = 0, 1, \ldots, n

This formula answers the question: if we have nn independent trials, each with success probability pp, what is the probability of getting exactly kk successes? The answer decomposes into two multiplicative components: the number of ways the successes can be arranged, and the probability of any particular arrangement.

Each component of this formula serves a specific purpose:

  • kk: the number of successes we're computing the probability for
  • nn: total number of independent trials
  • pp: probability of success on each trial
  • (nk)=n!k!(n−k)!\binom{n}{k} = \frac{n!}{k!(n-k)!}: the binomial coefficient, representing the number of ways to choose kk successes among nn trials
  • pkp^k: success-probability factor for kk specified success positions
  • (1−p)n−k(1-p)^{n-k}: failure-probability factor for n−kn-k specified failure positions

The formula multiplies the number of arrangements by the probability of each specific arrangement. Intuitively, if you flip a coin nn times with probability pp of heads, the PMF tells you the probability of getting exactly kk heads: you need kk successes (probability pkp^k), n−kn-k failures (probability (1−p)n−k(1-p)^{n-k}), and the binomial coefficient counts how many orderings produce that outcome. The key insight is that because the trials are independent, the probability of any specific sequence of outcomes is simply the product of the individual probabilities, and the binomial coefficient tells us how many such sequences lead to exactly kk successes.

The mean and variance have elegant formulas that reveal the distribution's structure:

E[X]=np,Var(X)=np(1−p)E[X] = np, \quad \text{Var}(X) = np(1-p)

These formulas encode intuitive relationships:

  • E[X]E[X]: the expected number of successes
  • Var(X)\text{Var}(X): the variance in the number of successes
  • nn: total number of trials
  • pp: probability of success on each trial
  • (1−p)(1-p): probability of failure, which appears in variance because maximum variance occurs when p=0.5p = 0.5 (maximum uncertainty)

The expected number of successes, npnp, is simply the number of trials times the success probability, exactly what intuition suggests. The variance formula np(1−p)np(1-p) is more subtle: it shows that variance is maximized when p=0.5p = 0.5 (complete uncertainty about each trial's outcome) and decreases toward zero as pp approaches 0 or 1 (near certainty). This makes sense because if pp is very close to 1, almost all trials succeed, leaving little room for variability in the total count.

Out[11]:
Visualization
Bar chart of binomial PMFs for n=20 at p=0.2, 0.5, and 0.8.
Binomial probability mass functions for fixed n=20 and varying success probability p, showing how the mean and skewness shift with p.
Line plot of binomial PMFs for p=0.3 at n=10, 30, and 50.
Binomial probability mass functions for fixed p=0.3 and varying n. On the count scale, the mean and absolute spread increase with n; the success proportion X/n becomes more concentrated around p.

Binomial Option Pricing Model

The Cox-Ross-Rubinstein (CRR) binomial model prices options by assuming that at each time step, the stock price either moves up by a factor uu or down by a factor dd. This discretization turns the continuous option pricing problem into a tree of possible outcomes that can be analyzed one step at a time. While this may seem like a crude approximation to the continuous world of actual markets, the model provides deep insights into option pricing and converges to the continuous Black-Scholes formula as the number of time steps increases.

After nn steps, the stock price has experienced kk up moves and n−kn-k down moves, giving:

ST=S0ukdn−kS_T = S_0 u^k d^{n-k}

This formula shows how the final price depends entirely on the number of up moves, not their specific timing. The components are:

  • STS_T: stock price at maturity after nn steps
  • S0S_0: initial stock price
  • uu: up factor (price multiplier for an up move)
  • dd: down factor (price multiplier for a down move)
  • kk: number of up moves
  • n−kn-k: number of down moves

Whether the stock rises first and then falls, or falls first and then rises, the final price is the same as long as the total number of up and down moves matches. This recombination supports efficient pricing for European options, whose payoff depends only on the final price. Path-dependent claims such as Asian options require the tree state to be augmented with path information; the simple terminal-price state is not sufficient.

The number of up moves follows a binomial distribution under the risk-neutral measure, with probability qq of an up move:

q=erΔt−du−dq = \frac{e^{r\Delta t} - d}{u - d}

This risk-neutral probability is central to derivatives pricing. Notice that qq is not the actual probability that the stock will go up. It is the artificial probability that makes the expected return on the stock equal to the risk-free rate. The components are:

  • qq: risk-neutral probability of an up move
  • rr: risk-free interest rate (annualized)
  • Δt\Delta t: length of each time step in years
  • uu: up factor (stock price multiplier for an up move)
  • dd: down factor (stock price multiplier for a down move)
  • erΔte^{r\Delta t}: the growth factor for a risk-free investment over one time step

This formula is derived from the no-arbitrage condition: the expected stock return under the risk-neutral measure must equal the risk-free rate. Setting q⋅u+(1−q)⋅d=erΔtq \cdot u + (1-q) \cdot d = e^{r\Delta t} and solving for qq yields this result. In this complete one-period binomial model, a replicable option is priced from the risk-neutral probability rather than the physical up-move probability. This conclusion depends on no arbitrage, market completeness, and the ability to form the replicating portfolio; it is not a universal statement about every derivative market.

Assuming d<ud<u, this qq is a strict probability and the one-period stock-and-bond model is arbitrage-free only when

d<erΔt<u,d < e^{r\Delta t} < u,

which is equivalent to 0<q<10<q<1. Inputs outside this range do not define valid CRR risk-neutral probabilities.

Out[12]:
Visualization
Recombining binomial tree diagram showing stock prices at each node.
A 4-step binomial tree showing possible stock price paths. The tree recombines because an up move followed by a down move gives the same price as a down move followed by an up move. Node labels show the stock price at each node.
In[13]:
Code
import numpy as np
from scipy.special import comb


def binomial_option_price(S0, K, T, r, sigma, n_steps, option_type="call"):
    """
    Price a European option using the CRR binomial model.

    Parameters:
        S0: Initial stock price
        K: Strike price
        T: Time to expiration (years)
        r: Risk-free rate (annualized)
        sigma: Volatility (annualized)
        n_steps: Number of time steps
        option_type: 'call' or 'put'

    Returns:
        Option price
    """
    if option_type not in {"call", "put"}:
        raise ValueError("option_type must be 'call' or 'put'")

    dt = T / n_steps

    # CRR parameters
    u = np.exp(sigma * np.sqrt(dt))
    d = 1 / u
    q = (np.exp(r * dt) - d) / (u - d)
    if not 0 < q < 1:
        raise ValueError(
            "CRR no-arbitrage condition requires d < exp(r*dt) < u"
        )

    # Calculate all possible final stock prices
    k = np.arange(n_steps + 1)
    final_prices = S0 * (u**k) * (d ** (n_steps - k))

    # Calculate payoffs
    if option_type == "call":
        payoffs = np.maximum(final_prices - K, 0)
    else:
        payoffs = np.maximum(K - final_prices, 0)

    # Risk-neutral probabilities
    probs = comb(n_steps, k) * (q**k) * ((1 - q) ** (n_steps - k))

    # Expected payoff under risk-neutral measure, discounted
    option_price = np.exp(-r * T) * np.sum(probs * payoffs)

    return option_price


# Price a call option
S0 = 100
K = 100
T = 0.5  # 6 months
r = 0.05
sigma = 0.20
Out[14]:
Console
Binomial Model Call Option Prices
-----------------------------------
Steps:   10  |  Price: $6.7494
Steps:   50  |  Price: $6.8605
Steps:  100  |  Price: $6.8746
Steps:  200  |  Price: $6.8817
Steps:  500  |  Price: $6.8859
-----------------------------------
Black-Scholes:  |  Price: $6.8887

As the number of steps increases, the CRR binomial price converges to the Black-Scholes price under the model's regularity conditions. The error can oscillate with the number of steps because the strike sits differently relative to successive lattices. For sufficiently regular European payoffs, suitably controlled sequences often have first-order error on the scale of 1/n1/n, but that rate is not uniform across payoffs and parameter choices. The convergence reflects a mathematical connection between the binomial random walk and Brownian motion, formalized in Donsker's theorem.

Out[15]:
Visualization
Line plot showing binomial price converging to Black-Scholes price with increasing steps.
Convergence of the binomial option pricing model to the Black-Scholes benchmark under their shared model assumptions. The oscillation around the continuous-model limit decreases as n grows.

The Poisson Distribution

The Poisson distribution models the number of events occurring in a fixed interval when events happen at a constant average rate and independently of each other. Unlike the binomial distribution, which assumes a fixed number of trials, the Poisson distribution allows for any number of events, making it useful for modeling rare occurrences over a continuous interval.

Poisson Distribution

A discrete probability distribution expressing the probability of a given number of events occurring in a fixed interval, when events occur at a constant average rate λ\lambda independently of each other. Used for modeling defaults, jumps, and arrival processes.

If X∼Poisson(λ)X \sim \text{Poisson}(\lambda), the probability mass function is:

P(X=k)=λke−λk!,k=0,1,2,…P(X = k) = \frac{\lambda^k e^{-\lambda}}{k!}, \quad k = 0, 1, 2, \ldots

This formula has a clear structure when we trace its derivation from first principles. Imagine dividing the interval into many tiny subintervals, each with a small probability of containing an event. As the number of subintervals goes to infinity while keeping the expected total number of events fixed at λ\lambda, we obtain the Poisson distribution.

The components of this formula are:

  • kk: the number of events we're computing the probability for
  • λ\lambda: the average rate of events (expected number of events in the interval)
  • e−λe^{-\lambda}: the probability of zero events, which is a normalization factor
  • λk\lambda^k: the intensity factor for kk events; together with e−λe^{-\lambda}, it makes P(X=k)P(X=k) largest at λ=k\lambda=k for k>0k>0
  • k!k!: accounts for the indistinguishability of event orderings

The Poisson distribution arises as the limit of a binomial distribution when n→∞n \to \infty and p→0p \to 0 while np=λnp = \lambda remains constant. More generally, a Poisson approximation requires independent rare-event opportunities whose largest individual probability vanishes while the sum of their probabilities approaches λ\lambda. A finite expected count alone is not enough: dependence or a non-negligible individual probability can produce a different count distribution. Credit portfolios with small independent default probabilities and jump models with independent small-interval arrivals can satisfy the approximation; clustered defaults and jumps require a different model.

Both the mean and variance equal λ\lambda:

E[X]=Var(X)=λE[X] = \text{Var}(X) = \lambda

This property deserves emphasis:

  • E[X]E[X]: the expected number of events
  • Var(X)\text{Var}(X): the variance in the number of events
  • λ\lambda: the average rate parameter

The equality of mean and variance is a distinctive property of the Poisson distribution. When analyzing count data, if the observed variance significantly exceeds the mean (overdispersion), this suggests the Poisson model may be inadequate. Dependence is one possible cause, but rate heterogeneity, omitted covariates, and excess zeros can also produce overdispersion. For instance, defaults may cluster because borrowers share economic conditions even when individual default mechanisms differ.

Out[16]:
Visualization
Line plot of Poisson PMFs for λ equal to 1, 3, 7, and 15.
Poisson probability mass functions for different rate parameters λ, showing how the distribution becomes more symmetric as λ increases.
Bar chart comparing Poisson(5) probabilities to binomial approximations with n equal to 10, 50, and 200.
Poisson distribution as the rare-event limit of the binomial distribution, comparing Poisson(λ=5) to Binomial(n, λ/n) for increasing n.

Applications of the Poisson Distribution

The Poisson distribution is particularly useful for modeling:

  • Jump processes in asset prices: The Merton jump-diffusion model adds jumps whose arrivals follow a Poisson process and whose sizes follow a separate distribution to geometric Brownian motion. The stock price follows GBM between jumps and moves discontinuously at jump times. This combination represents continuous diffusion-like fluctuations together with occasional discontinuous price moves.

  • Credit defaults in a portfolio: If defaults occur independently with a small probability, the number of defaults in a large portfolio is approximately Poisson distributed. This approximation underlies many credit risk models, although the assumption of independence often breaks down during systemic crises, when many borrowers default simultaneously.

  • Order arrivals: In market microstructure, the arrival of orders is often modeled as a Poisson process. This assumption implies that the time between consecutive orders follows an exponential distribution, without clustering or periodicity in order flow. While real markets show more complex arrival patterns, the Poisson model provides a useful baseline.

In[17]:
Code
import numpy as np


def simulate_jump_diffusion(
    S0, mu, sigma, lambda_jump, mu_jump, sigma_jump, T, n_steps, n_paths
):
    """
    Simulate Merton jump-diffusion price paths.

    Parameters:
        S0: Initial stock price
        mu: Drift of the diffusion component
        sigma: Volatility of the diffusion component
        lambda_jump: Average number of jumps per year (Poisson intensity)
        mu_jump: Mean of log-jump size
        sigma_jump: Std dev of log-jump size
        T: Time horizon in years
        n_steps: Number of time steps
        n_paths: Number of paths to simulate

    Returns:
        Array of simulated price paths
    """
    dt = T / n_steps

    np.random.seed(42)

    # Diffusion component
    Z_diffusion = np.random.standard_normal((n_paths, n_steps))
    diffusion = (mu - 0.5 * sigma**2) * dt + sigma * np.sqrt(dt) * Z_diffusion

    # Jump component
    # Number of jumps in each interval (Poisson distributed)
    n_jumps = np.random.poisson(lambda_jump * dt, (n_paths, n_steps))

    # Log jump sizes (normal)
    jump_sizes = np.zeros((n_paths, n_steps))
    for i in range(n_paths):
        for j in range(n_steps):
            if n_jumps[i, j] > 0:
                jumps = np.random.normal(mu_jump, sigma_jump, n_jumps[i, j])
                jump_sizes[i, j] = np.sum(jumps)

    # Combine diffusion and jumps
    log_returns = diffusion + jump_sizes

    # Calculate prices
    log_prices = np.zeros((n_paths, n_steps + 1))
    log_prices[:, 0] = np.log(S0)
    log_prices[:, 1:] = np.log(S0) + np.cumsum(log_returns, axis=1)

    return np.exp(log_prices)


# Compare GBM vs Jump-Diffusion
S0 = 100
mu = 0.10
sigma = 0.20
T = 1.0
n_steps = 252
n_paths = 10000

# Pure GBM paths
gbm_paths = simulate_gbm_paths(S0, mu, sigma, T, n_steps, n_paths)

# Jump-diffusion paths
jd_paths = simulate_jump_diffusion(
    S0,
    mu,
    sigma,
    lambda_jump=2,  # 2 jumps per year on average
    mu_jump=-0.05,  # Jumps tend to be negative
    sigma_jump=0.10,
    T=T,
    n_steps=n_steps,
    n_paths=n_paths,
)

gbm_returns = np.diff(np.log(gbm_paths), axis=1).flatten()
jd_returns = np.diff(np.log(jd_paths), axis=1).flatten()
Out[18]:
Console
Return Distribution Comparison
==================================================
Statistic                     GBM  Jump-Diffusion
--------------------------------------------------
Mean                     0.000312       -0.000083
Std Dev                  0.012609        0.016096
Skewness                  -0.0006         -3.1474
Excess Kurtosis           -0.0002         55.5761
Min                       -0.0605         -0.4391
Max                        0.0611          0.3572

With the negatively centered jump-size distribution used here, the simulated jump-diffusion returns have a more negative sample skewness and a larger fourth standardized moment than the simulated GBM returns. Changing the jump-size distribution can change the sign of skewness. Likewise, higher kurtosis alone does not determine a unique allocation of probability among the center, shoulders, and tails; the histogram and quantiles must be examined directly.

Out[19]:
Visualization
Overlaid full-range histograms on a logarithmic density axis, including simulated jump-diffusion daily log-return extrema near -43.9% and 35.7%.
Full-range histograms of simulated daily log returns from geometric Brownian motion and Merton jump-diffusion on a logarithmic density axis. All simulated observations are shown: the GBM sample ranges from about -6.1% to 6.1%, while the jump-diffusion sample ranges from about -43.9% to 35.7%.
Q-Q scatter plot of simulated GBM and jump-diffusion returns versus theoretical normal quantiles.
Q-Q plot against the normal distribution for the simulated GBM and jump-diffusion returns, showing their different departures from normality.

Fat-Tailed Distributions

A normal distribution is only one possible model for returns. When an observed or simulated return distribution places more probability in sufficiently remote tails than a fitted normal model, normal tail calculations can misstate extreme-event frequencies and downstream risk measures. The examples below use a seeded Student's t sample to show how to diagnose that mismatch without presenting synthetic observations as market data.

Evidence of Fat Tails

The term "fat tails" refers to distributions where extreme events occur more frequently than a normal distribution would suggest. A distribution with fat tails places more probability mass in the regions far from the mean, meaning that large positive or negative values are more common than the normal distribution's rapid exponential decay would predict. Several ways to measure this phenomenon:

Kurtosis: The normal distribution has a kurtosis of 3 (or excess kurtosis of 0). Kurtosis is the fourth central moment divided by the variance squared. A higher value signals more weight on large standardized deviations overall, but it does not by itself order every tail probability or identify where probability moved between the center and shoulders.

Tail exponents: For power-law-tailed distributions, the probability of extreme events decays polynomially rather than exponentially. The tail exponent α\alpha characterizes how fast this decay occurs. A smaller α\alpha means slower polynomial decay and fatter tails, while a larger α\alpha means faster polynomial decay and thinner tails within the power-law family. For every finite α\alpha, however, the asymptotic tail remains heavier than a normal tail.

Out[21]:
Console
Simulated Daily Returns (Heavy-Tailed, t-distribution df=4)
==================================================
Number of observations: 6,000
Mean daily return: 0.0079%
Daily volatility: 1.3944%
Skewness: 0.086
Excess Kurtosis: 5.719

Extreme Event Frequency
--------------------------------------------------
Event             Observed  Expected (Normal)      Ratio
--------------------------------------------------
|r| > 3σ                80               16.2        4.9x
|r| > 4σ                30                0.4       78.9x
|r| > 5σ                 9                0.0     2616.4x

The seeded sample was generated from a heavy-tailed distribution, so its departure from the normal benchmark is expected rather than empirical evidence about markets. For this seed, absolute moves beyond four sample standard deviations occur 30 times, compared with about 0.38 expected in 6,000 independent normal draws, a ratio of about 79. Nine observations exceed five sample standard deviations, compared with about 0.0034 under that same normal benchmark. The exercise demonstrates how a known heavy-tailed data-generating process changes extreme-event counts; it does not estimate their frequency in an observed market.

The Student's t Distribution

The Student's t distribution provides a simple way to model fat tails while retaining many convenient properties of the normal distribution. William Sealy Gosset originally developed the t distribution (writing under the pseudonym "Student") for quality control at Guinness brewery. It is now widely used in finance for modeling the heavy tails observed in return data.

Student's t Distribution

A probability distribution with heavier tails than the normal distribution, characterized by its degrees of freedom parameter ν\nu. As ν→∞\nu \to \infty, it converges to the normal distribution. For small ν\nu, it assigns higher probability to extreme events.

The probability density function of the standard t distribution with ν\nu degrees of freedom is:

f(x)=Γ(ν+12)νπ Γ(ν2)(1+x2ν)−ν+12f(x) = \frac{\Gamma\left(\frac{\nu+1}{2}\right)}{\sqrt{\nu\pi}\,\Gamma\left(\frac{\nu}{2}\right)} \left(1 + \frac{x^2}{\nu}\right)^{-\frac{\nu+1}{2}}

This formula may look complex, but its structure shows the key difference between the t distribution and the normal distribution. While the normal distribution uses the exponential function e−x2/2e^{-x^2/2} which decays very rapidly, the t distribution uses a power function (1+x2/ν)−(ν+1)/2(1 + x^2/\nu)^{-(\nu+1)/2} which decays much more slowly. This slower decay is precisely what produces the heavier tails.

Each component:

  • xx: the value at which we evaluate the density
  • ν\nu: degrees of freedom parameter (controls tail thickness; lower ν\nu means fatter tails)
  • Γ(⋅)\Gamma(\cdot): the gamma function, generalizing the factorial to non-integer values (Γ(n)=(n−1)!\Gamma(n) = (n-1)! for positive integers)
  • Γ(ν+12)νπ Γ(ν2)\frac{\Gamma\left(\frac{\nu+1}{2}\right)}{\sqrt{\nu\pi}\,\Gamma\left(\frac{\nu}{2}\right)}: a normalization constant ensuring the density integrates to 1
  • (1+x2/ν)−(ν+1)/2(1 + x^2/\nu)^{-(\nu+1)/2}: the core term that produces polynomial rather than exponential tail decay

The polynomial decay (1+x2/ν)−(ν+1)/2(1 + x^2/\nu)^{-(\nu+1)/2} versus the normal's exponential decay e−x2/2e^{-x^2/2} is why the t distribution has heavier tails. To understand this intuitively, consider what happens as xx becomes large. For the normal distribution, doubling xx quadruples the magnitude of the exponent because (2x)2=4x2(2x)^2 = 4x^2, so the density decays extremely rapidly. For the t distribution, the density is asymptotically proportional to ∣x∣−(ν+1)|x|^{-(\nu+1)}, so doubling a large threshold reduces the density by about 2−(ν+1)2^{-(\nu+1)}. Its one-sided survival probability is proportional to x−νx^{-\nu}, so doubling a large threshold reduces that probability by about 2−ν2^{-\nu}. As ν→∞\nu \to \infty, the polynomial term converges to the exponential, and the t distribution approaches the normal.

Key properties of the t distribution:

  • Degrees of freedom (ν\nu): Controls tail thickness. Lower values mean fatter tails.
  • Convergence: As ν→∞\nu \to \infty, the t distribution approaches the standard normal.
  • Variance: For ν>2\nu > 2, the variance is ν/(ν−2)\nu/(\nu-2), which is always greater than 1. For 1<ν≤21<\nu\leq2, the mean exists but the variance is infinite. For 0<ν≤10<\nu\leq1, the mean and therefore the variance are undefined. The second raw moment diverges throughout 0<ν≤20<\nu\leq2.
  • Kurtosis: For ν>4\nu > 4, excess kurtosis is 6/(ν−4)6/(\nu-4), always positive and decreasing in ν\nu. For ν≤4\nu \leq 4, the kurtosis is undefined because the fourth moment does not exist.

Degrees of freedom must be estimated for the dataset and model at hand; there is no universal range for financial returns. Values near 4 are nevertheless useful for understanding the boundary behavior of the family: for ν≤4\nu \leq 4, the fourth moment is undefined, while for ν>4\nu > 4 it is finite. The seeded example below uses ν=4\nu=4 deliberately to demonstrate a particularly heavy-tailed case.

Out[22]:
Visualization
Line plot of normal and t distribution PDFs for multiple degrees of freedom on a linear y-axis.
Comparison of normal and Student's t distributions with varying degrees of freedom. Lower degrees of freedom produce heavier tails, assigning higher probability to extreme events.
Semilog-y plot of the same PDFs to emphasize tail differences.
Tail behavior on a log scale, which shows how the Student's t distribution assigns higher probability to extreme events than the normal distribution.

The log-scale plot on the right reveals how dramatically the distributions differ in their tails. At 5 standard deviations, the t distribution with 3 degrees of freedom assigns probability thousands of times higher than the normal distribution. This difference, while invisible on a linear scale where both probabilities appear to be essentially zero, has enormous practical significance for risk management. The one-sided probability of a 5-sigma event under a normal distribution is about 1 in 3.5 million observations; under a t distribution with 3 degrees of freedom scaled to have unit variance, it is about 1 in 600 observations. If each observation is treated as an independent daily return and a year contains 252 trading days, those probabilities imply expected waiting times of roughly 13,900 trading years and 2.4 trading years, respectively. These are model-based average waiting times, not schedules for when events will occur.

Fitting the t Distribution to Simulated Returns

We can fit a scaled t distribution to the seeded heavy-tailed sample and compare its maximized log-likelihood with that of a fitted normal distribution. Because the sample was generated from a t distribution, this is a reproducibility check and model-comparison demonstration, not evidence about observed market returns.

In[23]:
Code
import numpy as np
from scipy import stats

# Fit a scaled t distribution to the simulated returns
# The t distribution has location (mu), scale (sigma), and degrees of freedom (nu)
# Use the returns array from the fat-tails-evidence cell
# Ensure returns is available and convert to numpy array
# Regenerate data to ensure availability (matching fat-tails-evidence cell)
np.random.seed(42)
n_days = 6000
daily_returns = np.random.standard_t(df=4, size=n_days) * 0.01 + 0.0003
returns_array = daily_returns

# Maximum likelihood estimation
t_params = stats.t.fit(returns_array)
t_nu, t_loc, t_scale = t_params

# Fit normal for comparison
norm_loc, norm_scale = stats.norm.fit(returns_array)

# Calculate log-likelihoods
t_loglik = np.sum(stats.t.logpdf(returns_array, t_nu, t_loc, t_scale))
norm_loglik = np.sum(stats.norm.logpdf(returns_array, norm_loc, norm_scale))
Out[24]:
Console
Distribution Fitting Results
==================================================

Normal Distribution:
  Location (μ): 0.000079
  Scale (σ): 0.013944
  Log-likelihood: 17,123

Student's t Distribution:
  Degrees of freedom (ν): 4.23
  Location: 0.000055
  Scale: 0.010172
  Log-likelihood: 17,526

Log-likelihood improvement: 403
The t distribution provides a substantially better fit to the data.

For this seed, the fitted t distribution has about 4.23 degrees of freedom and improves the maximized log-likelihood by about 403 relative to the fitted normal model. That large likelihood difference is expected because the sample was generated from a t distribution. Exponentiating the difference would produce a maximized likelihood ratio, not posterior odds; posterior model odds would also require priors and marginal likelihoods.

Out[25]:
Visualization
Full-range histogram of all 6,000 simulated returns with overlaid normal and t density curves.
Fitted distributions compared with all 6,000 simulated heavy-tailed returns, including the seeded sample extremes. The Student's t distribution better captures the peaked center and fat tails.
Q-Q scatter plots of all 6,000 sample values versus fitted-normal and fitted-t theoretical quantiles with a 45-degree reference line.
Q-Q plots of all 6,000 simulated returns against the fitted normal and fitted t distributions, showing where the normal fit breaks down in the tails.

Power-Law Tails

Power-law notation provides a general way to describe polynomial tail decay. Student's t distributions already have power-law tails, but other regularly varying models can use a different exponent or slowly varying factor:

P(X>x)∼Cx−αas x→∞,C>0P(X > x) \sim Cx^{-\alpha} \quad \text{as } x \to \infty, \qquad C>0

This asymptotic relationship describes the right-tail survival function for large thresholds. It includes the polynomial right tail of a Student's t distribution as a special case. For a nonnegative variable, the exponent controls positive moments; for a signed variable, existence of a full moment also depends on the left tail, which this one-sided relationship does not specify.

The components of this relationship are:

  • P(X>x)P(X > x): the probability that the random variable exceeds value xx (the survival function)
  • xx: a large threshold value
  • α\alpha: the tail exponent, which controls how fast probabilities decay. Some authors also call α\alpha a tail index; in another common extreme-value-theory convention, the tail index is ξ=1/α\xi=1/\alpha
  • ∼\sim: denotes asymptotic equivalence. The ratio of the left side to Cx−αCx^{-\alpha} approaches 1 as x→∞x \to \infty

A larger α\alpha means faster polynomial decay and a smaller α\alpha means slower decay. As an illustration, if α=3\alpha=3 and the asymptotic approximation is appropriate at both thresholds, doubling a sufficiently large threshold reduces its exceedance probability by approximately 2−3=1/82^{-3}=1/8. This is much slower than the exponential falloff of normal tails.

For comparison, the one-sided standard-normal survival probability falls by a factor of about 5.1×10105.1\times 10^{10} from 4σ to 8σ, roughly eleven orders of magnitude. Under the illustrative cubic law, the corresponding asymptotic reduction is only a factor of 8. Which approximation is appropriate must be established from a specified model or dataset rather than inferred from this contrast alone.

Power Law Distribution

A distribution where the probability of observing a value greater than xx decays as a power of xx: P(X>x)∝x−αP(X > x) \propto x^{-\alpha}. The tail exponent α\alpha determines how fat the tails are. Lower α\alpha means fatter tails.

Out[26]:
Visualization
Log-log plot comparing normal and unit-variance t survival functions with illustrative power-law tails aligned at a unitless threshold of one.
Tail decay on a common unitless threshold axis. The normal and unit-variance t curves are survival functions; the illustrative power-law tails are aligned to the normal survival probability at x=1 and do not define complete standardized distributions. Power-law tails appear as straight lines on the log-log scale, while the normal tail curves downward.

The implications of power-law tails are serious for risk management:

  • Moment thresholds require both tails: For p>0p>0 and a nonnegative variable with P(X>x)∼Cx−αP(X>x)\sim Cx^{-\alpha}, the right-tail contribution to E[Xp]E[X^p] is finite for p<αp<\alpha and diverges for p≥αp\geq\alpha. For a signed return, the displayed right-tail relation alone cannot guarantee that a full mean or variance exists because the left tail is unspecified. If both tails have comparable exponents near 3, absolute moments of orders below 3, including the mean and variance, are finite, while the fourth moment is not.
  • Slow convergence: Under i.i.d. sampling with finite variance (α>2\alpha > 2 in this tail model), the Central Limit Theorem applies, but its normal approximation can remain poor at practically available sample sizes.
  • Tail concentration and dependence: With power-law marginal tails, a small number of observations can dominate a sample risk estimate. Diversification may be weaker when those tails combine with strong common or tail dependence, but marginal tail thickness alone does not determine diversification benefits.

Implications for Risk Management

Fat tails change how we should think about risk:

VaR model error: A fitted normal model can underestimate sufficiently extreme heavy left-tail quantiles, but it need not underestimate VaR at every confidence level. The direction and magnitude depend on the fitted distributions, skewness, and confidence level. The seeded example below slightly overestimates 95% VaR and underestimates 99% VaR.

Expected shortfall: Under the continuous-loss convention used here, expected shortfall is the mean loss conditional on being at or beyond the VaR threshold. Terminology such as "Conditional VaR" varies across sources, and distributions with atoms require a quantile-integral definition to handle the boundary consistently. For continuous heavy-tailed models, ES can be substantially larger than VaR.

In[27]:
Code
import numpy as np
from scipy import stats


def calculate_var_es(returns, confidence_levels=[0.95, 0.99]):
    """
    Calculate VaR and Expected Shortfall from a return sample.

    Returns:
        Dictionary with VaR and ES for each confidence level
    """
    results = {}
    for cl in confidence_levels:
        # VaR is the negative of the quantile at (1 - confidence level)
        var = -np.percentile(returns, (1 - cl) * 100)

        # ES is the mean of returns worse than VaR
        losses = -returns
        es = np.mean(losses[losses >= var])

        results[cl] = {"VaR": var, "ES": es, "ES/VaR": es / var}
    return results


# Regenerate data to ensure availability
np.random.seed(42)
n_days = 6000
returns_array = np.random.standard_t(df=4, size=n_days) * 0.01 + 0.0003
norm_loc, norm_scale = stats.norm.fit(returns_array)

# Compare the simulated sample with a fitted normal model
simulated_risk = calculate_var_es(returns_array)

# Normal assumption
norm_var_95 = -stats.norm.ppf(0.05, norm_loc, norm_scale)
norm_var_99 = -stats.norm.ppf(0.01, norm_loc, norm_scale)


# For normal (reported as a positive loss), ES = -mu + sigma * phi(z) / alpha where phi is the pdf
def normal_es(alpha, mu, sigma):
    z = stats.norm.ppf(alpha)
    return -mu + sigma * stats.norm.pdf(z) / alpha


norm_es_95 = normal_es(0.05, norm_loc, norm_scale)
norm_es_99 = normal_es(0.01, norm_loc, norm_scale)
Out[28]:
Console
Risk Measure Comparison: Simulated Sample vs Normal Assumption
=================================================================
Measure                 Simulated       Normal        Ratio
-----------------------------------------------------------------
95% VaR                   2.1177%      2.2856%         0.93x
95% ES                    3.1541%      2.8682%         1.10x
99% VaR                   3.6798%      3.2358%         1.14x
99% ES                    5.0384%      3.7084%         1.36x
-----------------------------------------------------------------

The simulated 95% VaR is lower than the normal estimate,
while ES and the 99% measures are higher for this seed.

The table does not show uniform underestimation. For this seed, the simulated 95% VaR is about 7% lower than the fitted-normal estimate, while simulated 95% ES is about 10% higher. At 99%, simulated VaR is about 14% higher and simulated ES about 36% higher. The example therefore supports checking each risk measure and confidence level separately; it does not establish a universal multiplier or a claim about regulatory capital.

Out[29]:
Visualization
Grouped bar chart comparing simulated-sample and fitted-normal VaR and ES at 95% and 99% confidence levels.
Comparison of VaR and Expected Shortfall for a simulated heavy-tailed sample and its fitted normal model. The normal model slightly overestimates 95% VaR but underestimates ES and both 99% measures for this seed.

Choosing a Distribution

Distribution choice begins with the variable being modeled. A normal or Student's t model can describe signed continuous returns, a lognormal model can describe a positive level generated by exponentiated normal log changes, a binomial model counts successes in a fixed number of trials, and a Poisson model counts events over an interval. A power-law tail specification is a statement about remote-tail decay rather than a complete model for the center of a distribution.

The horizon and decision also matter. A model that is adequate for central daily-return behavior may be poor for multi-day dependence, extreme quantiles, or expected shortfall. Fit diagnostics should therefore match the intended use: inspect residual dependence and stability across regimes, compare empirical and fitted quantiles, and check the particular risk measure or payoff rather than relying on one summary statistic.

Summary

The normal distribution provides a tractable baseline, while the lognormal distribution connects normal log changes to positive modeled prices. Binomial and Poisson distributions address different discrete questions: a fixed number of trials versus event counts over an interval. Student's t and power-law models make slower tail decay explicit, but their parameters and moment conditions must be checked rather than assumed.

No distribution is selected by its name alone. The relevant variable, horizon, dependence structure, evidence, and decision determine whether its assumptions are useful.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about probability distributions in finance.

Probability Distributions in Finance

Question 1 of 80 of 8 completed
According to the 68-95-99.7 rule, approximately what percentage of observations from a normal distribution fall within two standard deviations of the mean?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025probabilitydistributions-2, author = {Michael Brenndoerfer}, title = {Probability Distributions in Finance: Normal, Lognormal}, year = {2025}, url = {https://mbrenndoerfer.com/writing/probability-distributions-quantitative-finance}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2025). Probability Distributions in Finance: Normal, Lognormal. Retrieved from https://mbrenndoerfer.com/writing/probability-distributions-quantitative-finance
MLAAcademic
Michael Brenndoerfer. "Probability Distributions in Finance: Normal, Lognormal." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/probability-distributions-quantitative-finance>.
CHICAGOAcademic
Michael Brenndoerfer. "Probability Distributions in Finance: Normal, Lognormal." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/probability-distributions-quantitative-finance.
HARVARDAcademic
Michael Brenndoerfer (2025) 'Probability Distributions in Finance: Normal, Lognormal'. Available at: https://mbrenndoerfer.com/writing/probability-distributions-quantitative-finance (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2025). Probability Distributions in Finance: Normal, Lognormal. https://mbrenndoerfer.com/writing/probability-distributions-quantitative-finance

About the author

Continue with the full handbook

This chapter is part of Quantitative Finance. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Quantitative Finance
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.