Part of Quantitative Finance
Covers probability distributions essential for quantitative finance: normal, lognormal, binomial, Poisson, and fat-tailed distributions with Python examples.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Common Probability Distributions in Finance
Probability distributions are a core modeling tool in quantitative finance. Many portfolio-risk estimates, option-pricing models, and default-likelihood models require assumptions about how uncertain outcomes are distributed. A misspecified distribution can materially change estimated tail probabilities, option values, and risk limits.
The choice becomes especially visible in the tails. Later in this chapter, the same simulated heavy-tailed sample produces a lower 95% VaR but higher expected shortfall and higher 99% risk measures than a fitted normal model. That mixed result is the practical reason to inspect a distribution's assumptions rather than relying on a blanket rule about normal models.
This chapter covers several common models in quantitative finance: the normal distribution as a baseline for returns and mean-variance analysis, the lognormal distribution for positive quantities such as modeled asset prices, discrete distributions like the binomial and Poisson for counts and event models, and heavier-tailed alternatives. Whether any of them fits a particular dataset or decision well must be checked against the relevant horizon and evidence. For each distribution, we examine both the mathematical properties and the financial intuition behind its use.
The Normal Distribution
The normal distribution, also called the Gaussian distribution, is a common baseline in financial modeling. One mathematical reason is the classical Central Limit Theorem: standardized sums of independent, identically distributed random variables with finite variance converge in distribution to a normal law. Thinking of a return as the aggregate effect of many small shocks can motivate a normal approximation, but market shocks need not be independent or identically distributed, so this is a modeling heuristic rather than a theorem about returns.
A continuous probability distribution characterized by its symmetric, bell-shaped curve. It is fully specified by two parameters: the mean (center) and variance (spread). Also known as the Gaussian distribution.
Mathematical Properties
To understand what it means for a random variable to follow a normal distribution, we must first grasp the concept of a probability density function. Unlike discrete random variables that take on specific values with certain probabilities, a continuous random variable like one following a normal distribution can take any value along the real line. Probability over a region is the integral of the density over that region; density at one point is not itself a probability. A higher density at a point means that sufficiently small, equal-width neighborhoods around that point have greater probability than neighborhoods around a lower-density point.
A random variable follows a normal distribution with mean and variance , written , if its probability density function is:
Each component serves a specific purpose. Let's examine each component:
- : the value at which we evaluate the density
- : the mean (center of the distribution)
- : the standard deviation (controls the spread)
- : a normalization constant ensuring the density integrates to 1
- : one half of the squared standardized distance from to the mean
The heart of the normal distribution lies in the exponential term. The quantity measures the squared distance between the point and the center of the distribution . Dividing by scales this distance relative to the spread of the distribution. A value two units away from the mean is much more unusual when than when , because the standardized distance is larger. The negative sign in the exponent ensures that the function decreases as we move away from the mean, and the exponential function turns this squared distance into a probability density that decays smoothly and symmetrically.
The exponential of a negative quadratic creates the characteristic bell shape: values near the mean have high probability density, while values far from the mean have exponentially decreasing density. For a standard normal variable, the two-sided probability beyond four standard deviations is about , and beyond five it is about . Whether those benchmark probabilities adequately describe a financial return series is an empirical question tied to the asset, horizon, and sample. The quadratic term in the exponent is the key: because squaring grows faster than linear growth, the tails of the normal distribution thin out quickly.
The normalization constant deserves special attention. This factor ensures that when we integrate the density function over all possible values from negative infinity to positive infinity, we obtain exactly 1. This is the total probability. The arises from the Gaussian integral, one of the most famous results in calculus, while the in the denominator accounts for how wider distributions (larger ) must have lower peak heights to maintain the same total area under the curve.
Key parameters:
- Mean (): The center of the distribution and the expected value of . In finance, this represents the expected return.
- Variance (): Measures the spread of the distribution. The square root, , is the standard deviation, which in finance we call volatility.
- Skewness: Zero for the normal distribution, meaning it is perfectly symmetric around the mean.
- Kurtosis: Equal to 3 for the normal distribution (or excess kurtosis of 0). Kurtosis is the fourth standardized moment; by itself, it does not uniquely determine how probability is divided among the center, shoulders, and tails.
The symmetry of the normal distribution is both its elegance and its limitation. Symmetry is a modeling assumption: it assigns equal probabilities to equal-sized deviations above and below the mean. Some financial return series exhibit asymmetry, with the direction and magnitude depending on the asset, sampling horizon, and market regime. This possibility is one reason to treat the normal distribution as a baseline rather than a universal model of returns.
The standard normal distribution is the special case where and . Any normal variable can be standardized:
This standardization process turns any normal random variable into the canonical standard normal form. The operation first centers the distribution by subtracting the mean (so the new center is at zero), then scales it by dividing by the standard deviation (so the new spread equals one). This step preserves the essential shape of the distribution while placing it in a standard coordinate system.
This standardization is central because it allows us to use a single table of probabilities for any normal distribution. The cumulative distribution function of the standard normal, denoted , gives the probability that a standard normal variable is less than . Before computers, statisticians relied on printed tables of values, and the standardization formula allowed them to answer questions about any normal distribution using this single reference. Even today, the standard normal is the universal reference point: when we speak of a "3-sigma event," we mean an outcome more than three standard deviations from the mean, regardless of what the actual mean and standard deviation are.
The 68-95-99.7 Rule
One of the most useful properties of the normal distribution is how probability mass is concentrated around the mean:
- Approximately 68% of observations fall within one standard deviation of the mean
- Approximately 95% fall within two standard deviations
- Approximately 99.7% fall within three standard deviations
These percentages emerge directly from integrating the normal density function, but their approximate values (68, 95, and 99.7) are worth committing to memory because they provide instant intuition for assessing how unusual any particular observation is. If someone tells you that an event was a "two-sigma move," you immediately know that such events occur roughly 5% of the time, or about one trading day in twenty.
This rule provides quick mental estimates for risk. If daily stock returns have a standard deviation of 1%, then under normality assumptions, the probability of an absolute move larger than 3% is about 0.27%. That is about once every 370 trading days, or roughly once every 1.5 trading years. This calculation reveals both the power and the danger of the normal assumption: it gives us precise quantitative predictions, but those predictions can be spectacularly wrong if the true distribution differs from normal.



Applications in Finance
The normal distribution appears throughout quantitative finance:
Portfolio theory
Portfolio theory often summarizes portfolios by expected return and variance. Normality is sufficient for those two moments to determine the full distribution of portfolio returns, but the mean-variance framework itself does not require normally distributed returns. When candidate portfolio returns are normal, expected utility for a fixed utility function can be evaluated from the mean and variance because those parameters determine the distribution.
Value at Risk (VaR)
VaR calculations often assume normal returns. Under that model, the signed VaR loss quantile at confidence level is:
This formula translates the abstract concept of a tail quantile into a concrete loss estimate. Under the assumption that returns are normally distributed with mean and standard deviation , we seek the threshold such that returns fall below this level only percent of the time. The standard normal quantile tells us how many standard deviations below the mean this threshold lies.
Each component of this formula plays a specific role:
- : the signed loss quantile at confidence level ; under a continuous model, losses exceed this threshold with probability
- : the expected return over the time horizon
- : the volatility (standard deviation of returns) over the time horizon
- : the -quantile of the standard normal distribution (e.g., for 95% confidence)
The formula works because under normality, the -percentile of returns is exactly . For example, the 95% VaR uses , meaning we expect losses to exceed this threshold only 5% of the time. Taking the negative of this return quantile yields the signed loss quantile . It is positive when , zero at equality, and negative when the modeled lower-tail return quantile is still a gain. A reporting policy that disallows negative VaR can use , but that floor is a reporting convention rather than the raw quantile.

Option pricing
The Black-Scholes model assumes that log-returns are normally distributed, as we will explore in a later chapter.
import numpy as np
from scipy import stats
def calculate_var_normal(mu, sigma, confidence_level=0.95, horizon=1):
"""
Calculate Value at Risk assuming normal returns.
Parameters:
mu: Expected return (daily)
sigma: Volatility (daily standard deviation)
confidence_level: Confidence level (e.g., 0.95 for 95% VaR)
horizon: Time horizon in days
Returns:
Signed VaR loss quantile; it is positive only when the modeled
lower-tail return quantile is negative
"""
# Scale parameters to the horizon
mu_scaled = mu * horizon
sigma_scaled = sigma * np.sqrt(horizon)
# Find the quantile
z = stats.norm.ppf(1 - confidence_level)
# VaR is the negative of the return quantile; the result may be negative
var = -(mu_scaled + z * sigma_scaled)
return var
# Example: Daily return of 0.05% with 1.5% volatility
daily_return = 0.0005
daily_vol = 0.015
var_95_1day = calculate_var_normal(daily_return, daily_vol, 0.95, 1)
var_99_1day = calculate_var_normal(daily_return, daily_vol, 0.99, 1)
var_95_10day = calculate_var_normal(daily_return, daily_vol, 0.95, 10)Portfolio with daily return 0.05% and volatility 1.50% 1-day 95% VaR: 2.42% 1-day 99% VaR: 3.44% 10-day 95% VaR: 7.30%
These VaR figures are loss thresholds, not maximum or expected losses. The 1-day 95% VaR of approximately 2.4% means that, under the fitted normal model, losses exceed 2.4% with 5% probability. The 10-day VaR is higher here because the volatility term grows with and outweighs the mean term, which grows with . Under the iid normal model, , so VaR need not increase with the horizon for every parameter choice. The square-root scaling of the volatility term follows from variance addition: if daily variances are , then the variance over days is , and the standard deviation is .
The Lognormal Distribution
Returns may be approximately normal, but a nonnegative stock-price level cannot follow an exact nondegenerate normal law because that law assigns some probability below zero. A local normal approximation may still be useful. A stock price can rise 50% but cannot fall more than 100%. This asymmetry is captured by the lognormal distribution, which constrains values to be positive. The lognormal distribution arises when multiplicative returns produce an accumulated log return that is normally distributed: exponentiating that normal log return gives a positive, lognormally distributed price. Multiplicative compounding by itself does not determine the distribution.
A probability distribution of a random variable whose logarithm is normally distributed. If , then follows a lognormal distribution. Asset prices are often modeled as lognormal.
From Returns to Prices
The connection between normal returns and lognormal prices is one of the most important relationships in quantitative finance. If log-returns are normally distributed, then prices are lognormally distributed. Consider a stock with initial price and continuously compounded return over some period:
This equation captures a key relationship in how prices evolve. The exponential function converts the additive world of log-returns into the multiplicative world of price changes. When , the stock has grown by a factor of , or about 10.5%. The slight difference between the log-return and the percentage return becomes more pronounced for larger moves.
The components of this price equation are:
- : stock price at time
- : initial stock price
- : continuously compounded return over the period
- : the growth factor, converting log-return to price ratio
If , then follows a lognormal distribution. This relationship is fundamental: the exponential function turns the symmetric, unbounded normal distribution into a right-skewed distribution bounded below by zero, which matches the behavior of asset prices that can rise without limit but cannot fall below zero. The step preserves the elegant properties of the normal distribution while mapping them onto the economically sensible domain of positive prices.
The probability density function of the lognormal distribution takes a form that reveals its connection to the normal:
Compare this to the normal density to see the main differences. The argument of the exponential now involves rather than , which reflects that we are measuring distance in the log-space. The presence of in front of the exponential arises from the chain rule when transforming from the normal to the lognormal through the change of variables.
Each element of this formula serves a distinct purpose:
- : the value at which we evaluate the density (must be positive)
- : the mean of , not the mean of itself
- : the standard deviation of , not of itself
- : this factor (compared to the normal PDF) accounts for the change of variables from to
The key insight is that and are parameters of the underlying normal distribution of log-values, not the lognormal variable itself. This distinction often causes confusion: when we say a stock has lognormal volatility of 20%, we mean the standard deviation of its log-returns, not of the returns themselves. The actual mean and variance of a lognormal variable are:
These formulas reveal how the parameters of the underlying normal distribution translate into the moments of the lognormal:
- : the expected value (mean) of the lognormal variable
- : the variance of
- : the mean of , the underlying normal distribution
- : the variance of
- : the mean is greater than when , and equal to it in the degenerate case
The term in the mean formula is a convexity adjustment: because the exponential function is convex, Jensen's inequality tells us that for any non-constant random variable . Specifically, the curvature of the exponential means that high values of contribute disproportionately to the mean of . This is why the expected value of a lognormal variable includes a variance correction rather than equaling , since higher variance increases this convexity effect. Holding the mean log-return fixed, greater therefore raises . In the GBM notation below, however, denotes arithmetic drift: holding that drift fixed, the log-mean falls by as variance rises, and does not depend on .
Geometric Brownian Motion
A standard baseline model for stock price dynamics is geometric Brownian motion (GBM), which produces lognormally distributed prices. This model has been central to quantitative finance since Samuelson's work in the 1960s and forms the foundation for the Black-Scholes option pricing formula. Understanding GBM requires familiarity with stochastic calculus, but the core idea is intuitive. Prices evolve continuously with both a deterministic trend and random fluctuations proportional to the current price.
Under GBM, the stock price evolves according to:
This stochastic differential equation describes infinitesimal changes in the stock price. Each term has a clear interpretation:
- : infinitesimal change in stock price
- : current stock price
- : drift rate (expected return per unit time)
- : diffusion volatility (annualized when time is measured in years); over an interval , the log-return standard deviation is
- : infinitesimal time increment
- : increment of a Wiener process (Brownian motion), with
The first term represents the deterministic trend, while the second term captures random fluctuations proportional to the current price. The key feature is that both terms are proportional to : a $100 stock and a $10 stock with the same parameters will have the same percentage return distribution, but the absolute dollar changes will be ten times larger for the more expensive stock. This proportionality is what makes GBM appropriate for modeling asset prices, which tend to grow (or shrink) in percentage terms.
Solving this stochastic differential equation yields:
This solution tells us exactly where the stock price will be at any future time, given the realized path of the Brownian motion. The formula has a clear structure:
- : stock price at time
- : initial stock price at time 0
- : drift rate (annualized expected return)
- : volatility (annualized standard deviation)
- : time horizon in years
- : value of the Wiener process at time , with
- : Itô correction term arising from the quadratic variation of Brownian motion
The Itô correction appears because we apply Itô's lemma to : the second derivative term contributes to the drift. This correction is one of the key differences between ordinary calculus and stochastic calculus. In ordinary calculus, if we know , we would conclude that . However, when follows a diffusion process, the additional term arises from the non-zero quadratic variation of Brownian motion. For , the expected GBM log-growth rate is below its arithmetic drift ; at , the two are equal.
This model forms the foundation of the Black-Scholes option pricing formula.


Simulating Price Paths
Monte Carlo simulation of stock prices under GBM is simple. The approach below draws exact GBM transition increments at each point on the chosen time grid and cumulatively sums them. Unlike an Euler approximation, this produces the exact GBM distribution at the grid points for constant parameters; it does not simulate between-grid extrema.
import numpy as np
def simulate_gbm_paths(S0, mu, sigma, T, n_steps, n_paths):
"""
Simulate geometric Brownian motion paths.
Parameters:
S0: Initial stock price
mu: Annualized drift (expected return)
sigma: Annualized volatility
T: Time horizon in years
n_steps: Number of time steps
n_paths: Number of simulation paths
Returns:
Array of shape (n_paths, n_steps + 1) containing price paths
"""
dt = T / n_steps
# Generate random increments
np.random.seed(42)
Z = np.random.standard_normal((n_paths, n_steps))
# Calculate log returns for each step
log_returns = (mu - 0.5 * sigma**2) * dt + sigma * np.sqrt(dt) * Z
# Cumulative sum to get log prices relative to initial
log_prices = np.zeros((n_paths, n_steps + 1))
log_prices[:, 0] = np.log(S0)
log_prices[:, 1:] = np.log(S0) + np.cumsum(log_returns, axis=1)
# Convert to prices
prices = np.exp(log_prices)
return prices
# Simulate 1 year of daily prices
S0 = 100
mu = 0.10 # 10% expected annual return
sigma = 0.25 # 25% annual volatility
T = 1.0
n_steps = 252 # Trading days
n_paths = 1000
paths = simulate_gbm_paths(S0, mu, sigma, T, n_steps, n_paths)
final_prices = paths[:, -1]Initial price: $100.00 Expected final price (theoretical): $110.52 Mean simulated final price: $110.44 Median simulated final price: $108.19 5th percentile: $71.15 95th percentile: $160.84
Notice that the mean simulated price is close to the theoretical expected value of , but the median is lower. This is typical of a non-degenerate lognormal distribution: its mean exceeds its median, and its mode is lower still. These quantities answer different questions. The median describes the middle terminal outcome across simulated paths, the mean describes ensemble expected terminal wealth, and long-run performance of repeatedly compounded wealth is governed by average log growth rather than by the arithmetic mean terminal price alone.


Discrete Distributions: Binomial and Poisson
Not all financial phenomena are continuous. Events like defaults, dividend announcements, or large jumps in prices are discrete occurrences. The binomial and Poisson distributions model these situations and give a framework for counting how many times something happens rather than measuring continuous quantities. These distributions are particularly important in credit risk, where we count defaults in a portfolio, and in derivatives pricing, where discrete lattice models approximate continuous price processes.
The Binomial Distribution
The binomial distribution models the number of successes in a fixed number of independent trials, each with the same probability of success. The classic example is flipping a coin multiple times and counting heads, but the same mathematics applies to questions such as: How many stocks in a portfolio will outperform? How many corporate bonds will default? How many days this month will have positive returns?
The probability distribution of the number of successes in independent Bernoulli trials, where each trial has success probability . Used in discrete option pricing models and credit risk modeling.
If , the probability mass function is:
This formula answers the question: if we have independent trials, each with success probability , what is the probability of getting exactly successes? The answer decomposes into two multiplicative components: the number of ways the successes can be arranged, and the probability of any particular arrangement.
Each component of this formula serves a specific purpose:
- : the number of successes we're computing the probability for
- : total number of independent trials
- : probability of success on each trial
- : the binomial coefficient, representing the number of ways to choose successes among trials
- : success-probability factor for specified success positions
- : failure-probability factor for specified failure positions
The formula multiplies the number of arrangements by the probability of each specific arrangement. Intuitively, if you flip a coin times with probability of heads, the PMF tells you the probability of getting exactly heads: you need successes (probability ), failures (probability ), and the binomial coefficient counts how many orderings produce that outcome. The key insight is that because the trials are independent, the probability of any specific sequence of outcomes is simply the product of the individual probabilities, and the binomial coefficient tells us how many such sequences lead to exactly successes.
The mean and variance have elegant formulas that reveal the distribution's structure:
These formulas encode intuitive relationships:
- : the expected number of successes
- : the variance in the number of successes
- : total number of trials
- : probability of success on each trial
- : probability of failure, which appears in variance because maximum variance occurs when (maximum uncertainty)
The expected number of successes, , is simply the number of trials times the success probability, exactly what intuition suggests. The variance formula is more subtle: it shows that variance is maximized when (complete uncertainty about each trial's outcome) and decreases toward zero as approaches 0 or 1 (near certainty). This makes sense because if is very close to 1, almost all trials succeed, leaving little room for variability in the total count.


Binomial Option Pricing Model
The Cox-Ross-Rubinstein (CRR) binomial model prices options by assuming that at each time step, the stock price either moves up by a factor or down by a factor . This discretization turns the continuous option pricing problem into a tree of possible outcomes that can be analyzed one step at a time. While this may seem like a crude approximation to the continuous world of actual markets, the model provides deep insights into option pricing and converges to the continuous Black-Scholes formula as the number of time steps increases.
After steps, the stock price has experienced up moves and down moves, giving:
This formula shows how the final price depends entirely on the number of up moves, not their specific timing. The components are:
- : stock price at maturity after steps
- : initial stock price
- : up factor (price multiplier for an up move)
- : down factor (price multiplier for a down move)
- : number of up moves
- : number of down moves
Whether the stock rises first and then falls, or falls first and then rises, the final price is the same as long as the total number of up and down moves matches. This recombination supports efficient pricing for European options, whose payoff depends only on the final price. Path-dependent claims such as Asian options require the tree state to be augmented with path information; the simple terminal-price state is not sufficient.
The number of up moves follows a binomial distribution under the risk-neutral measure, with probability of an up move:
This risk-neutral probability is central to derivatives pricing. Notice that is not the actual probability that the stock will go up. It is the artificial probability that makes the expected return on the stock equal to the risk-free rate. The components are:
- : risk-neutral probability of an up move
- : risk-free interest rate (annualized)
- : length of each time step in years
- : up factor (stock price multiplier for an up move)
- : down factor (stock price multiplier for a down move)
- : the growth factor for a risk-free investment over one time step
This formula is derived from the no-arbitrage condition: the expected stock return under the risk-neutral measure must equal the risk-free rate. Setting and solving for yields this result. In this complete one-period binomial model, a replicable option is priced from the risk-neutral probability rather than the physical up-move probability. This conclusion depends on no arbitrage, market completeness, and the ability to form the replicating portfolio; it is not a universal statement about every derivative market.
Assuming , this is a strict probability and the one-period stock-and-bond model is arbitrage-free only when
which is equivalent to . Inputs outside this range do not define valid CRR risk-neutral probabilities.

import numpy as np
from scipy.special import comb
def binomial_option_price(S0, K, T, r, sigma, n_steps, option_type="call"):
"""
Price a European option using the CRR binomial model.
Parameters:
S0: Initial stock price
K: Strike price
T: Time to expiration (years)
r: Risk-free rate (annualized)
sigma: Volatility (annualized)
n_steps: Number of time steps
option_type: 'call' or 'put'
Returns:
Option price
"""
if option_type not in {"call", "put"}:
raise ValueError("option_type must be 'call' or 'put'")
dt = T / n_steps
# CRR parameters
u = np.exp(sigma * np.sqrt(dt))
d = 1 / u
q = (np.exp(r * dt) - d) / (u - d)
if not 0 < q < 1:
raise ValueError(
"CRR no-arbitrage condition requires d < exp(r*dt) < u"
)
# Calculate all possible final stock prices
k = np.arange(n_steps + 1)
final_prices = S0 * (u**k) * (d ** (n_steps - k))
# Calculate payoffs
if option_type == "call":
payoffs = np.maximum(final_prices - K, 0)
else:
payoffs = np.maximum(K - final_prices, 0)
# Risk-neutral probabilities
probs = comb(n_steps, k) * (q**k) * ((1 - q) ** (n_steps - k))
# Expected payoff under risk-neutral measure, discounted
option_price = np.exp(-r * T) * np.sum(probs * payoffs)
return option_price
# Price a call option
S0 = 100
K = 100
T = 0.5 # 6 months
r = 0.05
sigma = 0.20Binomial Model Call Option Prices ----------------------------------- Steps: 10 | Price: $6.7494 Steps: 50 | Price: $6.8605 Steps: 100 | Price: $6.8746 Steps: 200 | Price: $6.8817 Steps: 500 | Price: $6.8859 ----------------------------------- Black-Scholes: | Price: $6.8887
As the number of steps increases, the CRR binomial price converges to the Black-Scholes price under the model's regularity conditions. The error can oscillate with the number of steps because the strike sits differently relative to successive lattices. For sufficiently regular European payoffs, suitably controlled sequences often have first-order error on the scale of , but that rate is not uniform across payoffs and parameter choices. The convergence reflects a mathematical connection between the binomial random walk and Brownian motion, formalized in Donsker's theorem.

The Poisson Distribution
The Poisson distribution models the number of events occurring in a fixed interval when events happen at a constant average rate and independently of each other. Unlike the binomial distribution, which assumes a fixed number of trials, the Poisson distribution allows for any number of events, making it useful for modeling rare occurrences over a continuous interval.
A discrete probability distribution expressing the probability of a given number of events occurring in a fixed interval, when events occur at a constant average rate independently of each other. Used for modeling defaults, jumps, and arrival processes.
If , the probability mass function is:
This formula has a clear structure when we trace its derivation from first principles. Imagine dividing the interval into many tiny subintervals, each with a small probability of containing an event. As the number of subintervals goes to infinity while keeping the expected total number of events fixed at , we obtain the Poisson distribution.
The components of this formula are:
- : the number of events we're computing the probability for
- : the average rate of events (expected number of events in the interval)
- : the probability of zero events, which is a normalization factor
- : the intensity factor for events; together with , it makes largest at for
- : accounts for the indistinguishability of event orderings
The Poisson distribution arises as the limit of a binomial distribution when and while remains constant. More generally, a Poisson approximation requires independent rare-event opportunities whose largest individual probability vanishes while the sum of their probabilities approaches . A finite expected count alone is not enough: dependence or a non-negligible individual probability can produce a different count distribution. Credit portfolios with small independent default probabilities and jump models with independent small-interval arrivals can satisfy the approximation; clustered defaults and jumps require a different model.
Both the mean and variance equal :
This property deserves emphasis:
- : the expected number of events
- : the variance in the number of events
- : the average rate parameter
The equality of mean and variance is a distinctive property of the Poisson distribution. When analyzing count data, if the observed variance significantly exceeds the mean (overdispersion), this suggests the Poisson model may be inadequate. Dependence is one possible cause, but rate heterogeneity, omitted covariates, and excess zeros can also produce overdispersion. For instance, defaults may cluster because borrowers share economic conditions even when individual default mechanisms differ.


Applications of the Poisson Distribution
The Poisson distribution is particularly useful for modeling:
-
Jump processes in asset prices: The Merton jump-diffusion model adds jumps whose arrivals follow a Poisson process and whose sizes follow a separate distribution to geometric Brownian motion. The stock price follows GBM between jumps and moves discontinuously at jump times. This combination represents continuous diffusion-like fluctuations together with occasional discontinuous price moves.
-
Credit defaults in a portfolio: If defaults occur independently with a small probability, the number of defaults in a large portfolio is approximately Poisson distributed. This approximation underlies many credit risk models, although the assumption of independence often breaks down during systemic crises, when many borrowers default simultaneously.
-
Order arrivals: In market microstructure, the arrival of orders is often modeled as a Poisson process. This assumption implies that the time between consecutive orders follows an exponential distribution, without clustering or periodicity in order flow. While real markets show more complex arrival patterns, the Poisson model provides a useful baseline.
import numpy as np
def simulate_jump_diffusion(
S0, mu, sigma, lambda_jump, mu_jump, sigma_jump, T, n_steps, n_paths
):
"""
Simulate Merton jump-diffusion price paths.
Parameters:
S0: Initial stock price
mu: Drift of the diffusion component
sigma: Volatility of the diffusion component
lambda_jump: Average number of jumps per year (Poisson intensity)
mu_jump: Mean of log-jump size
sigma_jump: Std dev of log-jump size
T: Time horizon in years
n_steps: Number of time steps
n_paths: Number of paths to simulate
Returns:
Array of simulated price paths
"""
dt = T / n_steps
np.random.seed(42)
# Diffusion component
Z_diffusion = np.random.standard_normal((n_paths, n_steps))
diffusion = (mu - 0.5 * sigma**2) * dt + sigma * np.sqrt(dt) * Z_diffusion
# Jump component
# Number of jumps in each interval (Poisson distributed)
n_jumps = np.random.poisson(lambda_jump * dt, (n_paths, n_steps))
# Log jump sizes (normal)
jump_sizes = np.zeros((n_paths, n_steps))
for i in range(n_paths):
for j in range(n_steps):
if n_jumps[i, j] > 0:
jumps = np.random.normal(mu_jump, sigma_jump, n_jumps[i, j])
jump_sizes[i, j] = np.sum(jumps)
# Combine diffusion and jumps
log_returns = diffusion + jump_sizes
# Calculate prices
log_prices = np.zeros((n_paths, n_steps + 1))
log_prices[:, 0] = np.log(S0)
log_prices[:, 1:] = np.log(S0) + np.cumsum(log_returns, axis=1)
return np.exp(log_prices)
# Compare GBM vs Jump-Diffusion
S0 = 100
mu = 0.10
sigma = 0.20
T = 1.0
n_steps = 252
n_paths = 10000
# Pure GBM paths
gbm_paths = simulate_gbm_paths(S0, mu, sigma, T, n_steps, n_paths)
# Jump-diffusion paths
jd_paths = simulate_jump_diffusion(
S0,
mu,
sigma,
lambda_jump=2, # 2 jumps per year on average
mu_jump=-0.05, # Jumps tend to be negative
sigma_jump=0.10,
T=T,
n_steps=n_steps,
n_paths=n_paths,
)
gbm_returns = np.diff(np.log(gbm_paths), axis=1).flatten()
jd_returns = np.diff(np.log(jd_paths), axis=1).flatten()Return Distribution Comparison ================================================== Statistic GBM Jump-Diffusion -------------------------------------------------- Mean 0.000312 -0.000083 Std Dev 0.012609 0.016096 Skewness -0.0006 -3.1474 Excess Kurtosis -0.0002 55.5761 Min -0.0605 -0.4391 Max 0.0611 0.3572
With the negatively centered jump-size distribution used here, the simulated jump-diffusion returns have a more negative sample skewness and a larger fourth standardized moment than the simulated GBM returns. Changing the jump-size distribution can change the sign of skewness. Likewise, higher kurtosis alone does not determine a unique allocation of probability among the center, shoulders, and tails; the histogram and quantiles must be examined directly.


Fat-Tailed Distributions
A normal distribution is only one possible model for returns. When an observed or simulated return distribution places more probability in sufficiently remote tails than a fitted normal model, normal tail calculations can misstate extreme-event frequencies and downstream risk measures. The examples below use a seeded Student's t sample to show how to diagnose that mismatch without presenting synthetic observations as market data.
Evidence of Fat Tails
The term "fat tails" refers to distributions where extreme events occur more frequently than a normal distribution would suggest. A distribution with fat tails places more probability mass in the regions far from the mean, meaning that large positive or negative values are more common than the normal distribution's rapid exponential decay would predict. Several ways to measure this phenomenon:
Kurtosis: The normal distribution has a kurtosis of 3 (or excess kurtosis of 0). Kurtosis is the fourth central moment divided by the variance squared. A higher value signals more weight on large standardized deviations overall, but it does not by itself order every tail probability or identify where probability moved between the center and shoulders.
Tail exponents: For power-law-tailed distributions, the probability of extreme events decays polynomially rather than exponentially. The tail exponent characterizes how fast this decay occurs. A smaller means slower polynomial decay and fatter tails, while a larger means faster polynomial decay and thinner tails within the power-law family. For every finite , however, the asymptotic tail remains heavier than a normal tail.
Simulated Daily Returns (Heavy-Tailed, t-distribution df=4) ================================================== Number of observations: 6,000 Mean daily return: 0.0079% Daily volatility: 1.3944% Skewness: 0.086 Excess Kurtosis: 5.719 Extreme Event Frequency -------------------------------------------------- Event Observed Expected (Normal) Ratio -------------------------------------------------- |r| > 3σ 80 16.2 4.9x |r| > 4σ 30 0.4 78.9x |r| > 5σ 9 0.0 2616.4x
The seeded sample was generated from a heavy-tailed distribution, so its departure from the normal benchmark is expected rather than empirical evidence about markets. For this seed, absolute moves beyond four sample standard deviations occur 30 times, compared with about 0.38 expected in 6,000 independent normal draws, a ratio of about 79. Nine observations exceed five sample standard deviations, compared with about 0.0034 under that same normal benchmark. The exercise demonstrates how a known heavy-tailed data-generating process changes extreme-event counts; it does not estimate their frequency in an observed market.
The Student's t Distribution
The Student's t distribution provides a simple way to model fat tails while retaining many convenient properties of the normal distribution. William Sealy Gosset originally developed the t distribution (writing under the pseudonym "Student") for quality control at Guinness brewery. It is now widely used in finance for modeling the heavy tails observed in return data.
A probability distribution with heavier tails than the normal distribution, characterized by its degrees of freedom parameter . As , it converges to the normal distribution. For small , it assigns higher probability to extreme events.
The probability density function of the standard t distribution with degrees of freedom is:
This formula may look complex, but its structure shows the key difference between the t distribution and the normal distribution. While the normal distribution uses the exponential function which decays very rapidly, the t distribution uses a power function which decays much more slowly. This slower decay is precisely what produces the heavier tails.
Each component:
- : the value at which we evaluate the density
- : degrees of freedom parameter (controls tail thickness; lower means fatter tails)
- : the gamma function, generalizing the factorial to non-integer values ( for positive integers)
- : a normalization constant ensuring the density integrates to 1
- : the core term that produces polynomial rather than exponential tail decay
The polynomial decay versus the normal's exponential decay is why the t distribution has heavier tails. To understand this intuitively, consider what happens as becomes large. For the normal distribution, doubling quadruples the magnitude of the exponent because , so the density decays extremely rapidly. For the t distribution, the density is asymptotically proportional to , so doubling a large threshold reduces the density by about . Its one-sided survival probability is proportional to , so doubling a large threshold reduces that probability by about . As , the polynomial term converges to the exponential, and the t distribution approaches the normal.
Key properties of the t distribution:
- Degrees of freedom (): Controls tail thickness. Lower values mean fatter tails.
- Convergence: As , the t distribution approaches the standard normal.
- Variance: For , the variance is , which is always greater than 1. For , the mean exists but the variance is infinite. For , the mean and therefore the variance are undefined. The second raw moment diverges throughout .
- Kurtosis: For , excess kurtosis is , always positive and decreasing in . For , the kurtosis is undefined because the fourth moment does not exist.
Degrees of freedom must be estimated for the dataset and model at hand; there is no universal range for financial returns. Values near 4 are nevertheless useful for understanding the boundary behavior of the family: for , the fourth moment is undefined, while for it is finite. The seeded example below uses deliberately to demonstrate a particularly heavy-tailed case.


The log-scale plot on the right reveals how dramatically the distributions differ in their tails. At 5 standard deviations, the t distribution with 3 degrees of freedom assigns probability thousands of times higher than the normal distribution. This difference, while invisible on a linear scale where both probabilities appear to be essentially zero, has enormous practical significance for risk management. The one-sided probability of a 5-sigma event under a normal distribution is about 1 in 3.5 million observations; under a t distribution with 3 degrees of freedom scaled to have unit variance, it is about 1 in 600 observations. If each observation is treated as an independent daily return and a year contains 252 trading days, those probabilities imply expected waiting times of roughly 13,900 trading years and 2.4 trading years, respectively. These are model-based average waiting times, not schedules for when events will occur.
Fitting the t Distribution to Simulated Returns
We can fit a scaled t distribution to the seeded heavy-tailed sample and compare its maximized log-likelihood with that of a fitted normal distribution. Because the sample was generated from a t distribution, this is a reproducibility check and model-comparison demonstration, not evidence about observed market returns.
import numpy as np
from scipy import stats
# Fit a scaled t distribution to the simulated returns
# The t distribution has location (mu), scale (sigma), and degrees of freedom (nu)
# Use the returns array from the fat-tails-evidence cell
# Ensure returns is available and convert to numpy array
# Regenerate data to ensure availability (matching fat-tails-evidence cell)
np.random.seed(42)
n_days = 6000
daily_returns = np.random.standard_t(df=4, size=n_days) * 0.01 + 0.0003
returns_array = daily_returns
# Maximum likelihood estimation
t_params = stats.t.fit(returns_array)
t_nu, t_loc, t_scale = t_params
# Fit normal for comparison
norm_loc, norm_scale = stats.norm.fit(returns_array)
# Calculate log-likelihoods
t_loglik = np.sum(stats.t.logpdf(returns_array, t_nu, t_loc, t_scale))
norm_loglik = np.sum(stats.norm.logpdf(returns_array, norm_loc, norm_scale))Distribution Fitting Results ================================================== Normal Distribution: Location (μ): 0.000079 Scale (σ): 0.013944 Log-likelihood: 17,123 Student's t Distribution: Degrees of freedom (ν): 4.23 Location: 0.000055 Scale: 0.010172 Log-likelihood: 17,526 Log-likelihood improvement: 403 The t distribution provides a substantially better fit to the data.
For this seed, the fitted t distribution has about 4.23 degrees of freedom and improves the maximized log-likelihood by about 403 relative to the fitted normal model. That large likelihood difference is expected because the sample was generated from a t distribution. Exponentiating the difference would produce a maximized likelihood ratio, not posterior odds; posterior model odds would also require priors and marginal likelihoods.


Power-Law Tails
Power-law notation provides a general way to describe polynomial tail decay. Student's t distributions already have power-law tails, but other regularly varying models can use a different exponent or slowly varying factor:
This asymptotic relationship describes the right-tail survival function for large thresholds. It includes the polynomial right tail of a Student's t distribution as a special case. For a nonnegative variable, the exponent controls positive moments; for a signed variable, existence of a full moment also depends on the left tail, which this one-sided relationship does not specify.
The components of this relationship are:
- : the probability that the random variable exceeds value (the survival function)
- : a large threshold value
- : the tail exponent, which controls how fast probabilities decay. Some authors also call a tail index; in another common extreme-value-theory convention, the tail index is
- : denotes asymptotic equivalence. The ratio of the left side to approaches 1 as
A larger means faster polynomial decay and a smaller means slower decay. As an illustration, if and the asymptotic approximation is appropriate at both thresholds, doubling a sufficiently large threshold reduces its exceedance probability by approximately . This is much slower than the exponential falloff of normal tails.
For comparison, the one-sided standard-normal survival probability falls by a factor of about from 4σ to 8σ, roughly eleven orders of magnitude. Under the illustrative cubic law, the corresponding asymptotic reduction is only a factor of 8. Which approximation is appropriate must be established from a specified model or dataset rather than inferred from this contrast alone.
A distribution where the probability of observing a value greater than decays as a power of : . The tail exponent determines how fat the tails are. Lower means fatter tails.

The implications of power-law tails are serious for risk management:
- Moment thresholds require both tails: For and a nonnegative variable with , the right-tail contribution to is finite for and diverges for . For a signed return, the displayed right-tail relation alone cannot guarantee that a full mean or variance exists because the left tail is unspecified. If both tails have comparable exponents near 3, absolute moments of orders below 3, including the mean and variance, are finite, while the fourth moment is not.
- Slow convergence: Under i.i.d. sampling with finite variance ( in this tail model), the Central Limit Theorem applies, but its normal approximation can remain poor at practically available sample sizes.
- Tail concentration and dependence: With power-law marginal tails, a small number of observations can dominate a sample risk estimate. Diversification may be weaker when those tails combine with strong common or tail dependence, but marginal tail thickness alone does not determine diversification benefits.
Implications for Risk Management
Fat tails change how we should think about risk:
VaR model error: A fitted normal model can underestimate sufficiently extreme heavy left-tail quantiles, but it need not underestimate VaR at every confidence level. The direction and magnitude depend on the fitted distributions, skewness, and confidence level. The seeded example below slightly overestimates 95% VaR and underestimates 99% VaR.
Expected shortfall: Under the continuous-loss convention used here, expected shortfall is the mean loss conditional on being at or beyond the VaR threshold. Terminology such as "Conditional VaR" varies across sources, and distributions with atoms require a quantile-integral definition to handle the boundary consistently. For continuous heavy-tailed models, ES can be substantially larger than VaR.
import numpy as np
from scipy import stats
def calculate_var_es(returns, confidence_levels=[0.95, 0.99]):
"""
Calculate VaR and Expected Shortfall from a return sample.
Returns:
Dictionary with VaR and ES for each confidence level
"""
results = {}
for cl in confidence_levels:
# VaR is the negative of the quantile at (1 - confidence level)
var = -np.percentile(returns, (1 - cl) * 100)
# ES is the mean of returns worse than VaR
losses = -returns
es = np.mean(losses[losses >= var])
results[cl] = {"VaR": var, "ES": es, "ES/VaR": es / var}
return results
# Regenerate data to ensure availability
np.random.seed(42)
n_days = 6000
returns_array = np.random.standard_t(df=4, size=n_days) * 0.01 + 0.0003
norm_loc, norm_scale = stats.norm.fit(returns_array)
# Compare the simulated sample with a fitted normal model
simulated_risk = calculate_var_es(returns_array)
# Normal assumption
norm_var_95 = -stats.norm.ppf(0.05, norm_loc, norm_scale)
norm_var_99 = -stats.norm.ppf(0.01, norm_loc, norm_scale)
# For normal (reported as a positive loss), ES = -mu + sigma * phi(z) / alpha where phi is the pdf
def normal_es(alpha, mu, sigma):
z = stats.norm.ppf(alpha)
return -mu + sigma * stats.norm.pdf(z) / alpha
norm_es_95 = normal_es(0.05, norm_loc, norm_scale)
norm_es_99 = normal_es(0.01, norm_loc, norm_scale)Risk Measure Comparison: Simulated Sample vs Normal Assumption ================================================================= Measure Simulated Normal Ratio ----------------------------------------------------------------- 95% VaR 2.1177% 2.2856% 0.93x 95% ES 3.1541% 2.8682% 1.10x 99% VaR 3.6798% 3.2358% 1.14x 99% ES 5.0384% 3.7084% 1.36x ----------------------------------------------------------------- The simulated 95% VaR is lower than the normal estimate, while ES and the 99% measures are higher for this seed.
The table does not show uniform underestimation. For this seed, the simulated 95% VaR is about 7% lower than the fitted-normal estimate, while simulated 95% ES is about 10% higher. At 99%, simulated VaR is about 14% higher and simulated ES about 36% higher. The example therefore supports checking each risk measure and confidence level separately; it does not establish a universal multiplier or a claim about regulatory capital.

Choosing a Distribution
Distribution choice begins with the variable being modeled. A normal or Student's t model can describe signed continuous returns, a lognormal model can describe a positive level generated by exponentiated normal log changes, a binomial model counts successes in a fixed number of trials, and a Poisson model counts events over an interval. A power-law tail specification is a statement about remote-tail decay rather than a complete model for the center of a distribution.
The horizon and decision also matter. A model that is adequate for central daily-return behavior may be poor for multi-day dependence, extreme quantiles, or expected shortfall. Fit diagnostics should therefore match the intended use: inspect residual dependence and stability across regimes, compare empirical and fitted quantiles, and check the particular risk measure or payoff rather than relying on one summary statistic.
Summary
The normal distribution provides a tractable baseline, while the lognormal distribution connects normal log changes to positive modeled prices. Binomial and Poisson distributions address different discrete questions: a fixed number of trials versus event counts over an interval. Student's t and power-law models make slower tail decay explicit, but their parameters and moment conditions must be checked rather than assumed.
No distribution is selected by its name alone. The relevant variable, horizon, dependence structure, evidence, and decision determine whether its assumptions are useful.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about probability distributions in finance.
Probability Distributions in Finance
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Quantitative Finance. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Quantitative FinanceStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!