Part of Quantitative Finance
Covers probability distributions, expected values, Bayes' theorem, and risk measures. Essential foundations for portfolio theory and derivatives pricing.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Probability Theory Fundamentals
Consequential finance decisions often involve uncertainty. Will the stock price rise or fall tomorrow? What's the likelihood a borrower defaults on their loan? How should we price an option when future volatility is unknown? Probability theory provides the mathematical language to quantify uncertainty and make decisions without pretending the outcome is known.
Probability theory turns vague notions like "likely" or "risky" into model- and data-specific numerical statements. For illustration, an analyst might report a hypothetical bond's estimated one-year default probability as 2.3%, or describe a specified return series and estimation window with 16% annualized volatility. The numbers are meaningful only with that scope; the same language supports portfolio analysis, derivatives pricing, and risk management.
The foundations we cover here, random variables, distributions, expected values, and conditional probabilities, appear throughout quantitative finance. Portfolio theory relies on expected returns and covariances. Option pricing depends on probability distributions of future asset prices. Credit risk models use conditional default probabilities. Many supervised-learning methods minimize an empirical loss as an estimate of expected risk. Master these fundamentals, and you'll have the building blocks for the quantitative techniques that follow.
Sample Spaces, Events, and Probability Axioms
Before we can assign probabilities, we need a probability space : a sample space , a sigma-algebra of measurable events, and a probability measure on those events. This framework begins with three foundational concepts for discussing randomness.
The sample space is the set of all possible outcomes of a random experiment. Each element represents one complete outcome that could occur.
The sample space is the foundation upon which all probability calculations rest. Think of it as the complete catalog of everything that could possibly happen when we conduct our random experiment. The key word here is "complete." The sample space must include every conceivable outcome, even those that seem unlikely or undesirable. If an outcome is missing from our sample space, we have no mathematical way to discuss its probability.
Consider rolling a six-sided die. The sample space is , listing every possible result. For a stock's daily return, the sample space might be , since returns cannot be less than -100% but can, in principle, be arbitrarily large (though extreme returns are rare). Notice how the nature of the random experiment determines the appropriate sample space: a finite set for the die, an infinite continuum for stock returns. Choosing the right sample space is the first modeling decision we make, and it shapes all subsequent analysis.
An event is a measurable subset of the sample space, . For a finite example such as a die, we often take , so every subset is an event. Events represent outcomes we care about, such as "rolling an even number" or "the stock return exceeds 5%."
Events emerge from our questions about uncertain situations. While the sample space contains raw outcomes, events package those outcomes into meaningful categories that correspond to questions we want to answer. The event "rolling an even number" is the subset , which collects all the individual outcomes where an even number appears. The event "stock return exceeds 5%" contains every possible return value greater than 0.05. By defining events as subsets, we gain access to the powerful machinery of set theory for combining and manipulating them.
Events can be combined using set operations. If is "stock rises" and is "volume is high," then:
- is "stock rises OR volume is high (or both)"
- is "stock rises AND volume is high"
- is "stock does not rise"
These set operations allow us to build complex events from simpler ones, mirroring how we naturally combine conditions in financial analysis. When a portfolio manager asks "what's the probability that either the market rallies or volatility spikes?" they're implicitly using the union operation. The intersection operation captures joint occurrences, while the complement captures the negation of an event.
The Kolmogorov Axioms
With sample spaces and events defined, we need rules for assigning probabilities to events. Andrey Kolmogorov formalized probability theory in 1933 with three standard axioms that any probability measure must satisfy. They capture basic consistency requirements while leaving room for many different probability measures.
- Non-negativity: for any event
- Normalization: (something must happen)
- Countable additivity: For mutually exclusive events ,
The first axiom says probabilities cannot be negative. The second normalizes the scale by assigning , while follows from the axioms. A nontrivial event can still have probability 1 or 0 without being logically certain or impossible; such events occur almost surely or have probability zero. The third axiom says that for mutually exclusive events, the probability of at least one occurring equals the sum of their individual probabilities.
Together with the measurable space and the surrounding mathematics, these axioms define a probability measure and yield the rules used in this chapter. Several useful properties follow immediately:
- Complement rule:
- Probability bounds:
- Addition rule:
The complement rule follows from the fact that and are mutually exclusive and their union is the entire sample space: . The probability bounds follow from the complement rule combined with non-negativity. The addition rule requires subtracting to avoid double-counting outcomes that belong to both events.
The addition rule accounts for double-counting when events overlap. If the probability of a stock rising is 0.55 and the probability of high volume is 0.30, we cannot simply add these to find the probability of either occurring; we must subtract the probability of both happening together. This correction is essential whenever we combine events that share common outcomes.
Random Variables
While sample spaces and events provide the conceptual foundation, we typically work with random variables that map outcomes to numbers. This mapping enables mathematical operations like addition, multiplication, and the computation of averages. The transition from abstract outcomes to numerical values makes probability theory quantitative; it lets us apply calculus, linear algebra, and optimization to uncertainty.
A real-valued random variable is a measurable function from a probability space to the real numbers: . In practical terms, measurability ensures that statements such as are events whose probabilities are defined. We use capital letters for random variables and lowercase letters for their realized values.
The definition of a random variable as a function might initially seem abstract, but it captures a natural idea: we take the raw outcomes of a random experiment and translate them into numbers that we can analyze mathematically. The word "variable" reflects that the numerical value varies depending on which outcome occurs. The word "random" indicates that we don't know in advance which outcome will occur.
When we say "let be the daily return of a stock," we mean is a function that, for each possible market scenario , produces a number representing that day's return. Before the market closes, is random because we do not know which will occur. After closing, we observe a specific realization (a 2.3% return). The distinction between and its realized value determines which tools apply. Before observation, we work with probabilities and expectations; after observation, we work with data.
Discrete Random Variables
A discrete random variable takes values from a countable set, such as integers or a finite list. Discrete random variables arise when outcomes fall into distinct categories or when we count occurrences of events. The probability mass function (PMF) specifies the probability of each possible value, giving us a complete description of how probability is distributed across the possible outcomes.
For a discrete random variable , the probability mass function is , giving the probability that equals each specific value .
The PMF answers the most direct question we can ask about a discrete random variable: "What's the probability that takes this particular value?" By specifying these probabilities for every possible value, the PMF provides complete information about the random variable's behavior. The subscript in reminds us which random variable we're describing, since different random variables have different PMFs.
Consider a simplified credit rating model where a bond can be in one of three states at year-end:
import numpy as np
# Credit rating transitions: Upgraded, Unchanged, Downgraded
states = ["Upgraded", "Unchanged", "Downgraded"]
probabilities = np.array([0.15, 0.70, 0.15])
# Verify this is a valid PMF (sums to 1)
assert np.isclose(probabilities.sum(), 1.0)P(Upgraded) = 0.15 P(Unchanged) = 0.70 P(Downgraded) = 0.15 Sum of probabilities: 1.00

The probabilities sum to 1 because exactly one outcome must occur: this bond will be upgraded, unchanged, or downgraded. The Kolmogorov axioms require the probabilities of this exhaustive list of mutually exclusive outcomes to sum to 1.
Continuous Random Variables
Financial quantities such as returns, prices, and interest rates are often approximated with continuous random variables, even though observed market values may lie on discrete ticks and empirical distributions may contain atoms. A continuous distribution assigns probability zero to each individual point; this property does not require every intermediate real value to belong to its support. When the distribution is absolutely continuous, a probability density function gives interval probabilities by integration.
For an absolutely continuous random variable , the probability density function (PDF) satisfies . Thus ; endpoint choices are equivalent because individual points have probability zero.
The PDF represents a fundamentally different concept than the PMF. Rather than giving the probability of exact values (which is zero for continuous variables), the PDF describes how probability is spread across the continuum of possible values. Higher density means probability is more concentrated in that region, while lower density means probability is more dispersed. The integral formula captures this: to find the probability of landing in an interval, we accumulate (integrate) the density across that interval.
The density itself is not a probability; it can exceed 1 for concentrated distributions. Only when we integrate over an interval do we obtain a probability between 0 and 1. Think of density as probability per unit length: a density of 2 at a point does not mean probability 2.
The Cumulative Distribution Function
Both discrete and continuous random variables have cumulative distribution functions, giving a unified way to describe probability distributions. The CDF answers the running question "what's the probability of being less than or equal to this value?" as we move along the number line.
The cumulative distribution function (CDF) of a random variable is:
where:
- : CDF evaluated at , giving a probability between 0 and 1
- : probability that takes a value less than or equal to
- As , (impossible to be below all values)
- As , (certain to be below infinity)
The CDF provides a universal framework for discrete and continuous random variables. It is nondecreasing and right-continuous, with limits 0 as and 1 as . These properties make it useful for computing interval probabilities: .
The CDF starts near 0 for very negative values and approaches 1 for large values. If the distribution is absolutely continuous, and almost everywhere. A continuous CDF need not have a density, so this derivative relationship requires the stronger absolute-continuity condition.


Key Probability Distributions in Finance
Certain probability distributions appear repeatedly in quantitative finance, and each encodes specific assumptions about how uncertainty manifests. Understanding their properties helps you recognize when to apply each one and interpret model outputs correctly. Choosing the right distribution is a fundamental modeling decision.
The Normal Distribution
The normal (Gaussian) distribution is common in financial modeling because it is tractable and because the Central Limit Theorem gives normal approximations to properly normalized sums under conditions such as independence or sufficiently weak dependence and finite variance. That theorem does not say that returns must be normal merely because trading aggregates many decisions. Normal models can approximate the center of some return distributions, while their tails and dependence often require separate treatment.
A normal random variable has density:
where:
- : probability density at value , representing the relative likelihood of observing a value near
- : mean (center) of the distribution, where the peak occurs
- : standard deviation (spread), controlling the width of the bell curve
- : variance, the square of the standard deviation
- : exponential function ( raised to the given power)
- : mathematical constant (≈ 3.14159)
- The term is one-half the squared standardized distance
- The normalization factor ensures the total area under the curve equals 1
The formula's structure reveals important intuition: the exponential of a negative quadratic creates the bell shape. Values near the mean make the exponent close to zero, while the exponent's magnitude grows quadratically with standardized distance and drives the density toward zero symmetrically. The density's slope magnitude does not increase forever; it peaks one standard deviation from the mean and then tends back toward zero in the tails.
The parameters are:
- : Mean, the center of the distribution. In finance, this represents expected return.
- : Variance, measuring spread. Its square root is the standard deviation, often called volatility.

The standard normal distribution (, ) is especially important. Any normal variable can be standardized:
where:
- : standardized random variable (also called z-score)
- : original normal random variable
- : mean of
- : standard deviation of
- : standard normal distribution with mean 0 and variance 1
The standardization formula shifts the distribution so the mean becomes zero (by subtracting ) and scales it so the standard deviation becomes one (by dividing by ). This transformation preserves shape while relocating the distribution. The standardized value tells us how many standard deviations lies from its mean; a of 2 means two standard deviations above average.
This transformation allows us to use standard normal tables or functions to compute probabilities for any normal distribution. We only need to tabulate one distribution rather than infinitely many. When a financial analyst computes that a return is "two sigma below average," they're implicitly using this standardization: they've converted the return to a z-score to assess how unusual it is.
The Log-Normal Distribution
While returns may be approximated as normal in some settings, a normal price model assigns positive probability to negative prices. A log-normal model avoids that result by restricting the modeled price to strictly positive values. This is a property of the model, not an exceptionless market fact: common stock can become worthless or be cancelled. Modeling the logarithm of a positive price as normal produces the log-normal distribution.
If , then follows a log-normal distribution. Under the common model in which period log returns are jointly normal (including iid normal log returns), their sum is normal and the compounded price ratio is log-normal. Marginal normality of each period alone is not sufficient.

Under the Black-Scholes model, a stock price at a fixed future time is log-normal, which keeps the modeled price positive while preserving normal log-return calculations. For derivative valuation, the relevant distribution is the model's risk-neutral distribution; it should not be confused with an estimate of the stock's real-world return distribution. Within the model, percentage gains are unbounded above while simple-return losses are bounded at 100%.
Other Important Distributions
Several other distributions appear frequently in quantitative finance, each suited to particular modeling contexts:
- Uniform distribution: Has constant density on a bounded interval, so subintervals of equal length have equal probability. Each individual point still has probability zero. It is used for random number generation and some Monte Carlo methods.
- Exponential distribution: Models waiting times between events, such as time until the next trade or default. Its "memoryless" property means the probability of an event occurring in the next moment does not depend on how long we have already waited.
- Poisson distribution: Counts events in fixed intervals, like number of trades per minute or credit events per year. Appropriate when events occur randomly and independently at a constant average rate.
- Student's t-distribution: Similar to normal but with heavier tails. A fitted t model can represent heavier marginal tails than a fitted normal model; whether it fits extreme returns better depends on the asset, horizon, parameters, dependence structure, and validation data. The t-distribution has a parameter (degrees of freedom) that controls tail heaviness, allowing it to interpolate between normal-like behavior and fat-tailed behavior.
- Chi-squared distribution: Under independent normal observations, a scaled sample variance has an exact chi-squared distribution. Other chi-squared tests rely on their own finite-sample or asymptotic assumptions, so this is not a generic law for volatility uncertainty.
Expected Value
When it exists, expected value is the distribution's probability-weighted mean. A long-run sample average converges to it only under law-of-large-numbers conditions such as iid repetitions with finite expectation, or an appropriate ergodic/dependence condition.
For a discrete random variable with PMF , the expected value is:
For a continuous random variable with PDF :
where:
- : expected value (also called expectation or mean) of random variable
- : possible values that can take
- : probability mass function giving for discrete
- : probability density function for continuous
- : sum over all possible values of
- : integral over the entire real line
The expected value is a probability-weighted average. Each possible value contributes according to both its magnitude and its probability (or density), and these contributions are summed (or integrated). A rare outcome with sufficiently large magnitude can therefore have a material effect on the mean.
The parallel structure between the discrete and continuous formulas reflects a deeper unity: in both cases, we're computing a weighted average where the weights come from the probability distribution. The sum becomes an integral when the set of possible values becomes continuous, but the underlying logic remains the same.
Financial Interpretation
For iid repetitions with finite expectation, the sample average converges to expected value. That statement concerns repetitions under a stable distribution; it does not imply that one stock's time average must approach a fixed annual expectation when its return process changes over time.
Consider a simplified binary stock model:
import numpy as np
# Binary model: stock either goes up 20% or down 10%
outcomes = np.array([0.20, -0.10]) # Possible returns
probabilities = np.array([0.60, 0.40]) # Probabilities
# Expected return
expected_return = np.sum(outcomes * probabilities)Possible outcomes: up 20% with prob 60%, down -10% with prob 40% Expected return: E[R] = 20% × 60% + (-10%) × 40% = 8.0%
The expected return of 8% is not a possible outcome of one trial. For iid repetitions of this binary experiment, the sample average converges to 8% by the law of large numbers; any individual trial still produces either a 20% gain or a 10% loss.
Properties of Expected Value
Expected value satisfies several useful properties that make it indispensable for financial calculations:
Linearity: For any random variables and and constants :
where:
- : constants (real numbers)
- : random variables (can be dependent or independent)
- : expected value operator
- : linear combination of the random variables
- : the same linear combination applied to the expected values
This property holds even if and are dependent. The expectation "passes through" the linear combination regardless of any relationship between the variables. This is not true for most other summary measures. The variance of a sum, for example, depends critically on the correlation between variables. Linearity makes portfolio expected return calculations straightforward: the expected return of a portfolio is the weighted average of individual expected returns. If you hold 60% stocks with expected return 10% and 40% bonds with expected return 4%, the portfolio expected return is simply , regardless of how stocks and bonds move together.
Expectation of a function: For a function whose expectation exists:
where:
- : any function applied to the random variable
- : the function evaluated at a specific value
- : probability mass function (for discrete )
- : probability density function (for continuous )
This result, known as the Law of the Unconscious Statistician (LOTUS), is useful when the stated expectation is finite: we weight by the original distribution of without first deriving the distribution of . It supports moment calculations and model-based derivative payoffs subject to their integrability conditions.
Variance and Standard Deviation
While expected value describes the center of a distribution, it tells us nothing about spread. Two investments might have the same expected return but vastly different risk profiles. Consider two stocks, both with 10% expected return. One might reliably return between 8% and 12%, while the other swings between -30% and +50%. Expected value alone cannot distinguish these dramatically different risk profiles. Variance quantifies this dispersion.
The variance of a random variable is the expected squared deviation from the mean:
where:
- : variance of random variable , measuring the spread or dispersion of values around the mean
- : expected value operator
- : mean of
- : squared deviation from the mean; squaring makes both signs contribute positively and emphasizes larger deviations
- : expected value of squared (the "mean of the square")
- : square of the expected value (the "square of the mean")
The standard deviation is , which returns the measure to the original units of .
Variance squares deviations to remove signs and give larger deviations more weight. This symmetric quadratic penalty is a modeling choice; general risk aversion does not imply that gains and losses of equal size should receive identical quadratic treatment.
The second formula, , is computationally convenient. It says variance equals the "mean of the square minus the square of the mean." This formula is often easier to apply because it doesn't require centering the data before squaring. We can compute and separately and then combine them.
To see why these formulas are equivalent, expand the definition:
Since , this gives us . The derivation uses linearity of expectation: we can pass the expectation through the sum and pull constants outside, reducing the problem to computing and .
Financial Interpretation
Standard deviation is a standard measure of return volatility, not a complete measure of financial risk. A stock with 20% annualized return volatility has greater dispersion around its mean than one with 10% volatility. Under normality, roughly 68% of returns fall within one standard deviation of the mean, and 95% fall within two standard deviations. These percentages provide benchmarks for interpreting volatility: a 20% volatility stock with 10% expected return will, about two-thirds of the time, return between -10% and +30% in a given year.
# Using the binary model from before
# Expected return already calculated
expected_return_squared = np.sum(outcomes**2 * probabilities)
variance = expected_return_squared - expected_return**2
std_dev = np.sqrt(variance)E[R] = 8.00% E[R²] = 0.0280 Var(R) = E[R²] - (E[R])² = 0.0280 - (0.0800)² = 0.0216 Standard deviation = 14.70%
The 14.7% standard deviation indicates substantial uncertainty around the 8% expected return. This volatility measure helps investors understand the range of outcomes they face and calibrate position sizes appropriately. A risk-averse investor might demand higher expected returns from this volatile investment compared to a more stable alternative.
Properties of Variance
Variance has different properties than expected value, which reflects its role in measuring spread rather than center:
Scaling: . Doubling a position quadruples variance, while standard deviation scales by and expected-return exposure scales by . The quadratic statement applies to variance specifically, not to every risk measure.
Sum of independent variables: If and are independent:
This result follows from independence: when and are independent, their covariance is zero, eliminating the cross term that would otherwise appear. Independence means knowing the value of one variable provides no information about the other, so their variances add without reinforcement or cancellation.
Sum of dependent variables: In general:
where:
- : variance of the sum of and
- , : individual variances
- : covariance between and
- The factor of 2 arises because when expanding , the cross term appears, and its expectation is
The covariance term determines how asset interactions affect portfolio variance. For a two-asset combination with fixed marginal variances and weights of the same sign, negative covariance lowers variance relative to zero covariance. With opposite-sign weights, the covariance contribution reverses sign. Whether the portfolio is less volatile than every constituent also depends on weights, individual volatilities, and constraints.
Covariance and Correlation
When analyzing multiple assets, we need to understand how they move together. Covariance and correlation quantify this relationship and provide the mathematical foundation for portfolio construction and risk management. A portfolio's risk depends on each component's volatility and on how those components interact.
The covariance between random variables and measures their joint variability:
where:
- : covariance between and
- : mean of
- : mean of
- : product of deviations from respective means
- : expected value of the product
- The units of covariance are the product of the units of and
The covariance definition captures co-movement through products of deviations. When both variables are above their means, or both are below, the product is positive; opposite-direction deviations produce a negative product.
Positive covariance means that same-sign deviations make a positive net contribution on average, while negative covariance means opposite-sign deviations dominate on average. Zero covariance means no linear association; it does not rule out nonlinear dependence, so observing can still provide information about .
The equivalence of the two covariance formulas follows from expanding the definition:
The second formula, , often simplifies calculations: it says covariance equals the expected product minus the product of expectations. If and are independent, , so the covariance is zero. Independence implies zero covariance, though the converse is not generally true.
For finite, nonzero standard deviations, the correlation coefficient normalizes covariance to lie between -1 and 1:
where:
- : correlation coefficient between and (bounded between -1 and 1)
- : covariance between and
- : standard deviation of
- : standard deviation of
- : product of standard deviations, which has the same units as covariance, making dimensionless
Dividing by the product of standard deviations converts covariance to a standardized scale. A correlation of means a perfect positive linear relationship where knowing exactly determines . A correlation of means a perfect negative linear relationship, and means no linear relationship.
Correlation is easier to interpret than covariance because it's unitless and bounded. A correlation of 1 indicates a perfect positive linear relationship, -1 indicates a perfect negative linear relationship, and 0 indicates no linear relationship. The normalization by standard deviations removes the scale dependence that makes covariance hard to interpret. Measuring the same returns in percentages rather than decimals is a fixed positive rescaling and leaves correlation unchanged. Currency translation is different: exchange rates generally vary over time, so converting asset values or returns between currencies introduces FX variation and can change correlations unless the conversion is a constant scaling.
The bounds of -1 and 1 follow from the Cauchy-Schwarz inequality, a fundamental result in mathematics. This bound provides a universal scale for comparison: a correlation of 0.8 between stocks A and B means they move together more strongly than if the correlation were 0.5, regardless of what A and B are or how volatile they are individually.



Low or negative correlation can reduce portfolio volatility relative to concentrated exposure, but correlation alone does not guarantee risk below every constituent. The two-asset variance also depends on both volatilities, the portfolio weights, and constraints.
Higher Moments: Skewness and Kurtosis
Mean and variance describe the first two moments of a distribution, but they don't fully characterize it. Two distributions can have identical means and variances yet differ dramatically in their shapes. Higher moments capture asymmetry and tail behavior, which are both critical for understanding financial risk. Investors care about average outcomes, volatility, and the likelihood of extreme gains versus extreme losses.
Skewness
When and the third absolute central moment is finite, skewness measures asymmetry around the mean:
where:
- : random variable
- : mean of
- : standard deviation of
- : expected value operator
- The cubic power captures asymmetry: positive deviations contribute positively () while negative deviations contribute negatively (), so a distribution with more extreme positive deviations has positive skewness
- Dividing by makes skewness dimensionless and scale-invariant, allowing comparison across different variables
The choice of the cubic power is not arbitrary. It's the lowest odd power that captures asymmetry. The first power (plain deviations) would sum to zero by definition of the mean. The second power (squared deviations) gives variance but loses sign information. The third power preserves the sign of each deviation, allowing positive and negative extremes to contribute differently. When the distribution has a long right tail with occasional extreme positive values, these large positive cubes dominate and yield positive skewness. A long left tail with extreme negative values yields negative skewness.
A distribution's skewness reveals the direction and magnitude of its asymmetry:
- Positive skewness: In the plotted skew-normal example, the right tail is longer and rare positive values dominate the third moment.
- Negative skewness: In the plotted skew-normal example, the left tail is longer and rare negative values dominate the third moment. Some equity-return samples and horizons show this pattern, but it is not universal across assets or regimes.
- Zero skewness: The third standardized moment is zero. A symmetric distribution with a finite third moment has zero skewness, but zero skewness does not by itself imply symmetry.

Skewness can matter even when two investments have the same mean and variance. Positive skew offers rare large gains, while negative skew exposes the investor to rare large losses. Whether an investor prefers one profile depends on the investor's utility, constraints, and the rest of the return distribution; skewness alone does not determine the choice.
Kurtosis
When and the fourth central moment is finite, kurtosis is the fourth standardized moment:
where:
- : random variable
- : mean of
- : standard deviation of
- The fourth power emphasizes extreme deviations because raising to the fourth power magnifies large values much more than small ones (e.g., but ), making kurtosis highly sensitive to tail behavior
- Unlike the cubic power in skewness, the fourth power is always positive, so both tails contribute positively
Excess kurtosis subtracts 3 (the kurtosis of a normal distribution): . This normalization sets the normal distribution as the baseline with excess kurtosis of zero.
Kurtosis uses fourth powers and therefore requires a finite fourth moment as a population quantity. A Student's t distribution with four degrees of freedom has no finite population fourth moment, so its population kurtosis is undefined; a finite sample still produces a numerical sample statistic, but that statistic is not estimating a finite t(4) kurtosis parameter.
Excess kurtosis compares a distribution's fourth standardized moment with the normal benchmark:
- Excess kurtosis > 0 (leptokurtic): A fourth standardized moment above the normal benchmark. Many financial-return samples exhibit positive excess kurtosis, but this alone does not guarantee larger tail probability at every cutoff.
- Excess kurtosis = 0 (mesokurtic): Equal to the normal distribution's fourth-moment benchmark; this does not imply the distribution is normal.
- Excess kurtosis < 0 (platykurtic): A fourth standardized moment below the normal benchmark; it does not guarantee uniformly lighter tails.
Many financial-return series have tails that a fitted normal model understates. Positive excess kurtosis is one warning sign, but tail probabilities should be checked directly because a fourth moment is not a complete description of tail behavior. Market moves during stressed periods can therefore be much more frequent than a normal calibration suggests without supporting a universal "once-in-a-century" label.
from scipy.stats import kurtosis, t
# Generate samples from normal and t-distributions
np.random.seed(42)
n_samples = 100000
normal_samples = np.random.randn(n_samples)
t_samples = t.rvs(df=4, size=n_samples) * np.sqrt(
(4 - 2) / 4
) # t(4) scaled to unit variance
# Calculate finite-sample excess-kurtosis statistics; t(4) has no finite population kurtosis
normal_kurtosis = kurtosis(normal_samples)
t_kurtosis = kurtosis(t_samples)Normal distribution excess kurtosis: -0.008 Student's t (df=4) finite-sample excess-kurtosis statistic: 17.487 Probability of |X| > 4: Normal distribution: 0.0050% Student's t (df=4): 0.4960%


After scaling both models to unit variance, the t-distribution with 4 degrees of freedom has about 4.9 times the standard normal's upper-tail probability at . The densities cross within the displayed tail region, so the t density is not higher at every threshold. This is a comparison between two specified models, not an estimate of a single "true" market probability or proof of a particular loss outcome. For risk management, the practical lesson is narrower: a model with lighter tails can assign a much smaller exceedance probability than a heavier-tailed alternative sufficiently far into the tail.
Conditional Probability
Conditional probability answers the question: "Given that event occurred, what's the probability of event ?" This concept is essential in finance, where we constantly update beliefs based on new information. Markets incorporate news continuously, so our probability assessments must evolve accordingly.
The conditional probability of given is:
where:
- : probability of event occurring given that event has occurred
- : probability that both and occur (joint probability)
- : probability of event (must be greater than 0) The formula captures a simple idea: within the probability mass assigned to , what share is also assigned to ? The denominator is the total probability mass of , and the numerator is the joint probability mass of .
The conditional probability formula implements a natural idea: conditioning on restricts the probability measure to the event and renormalizes it. Within this restricted measure, we ask what share of the probability mass also belongs to . The denominator is the total mass assigned to , while the numerator is the mass assigned to both events.
The formula says: to find the probability of given , we restrict attention to and divide the mass of by the mass of . This restriction and renormalization is the essence of conditioning. For finite equiprobable outcomes, the ratio also equals a fraction of outcome counts, but that counting interpretation does not hold in general.
Financial Example: Credit Risk
Suppose we're analyzing a corporate bond portfolio. Define:
- = bond defaults within one year
- = bond is rated BB (below investment grade)
If 5% of the portfolio are BB-rated bonds, 1% of all bonds default, and 0.3% of bonds are both BB-rated and default, then:
where:
- : probability of default given the bond is BB-rated
- : probability a bond is both BB-rated and defaults (0.3%)
- : probability a bond is BB-rated (5%)
BB-rated bonds have a 6% default probability, compared to 1% for the overall portfolio. Holding recovery, maturity, liquidity, taxes, and other relevant terms fixed, a higher default probability increases expected credit loss and can support a larger credit spread. The probability ratio alone does not determine a bond's yield. The calculation illustrates how conditioning on information, specifically the rating, changes our probability assessment.
Independence
Two events are independent if knowing that one occurred tells us nothing about the other. Independence is a useful simplifying assumption that, when valid, greatly simplifies probability calculations.
Events and are independent if and only if:
When , this is equivalent to .
where:
- : the joint probability factors into the product of marginals
- : when , learning occurred doesn't change the probability of
When , the two conditions are equivalent: substituting and multiplying by yields the product criterion. If , the elementary conditional ratio is undefined, while the product criterion remains valid.
The first characterization captures the intuitive meaning of independence: observing provides no information about , so our probability assessment for remains unchanged. The second characterization, the multiplication rule, provides a computational criterion: if the joint probability equals the product of marginal probabilities, the events are independent.
Independence is a powerful simplifying assumption. Many liquid-asset return series have small linear autocorrelations, but that does not establish independence: absolute and squared returns can remain dependent through volatility clustering. If returns are independent under a stated model, multi-day joint probabilities factor into products. The assumption should be tested against the dependence relevant to the calculation.
Bayes' Theorem
Bayes' Theorem provides a systematic method for updating probabilities when new information arrives, establishing the mathematical basis for learning from data and adapting beliefs. These skills are essential in finance, where conditions change constantly.
where:
- : posterior probability of given evidence
- : likelihood of observing if is true
- : prior probability of before observing
- : marginal probability of (normalizing constant)
In words: the probability of given equals the probability of given , times the prior probability of , divided by the probability of .
Bayes' Theorem follows directly from the definition of conditional probability applied twice. Starting from and noting that (by rearranging the conditional probability definition for given ), we obtain the theorem. Despite this simple derivation, the theorem has practical implications for reasoning under uncertainty.
The components have specific names that reflect their roles in the updating process:
- : Prior probability, representing our belief about before observing
- : Likelihood, measuring how probable the evidence is if is true
- : Posterior probability, representing our updated belief about after observing
- : Evidence probability, which is a normalizing constant
Once a prior and likelihood model are specified, the theorem provides the rule for computing the posterior after observing evidence. The posterior can then become the prior for the next round of updating as more evidence arrives.
Worked Example: Updating Default Probabilities
A bank is evaluating a corporate borrower. Based on the firm's industry and size, the bank initially estimates a 3% probability of default within one year. This is the prior probability.
The bank then receives the firm's quarterly financial statements showing deteriorating profit margins. Historical estimates from a comparable borrower cohort, measured at the same observation point before a one-year forecast horizon, show:
- 40% of companies that defaulted within the following year showed deteriorating margins at the observation date
- 10% of companies that did not default within that same year showed deteriorating margins at the observation date
The bank wants to update its default probability given this new information.
Step 1: Define events
- = company defaults within one year
- = company shows deteriorating margins
Step 2: List known probabilities
- (prior default probability)
- (prior probability of no default)
- (likelihood of margin deterioration given default)
- (likelihood of margin deterioration given no default)
Step 3: Calculate using the law of total probability
Before applying Bayes' Theorem, we need the total probability of observing deteriorating margins, regardless of whether default occurs. The law of total probability provides this by considering all mutually exclusive ways the evidence could arise:
where:
- : total probability of observing deteriorating margins
- : probability of deteriorating margins given the company defaults
- : prior probability of default
- : probability of deteriorating margins given no default
- : probability of no default
This formula works because every company either defaults or doesn't, so we can decompose the probability of deteriorating margins into these two mutually exclusive cases. We weight each conditional probability by the probability of that case occurring. The law of total probability essentially averages the conditional probabilities, weighted by the probability of each conditioning event.
# Prior probabilities
p_default = 0.03
p_no_default = 0.97
# Likelihoods
p_margin_given_default = 0.40
p_margin_given_no_default = 0.10
# Calculate P(M) using law of total probability
p_margin = (
p_margin_given_default * p_default
+ p_margin_given_no_default * p_no_default
)P(Deteriorating Margins) = 0.4 × 0.03 + 0.1 × 0.97 P(Deteriorating Margins) = 0.1090
Step 4: Apply Bayes' Theorem
Now we have all the ingredients to compute the posterior probability:
where:
- : posterior probability of default given deteriorating margins (what we want to find)
- : likelihood of deteriorating margins among defaulting companies (0.40)
- : prior default probability (0.03)
- : probability of deteriorating margins (calculated in Step 3)
The numerator represents the probability of both defaulting and showing deteriorating margins. The denominator normalizes this by dividing by the total probability of the evidence.
# Bayes' Theorem
p_default_given_margin = (p_margin_given_default * p_default) / p_marginP(Default | Deteriorating Margins) = (0.4 × 0.03) / 0.1090 P(Default | Deteriorating Margins) = 0.1101 Updated default probability: 11.01% Prior default probability: 3.00% Increase factor: 3.7x
The deteriorating margins more than tripled our assessed default probability from 3% to about 11%. This updated probability should inform lending decisions, pricing, and reserve requirements. The magnitude of this update reflects both the strength of the evidence, since deteriorating margins are much more common among defaulters, and the prior belief that default was already considered unlikely.
Visualizing Bayes' Theorem
The following visualization shows how the prior probability updates to the posterior as we incorporate new evidence.

Practical Implementation: Simulating Returns and Risk
This section brings together the concepts from this chapter by simulating stock returns and computing risk measures. This exercise demonstrates how probability distributions translate to financial risk analysis.
Simulating Daily Returns
We'll simulate iid normal daily simple returns as a local approximation and compute key statistics. Because a normal law is unbounded, the model assigns positive probability to returns at or below -100%; the seeded draw is validated before applying .
import numpy as np
np.random.seed(42)
# Parameters for daily returns
annual_return = 0.10 # Annualized arithmetic mean of daily simple returns
annual_volatility = 0.20 # 20% annual volatility
# Convert under iid daily arithmetic-mean and zero-autocovariance scaling
trading_days = 252
daily_return = annual_return / trading_days
daily_volatility = annual_volatility / np.sqrt(trading_days)
# Simulate 5 years of daily returns
n_years = 5
n_days = trading_days * n_years
returns = np.random.normal(daily_return, daily_volatility, n_days)
assert np.all(returns > -1), "simple returns must exceed -100% before log1p"Simulation Parameters: Annualized arithmetic daily mean: 10.0% Annual volatility: 20.0% Daily expected return: 0.0397% Daily volatility: 1.2599% Simulated 1260 trading days (5 years)
Computing Sample Statistics
Now we calculate the sample mean, standard deviation, skewness, and kurtosis.
from scipy import stats
# Sample statistics
sample_mean = np.mean(returns)
sample_std = np.std(returns, ddof=1) # ddof=1 for sample std
sample_skew = stats.skew(returns)
sample_kurtosis = stats.kurtosis(returns) # Excess kurtosis
# Annualize sample statistics
annualized_mean = sample_mean * trading_days
annualized_std = sample_std * np.sqrt(trading_days)Sample Statistics (Daily): Mean: 0.000877 Standard Deviation: 0.012467 Skewness: 0.0792 Excess Kurtosis: 0.0151 Annualized: Mean Return: 22.10% (target: 10.00%) Volatility: 19.79% (target: 20.00%)
For this seed, annualized volatility is about 19.79%, close to the 20% input, while the annualized sample mean is about 22.10%, well above the 10% input. The gap illustrates how noisy a mean estimate can remain with 1,260 daily observations. Near-zero sample skewness and excess kurtosis are consistent with normal draws, but those two statistics alone do not confirm normality; here the generating call establishes the normal model.
Visualizing the Return Distribution

Computing Value at Risk
Value at Risk (VaR) is a quantile threshold for loss over a specified horizon. Under a continuous calibrated distribution, a 95% VaR has 5% probability beyond the threshold: 5% is one period in twenty in probability terms, not a fixed schedule of exceedances. With discrete or empirical distributions, ties and the chosen quantile convention can change the exact fraction. VaR is neither the worst possible loss nor the expected loss beyond that threshold. The code below reports signed lower-tail return cutoffs. If loss is defined as , conventional positive loss VaR is the negative of the corresponding return cutoff.
# Signed lower-tail return cutoffs at 95% and 99% confidence levels
var_95 = np.percentile(returns, 5) # positive loss VaR = -var_95
var_99 = np.percentile(returns, 1) # positive loss VaR = -var_99
# Parametric VaR using normal distribution
var_95_parametric = sample_mean + stats.norm.ppf(0.05) * sample_std
var_99_parametric = sample_mean + stats.norm.ppf(0.01) * sample_stdDaily signed return cutoffs (positive loss VaR is their negative): 95% confidence (5th-percentile return cutoff): Empirical sample percentile: -1.9143% Fitted normal: -1.9629% 99% confidence (1st-percentile return cutoff): Empirical sample percentile: -2.6514% Fitted normal: -2.8125% Estimated 5th-percentile daily return: -1.91% This is the empirical sample percentile; the fitted-normal threshold is listed separately above.
For this seeded sample, the empirical 5th- and 1st-percentile daily return cutoffs are -1.9143% and -2.6514%. They correspond to positive loss VaR estimates of 1.9143% and 2.6514%. The fitted-normal return cutoffs are -1.9629% and -2.8125%, corresponding to positive loss VaR estimates of 1.9629% and 2.8125%.
In the signed-return presentation used here, historical VaR directly uses a lower-tail sample percentile, while parametric VaR calculates the corresponding percentile from a fitted distribution, normal in this example. Negating either cutoff gives the positive loss VaR. When the model fits and the sample is sufficiently informative, the estimates may be similar. In a finite real-data sample, historical VaR can be higher or lower than normal parametric VaR; its main distinction is that it does not impose a normal shape on the observations in the estimation window.

Converting Price Path from Returns
Starting from an initial price, we can construct a price path using cumulative returns.

Initial price: $100.00 Final price: $273.74 Total return: 173.74% Annualized return: 22.31%
The simulated price path is one realization under the model. The 10% input is an annualized arithmetic mean of daily simple returns, whereas 22.31% is the geometric CAGR of this realized path, so equality is not a model target. Sampling variation explains why this seed's annualized sample arithmetic mean is 22.10% rather than 10%. The separate move from that sample mean to the 22.31% CAGR comes from geometrically compounding the realized daily sequence; it is not another estimate of the arithmetic-mean input.
Key Parameters
The key parameters for simulating and analyzing financial returns are:
- annual_return: Annualized arithmetic mean of iid daily simple returns (10% here), so the daily mean is . It is not an exact target for the one-year compounded return, which under this model has expectation .
- annual_volatility: Annualized standard deviation under iid or zero-autocovariance scaling; daily volatility is the annual value divided by .
- trading_days: Modeling convention for trading days per year (commonly 252). Exact counts vary by market and calendar year.
- confidence_level: For VaR calculations, typically 95% or 99%. Higher confidence means more conservative risk estimates.
- initial_price: Starting price for the simulation. All subsequent prices are computed relative to this value.
Limitations and Impact
Probability theory provides powerful tools for financial modeling, but practitioners must understand its limitations.
The assumption of known, stable probability distributions rarely holds perfectly in financial markets. Market regimes shift, causing parameters such as volatility and dependence to change over time. In stressed periods, correlations among many risky assets can rise, weakening some diversification benefits, although the pattern is heterogeneous and not every pair approaches one. A model calibrated only on a calm historical window may therefore represent a later regime poorly.
Tail misspecification presents another fundamental challenge. Many financial-return series have extreme observations that a fitted normal model assigns very little probability. In those settings, a normal model understates tail risk and can produce risk limits that are too low unless it is supplemented with stress scenarios or a better tail model. That statistical limitation does not, by itself, establish a single causal explanation for any financial crisis.
The gap between sample statistics and true parameters also matters. With finite data, our estimates of means, variances, and correlations contain sampling error. Optimization procedures that treat estimated parameters as true values can generate portfolios that overfit to noise and perform poorly out of sample. For an independent sample with finite variance, the standard error of a sample mean shrinks only at the square-root rate, so the history required depends on the process, estimator, horizon, and desired precision.
Despite these limitations, probability theory supplies a consistent language for specified financial tasks. Quantiles such as VaR summarize loss thresholds under a model or empirical distribution. Risk-neutral expectations support derivative valuation under stated assumptions. Expected returns and covariance matrices make portfolio trade-offs explicit. These methods influence practice, but their usefulness depends on estimation quality, model fit, and the decision constraints around them.
The key is using these tools while remaining aware of their assumptions. Supplement parametric models with stress tests. Use longer-tailed distributions when appropriate. Update beliefs as new data arrives, following Bayes' theorem. All models are approximations: useful, but not reality.
Summary
This chapter covered the probability theory foundations essential for quantitative finance.
We began with sample spaces, measurable events, and the Kolmogorov axioms that define probability measures. Discrete variables use probability mass functions; absolutely continuous variables can be represented with densities, while every random variable has a CDF.
Expected value provides a probability-weighted average, representing the theoretical mean outcome. Variance measures dispersion around this mean, with standard deviation serving as a standard measure of return volatility rather than a complete measure of financial risk. Higher moments (skewness and kurtosis) capture asymmetry and tail behavior, revealing patterns that mean and variance miss.
Conditional probability quantifies how learning new information updates our beliefs. Bayes' Theorem formalizes this updating process, and gives a systematic method for incorporating evidence. The credit risk example demonstrated how new financial data can materially revise default probability estimates.
Key distributions in finance include normal models for returns and log-normal models for prices. However, many empirical return series have heavier tails than a fitted normal distribution, and some asset classes and horizons exhibit negative skewness. These departures from normality matter for risk management.
The practical simulation tied these concepts together: generating returns from probability distributions, computing sample statistics, calculating risk measures like Value at Risk, and constructing price paths. These techniques provide the computational foundation for more advanced quantitative methods.
With these probability foundations, you can tackle portfolio theory, time series analysis, and derivative pricing in the chapters ahead. Each builds directly on the concepts of distributions, expected values, and conditional probabilities introduced here.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about probability theory fundamentals.
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Quantitative Finance. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Quantitative FinanceStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!