Part of Quantitative Finance
Covers moments of returns, hypothesis testing, and confidence intervals. Essential statistical techniques for analyzing financial data and quantifying risk.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Statistical Data Analysis and Inference
Before building trading strategies, pricing derivatives, or constructing portfolios, you need to understand the data you're working with. Statistical analysis informs many quantitative decisions in finance. When you examine a stock's historical returns, you're implicitly asking: What return can I expect? How much risk am I taking? Are extreme losses more likely than a simple model would suggest?
This chapter equips you with the tools to answer these questions rigorously. We begin with descriptive statistics, the numerical summaries that characterize return distributions. Mean and volatility are familiar concepts. Skewness records a third standardized moment, while kurtosis records a fourth; neither statistic alone determines tail probabilities. We then move to statistical inference: how do you draw conclusions about a population, such as all possible future returns, from a sample of historical data? The Central Limit Theorem provides the theoretical foundation, while hypothesis testing and confidence intervals give you practical frameworks for making decisions under uncertainty.
These techniques support systematic analysis in portfolio optimization, risk management, and strategy evaluation.
Descriptive Statistics for Financial Returns
Descriptive statistics summarize the essential features of a dataset. For financial returns, we focus on four moment-based summaries: mean, volatility, skewness, and kurtosis. Together they describe complementary aspects of a return distribution.
Returns: The Foundation of Financial Analysis
Raw prices are difficult to compare across assets and time periods. A $10 move in a $100 stock is very different from a $10 move in a $1000 stock. The first represents a 10% change, while the second is only 1%. Returns normalize price changes, making relative moves comparable. Returns are often closer to stationary than price levels, but the transformation does not guarantee stationarity; that property must be checked for the series, horizon, and sample in question.
The choice to work with returns rather than prices can help with a basic statistical requirement: stationarity. Price levels often contain trends, so their distribution changes over time. Return means, variances, and dependence can also change across regimes, however. The techniques in this chapter therefore require diagnostics for structural breaks, serial dependence, and time-varying volatility rather than an automatic assumption that returns are stationary.
Simple returns measure the percentage change in price:
where:
- : simple return at time (expressed as a decimal, e.g., 0.05 for 5%)
- : price at time
- : price at time
Log returns (continuously compounded returns) are defined as:
where is the log return at time . The logarithm turns the price ratio into an additive measure.
Log returns are additive over time and approximately equal to simple returns for small values. Their additivity is convenient for time aggregation, while simple returns remain natural for portfolio accounting and cross-sectional aggregation. The appropriate convention depends on the analysis.
The distinction between simple and log returns deserves careful attention because it affects how we aggregate returns over time. With simple returns, compounding over multiple periods requires multiplication: a 10% gain followed by a 10% loss does not return you to your starting point (you end up at 99% of your original value). Log returns, however, are additive. The two-day log return is simply the sum of the individual daily log returns. Both definitions are widely used; the appropriate choice depends on whether the task emphasizes time aggregation, portfolio accounting, or another convention.
Let's calculate returns from price data:
import numpy as np
# Simulated daily closing prices
np.random.seed(42)
prices = 100 * np.cumprod(1 + np.random.normal(0.0005, 0.02, 252))
# Calculate simple and log returns
simple_returns = np.diff(prices) / prices[:-1]
log_returns = np.diff(np.log(prices))Number of daily returns: 251 Simple return (first day): -0.0023 Log return (first day): -0.0023 Difference: 0.000003
For small percentage moves, simple and log returns are nearly identical. The difference becomes material for larger moves: a 20% simple gain corresponds to a log return of about 18.2%, while a 20% simple loss corresponds to about -22.3%. The appropriate convention depends on the calculation rather than on a universal frequency threshold.
The First Moment: Mean Return
The arithmetic mean of returns measures the average performance over the sample period. This is the most intuitive statistic: add up all the returns and divide by how many there are. It summarizes the sample's first moment, but it need not equal an observed return, the most common return, or a value representative of a skewed or multimodal distribution.
where:
- : sample mean return
- : number of observations in the sample
- : individual return observation at time
The summation notation indicates that we add up all returns from the first observation () through the last (). The factor in front converts this sum into an average. This formula produces a single measure of central tendency, but observations need not cluster around it.
Using the common 252-trading-day convention, we annualize daily mean log returns by multiplying by 252:
This annualization assumes that expected returns accumulate linearly over time. If you expect to earn 0.04% per day on average, then over 252 trading days, you expect to earn approximately 252 × 0.04% = 10.08%. This linear scaling is exact for log returns (due to their additive property) and a reasonable approximation for simple returns when daily values are small.
# Calculate mean returns
mean_daily = np.mean(log_returns)
mean_annual = mean_daily * 252Mean daily return: 0.000199 (0.0199%) Annualized return: 0.0500 (5.00%)
The sample mean estimates the population mean under a specified return model; it is not automatically an investor's forward expected reward. Historical sample means are unstable, and a few extreme days can shift the estimate significantly. This instability arises because the mean is computed by adding all observations, giving equal weight to every day including the outliers. A single day with a 10% move (positive or negative) can shift the annual mean estimate by several percentage points.
The Second Moment: Volatility
Volatility measures the dispersion of returns around the mean. It quantifies risk in the most commonly used sense: how much do returns deviate from what we expect? A high-volatility asset tends to have larger deviations around its mean, while a low-volatility asset tends to have smaller deviations. Dispersion alone does not determine whether returns are forecastable. This concept of risk as variability is central to modern portfolio theory.
The sample variance is:
where:
- : sample variance
- : degrees of freedom; for i.i.d. observations with a common finite variance, Bessel's correction gives the usual unbiased sample-variance estimator
- : deviation of each return from the sample mean
- : number of observations in the sample
The logic behind variance is straightforward. First, we calculate how far each return deviates from the mean by computing . Some deviations are positive (returns above average) and some are negative (returns below average). If we simply averaged these deviations, the positives and negatives would cancel out, giving us zero, which is not useful as a dispersion measure. By squaring each deviation, we ensure all terms are nonnegative, and larger deviations receive disproportionately more weight.
For i.i.d. observations with a common finite variance, we use in the denominator rather than to obtain the usual unbiased estimator of that population variance. This adjustment, known as Bessel's correction, accounts for the fact that we estimated the mean from the same data. When we calculate deviations from the sample mean, the deviations are constrained to sum to zero, removing one algebraic degree of freedom. Under those sampling conditions, dividing by accounts for that lost degree of freedom and makes the sample variance unbiased for the common population variance.
Volatility is the square root of variance, annualized by multiplying by :
where:
- : annualized volatility
- : daily volatility (standard deviation of daily returns)
- : scaling factor based on approximately 252 trading days per year (the square root arises because variance scales linearly with time, so standard deviation scales with the square root)
The square root scaling deserves explanation because it differs from how we annualize means. When we sum independent random variables, their variances add. If daily variance is , then annual variance (summing 252 independent daily variances) is . Taking the square root to convert back to standard deviation gives us . This is why volatility scales with the square root of time rather than linearly.
# Calculate volatility
std_daily = np.std(log_returns, ddof=1) # ddof=1 for sample std
std_annual = std_daily * np.sqrt(252)Daily volatility: 0.019319 (1.9319%) Annualized volatility: 0.3067 (30.67%)
Volatility is central to finance: it appears in option pricing formulas, portfolio optimization, and risk metrics like Value at Risk. It is often more forecastable than mean return, but it remains time-varying and can shift sharply across regimes. Historical volatility is therefore an estimate with a horizon and model attached, not a stable constant.
The Third Moment: Skewness
Skewness measures one form of asymmetry in a distribution. Mean and variance describe center and spread but do not describe asymmetry, even though both are also defined for many asymmetric distributions. The third standardized moment records whether large standardized deviations contribute more on the positive or negative side.
Define the sample central moments with denominator as
The moment coefficient of sample skewness used by scipy.stats.skew when bias=True is
where:
- : second central sample moment using denominator
- : third central sample moment using denominator
- : dimensionless moment coefficient; SciPy can apply a finite-sample bias correction with
bias=False
Understanding this formula requires recognizing the role of each component. Subtracting the mean centers each return, and dividing by expresses the deviation on a dimensionless scale. Cubing those standardized deviations preserves their signs and gives large deviations more influence.
The key step is raising these standardized returns to the third power. Unlike squaring (which makes everything positive), cubing preserves the sign: negative values remain negative after cubing, while positive values remain positive. Cubing also amplifies extreme values more than moderate ones. A z-score of 3 contributes 27 to the sum, while a z-score of 2 contributes only 8. This makes skewness heavily influenced by the distribution's tails, especially which tail extends farther.
The sign of the third standardized moment is interpreted as follows, with the tail descriptions serving as heuristics rather than complete tail orderings:
- Zero skewness: The third central moment is zero. Symmetric distributions with a finite third moment have zero skewness, but zero skewness does not imply symmetry.
- Negative skewness: The third standardized moment is negative; this often accompanies a more influential left tail.
- Positive skewness: The third standardized moment is positive; this often accompanies a more influential right tail.
from scipy import stats
skewness = stats.skew(log_returns)Skewness: 0.2353 Interpretation: Sample skewness is within the chosen ±0.5 threshold



Some equity return samples exhibit negative skewness, but its sign and magnitude vary by asset, horizon, estimator, and market regime. A symmetric model cannot represent an asymmetric fitted loss distribution, so risk analysis should compare the relevant left-tail probabilities or quantiles directly rather than infer every tail probability from skewness alone.
The Fourth Moment: Kurtosis
Kurtosis is the fourth standardized moment and is especially sensitive to large standardized deviations. While skewness records one form of asymmetry, kurtosis summarizes how strongly those large deviations influence the fourth moment. It does not, by itself, rank tail probabilities against a normal distribution at every threshold.
Using the same central moments, the raw moment coefficient of kurtosis is
where:
- : fourth central sample moment using denominator
- : square of the second central sample moment
- : raw sample moment coefficient; finite-sample bias corrections use a different formula
The fourth power creates an even more extreme amplification of outliers than the cube used in skewness. A standardized return of 3 (three standard deviations from the mean) contributes 81 to the sum, while a standardized return of 4 contributes 256. A single extreme observation can therefore increase the kurtosis estimate dramatically. This sensitivity is relevant when studying extremes, but the fourth moment is not a complete measure of tail behavior and has considerable sampling variability.
A normal distribution has a kurtosis of 3. Excess kurtosis subtracts 3 to make the normal distribution the baseline:
where:
- : the raw kurtosis value computed from the sample
- : the kurtosis of the standard normal distribution, used as the reference point (subtracting it makes the normal distribution have excess kurtosis of zero)
By subtracting 3, we create a measure where zero is the normal-distribution benchmark. Positive excess kurtosis means a larger fourth standardized moment than the normal benchmark; it does not, by itself, guarantee larger tail probability at every threshold. scipy.stats.kurtosis reports excess kurtosis by default because its fisher argument defaults to True; other software may use a different convention.
The usual labels summarize the fourth-moment comparison, but they do not replace a direct tail-probability comparison:
- Zero: Same raw fourth standardized moment as the normal benchmark (mesokurtic)
- Positive: Larger fourth standardized moment than the normal benchmark (leptokurtic)
- Negative: Smaller fourth standardized moment than the normal benchmark (platykurtic)
kurtosis = stats.kurtosis(
log_returns
) # scipy returns excess kurtosis by defaultExcess kurtosis: 0.4562 Interpretation: Sample excess kurtosis is within the chosen ±1 descriptive threshold

Many return samples have positive sample excess kurtosis, but the estimate depends on the asset, frequency, window, and estimator. A large estimate is evidence that the sample's fourth standardized moment differs from the normal benchmark. Establishing how often specified market moves occur requires a named dataset, fitted model, horizon, and dependence assumptions; the synthetic sample used in this chapter cannot establish those empirical frequencies.
Visualizing the Distribution
Let's visualize these statistics with a histogram and overlay a normal distribution for comparison:

Summary Statistics Table
Let's create a function that computes all descriptive statistics in one place:
def compute_return_statistics(
returns, periods_per_year=252, return_convention="simple"
):
"""Compute comprehensive descriptive statistics for returns."""
result = {
"Count": len(returns),
"Mean (daily)": np.mean(returns),
"Mean (annual)": np.mean(returns) * periods_per_year,
"Std Dev (daily)": np.std(returns, ddof=1),
"Std Dev (annual)": np.std(returns, ddof=1) * np.sqrt(periods_per_year),
"Skewness": stats.skew(returns),
"Excess Kurtosis": stats.kurtosis(returns),
"Min": np.min(returns),
"Max": np.max(returns),
f"Sharpe Ratio ({return_convention} returns, zero reference)": (
np.mean(returns) * periods_per_year
)
/ (np.std(returns, ddof=1) * np.sqrt(periods_per_year)),
}
return result
summary = compute_return_statistics(log_returns, return_convention="log")Return Statistics Summary ---------------------------------------- Count: 251 Mean (daily): 0.000199 Mean (annual): 0.050049 Std Dev (daily): 0.019319 Std Dev (annual): 0.306673 Skewness: 0.235344 Excess Kurtosis: 0.456209 Min: -0.053290 Max: 0.074694 Sharpe Ratio (log returns, zero reference): 0.163199
The ratio included above is explicitly based on log returns with a zero reference. The conventional ex-post Sharpe ratio uses arithmetic excess returns relative to a matching risk-free return; the worked example later uses simulated simple returns and a zero reference. In either convention, the ratio divides the selected return measure by its volatility:
where:
- : annualized mean return of the portfolio or strategy
- : risk-free rate (often assumed to be zero for simplicity, as in the code above)
- : annualized volatility
The Sharpe ratio addresses a fundamental question in investing: how much return are you earning per unit of risk taken? A strategy with high returns might simply be taking excessive risk. The Sharpe ratio normalizes performance by volatility, allowing comparisons across strategies when the return horizon and conventions match. A Sharpe ratio of 1.0 is dimensionless: excess return equals volatility numerically at the same horizon. Whether that value is attractive depends on the horizon, estimation method, peer set, costs, and constraints.
Statistical Inference: From Samples to Populations
Descriptive statistics summarize historical data, but we care about the future. Statistical inference bridges this gap by using sample data to make probabilistic statements about population parameters. This transition from description to inference is perhaps the most important conceptual leap in quantitative analysis.
Sample vs. Population
The distinction between sample and population is fundamental to understanding everything that follows in this section. When we compute a mean return from historical data, we're not calculating the "true" expected return; we're estimating it from limited information.
The population is the distribution of observations defined by a specified stochastic process, asset, sampling frequency, and regime. Historical analysis often treats observations within a chosen horizon as draws from that population, an assumption that must be examined rather than inferred from an infinite time horizon.
A sample is a subset of observations we collect. Historical returns form our sample, typically covering a limited time period.
We observe sample statistics (denoted with bars or hats: , ) and use them to estimate population parameters (denoted with Greek letters: , ). This notational convention is a constant reminder that what we calculate from data (sample statistics) and what we're trying to learn about (population parameters) are different quantities.
The key challenge: sample statistics vary from sample to sample. If you calculate the mean return from 2010-2020 versus 2015-2025, you'll get different numbers even for the same stock. This variability is called sampling error. It's not a mistake or a flaw in your analysis; it is an inherent feature of drawing conclusions from incomplete information. Understanding and quantifying this variability is the core task of statistical inference.
The Standard Error
The standard error quantifies how much a sample statistic varies across different samples. It answers: if I collected a new sample of the same size, how different might my estimate be? For independent observations with a common finite variance , the sample-mean variance is , so:
where:
- : standard error of the sample mean, measuring how much would vary across repeated samples
- : population standard deviation (true variability of individual returns)
- : sample size (number of observations)
- The in the denominator reflects that averaging reduces variability, but with diminishing returns. To halve the standard error, you need four times as much data
This formula captures one of the most important relationships in statistics under that sampling model. In general,
so dependent returns require covariance or long-run-variance terms. When the independence and common-variance conditions apply, more variable data produces less precise estimates and larger samples produce more precise estimates. The square root relationship means precision improves slowly with additional data.
In practice, we estimate with the sample standard deviation :
For i.i.d. normal observations, replacing with leads to exact Student-t inference for the mean. For nonnormal i.i.d. observations, a t-based procedure may be a useful large-sample approximation under suitable regularity conditions. Large alone does not repair an invalid independence assumption or guarantee a good approximation for very heavy-tailed data.
n = len(log_returns)
standard_error = std_daily / np.sqrt(n)Sample size: 251 Sample mean: 0.000199 Standard error of mean: 0.001219 Ratio (SE/mean): 6.14

Under the independent common-variance model, the standard error decreases with the square root of the sample size. To halve that standard error, you need four times as much data. In finance, this slow improvement can leave mean return estimates imprecise even with long samples. In this seeded example, the printed ratio shows that the standard error is large relative to the sample mean, so the mean estimate is noisy.
The Central Limit Theorem
The Central Limit Theorem (CLT) is one of the most important results in statistics, and its implications extend beyond textbooks. It provides the theoretical justification for using normal distribution-based methods even when the data themselves are not normally distributed.
If are independent and identically distributed random variables with finite mean and finite, positive variance , then as :
where:
- : sample mean
- : population mean
- : standard error of the mean
- : convergence in distribution (meaning the distribution of the left side approaches the distribution on the right as grows)
- : standard normal distribution (mean 0, variance 1)
The left side of the equation is called the standardized sample mean or z-score of the sample mean. Under the stated i.i.d. and finite-variance conditions, the CLT tells us this standardized value becomes approximately normal as grows.
The sample mean is approximately normally distributed with mean and standard deviation .
This theorem is powerful. Under its i.i.d. and finite-variance conditions, it tells us that the standardized sample mean converges to a normal distribution regardless of the population's shape. A distribution can be skewed, bimodal, or heavy-tailed and still satisfy the theorem, but infinite variance or material dependence requires different conditions and may produce a different limit.
Averaging and the CLT describe two related but distinct effects. Averaging reduces the variance of the unscaled sample mean, which concentrates it around . Under the stated conditions, repeated convolution together with centering by and scaling by produces the normal limiting shape. Concentration by itself does not imply a bell curve.
Even if individual observations are not normally distributed, averages may be approximately normal when the sampling conditions and sample size make the CLT approximation adequate. Financial dependence, changing volatility, and very heavy tails can slow or invalidate that approximation, so normal-based inference for returns needs diagnostics or robust alternatives.
Let's demonstrate the CLT by repeatedly sampling and computing means:
# Generate a non-normal distribution (mixture of normals)
np.random.seed(123)
population = np.concatenate(
[
np.random.normal(-0.02, 0.01, 5000), # Crash regime
np.random.normal(0.001, 0.015, 45000), # Normal regime
]
)
# Sample means for different sample sizes
sample_sizes = [5, 30, 100, 500]
n_simulations = 1000
sample_means = {}
for size in sample_sizes:
means = [
np.mean(np.random.choice(population, size=size, replace=True))
for _ in range(n_simulations)
]
sample_means[size] = means



In this seeded simulation, the distribution of sample means is visibly bell-shaped by , despite the nonnormal mixture used as the population. The figure illustrates the progression: with only 5 observations, the sample-mean histogram retains visible irregularity; by 30, a bell-curve shape is evident; and by 500, the histogram is visually close to the theoretical normal curve. These sample sizes describe this example, not universal CLT thresholds.
Hypothesis Testing
Hypothesis testing provides a framework for making decisions based on data. In finance, you might ask: "Is this strategy's return significantly different from zero?" or "Does this factor predict returns?" These questions require a systematic approach that accounts for uncertainty in sample-based estimates.
The Framework
A hypothesis test has four components:
- Null hypothesis (): The default assumption, typically "no effect" or "no difference"
- Alternative hypothesis ( or ): What you're trying to demonstrate
- Test statistic: A number computed from data that measures evidence against
- P-value: Under the specified null model, the probability that the prespecified test statistic is at least as extreme as the observed value
Hypothesis testing starts by assuming the null model and asking how unusual the observed test statistic is under a prespecified one- or two-sided extremeness rule. A sufficiently small tail probability leads us to reject the null; otherwise, we fail to reject it. Note the asymmetry: we never "accept" the null hypothesis; we merely fail to find sufficient evidence against it.
The p-value is not the probability of the observed dataset or the probability that the null hypothesis is true. It is a tail probability for the chosen statistic under the null model. A p-value of 0.01, for example, means that the null model assigns 1% probability to statistic values at least as extreme as the observed value according to the prespecified test.
If the p-value is below a threshold (commonly ), we reject in favor of . The choice of is a convention, not a law of nature. The significance level controls the Type I error rate under the null. Type II error also depends on the specified alternative, effect size, sample size, variability, and test, so choosing alone does not determine a universal balance between the two errors.
Testing if Mean Return Differs from Zero
A common hypothesis test in strategy evaluation asks whether a modeled population mean return differs from zero under specified sampling assumptions. The test does not by itself establish persistent skill, a strategy edge, economic profitability, or tradability.
Under the null hypothesis that the true mean is zero:
The null hypothesis states that the modeled population mean return is zero. The alternative states that it is not zero (it could be positive or negative). This is a two-tailed test because we're interested in deviations from zero in either direction.
The test statistic is:
where:
- : test statistic
- : sample mean return
- : sample standard deviation (estimate of )
- : estimated standard error
The t-statistic measures how many standard errors the sample mean is from zero. A t-statistic of 2, for example, means the observed mean is 2 standard errors away from the hypothesized value of zero. This would be unusual if the true mean were zero. The larger the absolute value of the t-statistic, the stronger the evidence against the null hypothesis.
Under , this statistic has an exact t-distribution with degrees of freedom when the observations are i.i.d. normal. With nonnormal i.i.d. data, the test is generally approximate for sufficiently large samples. Serial dependence or heteroskedasticity can invalidate the standard error ; long-run-variance estimators or an appropriate bootstrap are then needed.
# Test if mean return is significantly different from zero
# Variables mean_daily, standard_error, and n were computed in "The First Moment" and "The Standard Error" sections
t_statistic = mean_daily / standard_error
degrees_freedom = n - 1
# Two-tailed p-value
p_value = 2 * (1 - stats.t.cdf(abs(t_statistic), degrees_freedom))
# Using scipy's ttest_1samp for verification
t_stat_scipy, p_value_scipy = stats.ttest_1samp(log_returns, 0)Hypothesis Test: Is mean return different from zero? -------------------------------------------------- Mean daily return: 0.000199 Standard error: 0.001219 t-statistic: 0.1629 Degrees of freedom: 250 P-value (two-tailed): 0.8707 Conclusion: Fail to reject H₀ at α=0.05. Insufficient evidence that mean differs from zero.
One-Tailed vs. Two-Tailed Tests
The test above was two-tailed, asking whether the mean differs from zero in either direction. Often in finance, you care specifically about whether returns are positive:
This formulation changes the question from "is there an effect?" to "is there a positive effect?" The null hypothesis now includes all non-positive values, and we only reject it if we find strong evidence of positive returns.
For a symmetric test statistic, a one-tailed p-value is half the two-tailed p-value only when the direction was specified in advance and the observed statistic points in that direction. Here we compute the upper-tail probability directly:
# One-tailed test: is mean return significantly positive?
p_value_one_tailed = 1 - stats.t.cdf(t_statistic, degrees_freedom)One-Tailed Test: Is mean return significantly positive? -------------------------------------------------- t-statistic: 0.1629 P-value (one-tailed): 0.4354 Conclusion: Insufficient evidence that returns are significantly positive.
The Multiple Testing Problem
In practice, quantitative researchers test many strategies, factors, or parameters. If all 100 null hypotheses are true and each test rejects with probability exactly under its null, the expected number of false rejections is 5. If each test is merely valid at level 0.05, the expectation is at most 5. This is the multiple testing problem, and it represents one of the most significant challenges in quantitative finance research.
The mathematics is straightforward, but its implications are serious. With a 5% false positive rate per test, the probability of at least one false positive across independent tests is . For 100 tests, this exceeds 99%. In other words, finding at least one "significant" result is virtually guaranteed even when no true effects exist.
Bonferroni correction: Divide by the number of tests. Testing 100 strategies? Use .
Benjamini-Hochberg: Under independent tests, or under certain positive-dependence conditions, the standard step-up procedure controls the false discovery rate (FDR) rather than the family-wise error rate.
The Bonferroni correction is simple but can be conservative; how conservative it is depends on the joint dependence structure of the tests. The Benjamini-Hochberg procedure offers a different approach: under independence or suitable positive dependence, it controls the expected false-discovery proportion at the chosen level (using zero when there are no discoveries), rather than the probability of any false discovery. Under arbitrary dependence, use an adjustment or procedure with a guarantee that matches the dependence structure.
# Simulate multiple testing problem
np.random.seed(42)
n_strategies = 100
n_days = 252
# All strategies have zero true mean (null is true for all)
strategy_returns = np.random.normal(0, 0.02, (n_strategies, n_days))
# Test each strategy
p_values = []
for i in range(n_strategies):
_, pval = stats.ttest_1samp(strategy_returns[i], 0)
p_values.append(pval)
p_values = np.array(p_values)Multiple Testing Simulation (all strategies have zero true mean) ------------------------------------------------------------ Strategies tested: 100 Strategies 'significant' at α=0.05: 4 Expected false positives: 5 After Bonferroni correction (α=0.0005): Strategies 'significant': 0

This simulation shows why backtesting results should be viewed skeptically. When you've tested many variations and found one that "works," it may simply be a false positive. Selecting the maximum from 100 null strategies can make the winner appear stronger than a typical null result because it was selected for favorable noise. This selection effect is one form of data dredging and requires careful attention.
Confidence Intervals
Hypothesis tests give binary decisions, while confidence intervals additionally communicate the precision of an estimate by providing a range of parameter values compatible with the procedure and assumptions.
For i.i.d. normal observations, the following interval has exact Student-t coverage. For a nonnormal sample, it is a large-sample approximation only when the standard error and regularity conditions justify it; dependent returns require an appropriate long-run-variance estimate. The interval is:
where:
- : sample mean (center of the interval)
- : positive upper critical value from the t-distribution with degrees of freedom (for 95% confidence with large , this is approximately 1.96)
- : sample size
- : significance level (for 95% confidence, )
- : estimated standard error of the mean, valid for the stated sampling model
- : margin of error
The construction follows from the sampling distribution of the standardized mean under those conditions. Across repeated samples, the procedure uses the positive critical magnitude so that 95% of the resulting intervals cover the fixed population mean when .
# Calculate 95% confidence interval for mean return
# Uses mean_daily, standard_error, and degrees_freedom computed in earlier sections
confidence_level = 0.95
alpha = 1 - confidence_level
# Critical value from t-distribution
t_critical = stats.t.ppf(1 - alpha / 2, degrees_freedom)
# Confidence interval
margin_of_error = t_critical * standard_error
ci_lower = mean_daily - margin_of_error
ci_upper = mean_daily + margin_of_error
# Annualized confidence interval
ci_lower_annual = ci_lower * 252
ci_upper_annual = ci_upper * 25295% Confidence Interval for Mean Return -------------------------------------------------- Daily: [-0.002203, 0.002600] Annualized: [-55.51%, 65.52%] Point estimate (annualized): 5.00% Width of CI (annualized): 121.04%

A 95% confidence interval means: if we repeated this sampling procedure many times, 95% of the constructed intervals would contain the true population mean. It does not mean there is a 95% probability the true mean lies in this specific interval. The true mean is fixed; our interval either contains it or it doesn't.
This distinction matters conceptually. The frequentist interpretation views probability as a long-run frequency. Across many repetitions of the sampling procedure, 95% of intervals would capture the true parameter. The Bayesian interpretation, which allows probability statements about parameters, requires additional assumptions about prior beliefs.
Confidence Intervals and Risk-Adjusted Returns
Confidence intervals are particularly useful for Sharpe ratios, which are notoriously imprecise. A strategy might show a Sharpe ratio of 1.0, but the confidence interval could span from 0.2 to 1.8, which would change our assessment of the strategy's quality.
def sharpe_confidence_interval(
returns, reference_return=0.0, confidence=0.95, periods_per_year=252
):
"""
Compute an approximate annualized Sharpe-ratio interval relative to a
matching-period reference return.
returns and reference_return must use the same period and return
convention. The default reference_return=0.0 makes the zero-reference
assumption explicit. Under an iid normal approximation, first compute the
single-period ratio and its standard error, then annualize both by
multiplying by sqrt(periods_per_year). Serial dependence requires a
different long-run-variance adjustment.
"""
n = len(returns)
excess_returns = returns - reference_return
std_return = np.std(excess_returns, ddof=1)
sharpe_period = np.mean(excess_returns) / std_return
sharpe = sharpe_period * np.sqrt(periods_per_year)
# IID normal approximation, kept on a consistent annualized scale
se_sharpe_period = np.sqrt((1 + 0.5 * sharpe_period**2) / n)
se_sharpe = se_sharpe_period * np.sqrt(periods_per_year)
z_critical = stats.norm.ppf(1 - (1 - confidence) / 2)
ci_lower = sharpe - z_critical * se_sharpe
ci_upper = sharpe + z_critical * se_sharpe
return sharpe, ci_lower, ci_upper, se_sharpe
sharpe, sr_lower, sr_upper, se_sr = sharpe_confidence_interval(
log_returns, reference_return=0.0
)95% Confidence Interval for Sharpe Ratio -------------------------------------------------- Sharpe Ratio: 0.1632 Standard Error: 1.0020 95% CI: [-1.8007, 2.1271] Under the IID normal approximation, the CI includes zero.
Worked Example: Analyzing Simulated Strategy Returns
Let's walk through an illustrative analysis of a seeded synthetic return series with small gains and occasional larger losses.
# Simulate an illustrative i.i.d. return mixture with a heavy left tail
np.random.seed(42)
n_days = 504 # Two years of trading
# Base simulated returns with slight positive drift
base_returns = np.random.normal(0.0003, 0.012, n_days)
# Inject occasional large losses into the synthetic series
crash_mask = np.random.random(n_days) < 0.02 # 2% chance of crash day
base_returns[crash_mask] = np.random.normal(-0.04, 0.015, crash_mask.sum())
strategy_returns = base_returnsStrategy Performance Analysis ================================================== Count: 504 Mean (daily): -0.000211 Mean (annual): -0.053113 Std Dev (daily): 0.012655 Std Dev (annual): 0.200888 Skewness: -0.247155 Excess Kurtosis: 1.341945 Min: -0.060398 Max: 0.046533 Sharpe Ratio (simple returns, zero reference): -0.264389


In this simulation, the deliberately injected loss days produce negative sample skewness and higher sample kurtosis. Those results describe only the specified synthetic data-generating process.
# Test whether the modeled gross mean return differs from zero
t_stat, p_val = stats.ttest_1samp(strategy_returns, 0)
# Confidence intervals
n = len(strategy_returns)
se = np.std(strategy_returns, ddof=1) / np.sqrt(n)
t_crit = stats.t.ppf(0.975, n - 1)
mean_return = np.mean(strategy_returns)
ci_daily = (mean_return - t_crit * se, mean_return + t_crit * se)
ci_annual = (ci_daily[0] * 252, ci_daily[1] * 252)Statistical Inference ================================================== t-statistic: -0.3739 p-value (two-tailed): 0.7086 95% CI for mean daily return: [-0.001318, 0.000897] 95% CI for annualized return: [-33.22%, 22.60%] Conclusion: Cannot conclude strategy returns differ from zero (p ≥ 0.05) Sharpe Ratio: -0.264 (95% CI: [-1.650, 1.122])
This raw mean-return test is not a test of net profitability. Profitability additionally depends on transaction costs, implementation, benchmark or factor exposures, and the return convention; persistence and selection effects require out-of-sample or multiple-testing analysis.
Key Parameters
The key parameters for statistical analysis of returns:
- n: Sample size (number of observations). Larger samples reduce standard errors and narrow confidence intervals, improving estimation precision.
- α (alpha): Significance level for hypothesis tests (commonly 0.05). Lower values reduce Type I error under the null and, with the design and alternative held fixed, reduce power.
- ddof: Degrees of freedom adjustment for variance estimation. For i.i.d. observations,
ddof=1gives the usual unbiased estimator of population variance; taking its square root does not make sample standard deviation exactly unbiased. - periods_per_year: Annualization factor (252 for daily trading data). The code multiplies the mean by this factor and volatility and the Sharpe ratio by its square root; skewness and kurtosis remain unchanged.
- confidence_level: Probability coverage for confidence intervals (commonly 0.95). Higher levels produce wider intervals.
Limitations and Practical Considerations
Statistical inference in finance faces several challenges that don't arise in controlled experiments.
A major issue is non-stationarity. Financial markets evolve over time. When you estimate a mean from historical data, you assume that the fitted historical process is informative about the forecast horizon. The strength of that assumption depends on the parameter, asset, horizon, and regime, and should be evaluated with structural-break diagnostics and out-of-sample performance. Regime changes and evolving market microstructure can violate the identical-distribution assumption underlying classical statistics.
A second challenge is the low signal-to-noise ratio in many financial return series. The ratio of daily volatility to mean return depends on the asset, period, and estimator, but mean-return uncertainty often remains large even with long samples. A useful analysis should report the sample-specific standard error or confidence interval rather than rely on a universal volatility-to-mean ratio.
Serial dependence and volatility clustering can further complicate inference. Return autocorrelation varies by asset, frequency, and market microstructure, while squared or absolute returns often show persistence. Standard errors computed under independence can be biased in either direction depending on the dependence pattern; heteroskedasticity- and autocorrelation-consistent estimators or suitable bootstraps estimate the relevant long-run variance.
These limitations do not render statistical analysis useless. Instead, they require humility. Treat point estimates with skepticism. Prefer wide confidence intervals to false precision. Consider robust statistics that are less sensitive to outliers. And always remember that statistical significance is not the same as economic significance or practical tradability.
Summary
This chapter covered the statistical foundation for analyzing financial data. We covered the key concepts that you'll use throughout quantitative finance.
Descriptive statistics capture complementary features of return distributions. The mean measures average return, while volatility quantifies dispersion. Skewness records the third standardized moment and kurtosis the fourth; their estimates vary across assets, horizons, and regimes, and neither statistic alone determines every tail probability.
Statistical inference bridges the gap between historical samples and population parameters. The standard error quantifies uncertainty in estimates, with precision improving as the square root of sample size under the stated sampling model. The classical Central Limit Theorem supports normal approximations for sample means under i.i.d. finite-variance sampling; dependence and non-stationarity require additional checks or robust methods.
Hypothesis testing provides a framework for decision-making under uncertainty. This chapter demonstrated how to test whether a modeled population mean return differs from zero using t-statistics and p-values. That test does not establish skill or net profitability. The multiple testing problem warns against data mining: testing many hypotheses inflates false positive rates.
Confidence intervals offer more information than binary hypothesis tests. For instance, a 95% confidence interval for the Sharpe ratio shows the Sharpe-ratio values compatible with the data under the stated model and procedure. Wide intervals signal imprecision and should temper confidence in conclusions.
These tools form the basis for portfolio construction, risk management, and strategy evaluation. As you progress to more advanced topics, you'll encounter sophisticated extensions of these basic ideas. The principles of estimation, uncertainty quantification, and statistical inference remain the same.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about statistical data analysis and inference in finance.
Statistical Data Analysis and Inference
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Quantitative Finance. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Quantitative FinanceStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!