Statistical Data Analysis & Inference in Finance

Michael BrenndoerferOctober 22, 202543 min read

Part of Quantitative Finance

Covers moments of returns, hypothesis testing, and confidence intervals. Essential statistical techniques for analyzing financial data and quantifying risk.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Statistical Data Analysis and Inference

Before building trading strategies, pricing derivatives, or constructing portfolios, you need to understand the data you're working with. Statistical analysis informs many quantitative decisions in finance. When you examine a stock's historical returns, you're implicitly asking: What return can I expect? How much risk am I taking? Are extreme losses more likely than a simple model would suggest?

This chapter equips you with the tools to answer these questions rigorously. We begin with descriptive statistics, the numerical summaries that characterize return distributions. Mean and volatility are familiar concepts. Skewness records a third standardized moment, while kurtosis records a fourth; neither statistic alone determines tail probabilities. We then move to statistical inference: how do you draw conclusions about a population, such as all possible future returns, from a sample of historical data? The Central Limit Theorem provides the theoretical foundation, while hypothesis testing and confidence intervals give you practical frameworks for making decisions under uncertainty.

These techniques support systematic analysis in portfolio optimization, risk management, and strategy evaluation.

Descriptive Statistics for Financial Returns

Descriptive statistics summarize the essential features of a dataset. For financial returns, we focus on four moment-based summaries: mean, volatility, skewness, and kurtosis. Together they describe complementary aspects of a return distribution.

Returns: The Foundation of Financial Analysis

Raw prices are difficult to compare across assets and time periods. A $10 move in a $100 stock is very different from a $10 move in a $1000 stock. The first represents a 10% change, while the second is only 1%. Returns normalize price changes, making relative moves comparable. Returns are often closer to stationary than price levels, but the transformation does not guarantee stationarity; that property must be checked for the series, horizon, and sample in question.

The choice to work with returns rather than prices can help with a basic statistical requirement: stationarity. Price levels often contain trends, so their distribution changes over time. Return means, variances, and dependence can also change across regimes, however. The techniques in this chapter therefore require diagnostics for structural breaks, serial dependence, and time-varying volatility rather than an automatic assumption that returns are stationary.

Simple Returns vs. Log Returns

Simple returns measure the percentage change in price:

Rt=Pt−Pt−1Pt−1R_t = \frac{P_t - P_{t-1}}{P_{t-1}}

where:

  • RtR_t: simple return at time tt (expressed as a decimal, e.g., 0.05 for 5%)
  • PtP_t: price at time tt
  • Pt−1P_{t-1}: price at time t−1t-1

Log returns (continuously compounded returns) are defined as:

rt=ln⁡(PtPt−1)r_t = \ln\left(\frac{P_t}{P_{t-1}}\right)

where rtr_t is the log return at time tt. The logarithm turns the price ratio into an additive measure.

Log returns are additive over time and approximately equal to simple returns for small values. Their additivity is convenient for time aggregation, while simple returns remain natural for portfolio accounting and cross-sectional aggregation. The appropriate convention depends on the analysis.

The distinction between simple and log returns deserves careful attention because it affects how we aggregate returns over time. With simple returns, compounding over multiple periods requires multiplication: a 10% gain followed by a 10% loss does not return you to your starting point (you end up at 99% of your original value). Log returns, however, are additive. The two-day log return is simply the sum of the individual daily log returns. Both definitions are widely used; the appropriate choice depends on whether the task emphasizes time aggregation, portfolio accounting, or another convention.

Let's calculate returns from price data:

In[2]:
Code
import numpy as np

# Simulated daily closing prices
np.random.seed(42)
prices = 100 * np.cumprod(1 + np.random.normal(0.0005, 0.02, 252))

# Calculate simple and log returns
simple_returns = np.diff(prices) / prices[:-1]
log_returns = np.diff(np.log(prices))
Out[3]:
Console
Number of daily returns: 251
Simple return (first day): -0.0023
Log return (first day): -0.0023
Difference: 0.000003

For small percentage moves, simple and log returns are nearly identical. The difference becomes material for larger moves: a 20% simple gain corresponds to a log return of about 18.2%, while a 20% simple loss corresponds to about -22.3%. The appropriate convention depends on the calculation rather than on a universal frequency threshold.

The First Moment: Mean Return

The arithmetic mean of returns measures the average performance over the sample period. This is the most intuitive statistic: add up all the returns and divide by how many there are. It summarizes the sample's first moment, but it need not equal an observed return, the most common return, or a value representative of a skewed or multimodal distribution.

rˉ=1n∑i=1nri\bar{r} = \frac{1}{n}\sum_{i=1}^{n} r_i

where:

  • rˉ\bar{r}: sample mean return
  • nn: number of observations in the sample
  • rir_i: individual return observation at time ii

The summation notation ∑i=1n\sum_{i=1}^{n} indicates that we add up all returns from the first observation (i=1i=1) through the last (i=ni=n). The factor 1n\frac{1}{n} in front converts this sum into an average. This formula produces a single measure of central tendency, but observations need not cluster around it.

Using the common 252-trading-day convention, we annualize daily mean log returns by multiplying by 252:

rˉannual=252×rˉdaily\bar{r}_{annual} = 252 \times \bar{r}_{daily}

This annualization assumes that expected returns accumulate linearly over time. If you expect to earn 0.04% per day on average, then over 252 trading days, you expect to earn approximately 252 × 0.04% = 10.08%. This linear scaling is exact for log returns (due to their additive property) and a reasonable approximation for simple returns when daily values are small.

In[4]:
Code
# Calculate mean returns
mean_daily = np.mean(log_returns)
mean_annual = mean_daily * 252
Out[5]:
Console
Mean daily return: 0.000199 (0.0199%)
Annualized return: 0.0500 (5.00%)

The sample mean estimates the population mean under a specified return model; it is not automatically an investor's forward expected reward. Historical sample means are unstable, and a few extreme days can shift the estimate significantly. This instability arises because the mean is computed by adding all observations, giving equal weight to every day including the outliers. A single day with a 10% move (positive or negative) can shift the annual mean estimate by several percentage points.

The Second Moment: Volatility

Volatility measures the dispersion of returns around the mean. It quantifies risk in the most commonly used sense: how much do returns deviate from what we expect? A high-volatility asset tends to have larger deviations around its mean, while a low-volatility asset tends to have smaller deviations. Dispersion alone does not determine whether returns are forecastable. This concept of risk as variability is central to modern portfolio theory.

The sample variance is:

s2=1n−1∑i=1n(ri−rˉ)2s^2 = \frac{1}{n-1}\sum_{i=1}^{n} (r_i - \bar{r})^2

where:

  • s2s^2: sample variance
  • n−1n-1: degrees of freedom; for i.i.d. observations with a common finite variance, Bessel's correction gives the usual unbiased sample-variance estimator
  • ri−rˉr_i - \bar{r}: deviation of each return from the sample mean
  • nn: number of observations in the sample

The logic behind variance is straightforward. First, we calculate how far each return deviates from the mean by computing (ri−rˉ)(r_i - \bar{r}). Some deviations are positive (returns above average) and some are negative (returns below average). If we simply averaged these deviations, the positives and negatives would cancel out, giving us zero, which is not useful as a dispersion measure. By squaring each deviation, we ensure all terms are nonnegative, and larger deviations receive disproportionately more weight.

For i.i.d. observations with a common finite variance, we use n−1n-1 in the denominator rather than nn to obtain the usual unbiased estimator of that population variance. This adjustment, known as Bessel's correction, accounts for the fact that we estimated the mean from the same data. When we calculate deviations from the sample mean, the deviations are constrained to sum to zero, removing one algebraic degree of freedom. Under those sampling conditions, dividing by n−1n-1 accounts for that lost degree of freedom and makes the sample variance unbiased for the common population variance.

Volatility is the square root of variance, annualized by multiplying by 252\sqrt{252}:

σannual=σdaily×252\sigma_{annual} = \sigma_{daily} \times \sqrt{252}

where:

  • σannual\sigma_{annual}: annualized volatility
  • σdaily\sigma_{daily}: daily volatility (standard deviation of daily returns)
  • 252≈15.87\sqrt{252} \approx 15.87: scaling factor based on approximately 252 trading days per year (the square root arises because variance scales linearly with time, so standard deviation scales with the square root)

The square root scaling deserves explanation because it differs from how we annualize means. When we sum independent random variables, their variances add. If daily variance is σdaily2\sigma_{daily}^2, then annual variance (summing 252 independent daily variances) is 252×σdaily2252 \times \sigma_{daily}^2. Taking the square root to convert back to standard deviation gives us 252×σdaily\sqrt{252} \times \sigma_{daily}. This is why volatility scales with the square root of time rather than linearly.

In[6]:
Code
# Calculate volatility
std_daily = np.std(log_returns, ddof=1)  # ddof=1 for sample std
std_annual = std_daily * np.sqrt(252)
Out[7]:
Console
Daily volatility: 0.019319 (1.9319%)
Annualized volatility: 0.3067 (30.67%)

Volatility is central to finance: it appears in option pricing formulas, portfolio optimization, and risk metrics like Value at Risk. It is often more forecastable than mean return, but it remains time-varying and can shift sharply across regimes. Historical volatility is therefore an estimate with a horizon and model attached, not a stable constant.

The Third Moment: Skewness

Skewness measures one form of asymmetry in a distribution. Mean and variance describe center and spread but do not describe asymmetry, even though both are also defined for many asymmetric distributions. The third standardized moment records whether large standardized deviations contribute more on the positive or negative side.

Define the sample central moments with denominator nn as

mk=1n∑i=1n(ri−rˉ)k.m_k = \frac{1}{n}\sum_{i=1}^{n}(r_i-\bar r)^k.

The moment coefficient of sample skewness used by scipy.stats.skew when bias=True is

g1=m3m23/2.g_1 = \frac{m_3}{m_2^{3/2}}.

where:

  • m2m_2: second central sample moment using denominator nn
  • m3m_3: third central sample moment using denominator nn
  • g1g_1: dimensionless moment coefficient; SciPy can apply a finite-sample bias correction with bias=False

Understanding this formula requires recognizing the role of each component. Subtracting the mean centers each return, and dividing by m2\sqrt{m_2} expresses the deviation on a dimensionless scale. Cubing those standardized deviations preserves their signs and gives large deviations more influence.

The key step is raising these standardized returns to the third power. Unlike squaring (which makes everything positive), cubing preserves the sign: negative values remain negative after cubing, while positive values remain positive. Cubing also amplifies extreme values more than moderate ones. A z-score of 3 contributes 27 to the sum, while a z-score of 2 contributes only 8. This makes skewness heavily influenced by the distribution's tails, especially which tail extends farther.

The sign of the third standardized moment is interpreted as follows, with the tail descriptions serving as heuristics rather than complete tail orderings:

  • Zero skewness: The third central moment is zero. Symmetric distributions with a finite third moment have zero skewness, but zero skewness does not imply symmetry.
  • Negative skewness: The third standardized moment is negative; this often accompanies a more influential left tail.
  • Positive skewness: The third standardized moment is positive; this often accompanies a more influential right tail.
In[8]:
Code
from scipy import stats

skewness = stats.skew(log_returns)
Out[9]:
Console
Skewness: 0.2353
Interpretation: Sample skewness is within the chosen ±0.5 threshold
Out[10]:
Visualization
Density plot of a negatively skewed distribution with a longer left tail.
Negative-skew example with a negative third standardized moment and a longer left side.
Density plot of a symmetric normal distribution.
Zero-skew example using a symmetric normal distribution.
Density plot of a positively skewed distribution with a longer right tail.
Positive-skew example with a positive third standardized moment and a longer right side.

Some equity return samples exhibit negative skewness, but its sign and magnitude vary by asset, horizon, estimator, and market regime. A symmetric model cannot represent an asymmetric fitted loss distribution, so risk analysis should compare the relevant left-tail probabilities or quantiles directly rather than infer every tail probability from skewness alone.

The Fourth Moment: Kurtosis

Kurtosis is the fourth standardized moment and is especially sensitive to large standardized deviations. While skewness records one form of asymmetry, kurtosis summarizes how strongly those large deviations influence the fourth moment. It does not, by itself, rank tail probabilities against a normal distribution at every threshold.

Using the same central moments, the raw moment coefficient of kurtosis is

g2raw=m4m22.g_2^{\mathrm{raw}} = \frac{m_4}{m_2^2}.

where:

  • m4m_4: fourth central sample moment using denominator nn
  • m22m_2^2: square of the second central sample moment
  • g2rawg_2^{\mathrm{raw}}: raw sample moment coefficient; finite-sample bias corrections use a different formula

The fourth power creates an even more extreme amplification of outliers than the cube used in skewness. A standardized return of 3 (three standard deviations from the mean) contributes 81 to the sum, while a standardized return of 4 contributes 256. A single extreme observation can therefore increase the kurtosis estimate dramatically. This sensitivity is relevant when studying extremes, but the fourth moment is not a complete measure of tail behavior and has considerable sampling variability.

A normal distribution has a kurtosis of 3. Excess kurtosis subtracts 3 to make the normal distribution the baseline:

Excess Kurtosis=Kurtosis−3\text{Excess Kurtosis} = \text{Kurtosis} - 3

where:

  • Kurtosis\text{Kurtosis}: the raw kurtosis value computed from the sample
  • 33: the kurtosis of the standard normal distribution, used as the reference point (subtracting it makes the normal distribution have excess kurtosis of zero)

By subtracting 3, we create a measure where zero is the normal-distribution benchmark. Positive excess kurtosis means a larger fourth standardized moment than the normal benchmark; it does not, by itself, guarantee larger tail probability at every threshold. scipy.stats.kurtosis reports excess kurtosis by default because its fisher argument defaults to True; other software may use a different convention.

The usual labels summarize the fourth-moment comparison, but they do not replace a direct tail-probability comparison:

  • Zero: Same raw fourth standardized moment as the normal benchmark (mesokurtic)
  • Positive: Larger fourth standardized moment than the normal benchmark (leptokurtic)
  • Negative: Smaller fourth standardized moment than the normal benchmark (platykurtic)
In[11]:
Code
kurtosis = stats.kurtosis(
    log_returns
)  # scipy returns excess kurtosis by default
Out[12]:
Console
Excess kurtosis: 0.4562
Interpretation: Sample excess kurtosis is within the chosen ±1 descriptive threshold
Out[13]:
Visualization
Line chart comparing a normal distribution and a fat-tailed t-distribution, with the tails highlighted to emphasize higher probability of extreme values.
Comparison of a normal distribution (excess kurtosis = 0) with a variance-matched Student's t distribution with five degrees of freedom (finite excess kurtosis = 6). The t distribution has a higher center and heavier far tails for this parameterization.

Many return samples have positive sample excess kurtosis, but the estimate depends on the asset, frequency, window, and estimator. A large estimate is evidence that the sample's fourth standardized moment differs from the normal benchmark. Establishing how often specified market moves occur requires a named dataset, fitted model, horizon, and dependence assumptions; the synthetic sample used in this chapter cannot establish those empirical frequencies.

Visualizing the Distribution

Let's visualize these statistics with a histogram and overlay a normal distribution for comparison:

Out[14]:
Visualization
Histogram of daily returns with overlaid normal distribution curve.
Distribution of simulated daily log returns compared with a normal distribution having the same sample mean and standard deviation. Both the histogram and fitted curve come from the chapter's seeded synthetic example.

Summary Statistics Table

Let's create a function that computes all descriptive statistics in one place:

In[15]:
Code
def compute_return_statistics(
    returns, periods_per_year=252, return_convention="simple"
):
    """Compute comprehensive descriptive statistics for returns."""
    result = {
        "Count": len(returns),
        "Mean (daily)": np.mean(returns),
        "Mean (annual)": np.mean(returns) * periods_per_year,
        "Std Dev (daily)": np.std(returns, ddof=1),
        "Std Dev (annual)": np.std(returns, ddof=1) * np.sqrt(periods_per_year),
        "Skewness": stats.skew(returns),
        "Excess Kurtosis": stats.kurtosis(returns),
        "Min": np.min(returns),
        "Max": np.max(returns),
        f"Sharpe Ratio ({return_convention} returns, zero reference)": (
            np.mean(returns) * periods_per_year
        )
        / (np.std(returns, ddof=1) * np.sqrt(periods_per_year)),
    }
    return result


summary = compute_return_statistics(log_returns, return_convention="log")
Out[16]:
Console
Return Statistics Summary
----------------------------------------
Count: 251
Mean (daily): 0.000199
Mean (annual): 0.050049
Std Dev (daily): 0.019319
Std Dev (annual): 0.306673
Skewness: 0.235344
Excess Kurtosis: 0.456209
Min: -0.053290
Max: 0.074694
Sharpe Ratio (log returns, zero reference): 0.163199

The ratio included above is explicitly based on log returns with a zero reference. The conventional ex-post Sharpe ratio uses arithmetic excess returns relative to a matching risk-free return; the worked example later uses simulated simple returns and a zero reference. In either convention, the ratio divides the selected return measure by its volatility:

Sharpe Ratio=rˉannual−rfσannual\text{Sharpe Ratio} = \frac{\bar{r}_{annual} - r_f}{\sigma_{annual}}

where:

  • rˉannual\bar{r}_{annual}: annualized mean return of the portfolio or strategy
  • rfr_f: risk-free rate (often assumed to be zero for simplicity, as in the code above)
  • σannual\sigma_{annual}: annualized volatility

The Sharpe ratio addresses a fundamental question in investing: how much return are you earning per unit of risk taken? A strategy with high returns might simply be taking excessive risk. The Sharpe ratio normalizes performance by volatility, allowing comparisons across strategies when the return horizon and conventions match. A Sharpe ratio of 1.0 is dimensionless: excess return equals volatility numerically at the same horizon. Whether that value is attractive depends on the horizon, estimation method, peer set, costs, and constraints.

Statistical Inference: From Samples to Populations

Descriptive statistics summarize historical data, but we care about the future. Statistical inference bridges this gap by using sample data to make probabilistic statements about population parameters. This transition from description to inference is perhaps the most important conceptual leap in quantitative analysis.

Sample vs. Population

The distinction between sample and population is fundamental to understanding everything that follows in this section. When we compute a mean return from historical data, we're not calculating the "true" expected return; we're estimating it from limited information.

Population and Sample

The population is the distribution of observations defined by a specified stochastic process, asset, sampling frequency, and regime. Historical analysis often treats observations within a chosen horizon as draws from that population, an assumption that must be examined rather than inferred from an infinite time horizon.

A sample is a subset of observations we collect. Historical returns form our sample, typically covering a limited time period.

We observe sample statistics (denoted with bars or hats: rˉ\bar{r}, σ^\hat{\sigma}) and use them to estimate population parameters (denoted with Greek letters: μ\mu, σ\sigma). This notational convention is a constant reminder that what we calculate from data (sample statistics) and what we're trying to learn about (population parameters) are different quantities.

The key challenge: sample statistics vary from sample to sample. If you calculate the mean return from 2010-2020 versus 2015-2025, you'll get different numbers even for the same stock. This variability is called sampling error. It's not a mistake or a flaw in your analysis; it is an inherent feature of drawing conclusions from incomplete information. Understanding and quantifying this variability is the core task of statistical inference.

The Standard Error

The standard error quantifies how much a sample statistic varies across different samples. It answers: if I collected a new sample of the same size, how different might my estimate be? For independent observations with a common finite variance σ2\sigma^2, the sample-mean variance is σ2/n\sigma^2/n, so:

SE(rˉ)=σnSE(\bar{r}) = \frac{\sigma}{\sqrt{n}}

where:

  • SE(rˉ)SE(\bar{r}): standard error of the sample mean, measuring how much rˉ\bar{r} would vary across repeated samples
  • σ\sigma: population standard deviation (true variability of individual returns)
  • nn: sample size (number of observations)
  • The n\sqrt{n} in the denominator reflects that averaging reduces variability, but with diminishing returns. To halve the standard error, you need four times as much data

This formula captures one of the most important relationships in statistics under that sampling model. In general,

Var⁡(rˉ)=1n2∑i=1n∑j=1nCov⁡(ri,rj),\operatorname{Var}(\bar r)=\frac{1}{n^2}\sum_{i=1}^{n}\sum_{j=1}^{n}\operatorname{Cov}(r_i,r_j),

so dependent returns require covariance or long-run-variance terms. When the independence and common-variance conditions apply, more variable data produces less precise estimates and larger samples produce more precise estimates. The square root relationship means precision improves slowly with additional data.

In practice, we estimate σ\sigma with the sample standard deviation ss:

SE^(rˉ)=sn\widehat{SE}(\bar{r}) = \frac{s}{\sqrt{n}}

For i.i.d. normal observations, replacing σ\sigma with ss leads to exact Student-t inference for the mean. For nonnormal i.i.d. observations, a t-based procedure may be a useful large-sample approximation under suitable regularity conditions. Large nn alone does not repair an invalid independence assumption or guarantee a good approximation for very heavy-tailed data.

In[17]:
Code
n = len(log_returns)
standard_error = std_daily / np.sqrt(n)
Out[18]:
Console
Sample size: 251
Sample mean: 0.000199
Standard error of mean: 0.001219
Ratio (SE/mean): 6.14
Out[19]:
Visualization
Line chart showing standard error percentage decreasing as sample size increases, with annotated points at common sample sizes such as 50, 252, 504, and 1000.
Under the independent common-variance model shown here, standard error of the mean decreases with the square root of sample size. Halving the standard error requires four times as much data.

Under the independent common-variance model, the standard error decreases with the square root of the sample size. To halve that standard error, you need four times as much data. In finance, this slow improvement can leave mean return estimates imprecise even with long samples. In this seeded example, the printed ratio shows that the standard error is large relative to the sample mean, so the mean estimate is noisy.

The Central Limit Theorem

The Central Limit Theorem (CLT) is one of the most important results in statistics, and its implications extend beyond textbooks. It provides the theoretical justification for using normal distribution-based methods even when the data themselves are not normally distributed.

Central Limit Theorem

If r1,r2,…r_1, r_2, \ldots are independent and identically distributed random variables with finite mean μ\mu and finite, positive variance σ2\sigma^2, then as n→∞n \to \infty:

rˉ−μσ/n→dN(0,1)\frac{\bar{r} - \mu}{\sigma/\sqrt{n}} \overset{d}{\to} N(0, 1)

where:

  • rˉ\bar{r}: sample mean
  • μ\mu: population mean
  • σ/n\sigma/\sqrt{n}: standard error of the mean
  • →d\overset{d}{\to}: convergence in distribution (meaning the distribution of the left side approaches the distribution on the right as nn grows)
  • N(0,1)N(0, 1): standard normal distribution (mean 0, variance 1)

The left side of the equation is called the standardized sample mean or z-score of the sample mean. Under the stated i.i.d. and finite-variance conditions, the CLT tells us this standardized value becomes approximately normal as nn grows.

The sample mean rˉ\bar{r} is approximately normally distributed with mean μ\mu and standard deviation σ/n\sigma/\sqrt{n}.

This theorem is powerful. Under its i.i.d. and finite-variance conditions, it tells us that the standardized sample mean converges to a normal distribution regardless of the population's shape. A distribution can be skewed, bimodal, or heavy-tailed and still satisfy the theorem, but infinite variance or material dependence requires different conditions and may produce a different limit.

Averaging and the CLT describe two related but distinct effects. Averaging reduces the variance of the unscaled sample mean, which concentrates it around μ\mu. Under the stated conditions, repeated convolution together with centering by μ\mu and scaling by σ/n\sigma/\sqrt{n} produces the normal limiting shape. Concentration by itself does not imply a bell curve.

Even if individual observations are not normally distributed, averages may be approximately normal when the sampling conditions and sample size make the CLT approximation adequate. Financial dependence, changing volatility, and very heavy tails can slow or invalidate that approximation, so normal-based inference for returns needs diagnostics or robust alternatives.

Let's demonstrate the CLT by repeatedly sampling and computing means:

In[20]:
Code
# Generate a non-normal distribution (mixture of normals)
np.random.seed(123)
population = np.concatenate(
    [
        np.random.normal(-0.02, 0.01, 5000),  # Crash regime
        np.random.normal(0.001, 0.015, 45000),  # Normal regime
    ]
)

# Sample means for different sample sizes
sample_sizes = [5, 30, 100, 500]
n_simulations = 1000

sample_means = {}
for size in sample_sizes:
    means = [
        np.mean(np.random.choice(population, size=size, replace=True))
        for _ in range(n_simulations)
    ]
    sample_means[size] = means
Out[21]:
Visualization
Histogram of sample means for samples of five observations with an overlaid theoretical normal curve.
Sample means for n = 5 retain visible irregularity from the non-normal population.
Histogram of sample means for samples of thirty observations with an overlaid theoretical normal curve.
In this seeded simulation, sample means for n = 30 form a visibly bell-shaped distribution.
Out[22]:
Visualization
Histogram of sample means for samples of one hundred observations with an overlaid theoretical normal curve.
In this seeded simulation, sample means for n = 100 visually track the theoretical normal curve.
Histogram of sample means for samples of five hundred observations with an overlaid theoretical normal curve.
In this seeded simulation, sample means for n = 500 are visually close to the theoretical normal approximation.

In this seeded simulation, the distribution of sample means is visibly bell-shaped by n=30n=30, despite the nonnormal mixture used as the population. The figure illustrates the progression: with only 5 observations, the sample-mean histogram retains visible irregularity; by 30, a bell-curve shape is evident; and by 500, the histogram is visually close to the theoretical normal curve. These sample sizes describe this example, not universal CLT thresholds.

Hypothesis Testing

Hypothesis testing provides a framework for making decisions based on data. In finance, you might ask: "Is this strategy's return significantly different from zero?" or "Does this factor predict returns?" These questions require a systematic approach that accounts for uncertainty in sample-based estimates.

The Framework

A hypothesis test has four components:

  1. Null hypothesis (H0H_0): The default assumption, typically "no effect" or "no difference"
  2. Alternative hypothesis (H1H_1 or HaH_a): What you're trying to demonstrate
  3. Test statistic: A number computed from data that measures evidence against H0H_0
  4. P-value: Under the specified null model, the probability that the prespecified test statistic is at least as extreme as the observed value

Hypothesis testing starts by assuming the null model and asking how unusual the observed test statistic is under a prespecified one- or two-sided extremeness rule. A sufficiently small tail probability leads us to reject the null; otherwise, we fail to reject it. Note the asymmetry: we never "accept" the null hypothesis; we merely fail to find sufficient evidence against it.

The p-value is not the probability of the observed dataset or the probability that the null hypothesis is true. It is a tail probability for the chosen statistic under the null model. A p-value of 0.01, for example, means that the null model assigns 1% probability to statistic values at least as extreme as the observed value according to the prespecified test.

If the p-value is below a threshold (commonly α=0.05\alpha = 0.05), we reject H0H_0 in favor of H1H_1. The choice of α=0.05\alpha = 0.05 is a convention, not a law of nature. The significance level controls the Type I error rate under the null. Type II error also depends on the specified alternative, effect size, sample size, variability, and test, so choosing α\alpha alone does not determine a universal balance between the two errors.

Testing if Mean Return Differs from Zero

A common hypothesis test in strategy evaluation asks whether a modeled population mean return differs from zero under specified sampling assumptions. The test does not by itself establish persistent skill, a strategy edge, economic profitability, or tradability.

Under the null hypothesis that the true mean is zero:

H0:μ=0vsH1:μ≠0H_0: \mu = 0 \quad \text{vs} \quad H_1: \mu \neq 0

The null hypothesis states that the modeled population mean return is zero. The alternative states that it is not zero (it could be positive or negative). This is a two-tailed test because we're interested in deviations from zero in either direction.

The test statistic is:

t=rˉ−0s/n=rˉSE(rˉ)t = \frac{\bar{r} - 0}{s/\sqrt{n}} = \frac{\bar{r}}{SE(\bar{r})}

where:

  • tt: test statistic
  • rˉ\bar{r}: sample mean return
  • ss: sample standard deviation (estimate of σ\sigma)
  • SE(rˉ)=s/nSE(\bar{r}) = s/\sqrt{n}: estimated standard error

The t-statistic measures how many standard errors the sample mean is from zero. A t-statistic of 2, for example, means the observed mean is 2 standard errors away from the hypothesized value of zero. This would be unusual if the true mean were zero. The larger the absolute value of the t-statistic, the stronger the evidence against the null hypothesis.

Under H0H_0, this statistic has an exact t-distribution with n−1n-1 degrees of freedom when the observations are i.i.d. normal. With nonnormal i.i.d. data, the test is generally approximate for sufficiently large samples. Serial dependence or heteroskedasticity can invalidate the standard error s/ns/\sqrt n; long-run-variance estimators or an appropriate bootstrap are then needed.

In[23]:
Code
# Test if mean return is significantly different from zero
# Variables mean_daily, standard_error, and n were computed in "The First Moment" and "The Standard Error" sections
t_statistic = mean_daily / standard_error
degrees_freedom = n - 1

# Two-tailed p-value
p_value = 2 * (1 - stats.t.cdf(abs(t_statistic), degrees_freedom))

# Using scipy's ttest_1samp for verification
t_stat_scipy, p_value_scipy = stats.ttest_1samp(log_returns, 0)
Out[24]:
Console
Hypothesis Test: Is mean return different from zero?
--------------------------------------------------
Mean daily return: 0.000199
Standard error: 0.001219
t-statistic: 0.1629
Degrees of freedom: 250
P-value (two-tailed): 0.8707

Conclusion: Fail to reject H₀ at α=0.05. Insufficient evidence that mean differs from zero.

One-Tailed vs. Two-Tailed Tests

The test above was two-tailed, asking whether the mean differs from zero in either direction. Often in finance, you care specifically about whether returns are positive:

H0:μ≤0vsH1:μ>0H_0: \mu \leq 0 \quad \text{vs} \quad H_1: \mu > 0

This formulation changes the question from "is there an effect?" to "is there a positive effect?" The null hypothesis now includes all non-positive values, and we only reject it if we find strong evidence of positive returns.

For a symmetric test statistic, a one-tailed p-value is half the two-tailed p-value only when the direction was specified in advance and the observed statistic points in that direction. Here we compute the upper-tail probability directly:

In[25]:
Code
# One-tailed test: is mean return significantly positive?
p_value_one_tailed = 1 - stats.t.cdf(t_statistic, degrees_freedom)
Out[26]:
Console
One-Tailed Test: Is mean return significantly positive?
--------------------------------------------------
t-statistic: 0.1629
P-value (one-tailed): 0.4354
Conclusion: Insufficient evidence that returns are significantly positive.

The Multiple Testing Problem

In practice, quantitative researchers test many strategies, factors, or parameters. If all 100 null hypotheses are true and each test rejects with probability exactly α=0.05\alpha = 0.05 under its null, the expected number of false rejections is 5. If each test is merely valid at level 0.05, the expectation is at most 5. This is the multiple testing problem, and it represents one of the most significant challenges in quantitative finance research.

The mathematics is straightforward, but its implications are serious. With a 5% false positive rate per test, the probability of at least one false positive across mm independent tests is 1−(1−0.05)m1 - (1-0.05)^m. For 100 tests, this exceeds 99%. In other words, finding at least one "significant" result is virtually guaranteed even when no true effects exist.

Multiple Testing Corrections

Bonferroni correction: Divide α\alpha by the number of tests. Testing 100 strategies? Use α=0.05/100=0.0005\alpha = 0.05/100 = 0.0005.

Benjamini-Hochberg: Under independent tests, or under certain positive-dependence conditions, the standard step-up procedure controls the false discovery rate (FDR) rather than the family-wise error rate.

The Bonferroni correction is simple but can be conservative; how conservative it is depends on the joint dependence structure of the tests. The Benjamini-Hochberg procedure offers a different approach: under independence or suitable positive dependence, it controls the expected false-discovery proportion at the chosen level (using zero when there are no discoveries), rather than the probability of any false discovery. Under arbitrary dependence, use an adjustment or procedure with a guarantee that matches the dependence structure.

In[27]:
Code
# Simulate multiple testing problem
np.random.seed(42)
n_strategies = 100
n_days = 252

# All strategies have zero true mean (null is true for all)
strategy_returns = np.random.normal(0, 0.02, (n_strategies, n_days))

# Test each strategy
p_values = []
for i in range(n_strategies):
    _, pval = stats.ttest_1samp(strategy_returns[i], 0)
    p_values.append(pval)

p_values = np.array(p_values)
Out[28]:
Console
Multiple Testing Simulation (all strategies have zero true mean)
------------------------------------------------------------
Strategies tested: 100
Strategies 'significant' at α=0.05: 4
Expected false positives: 5

After Bonferroni correction (α=0.0005):
Strategies 'significant': 0
Out[29]:
Visualization
Histogram of p-values from 100 tests under the null, with a vertical line at 0.05 and a shaded rejection region highlighting expected false positives.
Distribution of p-values when testing 100 strategies that all have zero true mean. Under the null hypothesis, p-values should be uniformly distributed between 0 and 1. The red shaded region shows the rejection area at α = 0.05, capturing approximately 5% of tests as false positives.

This simulation shows why backtesting results should be viewed skeptically. When you've tested many variations and found one that "works," it may simply be a false positive. Selecting the maximum from 100 null strategies can make the winner appear stronger than a typical null result because it was selected for favorable noise. This selection effect is one form of data dredging and requires careful attention.

Confidence Intervals

Hypothesis tests give binary decisions, while confidence intervals additionally communicate the precision of an estimate by providing a range of parameter values compatible with the procedure and assumptions.

For i.i.d. normal observations, the following interval has exact Student-t coverage. For a nonnormal sample, it is a large-sample approximation only when the standard error and regularity conditions justify it; dependent returns require an appropriate long-run-variance estimate. The interval is:

rˉ±t1−α/2,n−1×SE^(rˉ)\bar{r} \pm t_{1-\alpha/2, n-1} \times \widehat{SE}(\bar{r})

where:

  • rˉ\bar{r}: sample mean (center of the interval)
  • t1−α/2,n−1t_{1-\alpha/2, n-1}: positive upper critical value from the t-distribution with n−1n-1 degrees of freedom (for 95% confidence with large nn, this is approximately 1.96)
  • nn: sample size
  • α\alpha: significance level (for 95% confidence, α=0.05\alpha = 0.05)
  • SE^(rˉ)\widehat{SE}(\bar{r}): estimated standard error of the mean, valid for the stated sampling model
  • t1−α/2,n−1×SE^(rˉ)t_{1-\alpha/2, n-1} \times \widehat{SE}(\bar{r}): margin of error

The construction follows from the sampling distribution of the standardized mean under those conditions. Across repeated samples, the procedure uses the positive t1−α/2,n−1t_{1-\alpha/2,n-1} critical magnitude so that 95% of the resulting intervals cover the fixed population mean when α=0.05\alpha=0.05.

In[30]:
Code
# Calculate 95% confidence interval for mean return
# Uses mean_daily, standard_error, and degrees_freedom computed in earlier sections
confidence_level = 0.95
alpha = 1 - confidence_level

# Critical value from t-distribution
t_critical = stats.t.ppf(1 - alpha / 2, degrees_freedom)

# Confidence interval
margin_of_error = t_critical * standard_error
ci_lower = mean_daily - margin_of_error
ci_upper = mean_daily + margin_of_error

# Annualized confidence interval
ci_lower_annual = ci_lower * 252
ci_upper_annual = ci_upper * 252
Out[31]:
Console
95% Confidence Interval for Mean Return
--------------------------------------------------
Daily: [-0.002203, 0.002600]
Annualized: [-55.51%, 65.52%]

Point estimate (annualized): 5.00%
Width of CI (annualized): 121.04%
Out[32]:
Visualization
Horizontal confidence-interval bar for annualized mean return with a point estimate marker and a vertical line at zero return.
Visual representation of the 95% confidence interval for annualized mean return. The point estimate (orange dot) is surrounded by the confidence interval (blue bar). For the matching two-sided test of H0: μ = 0 at α = 0.05, using the same model and standard error, excluding zero corresponds to rejection and including zero to non-rejection.
Interpreting Confidence Intervals

A 95% confidence interval means: if we repeated this sampling procedure many times, 95% of the constructed intervals would contain the true population mean. It does not mean there is a 95% probability the true mean lies in this specific interval. The true mean is fixed; our interval either contains it or it doesn't.

This distinction matters conceptually. The frequentist interpretation views probability as a long-run frequency. Across many repetitions of the sampling procedure, 95% of intervals would capture the true parameter. The Bayesian interpretation, which allows probability statements about parameters, requires additional assumptions about prior beliefs.

Confidence Intervals and Risk-Adjusted Returns

Confidence intervals are particularly useful for Sharpe ratios, which are notoriously imprecise. A strategy might show a Sharpe ratio of 1.0, but the confidence interval could span from 0.2 to 1.8, which would change our assessment of the strategy's quality.

In[33]:
Code
def sharpe_confidence_interval(
    returns, reference_return=0.0, confidence=0.95, periods_per_year=252
):
    """
    Compute an approximate annualized Sharpe-ratio interval relative to a
    matching-period reference return.

    returns and reference_return must use the same period and return
    convention. The default reference_return=0.0 makes the zero-reference
    assumption explicit. Under an iid normal approximation, first compute the
    single-period ratio and its standard error, then annualize both by
    multiplying by sqrt(periods_per_year). Serial dependence requires a
    different long-run-variance adjustment.
    """
    n = len(returns)
    excess_returns = returns - reference_return
    std_return = np.std(excess_returns, ddof=1)
    sharpe_period = np.mean(excess_returns) / std_return
    sharpe = sharpe_period * np.sqrt(periods_per_year)

    # IID normal approximation, kept on a consistent annualized scale
    se_sharpe_period = np.sqrt((1 + 0.5 * sharpe_period**2) / n)
    se_sharpe = se_sharpe_period * np.sqrt(periods_per_year)

    z_critical = stats.norm.ppf(1 - (1 - confidence) / 2)

    ci_lower = sharpe - z_critical * se_sharpe
    ci_upper = sharpe + z_critical * se_sharpe

    return sharpe, ci_lower, ci_upper, se_sharpe


sharpe, sr_lower, sr_upper, se_sr = sharpe_confidence_interval(
    log_returns, reference_return=0.0
)
Out[34]:
Console
95% Confidence Interval for Sharpe Ratio
--------------------------------------------------
Sharpe Ratio: 0.1632
Standard Error: 1.0020
95% CI: [-1.8007, 2.1271]

Under the IID normal approximation, the CI includes zero.

Worked Example: Analyzing Simulated Strategy Returns

Let's walk through an illustrative analysis of a seeded synthetic return series with small gains and occasional larger losses.

In[35]:
Code
# Simulate an illustrative i.i.d. return mixture with a heavy left tail
np.random.seed(42)
n_days = 504  # Two years of trading

# Base simulated returns with slight positive drift
base_returns = np.random.normal(0.0003, 0.012, n_days)

# Inject occasional large losses into the synthetic series
crash_mask = np.random.random(n_days) < 0.02  # 2% chance of crash day
base_returns[crash_mask] = np.random.normal(-0.04, 0.015, crash_mask.sum())

strategy_returns = base_returns
Out[36]:
Console
Strategy Performance Analysis
==================================================
Count: 504
Mean (daily): -0.000211
Mean (annual): -0.053113
Std Dev (daily): 0.012655
Std Dev (annual): 0.200888
Skewness: -0.247155
Excess Kurtosis: 1.341945
Min: -0.060398
Max: 0.046533
Sharpe Ratio (simple returns, zero reference): -0.264389
Out[37]:
Visualization
Line chart showing the simulated strategy-performance path from an initial wealth index of one at trading day zero, with visible drawdowns.
Wealth index for the seeded simulated return mixture over two years, starting from $1. Injected loss days appear as sharp declines.
Histogram of seeded simulated daily returns with a fitted normal-distribution overlay and negative sample skewness.
Distribution of the simulated daily strategy returns compared with a fitted normal distribution. For this seeded mixture, the left tail contains the injected large losses.

In this simulation, the deliberately injected loss days produce negative sample skewness and higher sample kurtosis. Those results describe only the specified synthetic data-generating process.

In[38]:
Code
# Test whether the modeled gross mean return differs from zero
t_stat, p_val = stats.ttest_1samp(strategy_returns, 0)

# Confidence intervals
n = len(strategy_returns)
se = np.std(strategy_returns, ddof=1) / np.sqrt(n)
t_crit = stats.t.ppf(0.975, n - 1)
mean_return = np.mean(strategy_returns)

ci_daily = (mean_return - t_crit * se, mean_return + t_crit * se)
ci_annual = (ci_daily[0] * 252, ci_daily[1] * 252)
Out[39]:
Console

Statistical Inference
==================================================
t-statistic: -0.3739
p-value (two-tailed): 0.7086

95% CI for mean daily return: [-0.001318, 0.000897]
95% CI for annualized return: [-33.22%, 22.60%]

Conclusion: Cannot conclude strategy returns differ from zero (p ≥ 0.05)

Sharpe Ratio: -0.264 (95% CI: [-1.650, 1.122])

This raw mean-return test is not a test of net profitability. Profitability additionally depends on transaction costs, implementation, benchmark or factor exposures, and the return convention; persistence and selection effects require out-of-sample or multiple-testing analysis.

Key Parameters

The key parameters for statistical analysis of returns:

  • n: Sample size (number of observations). Larger samples reduce standard errors and narrow confidence intervals, improving estimation precision.
  • α (alpha): Significance level for hypothesis tests (commonly 0.05). Lower values reduce Type I error under the null and, with the design and alternative held fixed, reduce power.
  • ddof: Degrees of freedom adjustment for variance estimation. For i.i.d. observations, ddof=1 gives the usual unbiased estimator of population variance; taking its square root does not make sample standard deviation exactly unbiased.
  • periods_per_year: Annualization factor (252 for daily trading data). The code multiplies the mean by this factor and volatility and the Sharpe ratio by its square root; skewness and kurtosis remain unchanged.
  • confidence_level: Probability coverage for confidence intervals (commonly 0.95). Higher levels produce wider intervals.

Limitations and Practical Considerations

Statistical inference in finance faces several challenges that don't arise in controlled experiments.

A major issue is non-stationarity. Financial markets evolve over time. When you estimate a mean from historical data, you assume that the fitted historical process is informative about the forecast horizon. The strength of that assumption depends on the parameter, asset, horizon, and regime, and should be evaluated with structural-break diagnostics and out-of-sample performance. Regime changes and evolving market microstructure can violate the identical-distribution assumption underlying classical statistics.

A second challenge is the low signal-to-noise ratio in many financial return series. The ratio of daily volatility to mean return depends on the asset, period, and estimator, but mean-return uncertainty often remains large even with long samples. A useful analysis should report the sample-specific standard error or confidence interval rather than rely on a universal volatility-to-mean ratio.

Serial dependence and volatility clustering can further complicate inference. Return autocorrelation varies by asset, frequency, and market microstructure, while squared or absolute returns often show persistence. Standard errors computed under independence can be biased in either direction depending on the dependence pattern; heteroskedasticity- and autocorrelation-consistent estimators or suitable bootstraps estimate the relevant long-run variance.

These limitations do not render statistical analysis useless. Instead, they require humility. Treat point estimates with skepticism. Prefer wide confidence intervals to false precision. Consider robust statistics that are less sensitive to outliers. And always remember that statistical significance is not the same as economic significance or practical tradability.

Summary

This chapter covered the statistical foundation for analyzing financial data. We covered the key concepts that you'll use throughout quantitative finance.

Descriptive statistics capture complementary features of return distributions. The mean measures average return, while volatility quantifies dispersion. Skewness records the third standardized moment and kurtosis the fourth; their estimates vary across assets, horizons, and regimes, and neither statistic alone determines every tail probability.

Statistical inference bridges the gap between historical samples and population parameters. The standard error quantifies uncertainty in estimates, with precision improving as the square root of sample size under the stated sampling model. The classical Central Limit Theorem supports normal approximations for sample means under i.i.d. finite-variance sampling; dependence and non-stationarity require additional checks or robust methods.

Hypothesis testing provides a framework for decision-making under uncertainty. This chapter demonstrated how to test whether a modeled population mean return differs from zero using t-statistics and p-values. That test does not establish skill or net profitability. The multiple testing problem warns against data mining: testing many hypotheses inflates false positive rates.

Confidence intervals offer more information than binary hypothesis tests. For instance, a 95% confidence interval for the Sharpe ratio shows the Sharpe-ratio values compatible with the data under the stated model and procedure. Wide intervals signal imprecision and should temper confidence in conclusions.

These tools form the basis for portfolio construction, risk management, and strategy evaluation. As you progress to more advanced topics, you'll encounter sophisticated extensions of these basic ideas. The principles of estimation, uncertainty quantification, and statistical inference remain the same.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about statistical data analysis and inference in finance.

Statistical Data Analysis and Inference

Question 1 of 80 of 8 completed
Why do quantitative analysts typically prefer log returns over simple returns for statistical analysis?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025statisticaldata, author = {Michael Brenndoerfer}, title = {Statistical Data Analysis & Inference in Finance}, year = {2025}, url = {https://mbrenndoerfer.com/writing/statistical-data-analysis-inference-quantitative-finance}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2025). Statistical Data Analysis & Inference in Finance. Retrieved from https://mbrenndoerfer.com/writing/statistical-data-analysis-inference-quantitative-finance
MLAAcademic
Michael Brenndoerfer. "Statistical Data Analysis & Inference in Finance." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/statistical-data-analysis-inference-quantitative-finance>.
CHICAGOAcademic
Michael Brenndoerfer. "Statistical Data Analysis & Inference in Finance." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/statistical-data-analysis-inference-quantitative-finance.
HARVARDAcademic
Michael Brenndoerfer (2025) 'Statistical Data Analysis & Inference in Finance'. Available at: https://mbrenndoerfer.com/writing/statistical-data-analysis-inference-quantitative-finance (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2025). Statistical Data Analysis & Inference in Finance. https://mbrenndoerfer.com/writing/statistical-data-analysis-inference-quantitative-finance

About the author

Continue with the full handbook

This chapter is part of Quantitative Finance. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Quantitative Finance
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.