Descriptive Statistics with Python

Michael BrenndoerferUpdated December 25, 202526 min read

Part of Machine Learning from Scratch

Covers descriptive statistics fundamentals, including measures of central tendency (mean, median, mode), variability (variance, standard deviation, IQR).

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Descriptive Statistics: Summarizing and Understanding Data

Descriptive statistics help us summarize and understand data characteristics. These methods transform raw data into useful summaries that show patterns, typical values, and variability. They are a practical starting point for statistical analysis and machine learning because they help reveal structure, anomalies, and questionable assumptions before modeling [1].

Introduction

Before moving far into modeling or inference, it is good practice to examine the basic properties of your data. Descriptive statistics provide this early view by showing where distributions are centered, how widely values vary, and what shape they take. These summaries help us assess data quality, spot potential issues such as outliers or skewness, and communicate findings to readers who may not need the underlying technical detail.

Descriptive statistics provide different types of summary information about data, organized into several key categories:

  • Measures of central tendency (e.g., mean, median, mode): Identify typical or representative values in a dataset.
  • Measures of variability (e.g., variance, standard deviation, range, interquartile range): Quantify how spread out the data points are.
  • Measures of distribution shape (e.g., skewness, kurtosis): Describe the asymmetry and tail behavior of the data.

Understanding these descriptive measures matters for exploratory data analysis and for diagnosing potential problems in modeling. For instance, high skewness might suggest the need for data transformation, while extreme kurtosis could indicate the presence of outliers that warrant investigation. This chapter explores each of these fundamental concepts, giving the statistical basics needed for effective data analysis and modeling.

Measures of Central Tendency

Central tendency measures answer a key question: what is a typical value in this dataset? The mean and median answer this question for numerical data, while the mode identifies the most frequent value. The right choice depends on the data's properties and the analysis goals [2].

The Mean

The arithmetic mean, commonly called the average, is calculated by summing all values and dividing by the count of observations:

xˉ=1n∑i=1nxi\bar{x} = \frac{1}{n} \sum_{i=1}^n x_i

Where:

  • xˉ\bar{x}: Sample mean
  • xix_i: Individual observation
  • nn: Number of observations

The mean represents the balance point of a distribution and incorporates information from every data point. Under a normal model it is an efficient estimator of location, but symmetry alone is not enough to guarantee that property [2]. Its sensitivity to all values also makes it vulnerable to outliers: a single extremely large or small value can pull the mean far from where most observations cluster.

For example, consider household income in a neighborhood where most families earn between 40,000and40,000 and 80,000 annually, but one household earns 5million.Themeanincomemightbe5 million. The mean income might be 150,000, far above what most residents earn, while the median would remain around $60,000, better representing the typical household.

The Median

The median is the middle value when observations are arranged in order. For datasets with an odd number of observations, it is simply the center value; for even counts, it is the average of the two middle values:

Median={x(n+1)/2if n is oddx(n/2)+x(n/2)+12if n is even\text{Median} = \begin{cases} x_{(n+1)/2} & \text{if } n \text{ is odd} \\ \frac{x_{(n/2)} + x_{(n/2)+1}}{2} & \text{if } n \text{ is even} \end{cases}

Where x(i)x_{(i)} denotes the ii-th value in the sorted dataset.

For example, consider the sorted dataset: 12, 15, 18, 22, 25, 28, 31. With seven observations (odd count), the median is the fourth value: 22. If we add an eighth observation to get 12, 15, 18, 22, 25, 28, 31, 45, the median becomes the average of the fourth and fifth values: (22 + 25) / 2 = 23.5.

The median's main advantage is its resistance to the magnitude of extreme tail values. Moving an already extreme observation farther into the tail does not change its rank and usually leaves the median unchanged; changing observations near the middle can still change it. This makes the median particularly useful for skewed distributions or data with outliers.

The Mode

For a discrete sample, the mode is any value with the highest observed frequency. Two values tied for that maximum give two sample modes, as when exam scores of 75 and 90 each occur eight times and no other score occurs more than five times. For a continuous distribution, modes are peaks or local maxima of its density. Multiple peaks may suggest a mixture of subgroups, but they do not prove one, and the peaks seen in a histogram or KDE can change with bin width or smoothing bandwidth [3]. If every observed value has the same frequency, all values formally tie; many introductory summaries report that situation as having no unique mode.

The mode is especially useful for nominal categorical data, where mean and median are undefined, such as the most common color preference in a survey or the most frequent diagnosis in a medical dataset. Ordered categorical data may also support a median. For continuous numerical data, estimated density modes can help identify peaks, although their locations depend on the estimation method.

In practice, for continuous data, we often examine the mode through histograms or kernel density estimates rather than computing the exact most frequent value, especially when data have many unique values.

Kernel Density Estimation

Kernel density estimation (KDE) is a nonparametric method for estimating a probability density. It centers a kernel, often Gaussian, on each observation and averages the bandwidth-scaled kernels to form a smooth curve. KDE is useful for exploring continuous distributions and possible modes when exact repetitions are rare, but the selected bandwidth materially affects how many peaks appear [4].

Don't worry if this concept seems unclear at this point. It's not essential to understand right now. Feel free to continue and revisit it later if needed.

Measures of Spread

While central tendency measures locate the center of a distribution, measures of spread quantify how much data points vary around that center. Two datasets can have identical means but very different spreads, so center and scale should be examined together [5].

Variance and Standard Deviation

Population variance is the mean squared deviation from the population mean. For a random sample, the following corrected sample variance divides the sum of squared deviations by n−1n-1:

s2=1n−1∑i=1n(xi−xˉ)2s^2 = \frac{1}{n-1} \sum_{i=1}^n (x_i - \bar{x})^2

Where:

  • s2s^2: Sample variance
  • xix_i: Individual observation
  • xˉ\bar{x}: Sample mean
  • nn: Number of observations

For n≥2n \geq 2, division by n−1n-1 is commonly called Bessel's correction. Under an independent, identically distributed sampling model with finite variance, it makes s2s^2 an unbiased estimator of the population variance; it does not make ss an unbiased estimator of the population standard deviation [6]. For independent variables with finite variances, the variance of their sum equals the sum of their variances. Because variance has squared units, such as dollars squared, direct interpretation can be awkward.

Standard deviation addresses this interpretation issue by taking the square root of variance, returning to the original units:

s=s2=1n−1∑i=1n(xi−xˉ)2s = \sqrt{s^2} = \sqrt{\frac{1}{n-1} \sum_{i=1}^n (x_i - \bar{x})^2}

Standard deviation expresses dispersion in the data's original units. Its distance interpretation is clearest for distributions that are roughly symmetric and not heavy-tailed. For a normal distribution, about 68% of probability lies within one population standard deviation of the mean, 95% within two, and 99.7% within three [7]. These percentages are approximations for data whose distribution is only approximately normal, not a distribution-free outlier rule.

Population vs. Sample Standard Deviation: σ vs. s

You may notice that standard deviation is sometimes denoted as σ\sigma (sigma) and sometimes as ss. This distinction reflects whether we are working with a population or a sample:

  • Population standard deviation (σ\sigma): A parameter of the population distribution. If every member of a finite population is observed, it can be calculated directly with the NN denominator; otherwise it may be estimated from a sample.
  • Sample standard deviation (ss): A descriptive statistic calculated from sample observations. The familiar n−1n-1 form is the square root of the unbiased sample-variance estimator, but ss itself is generally biased for σ\sigma.

Choose the denominator according to the estimand. Use the NN denominator to describe the variability of the observations in hand as a complete finite population, and use the n−1n-1 variance estimator when estimating a common population variance from a random sample with unknown mean. State the convention when software defaults or the analytical goal could make it ambiguous [5], [6].

Range and Interquartile Range

The range is the simplest measure of spread, calculated as the difference between the maximum and minimum values:

Range=xmax−xmin\text{Range} = x_{\text{max}} - x_{\text{min}}

While easy to compute and interpret, the range uses only two data points and is extremely sensitive to outliers. A single erroneous or extreme value can inflate the range dramatically, making it an unreliable measure of typical spread.

The interquartile range (IQR) is less sensitive to extreme values because it measures the spread of the middle 50% of the data:

IQR=Q3−Q1\text{IQR} = Q_3 - Q_1

Where:

  • Q1Q_1: First quartile (25th percentile)
  • Q3Q_3: Third quartile (75th percentile)

Because the IQR focuses on the central half of the ordered values, changing an already extreme tail value usually leaves it unchanged. The conventional inner-fence rule flags values below Q1−1.5×IQRQ_1 - 1.5 \times \text{IQR} or above Q3+1.5×IQRQ_3 + 1.5 \times \text{IQR} as potential outliers. A flag is a prompt for investigation, not a reason to delete an observation automatically [8].

For example, consider a dataset of daily website visits over a week: 120, 135, 142, 128, 155, 130, 138. We can calculate several descriptive measures of center and spread:

  • Mean: (120 + 135 + 142 + 128 + 155 + 130 + 138) / 7 = 135.43
  • Median (middle sorted value): 135
  • Mode: No unique mode (all values occur once)
  • Range: 155 - 120 = 35
  • IQR: Q3 - Q1 = 142 - 128 = 14 (exclusive median-of-halves convention, omitting the overall median)
  • Standard Deviation: 11.22

The mean and median are close, although that alone does not establish symmetry. The range is 35, the IQR is 14 under the exclusive median-of-halves convention, and the sample standard deviation is 11.22. Whether that spread is large or small depends on the scale and application.

Measures of Distribution Shape

Beyond center and spread, distribution shape gives useful information about data behavior. Two common measures, skewness and kurtosis, summarize asymmetry and tail weight, respectively [9].

Skewness

One common sample measure of skewness uses central moments computed with a consistent 1/n1/n denominator:

mr=1n∑i=1n(xi−xˉ)r,g1=m3m23/2.m_r = \frac{1}{n}\sum_{i=1}^n (x_i-\bar{x})^r, \qquad g_1 = \frac{m_3}{m_2^{3/2}}.

For nonconstant data, m2>0m_2>0 and the statistic g1g_1 is defined and unitless. Some software instead reports a finite-sample bias-adjusted estimator, so comparisons should use the same convention.

A nondegenerate symmetric distribution with a finite third moment has skewness zero, but zero skewness does not by itself prove symmetry. Positive skewness, or right skew, often accompanies a longer or heavier right tail. In many familiar unimodal cases the mean then exceeds the median, but that ordering is not guaranteed for every distribution.

Negative skewness, or left skew, often accompanies a longer or heavier left tail. Test scores on an easy exam can show this pattern when many students score near the upper boundary and a smaller number receive much lower scores.

There is no universal cutoff that turns a skewness value into “moderate” or “high.” Interpret its magnitude relative to the estimator, sample size, application, and the accompanying distribution plot.

Kurtosis

Moment kurtosis standardizes the fourth central moment:

g2=m4m22.g_2 = \frac{m_4}{m_2^2}.

For nonconstant data, m2>0m_2>0 and g2g_2 is defined. Because fourth powers amplify large standardized deviations, kurtosis is sensitive to extreme observations. It is not a direct partition of probability between the center and tails, nor does it determine whether a density has a sharp or flat peak [10]. A normal distribution has population kurtosis 3. Many software packages instead report excess kurtosis, g2−3g_2-3, and may also apply a finite-sample bias correction; under the excess convention the normal reference is zero.

High kurtosis means that extreme standardized deviations contribute strongly to the fourth moment. It can be a warning that a normal model understates important extremes, but it does not guarantee a larger exceedance probability at every threshold. Risk analysis should therefore examine relevant tail probabilities and quantiles directly rather than rely on kurtosis alone.

Low kurtosis means that extreme standardized deviations make a smaller fourth-moment contribution than under the chosen reference. A continuous uniform distribution, for example, has raw kurtosis 1.8. This value does not by itself characterize the height or flatness of its density peak.

Kurtosis is one diagnostic for departures from a normal model, not a method-selection rule. When extremes matter, inspect distribution plots, tail quantiles, and model residuals and then choose methods whose assumptions fit the task.

Why Do Higher Moments Capture Shape?

You might wonder why the third and fourth powers in the skewness and kurtosis formulas capture asymmetry and tail behavior, respectively. The key lies in how exponentiation treats deviations from the mean:

  • Odd powers (3rd moment for skewness): Preserve the sign of deviations. Negative deviations remain negative and positive ones remain positive. In many right-skewed distributions, large positive deviations dominate the cubed sum; the reverse often occurs for left skew. The probabilities and magnitudes on both sides still determine the sign.

  • Even powers (4th moment for kurtosis): Make all deviations nonnegative, while amplifying large deviations. Raising a deviation to the fourth power makes extreme observations contribute disproportionately to the sum.

The standardization by powers of m2m_2 makes these measures unitless. Moments provide compact shape summaries, but each summary discards detail, so plots and application-specific tail measures remain important.

Visual Examples

Visualizing descriptive statistics makes these concepts clearer, helping us see patterns more easily. The following examples show how distributions can differ in their centers and spreads as well as in their overall shapes.

Example: Comparing Distributions with Different Properties

This visualization compares three distributions: a symmetric normal distribution, a right-skewed distribution, and a high-kurtosis distribution with heavy tails. Each plot includes reference lines showing the mean with a solid line and the median with a dashed line to illustrate how these measures respond to distributional shape.

Out[2]:
Visualization
Histogram of a symmetric normal distribution with nearly overlapping vertical lines for the mean and median.
A symmetric normal distribution keeps its mean and median nearly aligned at the center. The balanced shape has skewness near zero and moderate tails.
Histogram of a right-skewed distribution with a long tail to the right and the mean line positioned above the median.
A right-skewed distribution pulls the mean above the median as a small number of large values extend the right tail. This pattern is common in income, prices, and response times.
Histogram of a symmetric heavy-tailed distribution with observations extending far into both tails.
This heavy-tailed sample keeps its mean and median close while producing extreme observations on both sides. Its positive sample excess kurtosis reflects the strong contribution of those extremes.

Example: The 68-95-99.7 Rule

The empirical rule gives an easy way to understand standard deviation. Under a normal model, approximately 68% of probability lies within one population standard deviation of the mean, 95% within two, and 99.7% within three. Proportions in a finite random sample fluctuate around these population probabilities.

Out[3]:
Visualization
Normal distribution curve with nested blue, teal, and amber bands marking 68%, 95%, and 99.7% coverage within one, two, and three standard deviations of the mean.
Nested bands under a normal density show that about 68% of probability lies within one population standard deviation of the mean, 95% within two, and 99.7% within three. Probability beyond the outer band is correspondingly small.

Example: Box Plot Comparison

Box plots give a compact visual summary of distribution characteristics. This example compares the same three distributions using box plots to emphasize how the IQR captures central spread and the conventional fences flag potential outliers [11].

Out[4]:
Visualization
Three vertical box plots compare normal, right-skewed, and heavy-tailed samples. The skewed sample has upper potential outliers while the heavy-tailed sample has many potential outliers above and below the whiskers.
Box plots compress each distribution into its median, interquartile range, whiskers, and 1.5-IQR fliers. The normal sample is balanced, the right-skewed sample produces high-side potential outliers, and the heavy-tailed sample produces potential outliers in both directions.

Example: Measures of Spread Visualization

This example compares the range and IQR with standard deviation to show how each measure captures variability in data. It also shows how these measures behave when outliers are present.

Out[5]:
Visualization
Scatter plot of 100 observations with no points outside the 1.5-IQR fences, with horizontal lines marking the mean, one and two standard deviations, and a shaded interquartile range.
In a baseline sample with no 1.5-IQR potential outliers, the standard deviation bands and interquartile range describe the central spread.
Scatter plot after replacing two observations with low and high extremes, showing wider standard deviation bands while the shaded interquartile range changes much less.
Replacing two observations with low and high extremes expands the range and standard deviation while the interquartile range changes much less.

Choosing Appropriate Descriptive Statistics

Which descriptive statistics to report depends on the data's characteristics and your analysis goals. Consider these guidelines:

  • For roughly symmetric, unimodal distributions without influential extremes: Mean and standard deviation are often useful summaries. Their adequacy still depends on the analytical question and the distribution's tails.
  • For skewed distributions or data with influential extremes: Median and IQR are useful resistant summaries. Retain the mean when the target is an expectation, total, or cost, and consider reporting both conventional and resistant summaries.
  • For nominal categorical data: Report frequencies or proportions and, when useful, the mode. For discrete counts, the mean, median, quantiles, frequencies, or modes may be appropriate depending on the question. When a plot has several stable peaks, report and investigate them rather than compressing the distribution into one center.

Reporting several compatible statistics usually gives a more useful picture than relying on one measure. A difference between mean and median can suggest asymmetry, while a large difference between standard deviation and IQR can prompt investigation of tails or outliers; neither comparison proves the cause. Skewness and kurtosis add compact shape information, but plots and context are still needed to avoid misleading simplifications.

Practical Applications

Descriptive statistics are widely useful across data science and analytics:

  • Quality Control Manufacturing: Control charts monitor process statistics over time and flag patterns that warrant investigation [12]. Their control limits describe expected process variation and are not the same as engineering specification limits, so a statistically stable process can still produce items outside specification [13].
  • Healthcare Research: Survival quantiles can summarize treatment outcomes, but censoring must be handled with a survival estimator such as Kaplan–Meier rather than an ordinary sample median. A median is estimable only if the fitted survival curve falls to 0.5 or below during follow-up; restricted mean survival time may better match some clinical questions.
  • Financial Analysis: Return distributions are examined for asymmetry, heavy tails, and other departures from normal models. Skewness and kurtosis can summarize some of those features, but risk decisions also require horizon-specific tail probabilities, quantiles, and loss measures [14].
  • Marketing Analytics: Purchase amounts, time on site, and conversion rates can have quite different distributions. Analysts should inspect each metric rather than assume a common shape; when purchase amounts are strongly right-skewed, median and upper quantiles can complement the mean.
  • Climate Science: Temperature distributions are described using centers, spreads, and tail measures across stated locations, seasons, and reference periods. A trend in the mean can indicate a location shift. Changes in exceedance rates above fixed, reference-period thresholds address whether defined extremes have become more common, while changes in tail quantiles address whether their magnitudes have shifted. Variance alone cannot establish either conclusion [15].

These examples lead to several practical guidelines for working with descriptive statistics effectively.

Best Practices

Use visualization alongside summary statistics, iterating between them as the analysis develops. Histograms can expose multimodality, box plots can flag potential outliers, and scatter plots can reveal gaps and relationships. Anscombe's quartet demonstrates that datasets can share the same means, variances, correlation, and fitted regression line while having visibly different structures [16].

Report uncertainty alongside inferential point estimates when possible. A sample mean is more informative about a population mean when accompanied by a standard error or confidence interval; a standard deviation instead describes variation among observations [17]. For small samples, acknowledge that estimates can change substantially with additional data.

Document your choices about handling missing values and outliers. Different approaches, such as excluding missing data, imputing values, or trimming extremes, can produce substantially different summaries. Transparency about these decisions allows others to evaluate and reproduce your analysis.

Common Pitfalls

Choose the target estimand before combining percentages or ratios. If the goal is the pooled item-level return rate and one store has a 10% rate on 1,000 items while another has a 20% rate on 100 items, denominator weighting gives (100+20)/(1000+100)≈10.9%(100+20)/(1000+100) \approx 10.9\%, not 15%. An equal-store average answers a different question and may be appropriate when stores, rather than items, are the units of interest.

Standard deviation retains its mathematical meaning as the square root of variance, but a symmetric “mean ± one standard deviation” interpretation can mislead for strongly skewed data. The interval is neither a general coverage guarantee nor a bound on possible values. For nonnegative, right-skewed quantities such as many income measures, report the median and selected quantiles to show asymmetry and coverage; the IQR adds a resistant summary of central spread.

Beware of Simpson's paradox, where an association in aggregated data reverses or disappears after stratification by a third variable [18]. The grouping variable is not automatically a causal confounder; that interpretation needs a causal argument [19]. Computing summaries at relevant levels of aggregation helps reveal such reversals.

Summary

Descriptive statistics give us tools to understand and communicate data characteristics. The mean and median summarize numerical location, while frequencies and the mode summarize categorical prevalence. Variance and standard deviation quantify variability using every observation; the range and interquartile range describe it through ordered endpoints or quartiles. Skewness and kurtosis summarize aspects of asymmetry and extreme-deviation behavior, but they do not fully determine distribution shape or model suitability.

A single descriptive statistic rarely captures every feature relevant to an analysis. Effective summarization means choosing statistics that match the data and question, then pairing them with appropriate plots. Roughly symmetric, unimodal data without influential extremes are often summarized well by mean and standard deviation; skewed or outlier-prone data often benefit from median and IQR. Skewness and kurtosis are diagnostics, not automatic selectors of parametric methods or transformations.

Descriptive statistics involves calculation, but its real purpose is to develop the judgment needed to select appropriate summaries, recognize the patterns they reveal, and communicate findings clearly. These skills form the foundation for subsequent statistical analysis and machine learning, including data-driven decisions.

References

  1. Heckert, N. A., & Filliben, J. J. (2003). NIST/SEMATECH e-Handbook of Statistical Methods: Chapter 1, Exploratory Data Analysis. National Institute of Standards and Technology.
  2. NIST/SEMATECH. (2003a). Measures of location.
  3. Silverman, B. W. (1981). Using kernel density estimates to investigate multimodality. Journal of the Royal Statistical Society: Series B, 43(1), 97–99.
  4. Silverman, B. W. (1986). Density Estimation for Statistics and Data Analysis. Chapman and Hall.
  5. NIST/SEMATECH. (2003c). Measures of scale.
  6. Edelmann, D. (2026). The optimal standardization factor for the variance of the normal distribution: A historical perspective and recommendations for practice. Journal of Statistical Theory and Practice, 20(1), 49.
  7. NIST/SEMATECH. (2003b). Approximate intervals that contain most of the population values.
  8. NIST/SEMATECH. (2003d). What are outliers in the data?.
  9. NIST/SEMATECH. (2003e). Measures of skewness and kurtosis.
  10. Westfall, P. H. (2014). Kurtosis as peakedness, 1905–2014. R.I.P.. The American Statistician, 68(3), 191–195.
  11. NIST/SEMATECH. (2003h). Box plot.
  12. NIST/SEMATECH. (2003f). What are control charts?.
  13. NIST/SEMATECH. (2003i). What are variables control charts?.
  14. Cont, R. (2001). Empirical properties of asset returns: Stylized facts and statistical issues. Quantitative Finance, 1(2), 223–236.
  15. Intergovernmental Panel on Climate Change. (2012). Changes in climate extremes and their impacts on the natural physical environment. In Managing the Risks of Extreme Events and Disasters to Advance Climate Change Adaptation.
  16. Anscombe, F. J. (1973). Graphs in statistical analysis. The American Statistician, 27(1), 17–21.
  17. NIST/SEMATECH. (2003g). Confidence limits for the mean.
  18. Simpson, E. H. (1951). The interpretation of interaction in contingency tables. Journal of the Royal Statistical Society: Series B, 13(2), 238–241.
  19. Pearl, J. (2009). Simpson's paradox, confounding, and collapsibility. In Causality: Models, Reasoning, and Inference (2nd ed., pp. 173–200). Cambridge University Press.

Quiz

Ready to test your understanding of descriptive statistics? Take this quiz to reinforce what you've learned about measures of central tendency, spread, and distribution shape.

Descriptive Statistics Quiz

Question 1 of 100 of 10 completed
When should you use the median instead of the mean as a measure of central tendency?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025descriptivestatistics, author = {Michael Brenndoerfer}, title = {Descriptive Statistics with Python}, year = {2025}, url = {https://mbrenndoerfer.com/writing/descriptive-statistics-guide-python-data-analysis}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-10-06} }
APAAcademic
Michael Brenndoerfer (2025). Descriptive Statistics with Python. Retrieved from https://mbrenndoerfer.com/writing/descriptive-statistics-guide-python-data-analysis
MLAAcademic
Michael Brenndoerfer. "Descriptive Statistics with Python." 2026. Web. October 6, 2026. <https://mbrenndoerfer.com/writing/descriptive-statistics-guide-python-data-analysis>.
CHICAGOAcademic
Michael Brenndoerfer. "Descriptive Statistics with Python." Accessed October 6, 2026. https://mbrenndoerfer.com/writing/descriptive-statistics-guide-python-data-analysis.
HARVARDAcademic
Michael Brenndoerfer (2025) 'Descriptive Statistics with Python'. Available at: https://mbrenndoerfer.com/writing/descriptive-statistics-guide-python-data-analysis (Accessed: October 6, 2026).
SimpleBasic
Michael Brenndoerfer (2025). Descriptive Statistics with Python. https://mbrenndoerfer.com/writing/descriptive-statistics-guide-python-data-analysis

About the author

Continue with the full handbook

This chapter is part of Machine Learning from Scratch. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Machine Learning from Scratch
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.