Part of Machine Learning from Scratch
Classify data as quantitative or qualitative, discrete or continuous, and nominal, ordinal, interval, or ratio. Match each type to suitable analysis methods.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Types of Data: Understanding Data Classification
A column filled with numbers can record a measurement, a count, a ranking, or nothing more than an identifier. Treat those cases as interchangeable and you may calculate a meaningless average, erase a real ordering, or impose distances on a model that the data never contained. This chapter shows how to distinguish the common data classifications used here, which comparisons and operations their meanings support, and why the label is only the beginning of a defensible analysis.
Introduction
Data type influences your choice of statistical methods, visualization techniques, and machine learning algorithms. Classifying variables carefully helps you choose suitable techniques and interpret their results, but it does not replace checks of data quality, design, and modeling assumptions.
In this chapter, we'll use a common introductory hierarchy. It separates quantitative from qualitative data, then distinguishes discrete from continuous quantitative data and nominal from ordinal qualitative data. Terminology varies across fields, so the meaning of a variable is more important than the label alone.
Core Data Classification Systems
One common framework distinguishes quantitative variables, whose values represent numerical measurements or counts, from qualitative variables, whose values represent categories. In the hierarchy used here, quantitative variables can be discrete or continuous, while qualitative variables can be nominal or ordinal. Other texts sometimes use these terms differently, so document the operational meaning of each variable (Australian Bureau of Statistics, 2023a).
These classifications describe different aspects of a variable:
- Quantitative versus qualitative describes what its values represent.
- Discrete versus continuous describes its modeled support.
- Nominal, ordinal, interval, and ratio describe which empirical relations and transformations preserve meaning.
They do not form one universal partition. A count, for example, can be both discrete and ratio-scale, while a continuous measurement may be interval- or ratio-scale (Stevens, 1946; Velleman & Wilkinson, 1993).
The familiar nominal–ordinal–interval–ratio vocabulary was presented in psychologist S. S. Stevens's influential 1946 paper. Paraphrasing N. R. Campbell, Stevens made the rule behind a numerical representation part of the definition:
“[M]easurement, in the broadest sense, is defined as the assignment of numerals to objects or events according to rules.”
The phrase “according to rules” is the important part. It explains why two variables stored as numbers may support different comparisons and operations.
Quantitative vs. Qualitative Data
We begin with the distinction between quantitative (numerical) and qualitative (categorical) data.
Quantitative Data
Quantitative data consists of numerical values that represent measurements or counts. Examples include:
- height in centimeters
- annual income in dollars
- number of defects in a product
These are quantitative variables because their values represent measured amounts or counts. Numeric identifiers and numeric category codes are not quantitative merely because they are stored as numbers (Australian Bureau of Statistics, 2023b).
The operations that are meaningful depend on the measurement scale and context. Means and standard deviations are generally used for interval or ratio variables when their distributions make those summaries informative; medians and quantiles apply more broadly to ordered numerical values.
Qualitative Data
Qualitative data represents categories, labels, or attributes rather than numerical quantities. Category values are often stored as text, but they may also be stored as numeric codes. Examples include:
- a documented gender category
- a product brand or color
- a geographic region
These variables are qualitative when their values identify groups rather than magnitudes (Australian Bureau of Statistics, 2023b).
Arithmetic on category labels or their arbitrary codes is not meaningful: adding two region codes does not produce another interpretable region. Frequencies and proportions are meaningful, and categorical variables can also be used in inferential and predictive models through category-aware methods or suitable encodings.


Discrete vs. Continuous Data
Within the quantitative branch of the hierarchy used here, we can further distinguish between discrete and continuous data.
Discrete Data
Discrete quantitative data takes values from a finite or countably infinite set. Literal counts usually take non-negative whole-number values, but discrete numerical values need not be integers. Common discrete counts include:
- number of students in a classroom
- number of defects in a manufacturing process
- number of cars in a parking lot
Each observation records a count from a discrete set of permitted values (NIST/SEMATECH, n.d.).
Mathematical Definition:
A variable is discrete if it can only take on a finite or countably infinite set of values.
How to read this definition:
- is the random variable representing the data
- A finite set has a specific, limited number of possible values (e.g., )
- A countably infinite set can be put into one-to-one correspondence with the natural numbers (e.g., all non-negative integers )
In plain language, the permitted values can be listed one by one. A finite list eventually stops; a countably infinite list continues without end. The values do not have to be equally spaced or integer-valued.
The possible values of a discrete variable are distinct. For literal counts of indivisible units, values such as 2.5 students or 3.7 defects are not possible, although averages, rates, expected counts, and weighted counts can be fractional. Frequency tables and count models are often useful for discrete data, depending on the question and distribution.
Continuous Data
Continuous measurement data is modeled on a continuum, often an interval of real numbers. Common examples include:
- height
- weight
- temperature
- elapsed time
These quantities are commonly treated as continuous even though instruments record them at finite resolution. The continuum is a model of the underlying quantity, not a claim that a measuring device has unlimited resolution (Polyanskiy, 2018; JCGM, 2012).
Mathematical Definition:
A random variable is absolutely continuous if there is a probability density function such that, for any ,
How to read this equation:
The real numbers and are the left and right endpoints of an interval. The condition simply puts them in that order. On the left, is the probability that the random variable falls between those endpoints. On the right, is a placeholder for values along the horizontal axis, and the integral accumulates the density across the same interval. Geometrically, it is the area under the density curve from to .
The density is nonnegative, and its total area over the real line is 1. A density value at one point is not the probability of that exact value: for an absolutely continuous variable, for every individual . It therefore makes no difference whether the endpoints are written with strict or inclusive inequalities.
The equation also does not say that every interval must have positive probability. If throughout an interval, the integral over that interval is zero. An absolutely continuous distribution can therefore have gaps in its support, and its support may also be unbounded (Polyanskiy, 2018).
For a concrete example, suppose is uniform between 0 and 1, so on that interval. Then
The interval has width 0.3, so the corresponding area is 0.3. The probability is therefore 0.3.
Within any interval used to model possible measurements, infinitely many values lie between two distinct points. For example, the model permits values between 170.0 cm and 170.1 cm even when a particular instrument rounds to the nearest millimeter. This distinction between the underlying quantity and recorded resolution matters when choosing models and interpreting results (JCGM, 2012).


Data Type Hierarchy and Subcategories
Nominal vs. Ordinal Data
Within qualitative data, we can distinguish between nominal and ordinal data. The distinction changes which summaries and encodings preserve the available information.
Nominal Data
Nominal data consists of categories without an inherent order or ranking. The displayed order or numeric codes assigned to these categories are arbitrary. Examples include a documented set of gender categories, product colors, or sales regions. Whether a variable is nominal depends on its operational definition and the analysis context, not on how its values happen to be sorted.
Frederic Lord made this point memorable in a 1953 satire about football jersey numbers:
“Since the numbers don’t remember where they came from, they always behave just the same way, regardless.”
The joke is not permission to calculate anything with any numeric code. It warns that the stored values alone do not tell you what the numbers mean; you also need the question and the process that produced them.
Ordinal Data
Ordinal data consists of categories with an inherent order, but the distances between adjacent categories are not established as equal. We can rank the categories, but assigning codes 1, 2, and 3 does not by itself justify treating the gaps as equal. Education levels, satisfaction ratings, and letter grades are common examples when their ordering is defined for the task (Australian Bureau of Statistics, 2023a).
Interval vs. Ratio Data
Within quantitative data, we can distinguish between interval and ratio data based on the presence of a true zero point and the types of mathematical operations that are meaningful.
Interval Data
Interval data has equal units but an arbitrary zero point. Differences and comparisons of differences are meaningful, while ratios of raw values are not. Celsius temperature is a standard example: the difference from 20°C to 30°C equals the difference from 30°C to 40°C, but 20°C is not twice as hot as 10°C. Calendar years are also commonly treated as interval values. Some standardized scores, including certain IQ scores, are treated approximately as interval data only when their construction supports that interpretation (Stevens, 1946).
Ratio Data
Ratio data has equal units and a meaningful zero, so differences and ratios of measurements of the same attribute are interpretable. Height illustrates this property: 180 cm is twice 90 cm, and that ratio is unchanged after conversion to meters (Stevens, 1946). Weight and nonnegative elapsed time are also ratio variables. Income can be treated as ratio data only when observations share a compatible currency, period, and accounting definition; quantities such as net income may also be negative. Sums and products still require attention to units and domain meaning.
Practical Applications
Practical Implications
Data type constrains the analytical approaches that can represent a variable faithfully, but it rarely identifies one method by itself. Method choice also depends on:
- the research question
- whether a variable is an outcome or predictor
- the study design and dependence structure
- the distribution
- the model assumptions
Both quantitative and qualitative variables support descriptive, inferential, and predictive analysis when represented appropriately (Pennsylvania State University, n.d.-a).
Paul Velleman and Leland Wilkinson later summarized the limitation of rigid scale labels directly:
“…scale types are not fundamental attributes of the data…”
Their point is that scale type also depends on how a variable was measured, what else is known about it, and which question the analysis asks. The labels remain useful, but they are the beginning of method selection rather than the end.
Consider healthcare analytics. Pearson's correlation summarizes the strength and direction of a linear relationship between paired quantitative variables (Pennsylvania State University, n.d.-b). Poisson regression models occurrence counts under a Poisson conditional distribution; unequal observation windows require an exposure adjustment, and diagnostics should assess overdispersion (Pennsylvania State University, n.d.-c). A blood-pressure measurement or hospital-visit count does not determine either method on its own. Similarly, purchase amounts and ordinal satisfaction ratings need representations that preserve their meanings. In quality control, defect counts often call for count-aware process-control methods rather than methods that assume a continuous measurement, provided the method's distributional conditions are reasonable.
Best Practices
Start by identifying the semantic type and analytical role of each variable. At minimum, a useful data dictionary records:
- variable names and definitions
- units
- data types
- value codes
- missing-value codes
- derivations
For modeling work, it can also record measurement properties and whether a variable is an outcome, predictor, identifier, or weight (UK Data Service, 2021). For mixed datasets, consider whether the estimator handles each type natively or needs type-specific transformations.
For quantitative data, distinguish interval from ratio scales. Celsius temperatures support meaningful differences but not ratios of raw readings; ratios of weights measured in compatible units are meaningful. For qualitative data, preserve the absence or presence of order: nominal and ordinal variables often need different encodings. Exploratory analysis can reveal unexpected values and coding problems, but metadata and domain knowledge are needed to establish a variable's semantic measurement level.
Data Requirements and Preprocessing
Data quality checks should reflect the variable's meaning. For quantitative data, verify units and plausible ranges, then investigate unusual values instead of deleting them automatically (NIST/SEMATECH, n.d.-b). Convert measurements of the same quantity to common units before comparing their numerical values (NIST, 2010). For qualitative data, use consistent labels and distinguish missing, unknown, not applicable, and observed categories.
Preprocessing also depends on the model:
- For scale-sensitive methods, continuous features often benefit from standardization or another scale transformation fitted only on the training data; many tree-based methods are largely insensitive to feature scale (scikit-learn developers, n.d.-a).
- For count outcomes, diagnostics that show overdispersion warn that a basic Poisson model may fit poorly. When excess zeros are present and domain knowledge supports a separate zero-generating process, a zero-inflated model may be considered (Pennsylvania State University, n.d.-c; Lambert, 1992).
- For nominal variables, one-hot encoding gives each category a separate binary feature.
- For ordinal variables, ordinal encoding can preserve a defensible order, but numeric codes may make an estimator assume ordering or equal spacing that the data do not establish (scikit-learn developers, n.d.-b).
Document mappings, learned parameters, exceptions, and the rationale for each decision.
Common Pitfalls
Common pitfalls include:
- Treating an ordinal variable as nominal, which prevents the fitted analysis from exploiting its known order unless that order is restored.
- Treating arbitrary category codes as quantitative, which can impose unsupported distances.
- Assuming interval data has ratio properties, which leads to invalid multiplicative interpretations.
- Dichotomizing a continuous variable, which discards information and usually reduces statistical power; broader binning can also conceal gradients and nonlinear relationships (Altman & Royston, 2006). Age bins, for example, can hide relationships that a suitable continuous model could retain.
- Misclassifying qualitative measurement levels, which can produce unsuitable encodings or prevent a model from using known order.
Validate representations with domain knowledge, metadata, exploratory analysis, and model diagnostics.
Summary
Understanding data types is important in data science and AI. The distinction between quantitative and qualitative variables, together with discrete, continuous, nominal, ordinal, interval, and ratio properties, influences which analytical methods are appropriate. It does not replace attention to the question, variable roles, study design, data quality, and assumptions.
Quantitative data represents measurements or counts, and its meaningful arithmetic depends on measurement scale. Qualitative data calls for category-aware summaries and either category-aware models or suitable encodings. In the hierarchy used here, discrete quantitative data has finite or countably infinite support, whereas continuous measurement data is modeled on a continuum.
Verify storage formats against the variables' semantic meanings before analysis. These distinctions help guide choices throughout the workflow, from data collection and validation to preprocessing, modeling, and interpretation.
References
- Altman, D. G., & Royston, P. (2006). The cost of dichotomising continuous variables. BMJ, 332(7549), 1080.
- Australian Bureau of Statistics. (2023a). Variables.
- Australian Bureau of Statistics. (2023b). Quantitative and qualitative data.
- Joint Committee for Guides in Metrology. (2012). International Vocabulary of Metrology—Basic and General Concepts and Associated Terms (VIM), 3rd edition.
- Lambert, D. (1992). Zero-inflated Poisson regression, with an application to defects in manufacturing. Technometrics, 34(1), 1–14. DOI.
- Lord, F. M. (1953). On the statistical treatment of football numbers. American Psychologist, 8(12), 750–751.
- National Institute of Standards and Technology. (n.d.-a). What is a probability distribution?.
- National Institute of Standards and Technology. (n.d.-b). Detection of outliers.
- National Institute of Standards and Technology. (2010). Unit conversion.
- Pennsylvania State University. (n.d.-a). Choosing appropriate statistical methods.
- Pennsylvania State University. (n.d.-b). Correlation.
- Pennsylvania State University. (n.d.-c). Poisson regression.
- Polyanskiy, Y. (2018). Continuous random variables. MIT OpenCourseWare.
- scikit-learn developers. (n.d.-a). Importance of feature scaling.
- scikit-learn developers. (n.d.-b). Encoding categorical features.
- Stevens, S. S. (1946). On the theory of scales of measurement. Science, 103(2684), 677–680.
- UK Data Service. (2021). Data documentation: quantitative data.
- Velleman, P. F., & Wilkinson, L. (1993). Nominal, ordinal, interval, and ratio typologies are misleading. The American Statistician, 47(1), 65–72.
Quiz
Ready to test your understanding of data classification? Take this quiz to reinforce what you've learned about the different types of data and their characteristics.
Types of Data Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Machine Learning from Scratch. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Machine Learning from ScratchStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
1 comment
Thank you.
Hi Ebenezer - glad you find the article useful!