Volume 1 of 20 · PDF edition
Available nowLanguage Data and Classical NLP
For readers who need a rigorous, practical foundation in text processing, statistical language modeling, search, and distributional meaning.
Read every available chapter online for free. The paid edition is a focused, carefully typeset volume PDF with future chapters, updates, and errata included.
- Written chapters
- 18 written chapters
- Approximate pages
- ~570 pages
- Edition
- Version 2026.08.0
- Price
- $24 one-time

Author and edition details
About the author and Volume 1 PDF edition

Michael Brenndoerfer
Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.
- Edition
- Volume 1 PDF 2026.08.0
- Published
- Last reviewed
Focused learning path
What this volume covers
- Text as Data
- Classical Text Representations
- Distributional Semantics
Audience and prerequisites
Where this volume fits
For readers who need a rigorous, practical foundation in text processing, statistical language modeling, search, and distributional meaning.
Prerequisites: No NLP background required; basic Python and statistics help.
Free online preview
Start with “Character Encoding”
Covers ASCII origins and 7-bit limitations, Unicode code points and planes, UTF-8 variable-width encoding scheme, byte order marks and endianness, encoding detection heuristics, common encoding errors and mojibake, practical encoding/decoding in Python.
Exact contents
18 chapters available now
Part I: Text as Data
- 01Character Encoding
Covers ASCII origins and 7-bit limitations, Unicode code points and planes, UTF-8 variable-width encoding scheme, byte order marks and endianness, encoding detection heuristics, common encoding errors and mojibake, practical encoding/decoding in Python.
- 02Text Normalization
Covers Unicode normalization forms (NFC, NFD, NFKC, NFKD), case folding vs lowercasing, accent and diacritic handling, whitespace normalization, ligature expansion, full-width to half-width conversion, implementing a normalization pipeline.
- 03Regular Expressions
Covers regex syntax and metacharacters, character classes and quantifiers, grouping and backreferences, lookahead and lookbehind assertions, greedy vs lazy matching, common NLP patterns (emails, URLs, dates), regex performance considerations.
- 04Sentence Segmentation
Covers period disambiguation challenges, abbreviation handling, rule-based boundary detection, Punkt sentence tokenizer algorithm, evaluation metrics for segmentation, handling edge cases (quotes, parentheses, lists), multilingual segmentation issues.
- 05Word Tokenization
Covers whitespace tokenization limitations, punctuation handling rules, contractions and clitics, language-specific challenges (Chinese, Japanese, German compounds), Penn Treebank tokenization standard, building a rule-based tokenizer, tokenization evaluation.
Part II: Classical Text Representations
- 01Bag of Words
Covers document-term matrix construction, vocabulary building from corpus, word counting and frequency vectors, sparse matrix representation (CSR/CSC formats), vocabulary pruning (min_df, max_df), binary vs count representations, limitations of word order loss.
- 02N-grams
Covers bigram and trigram extraction, n-gram vocabulary explosion, n-gram frequency distributions, Zipf's law in n-grams, character n-grams for robustness, skip-grams and flexible windows, n-gram indexing for search.
- 03N-gram Language Models
Covers Markov assumption and chain rule, maximum likelihood estimation, probability calculation for sequences, handling unseen n-grams, start and end tokens, generating text from n-gram models, model storage and lookup efficiency.
- 04Smoothing Techniques
Covers add-one (Laplace) smoothing, add-k smoothing and tuning, Good-Turing smoothing derivation, Kneser-Ney smoothing intuition and formula, interpolation vs backoff, modified Kneser-Ney, comparing smoothing methods empirically.
- 05Perplexity
Covers cross-entropy definition and derivation, perplexity as branching factor, relationship to bits-per-character, held-out evaluation methodology, perplexity vs downstream performance, comparing models with perplexity, perplexity limitations and caveats.
- 06Term Frequency
Covers raw term frequency, log-scaled term frequency, boolean term frequency, augmented term frequency, L2-normalized frequency vectors, term frequency sparsity patterns, efficient term frequency computation.
- 07Inverse Document Frequency
Covers document frequency calculation, IDF formula derivation, IDF intuition (rare words matter more), smoothed IDF variants, IDF across corpus splits, relationship to information theory, implementing IDF efficiently.
- 08TF-IDF
Covers TF-IDF formula and variants, TF-IDF vector computation, TF-IDF normalization options, BM25 as TF-IDF extension, document similarity with TF-IDF, TF-IDF for feature extraction, sklearn TfidfVectorizer deep dive.
- 09BM25
Covers BM25 derivation from probabilistic IR, saturation parameter k1, length normalization parameter b, BM25+ and BM25L variants, field-weighted BM25, implementing BM25 scoring, BM25 vs TF-IDF empirically.
Part III: Distributional Semantics
- 01The Distributional Hypothesis
Covers Firth's "you shall know a word by the company it keeps," distributional similarity intuition, context window definitions, paradigmatic vs syntagmatic relations, word similarity from distributions, limitations of distributional semantics.
- 02Co-occurrence Matrices
Covers word-word co-occurrence matrices, word-document matrices, context window size effects, weighting by distance, symmetric vs directional contexts, matrix sparsity patterns, efficient construction algorithms.
- 03Pointwise Mutual Information
Covers PMI formula derivation, PMI interpretation as association, positive PMI (PPMI), shifted PPMI variants, PMI matrix properties, PMI vs raw counts comparison, PMI for collocation extraction.
- 04Singular Value Decomposition
Covers SVD mathematical formulation, truncated SVD for dimensionality reduction, LSA (Latent Semantic Analysis), choosing embedding dimensions, SVD computational complexity, randomized SVD for scale, interpreting SVD dimensions.
Volume 1 PDF
Own this focused edition
Get the carefully typeset PDF to keep, plus every future chapter, revision, and erratum for this volume.
- Carefully typeset standalone volume PDF
- Future chapters, updates, and errata included
- No oversized combined edition to render or download
Volume 1 of 20
$24
one-timeVersion 2026.08.0 · secure checkout via Stripe
Delivered by email · Free volume updates included
All-volume access
Every current and future volume
Get all 20 current volume slots, every future volume, and every update for $199. PDFs are always delivered as manageable individual volumes, never as one combined 12,500-page file.
$199 one-time
Get all-volume accessContinue through the library
Explore adjacent volumes
Version history
Kept current, not frozen in time
Each PDF purchase includes future editions. When the book changes, the updated copy appears in My books at no extra cost.
Current release
Edition 2026.08.0
Initial volume-library release with 19 written PDFs, a Volume 3 placeholder, and no combined edition.