Volume 1 of 20 · PDF edition

Available now

Language Data and Classical NLP

For readers who need a rigorous, practical foundation in text processing, statistical language modeling, search, and distributional meaning.

Read every available chapter online for free. The paid edition is a focused, carefully typeset volume PDF with future chapters, updates, and errata included.

Written chapters
18 written chapters
Approximate pages
~570 pages
Price
$24 one-time
Language AI Handbook, Volume 1: Language Data and Classical NLP cover

Author and edition details

About the author and Volume 1 PDF edition

Michael Brenndoerfer, author of Language AI Handbook

Michael Brenndoerfer

Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.

Edition
Volume 1 PDF 2026.08.0
Published
Last reviewed

Focused learning path

What this volume covers

  • Text as Data
  • Classical Text Representations
  • Distributional Semantics

Audience and prerequisites

Where this volume fits

For readers who need a rigorous, practical foundation in text processing, statistical language modeling, search, and distributional meaning.

Prerequisites: No NLP background required; basic Python and statistics help.

Free online preview

Start with “Character Encoding

Covers ASCII origins and 7-bit limitations, Unicode code points and planes, UTF-8 variable-width encoding scheme, byte order marks and endianness, encoding detection heuristics, common encoding errors and mojibake, practical encoding/decoding in Python.

Read the chapter

Exact contents

18 chapters available now

Part I: Text as Data

  1. 01
    Character Encoding

    Covers ASCII origins and 7-bit limitations, Unicode code points and planes, UTF-8 variable-width encoding scheme, byte order marks and endianness, encoding detection heuristics, common encoding errors and mojibake, practical encoding/decoding in Python.

  2. 02
    Text Normalization

    Covers Unicode normalization forms (NFC, NFD, NFKC, NFKD), case folding vs lowercasing, accent and diacritic handling, whitespace normalization, ligature expansion, full-width to half-width conversion, implementing a normalization pipeline.

  3. 03
    Regular Expressions

    Covers regex syntax and metacharacters, character classes and quantifiers, grouping and backreferences, lookahead and lookbehind assertions, greedy vs lazy matching, common NLP patterns (emails, URLs, dates), regex performance considerations.

  4. 04
    Sentence Segmentation

    Covers period disambiguation challenges, abbreviation handling, rule-based boundary detection, Punkt sentence tokenizer algorithm, evaluation metrics for segmentation, handling edge cases (quotes, parentheses, lists), multilingual segmentation issues.

  5. 05
    Word Tokenization

    Covers whitespace tokenization limitations, punctuation handling rules, contractions and clitics, language-specific challenges (Chinese, Japanese, German compounds), Penn Treebank tokenization standard, building a rule-based tokenizer, tokenization evaluation.

Part II: Classical Text Representations

  1. 01
    Bag of Words

    Covers document-term matrix construction, vocabulary building from corpus, word counting and frequency vectors, sparse matrix representation (CSR/CSC formats), vocabulary pruning (min_df, max_df), binary vs count representations, limitations of word order loss.

  2. 02
    N-grams

    Covers bigram and trigram extraction, n-gram vocabulary explosion, n-gram frequency distributions, Zipf's law in n-grams, character n-grams for robustness, skip-grams and flexible windows, n-gram indexing for search.

  3. 03
    N-gram Language Models

    Covers Markov assumption and chain rule, maximum likelihood estimation, probability calculation for sequences, handling unseen n-grams, start and end tokens, generating text from n-gram models, model storage and lookup efficiency.

  4. 04
    Smoothing Techniques

    Covers add-one (Laplace) smoothing, add-k smoothing and tuning, Good-Turing smoothing derivation, Kneser-Ney smoothing intuition and formula, interpolation vs backoff, modified Kneser-Ney, comparing smoothing methods empirically.

  5. 05
    Perplexity

    Covers cross-entropy definition and derivation, perplexity as branching factor, relationship to bits-per-character, held-out evaluation methodology, perplexity vs downstream performance, comparing models with perplexity, perplexity limitations and caveats.

  6. 06
    Term Frequency

    Covers raw term frequency, log-scaled term frequency, boolean term frequency, augmented term frequency, L2-normalized frequency vectors, term frequency sparsity patterns, efficient term frequency computation.

  7. 07
    Inverse Document Frequency

    Covers document frequency calculation, IDF formula derivation, IDF intuition (rare words matter more), smoothed IDF variants, IDF across corpus splits, relationship to information theory, implementing IDF efficiently.

  8. 08
    TF-IDF

    Covers TF-IDF formula and variants, TF-IDF vector computation, TF-IDF normalization options, BM25 as TF-IDF extension, document similarity with TF-IDF, TF-IDF for feature extraction, sklearn TfidfVectorizer deep dive.

  9. 09
    BM25

    Covers BM25 derivation from probabilistic IR, saturation parameter k1, length normalization parameter b, BM25+ and BM25L variants, field-weighted BM25, implementing BM25 scoring, BM25 vs TF-IDF empirically.

Part III: Distributional Semantics

  1. 01
    The Distributional Hypothesis

    Covers Firth's "you shall know a word by the company it keeps," distributional similarity intuition, context window definitions, paradigmatic vs syntagmatic relations, word similarity from distributions, limitations of distributional semantics.

  2. 02
    Co-occurrence Matrices

    Covers word-word co-occurrence matrices, word-document matrices, context window size effects, weighting by distance, symmetric vs directional contexts, matrix sparsity patterns, efficient construction algorithms.

  3. 03
    Pointwise Mutual Information

    Covers PMI formula derivation, PMI interpretation as association, positive PMI (PPMI), shifted PPMI variants, PMI matrix properties, PMI vs raw counts comparison, PMI for collocation extraction.

  4. 04
    Singular Value Decomposition

    Covers SVD mathematical formulation, truncated SVD for dimensionality reduction, LSA (Latent Semantic Analysis), choosing embedding dimensions, SVD computational complexity, randomized SVD for scale, interpreting SVD dimensions.

Volume 1 PDF

Own this focused edition

Get the carefully typeset PDF to keep, plus every future chapter, revision, and erratum for this volume.

  • Carefully typeset standalone volume PDF
  • Future chapters, updates, and errata included
  • No oversized combined edition to render or download

Volume 1 of 20

$24

one-time

Version 2026.08.0 · secure checkout via Stripe

Delivered by email · Free volume updates included

All-volume access

Every current and future volume

Get all 20 current volume slots, every future volume, and every update for $199. PDFs are always delivered as manageable individual volumes, never as one combined 12,500-page file.

Continue through the library

Explore adjacent volumes

Version history

Kept current, not frozen in time

Each PDF purchase includes future editions. When the book changes, the updated copy appears in My books at no extra cost.

Current release

Edition 2026.08.0

Initial volume-library release with 19 written PDFs, a Volume 3 placeholder, and no combined edition.