Volume 8 of 20 · PDF edition
In progressEncoder, Multilingual, and Translation Models
For readers building understanding, classification, extraction, multilingual, and translation systems rather than decoder-only chat products.
Read every available chapter online for free. The paid edition is a focused, carefully typeset volume PDF with future chapters, updates, and errata included.
- Written chapters
- 14 written chapters
- Planned chapters
- 14 planned chapters
- Approximate pages
- ~440 pages
- Edition
- Version 2026.08.0
- Price
- $24 one-time

Author and edition details
About the author and Volume 8 PDF edition

Michael Brenndoerfer
Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.
- Edition
- Volume 8 PDF 2026.08.0
- Published
- Last reviewed
Focused learning path
What this volume covers
- BERT and Variants
- Encoder-Decoder Models
- Multilingual Language Models and Cross-Lingual Transfer
- Machine Translation and Speech Translation
Audience and prerequisites
Where this volume fits
For readers building understanding, classification, extraction, multilingual, and translation systems rather than decoder-only chat products.
Prerequisites: Assumes Volumes 4–5; linguistic foundations from Volume 3 are recommended.
Free online preview
Start with “BERT Architecture”
Covers BERT model sizes, BERT layer configuration, BERT embedding layers, BERT attention patterns, BERT output representations.
Exact contents
14 chapters available now
The current PDF contains every linked chapter below. The remaining 14 planned chapters will be added through free volume updates.
Part XXIV: BERT and Variants
- 01BERT Architecture
Covers BERT model sizes, BERT layer configuration, BERT embedding layers, BERT attention patterns, BERT output representations.
- 02BERT Pre-training
Covers pre-training data preparation, MLM implementation, NSP task design, pre-training hyperparameters, pre-training duration.
- 03BERT Fine-tuning
Covers classification fine-tuning, sequence labeling fine-tuning, question answering fine-tuning, fine-tuning hyperparameters, catastrophic forgetting.
- 04BERT Representations
Covers [CLS] token usage, layer selection strategies, pooling strategies, BERT as feature extractor, frozen vs fine-tuned representations.
- 05RoBERTa
Covers dynamic masking, NSP removal, larger batches, more data, RoBERTa training recipe, RoBERTa vs BERT performance.
- 06ALBERT
Covers factorized embeddings, cross-layer parameter sharing, sentence order prediction, ALBERT efficiency, ALBERT performance trade-offs.
- 07ELECTRA
Covers generator training, discriminator training, RTD objective, ELECTRA sample efficiency, ELECTRA scaling, ELECTRA fine-tuning.
- 08DeBERTa
Covers disentangled attention formulation, enhanced mask decoder, DeBERTa position encoding, DeBERTa improvements, DeBERTa-v3 advances.
Part XXV: Encoder-Decoder Models
- 01T5 Architecture
Covers T5 encoder-decoder design, T5 attention patterns, T5 model sizes, T5 relative positions, T5 implementation.
- 02T5 Pre-training
Covers span corruption procedure, sentinel tokens, corruption rate, T5 pre-training data, T5 training scale.
- 03T5 Task Formatting
Covers task prefixes, classification as generation, NER as generation, QA as generation, task formatting examples.
- 04BART Architecture
Covers BART encoder-decoder, BART attention configuration, BART vs T5 comparison, BART model sizes.
- 05BART Pre-training
Covers token masking, token deletion, text infilling, sentence permutation, document rotation, objective combinations.
- 06mT5
Covers mT5 training data, language sampling, cross-lingual transfer, mT5 vs T5 performance, multilingual tokenization.
Part XXVI: Multilingual Language Models and Cross-Lingual Transfer
- 01Language Diversity, Typology, and ScriptsPlanned
Writing systems, morphology, word order, language families, and why English-centric assumptions fail.
- 02Multilingual Tokenization and Vocabulary AllocationPlanned
Fertility, script coverage, shared vocabularies, byte models, and unequal token costs across languages.
- 03Multilingual Pre-training and Data BalancingPlanned
Sampling temperatures, capacity allocation, transfer-interference trade-offs, and multilingual mixture design.
- 04Cross-Lingual Transfer and AlignmentPlanned
Shared representations, zero-shot transfer, alignment objectives, adapters, and transfer diagnostics.
- 05Low-Resource and Endangered LanguagesPlanned
Data scarcity, community participation, transliteration, active learning, and responsible evaluation.
- 06Code-Switching and Mixed-Language TextPlanned
Language identification, mixed scripts, code-switched generation, evaluation, and deployment failure modes.
- 07Cultural and Multilingual EvaluationPlanned
Translationese, construct validity, cultural knowledge, local harms, and evaluation led by native speakers.
Part XXVII: Machine Translation and Speech Translation
- 01Statistical Machine Translation FoundationsPlanned
Word alignment, phrase tables, language models, log-linear decoding, and the ideas inherited by neural MT.
- 02Neural Machine TranslationPlanned
Encoder-decoder translation, attention, transformer MT, training objectives, and decoding.
- 03Parallel Data Mining and Quality ControlPlanned
Bitext mining, alignment, filtering, back-translation, synthetic parallel data, and contamination checks.
- 04Many-to-Many and Low-Resource TranslationPlanned
Multilingual transfer, mixture-of-experts translation, zero-shot directions, and capacity bottlenecks.
- 05Translation Evaluation Beyond BLEUPlanned
Learned metrics, adequacy, fluency, terminology, human evaluation, and metric failure modes.
- 06Document-Level and Context-Aware TranslationPlanned
Terminology consistency, discourse context, pronouns, document memory, and long-form evaluation.
- 07Speech-to-Speech TranslationPlanned
Cascaded and end-to-end systems, latency, speaker preservation, prosody, and multilingual safety.
Volume 8 PDF
Own this focused edition
Get the carefully typeset PDF to keep, plus every future chapter, revision, and erratum for this volume.
- Carefully typeset standalone volume PDF
- Future chapters, updates, and errata included
- No oversized combined edition to render or download
Volume 8 of 20
$24
one-timeVersion 2026.08.0 · secure checkout via Stripe
Delivered by email · Free volume updates included
All-volume access
Every current and future volume
Get all 20 current volume slots, every future volume, and every update for $199. PDFs are always delivered as manageable individual volumes, never as one combined 12,500-page file.
$199 one-time
Get all-volume accessContinue through the library
Explore adjacent volumes
Version history
Kept current, not frozen in time
Each PDF purchase includes future editions. When the book changes, the updated copy appears in My books at no extra cost.
Current release
Edition 2026.08.0
Initial volume-library release with 19 written PDFs, a Volume 3 placeholder, and no combined edition.