Volume 8 of 20 · PDF edition

In progress

Encoder, Multilingual, and Translation Models

For readers building understanding, classification, extraction, multilingual, and translation systems rather than decoder-only chat products.

Read every available chapter online for free. The paid edition is a focused, carefully typeset volume PDF with future chapters, updates, and errata included.

Written chapters
14 written chapters
Planned chapters
14 planned chapters
Approximate pages
~440 pages
Price
$24 one-time
Language AI Handbook, Volume 8: Encoder, Multilingual, and Translation Models cover

Author and edition details

About the author and Volume 8 PDF edition

Michael Brenndoerfer, author of Language AI Handbook

Michael Brenndoerfer

Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.

Edition
Volume 8 PDF 2026.08.0
Published
Last reviewed

Focused learning path

What this volume covers

  • BERT and Variants
  • Encoder-Decoder Models
  • Multilingual Language Models and Cross-Lingual Transfer
  • Machine Translation and Speech Translation

Audience and prerequisites

Where this volume fits

For readers building understanding, classification, extraction, multilingual, and translation systems rather than decoder-only chat products.

Prerequisites: Assumes Volumes 4–5; linguistic foundations from Volume 3 are recommended.

Free online preview

Start with “BERT Architecture

Covers BERT model sizes, BERT layer configuration, BERT embedding layers, BERT attention patterns, BERT output representations.

Read the chapter

Exact contents

14 chapters available now

The current PDF contains every linked chapter below. The remaining 14 planned chapters will be added through free volume updates.

Part XXIV: BERT and Variants

  1. 01
    BERT Architecture

    Covers BERT model sizes, BERT layer configuration, BERT embedding layers, BERT attention patterns, BERT output representations.

  2. 02
    BERT Pre-training

    Covers pre-training data preparation, MLM implementation, NSP task design, pre-training hyperparameters, pre-training duration.

  3. 03
    BERT Fine-tuning

    Covers classification fine-tuning, sequence labeling fine-tuning, question answering fine-tuning, fine-tuning hyperparameters, catastrophic forgetting.

  4. 04
    BERT Representations

    Covers [CLS] token usage, layer selection strategies, pooling strategies, BERT as feature extractor, frozen vs fine-tuned representations.

  5. 05
    RoBERTa

    Covers dynamic masking, NSP removal, larger batches, more data, RoBERTa training recipe, RoBERTa vs BERT performance.

  6. 06
    ALBERT

    Covers factorized embeddings, cross-layer parameter sharing, sentence order prediction, ALBERT efficiency, ALBERT performance trade-offs.

  7. 07
    ELECTRA

    Covers generator training, discriminator training, RTD objective, ELECTRA sample efficiency, ELECTRA scaling, ELECTRA fine-tuning.

  8. 08
    DeBERTa

    Covers disentangled attention formulation, enhanced mask decoder, DeBERTa position encoding, DeBERTa improvements, DeBERTa-v3 advances.

Part XXV: Encoder-Decoder Models

  1. 01
    T5 Architecture

    Covers T5 encoder-decoder design, T5 attention patterns, T5 model sizes, T5 relative positions, T5 implementation.

  2. 02
    T5 Pre-training

    Covers span corruption procedure, sentinel tokens, corruption rate, T5 pre-training data, T5 training scale.

  3. 03
    T5 Task Formatting

    Covers task prefixes, classification as generation, NER as generation, QA as generation, task formatting examples.

  4. 04
    BART Architecture

    Covers BART encoder-decoder, BART attention configuration, BART vs T5 comparison, BART model sizes.

  5. 05
    BART Pre-training

    Covers token masking, token deletion, text infilling, sentence permutation, document rotation, objective combinations.

  6. 06
    mT5

    Covers mT5 training data, language sampling, cross-lingual transfer, mT5 vs T5 performance, multilingual tokenization.

Part XXVI: Multilingual Language Models and Cross-Lingual Transfer

  1. 01
    Language Diversity, Typology, and ScriptsPlanned

    Writing systems, morphology, word order, language families, and why English-centric assumptions fail.

  2. 02
    Multilingual Tokenization and Vocabulary AllocationPlanned

    Fertility, script coverage, shared vocabularies, byte models, and unequal token costs across languages.

  3. 03
    Multilingual Pre-training and Data BalancingPlanned

    Sampling temperatures, capacity allocation, transfer-interference trade-offs, and multilingual mixture design.

  4. 04
    Cross-Lingual Transfer and AlignmentPlanned

    Shared representations, zero-shot transfer, alignment objectives, adapters, and transfer diagnostics.

  5. 05
    Low-Resource and Endangered LanguagesPlanned

    Data scarcity, community participation, transliteration, active learning, and responsible evaluation.

  6. 06
    Code-Switching and Mixed-Language TextPlanned

    Language identification, mixed scripts, code-switched generation, evaluation, and deployment failure modes.

  7. 07
    Cultural and Multilingual EvaluationPlanned

    Translationese, construct validity, cultural knowledge, local harms, and evaluation led by native speakers.

Part XXVII: Machine Translation and Speech Translation

  1. 01
    Statistical Machine Translation FoundationsPlanned

    Word alignment, phrase tables, language models, log-linear decoding, and the ideas inherited by neural MT.

  2. 02
    Neural Machine TranslationPlanned

    Encoder-decoder translation, attention, transformer MT, training objectives, and decoding.

  3. 03
    Parallel Data Mining and Quality ControlPlanned

    Bitext mining, alignment, filtering, back-translation, synthetic parallel data, and contamination checks.

  4. 04
    Many-to-Many and Low-Resource TranslationPlanned

    Multilingual transfer, mixture-of-experts translation, zero-shot directions, and capacity bottlenecks.

  5. 05
    Translation Evaluation Beyond BLEUPlanned

    Learned metrics, adequacy, fluency, terminology, human evaluation, and metric failure modes.

  6. 06
    Document-Level and Context-Aware TranslationPlanned

    Terminology consistency, discourse context, pronouns, document memory, and long-form evaluation.

  7. 07
    Speech-to-Speech TranslationPlanned

    Cascaded and end-to-end systems, latency, speaker preservation, prosody, and multilingual safety.

Volume 8 PDF

Own this focused edition

Get the carefully typeset PDF to keep, plus every future chapter, revision, and erratum for this volume.

  • Carefully typeset standalone volume PDF
  • Future chapters, updates, and errata included
  • No oversized combined edition to render or download

Volume 8 of 20

$24

one-time

Version 2026.08.0 · secure checkout via Stripe

Delivered by email · Free volume updates included

All-volume access

Every current and future volume

Get all 20 current volume slots, every future volume, and every update for $199. PDFs are always delivered as manageable individual volumes, never as one combined 12,500-page file.

Continue through the library

Explore adjacent volumes

Version history

Kept current, not frozen in time

Each PDF purchase includes future editions. When the book changes, the updated copy appears in My books at no extra cost.

Current release

Edition 2026.08.0

Initial volume-library release with 19 written PDFs, a Volume 3 placeholder, and no combined edition.