PaLM: Pathways LLM: Large-Scale Training, Reasoning

Michael BrenndoerferAugust 6, 20257 min read

Part of History of Language AI

Covers Google's PaLM, the 540 billion parameter language model that demonstrated advance capabilities in complex reasoning, multilingual understanding.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

2022: PaLM

By 2022, language-model research was testing how performance changed with parameter count, training data, and compute. GPT-3 had reached 175 billion parameters in 2020. Google used its Pathways software and TPU infrastructure to train a substantially larger dense model across multiple TPU pods.

The resulting Pathways Language Model (PaLM) had 540 billion parameters, more than three times GPT-3's count. Google evaluated it on natural-language tasks, multilingual benchmarks, reasoning problems, and code generation.

PaLM improved over smaller comparison models on many reported evaluations. The paper highlighted few-shot results, multilingual transfer, arithmetic and reasoning tasks, and code generation. Performance still varied across tasks, so the results supported scaling trends rather than a general claim that the model reasoned reliably.

PaLM became a prominent data point in scaling research: a dense 540-billion-parameter model could be trained across thousands of accelerators and evaluated on a broad benchmark suite. Its cost and infrastructure requirements also showed how concentrated such experiments had become.

The Problem

GPT-3 and similar models produced fluent text but remained inconsistent on arithmetic, logical puzzles, and multi-step problems. Scaling experiments asked whether a larger pretrained model would improve few-shot performance on these evaluations without task-specific parameter updates.

Performance outside English was another concern. PaLM's training mixture included many languages, allowing the paper to measure cross-lingual and multilingual few-shot behavior in one model.

Code provided a separate test of transfer from a mixed text-and-code corpus. Generating a short solution from a prompt differed substantially from maintaining a multi-file software project, so benchmark gains did not establish general software-development competence.

Training 540 billion parameters exceeded the capacity of one accelerator pod. Pathways coordinated work across multiple TPU v4 pods. The training system also had to partition the model, manage communication, and recover from failures across thousands of chips.

Memory and communication constrained how parameters, optimizer state, and activations could be distributed. The engineering problem was therefore not simply to add chips, but to keep them supplied with work while coordinating model updates.

The question of how to best utilize computational resources also remained open. Previous models had demonstrated that larger models could perform better, but the optimal balance between model size, training data amount, and training procedures was not well understood. Researchers needed systematic approaches to training at unprecedented scales while ensuring that computational resources were used efficiently.

The Solution

Google combined a decoder-only transformer with distributed training through Pathways. The software coordinated a model whose parameters and computation exceeded the capacity of any single TPU pod.

The Pathways System

Pathways provided the software layer for training across multiple TPU pods. It coordinated distributed computation and parameter updates while handling failures in a large accelerator fleet.

Pathways Infrastructure

Pathways coordinated thousands of TPU v4 chips across multiple pods for one training run. This infrastructure made the 540-billion-parameter experiment feasible, though it did not make such training inexpensive or broadly accessible.

Architecture and Training Optimizations

PaLM was a dense decoder-only transformer. Its paper described architectural choices intended to improve training quality and serving efficiency, while Pathways handled distribution across the accelerator fleet.

The training stack partitioned model state and computation across devices and used lower-precision arithmetic where appropriate. These implementation choices reduced memory and communication costs but did not remove the large resource requirement.

Training Data and Procedures

PaLM was trained with next-token prediction on a mixture of text and code. The corpus drew from books, websites, and other text sources, plus code in multiple programming languages.

Filtering changed the composition of the training mixture but could not guarantee clean or unbiased data. The paper also evaluated toxicity and bias, documenting risks that remained after training.

Scaling to 540 Billion Parameters

The 540-billion-parameter run reflected a choice among model size, token count, and available compute. Later scaling work would examine whether different allocations of the same budget could train smaller models on more data.

PaLM's results therefore captured one point on the scaling curve, not a proof that parameter count alone was the optimal use of compute.

Applications and Impact

PaLM's evaluations covered arithmetic and reasoning tasks, multilingual transfer, and code generation. These were research results on a pretrained model rather than a released product or a guarantee of dependable problem solving.

The multilingual results showed that one training mixture could support few-shot tasks across many languages. Quality still varied by language and benchmark, especially where training data was scarce.

Code benchmarks tested completion and generation from natural-language prompts. They did not measure debugging or code review. Nor did they cover testing and long-term maintenance.

Google published the PaLM paper and evaluation results but did not openly release the 540-billion-parameter model weights. Researchers could study the reported methods and outputs without reproducing the complete system.

PaLM's broad evaluation table became a comparison point for later models. As with other benchmark suites, results depended on the prompts and datasets. Scoring procedures also affected the comparison, which did not summarize deployment reliability.

Limitations

PaLM's training and inference costs created a substantial access barrier. Reproducing the 540-billion-parameter run required infrastructure available to only a small number of organizations.

The model's performance varied across tasks and domains. While it demonstrated strong capabilities in many areas, performance on specific tasks could be inconsistent, and the model sometimes struggled with tasks that required very specialized knowledge or reasoning patterns. This variability meant that deploying the model in production applications required careful evaluation and potentially fine-tuning for specific use cases.

The model could generate harmful, biased, or incorrect text. A deployment therefore required evaluation on its intended use, output controls, and human review where errors carried material consequences.

The model's size and computational requirements made it difficult to deploy in resource-constrained environments. Applications requiring real-time responses, deployment on edge devices, or operation with limited computational budgets faced significant challenges. The efficiency improvements from techniques like sparse attention helped but did not eliminate these constraints.

The benchmark suite did not capture every deployment concern. Factual accuracy, behavior on edge cases, and reliability over longer interactions required additional testing.

Legacy and Looking Forward

PaLM demonstrated that Google could train a dense 540-billion-parameter transformer across thousands of TPU chips and obtain strong few-shot results on many reported tasks. It also made the resource concentration behind frontier scaling unusually clear.

Pathways was central to the result: the experiment depended on distributed systems that could keep a large TPU fleet coordinated through one run. That systems contribution was as important as the final parameter count.

The reasoning evaluations encouraged further work on arithmetic and multi-step tasks, while PaLM's own errors showed that scale had not made such behavior systematic.

Its multilingual results supplied evidence that a shared model could transfer across languages without a separate architecture for each one. Uneven data coverage and evaluation quality remained open problems.

The code results added to evidence that mixed text-and-code pretraining could support code understanding, generation, and debugging benchmarks. Later coding systems added tool use, execution feedback, and task-specific training.

PaLM is best read as a scaling and infrastructure result. It paired a very large dense transformer with Pathways and a broad evaluation suite. Questions about compute allocation and reproducibility remained, as did deployment reliability.

The chapter's main lesson is narrower than “bigger is always better.” Scaling improved many measured tasks, but the cost, data mixture, training design, and evaluation setup determined what those gains meant.

Quiz

The following questions review PaLM's parameter count, Pathways infrastructure, training objective, and limitations.

PaLM Quiz

Question 1 of 50 of 5 completed
What was the primary innovation that enabled training PaLM at 540 billion parameters?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025palmpathways, author = {Michael Brenndoerfer}, title = {PaLM: Pathways LLM: Large-Scale Training, Reasoning}, year = {2025}, url = {https://mbrenndoerfer.com/writing/palm-pathways-language-model-large-scale-training-reasoning}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-27} }
APAAcademic
Michael Brenndoerfer (2025). PaLM: Pathways LLM: Large-Scale Training, Reasoning. Retrieved from https://mbrenndoerfer.com/writing/palm-pathways-language-model-large-scale-training-reasoning
MLAAcademic
Michael Brenndoerfer. "PaLM: Pathways LLM: Large-Scale Training, Reasoning." 2026. Web. September 27, 2026. <https://mbrenndoerfer.com/writing/palm-pathways-language-model-large-scale-training-reasoning>.
CHICAGOAcademic
Michael Brenndoerfer. "PaLM: Pathways LLM: Large-Scale Training, Reasoning." Accessed September 27, 2026. https://mbrenndoerfer.com/writing/palm-pathways-language-model-large-scale-training-reasoning.
HARVARDAcademic
Michael Brenndoerfer (2025) 'PaLM: Pathways LLM: Large-Scale Training, Reasoning'. Available at: https://mbrenndoerfer.com/writing/palm-pathways-language-model-large-scale-training-reasoning (Accessed: September 27, 2026).
SimpleBasic
Michael Brenndoerfer (2025). PaLM: Pathways LLM: Large-Scale Training, Reasoning. https://mbrenndoerfer.com/writing/palm-pathways-language-model-large-scale-training-reasoning

About the author

Continue with the full handbook

This chapter is part of History of Language AI. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore History of Language AI
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.