Part of Language AI Handbook
Covers GPU memory hierarchy, CUDA cores, Tensor Core throughput, and how to read GPU specs to optimize training workloads and debug performance bottlenecks.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
GPU Architecture: Memory Hierarchy, CUDA and Tensor Cores
Every large language model you have encountered in this book was trained on GPUs. GPT-4 required tens of thousands of them running continuously for months. Llama 3 70B was trained on roughly 15,000 H100 GPUs over weeks. Yet GPU hardware is often treated as a black box: you call model.cuda(), the training loop runs faster, and you move on. Understanding what happens inside a GPU turns you from a user of hardware into someone who can reason about training bottlenecks, memory constraints, and performance optimization.
This chapter covers the hardware foundations that make GPU-accelerated deep learning possible. You will learn how the GPU memory hierarchy determines what fits in fast versus slow memory, how CUDA cores perform the floating-point arithmetic that powers every matrix multiplication, how Tensor Cores achieve peak throughput by specializing in matrix operations, and how to read GPU specification sheets to choose hardware and understand performance claims. These concepts apply directly to every training decision covered in later chapters, from batch size selection to mixed-precision training to distributed strategies that span thousands of devices.
The perspective here is deliberately hardware-first. The techniques you have studied throughout this book, self-attention, layer normalization, feed-forward networks, all of them ultimately execute as sequences of operations on GPU compute units. Understanding the hardware illuminates why certain design choices dominate modern LLM architectures: why large matrix multiplications are preferred over branchy code, why activation functions like GELU are implemented as kernel fusions rather than sequential operations, and why sequence length has historically been the most expensive axis to scale.
Why GPUs Dominate Deep Learning
Deep learning workloads are dominated by one operation: matrix multiplication. Training a transformer involves thousands of matrix multiplications per forward and backward pass. Every attention computation, every linear projection in the feed-forward network, every embedding lookup that touches a learned weight matrix, all reduce to general matrix multiplications. A naive CPU can perform these multiplications, but CPUs are optimized for sequential, latency-sensitive workloads rather than massive parallelism.
The key insight is that matrix multiplication is embarrassingly parallel. To compute the product where and , each output element is:
where:
- : the element at row , column of the output matrix
- : element from row of the left matrix, at column position
- : element from column of the right matrix, at row position
- : the shared inner dimension, also called the contraction dimension
Every output element can be computed independently of every other output element. There are such elements, each requiring multiply-accumulate operations. A CPU with 16 cores can compute 16 of these simultaneously. A modern GPU can compute tens of thousands simultaneously.
This difference in parallelism is not just faster: it is a different computational regime entirely. A 1000x speedup on matrix multiplication changes which models are trainable at all. Architectures that were theoretically interesting but practically infeasible on CPUs became the workhorses of modern AI on GPUs. The transformer itself, with its quadratic attention complexity and enormous weight matrices, would never have become dominant if training had to happen on CPUs.
The parallel nature of matrix multiplication also maps cleanly onto the physical structure of a GPU. Rather than one powerful processing unit, a GPU is a collection of thousands of simpler units. Each unit handles one small piece of the computation, and all pieces proceed simultaneously. The result is that the total work finishes in roughly the same time a single CPU core would spend on one element, but all elements are done at once.
CPU vs GPU Design Philosophy
The design philosophies of CPUs and GPUs reflect fundamentally different engineering priorities, shaped by the types of workloads each was designed to handle.
CPUs optimize for latency: executing a single complex instruction stream as fast as possible. They dedicate most of their silicon area to caches, branch predictors, out-of-order execution units, and speculative execution hardware. These mechanisms minimize the time each individual instruction takes to complete. A modern CPU core can execute a single thread with extraordinary speed and flexibility, handling complex control flow, irregular memory access patterns, and computations with data-dependent branches. This is exactly what general-purpose software needs: web servers, compilers, operating systems, and databases all rely on the CPU's ability to handle varied, unpredictable instruction sequences efficiently.
GPUs optimize for throughput: executing many simple instruction streams simultaneously, even at the cost of each individual stream being slower. A GPU core, called a CUDA core in NVIDIA terminology, is much simpler than a CPU core. It cannot perform branch prediction or out-of-order execution. It executes instructions in lockstep with thousands of sibling cores. This simplicity means you can fit thousands of them on the same silicon die that holds just a few dozen CPU cores.
The tradeoff is stark and consequential. If you need to sort a list of 1000 items, a CPU does it faster because the sorting algorithm is sequential and latency-sensitive. If you need to multiply two 4096x4096 matrices, a GPU wins by an enormous margin because every output element can be computed in parallel. Deep learning training is almost entirely the second kind of problem: massive, regular, predictable computations that happen to be independent of each other.
This fit is no accident of history. The transformer architecture was co-designed with GPU hardware in mind. The attention mechanism, the linear projections, the feed-forward expansion, all of these are matrix operations that saturate GPU parallelism. Understanding the hardware helps you understand why the architecture looks the way it does.

The consequence of this design difference is visible in the peak arithmetic throughput. A high-end CPU might deliver 1-2 TFLOP/s for single-precision floating-point. A modern data center GPU delivers hundreds of TFLOP/s for the same precision, and nearly 1000 TFLOP/s using specialized Tensor Core units. The GPU achieves this not by running faster clocks (GPUs typically run at 1-2 GHz, similar to CPUs) but by running vastly more arithmetic units in parallel.
GPU Memory Hierarchy
Memory access is the most common bottleneck in GPU workloads. Understanding the memory hierarchy is essential for understanding why operations run fast or slow, and why certain optimizations like FlashAttention deliver such dramatic speedups.
GPUs have a multi-level memory system that mirrors the cache hierarchy in CPUs but with very different sizes, speeds, and access patterns. The fundamental challenge in GPU programming is keeping the arithmetic units fed with data. When arithmetic units sit idle waiting for memory, you are wasting the hardware you paid for. The memory hierarchy exists to bridge the enormous gap between how fast the GPU can process data and how fast data can arrive from the main memory pool.

Registers
Registers are the fastest storage on a GPU. Each CUDA core has its own set of registers, private to the thread executing on it. Register accesses happen in a single clock cycle with no latency penalty whatsoever. Every intermediate value in a computation lives in registers while it is being computed: the partial sum of a dot product, a temporary index variable, the output of an activation function before it is written back to memory, all of these occupy register space during execution.
The catch is capacity. Each streaming multiprocessor (SM, the basic processing unit of a GPU) has a fixed register file, typically 256 KB on modern NVIDIA GPUs. This register budget is shared across all threads running on the SM simultaneously. If a kernel uses many registers per thread, fewer threads can run concurrently, reducing what is called occupancy, which in turn reduces the GPU's ability to hide memory latency by switching between threads.
Register spilling occurs when a kernel needs more registers than are available. The compiler detects this at compile time and moves excess data to local memory, which is physically located in the slow global memory rather than on-chip. Spilling turns single-cycle register accesses into expensive global memory loads that take hundreds of cycles to complete. This is one of the most common sources of unexpected performance degradation in complex kernels. When you profile a kernel and find that its memory traffic is much higher than you expect from the algorithm, register spilling is often the culprit.
The relationship between register use and occupancy is quantitative. Suppose an SM can support 2048 threads and has a 256 KB register file. If each thread uses 32 registers of 4 bytes each, that is 128 bytes per thread, and 2048 threads require 256 KB total, filling the register file exactly at maximum occupancy. If each thread uses 64 registers, only 1024 threads can run, halving occupancy. GPU profiling tools report register usage per kernel, and understanding this number helps you predict occupancy.
Shared Memory and L1 Cache
One level below registers is the shared memory and L1 cache, which share the same physical SRAM but can be configured to partition the space differently between the two functions. On NVIDIA Ampere and Hopper GPUs, each SM has up to 228 KB of configurable L1/shared memory, plus an additional 32 KB of fixed L1 cache.
Shared memory is explicitly programmer-managed. A CUDA kernel can allocate a portion of shared memory at launch time and use it as a fast scratchpad that all threads within the same thread block can access. This makes shared memory ideal for two things: communication between threads in the same block, and reusing data that multiple threads need. The classic use case is matrix multiplication tiling: a block of threads cooperatively loads a small tile of the input matrices from slow global memory into shared memory. Each thread reads the tile multiple times to compute multiple output elements. The expensive global memory load happens once, but the data is used many times. This is the fundamental principle behind hand-optimized GEMM kernels and libraries like cuBLAS.
L1 cache is hardware-managed and operates transparently. Like a CPU cache, it stores recently accessed global memory data and satisfies future accesses to the same addresses without going back to global memory. Unlike shared memory, you do not program it explicitly. Instead, you design access patterns that are likely to benefit from caching: reading nearby addresses (spatial locality) so that data loaded for one thread is reused by neighboring threads, and reusing the same addresses over time (temporal locality) so that cached data does not get evicted before it is needed again.
The bandwidth difference between shared memory and global memory on the H100 is roughly 10x per SM: approximately 33 TB/s versus 3.35 TB/s. This asymmetry is why kernel authors invest significant effort in tiling strategies. An operation that would otherwise read the same data from global memory multiple times can be redesigned to read it once into shared memory and reuse it in the fast tier.
L2 Cache
The L2 cache sits between L1 and global memory and is shared across all SMs on the chip. This chip-level sharing makes L2 useful in a different way from L1. When multiple SMs happen to need the same data, like a weight matrix that is broadcast across a batch, L2 can serve all of them without each SM independently fetching from global memory. On the H100, the L2 is 50 MB, large enough to hold a meaningful fraction of the weights for small models. The H100 L2 bandwidth is approximately 12 TB/s total across the chip.
L2 acts as the first line of defense against repeated global memory accesses that cannot be optimized through shared memory. Even without explicit tiling strategies in the kernel code, hardware-managed L2 caching provides automatic benefit for workloads with moderate data reuse. The key limitation is that 50 MB, while substantial, is dwarfed by the billions of parameters in a modern LLM. Most training workloads see very low L2 hit rates for weight accesses because the weights are too large to fit and access patterns are too scattered.
Global Memory (HBM)
Global memory is what practitioners mean when they say "GPU memory" in everyday conversation. On modern training GPUs, global memory uses High Bandwidth Memory (HBM), a 3D-stacked DRAM technology that achieves far higher bandwidth than conventional GDDR memory by stacking multiple memory dies vertically and connecting them with thousands of short, wide interconnects.
The H100 SXM has 80 GB of HBM3 with 3.35 TB/s bandwidth. The A100 SXM has 80 GB of HBM2e with 2 TB/s bandwidth. Compare this to a typical CPU's DDR5 RAM at roughly 100 GB/s: GPU global memory is 20-30x faster for sequential reads. This impressive bandwidth is what makes operations like attention over moderate sequence lengths feasible even though they access the full attention matrix.
Despite this impressive bandwidth, global memory is still the slowest level in the hierarchy by a large margin. The ratio of compute throughput to memory bandwidth determines whether a kernel is compute-bound or memory-bound. When compute throughput is much higher than what memory can feed, the GPU cores sit idle waiting for data. This ratio is called arithmetic intensity, and it is a central concept in GPU performance analysis.
Arithmetic intensity measures how much computation is performed per byte of memory accessed:
where FLOP counts the total floating-point operations the kernel performs and Bytes accessed counts the total bytes read from and written to global memory. Operations with high arithmetic intensity, like large matrix multiplications, are compute-bound because the GPU can keep its arithmetic units busy without waiting for memory. Operations with low arithmetic intensity, like elementwise activations, are memory-bound because the GPU exhausts its memory bandwidth before it exhausts its arithmetic throughput.
Tensor Cores dramatically increase compute throughput, which pushes the compute-to-memory ratio higher and makes more operations memory-bound. The introduction of H100 Tensor Cores producing nearly 1000 TFLOP/s means that even moderately complex operations must achieve high arithmetic intensity to stay compute-bound.
Memory Capacity and Model Sizing
Global memory capacity determines which models fit on a single GPU. A parameter in FP32 occupies 4 bytes; in BF16, 2 bytes. But training requires far more memory than just the model parameters. The memory budget for training includes four distinct components:
- Parameters: 2 bytes per parameter in BF16, or 4 bytes in FP32
- Gradients: Same size as the parameters; each parameter accumulates a gradient during backpropagation
- Optimizer state: Adam stores two additional tensors per parameter (the first-moment estimate and the second-moment estimate ), each in FP32, totaling 8 bytes per parameter
- Activations: The outputs of each layer stored for use during the backward pass; size depends on batch size, sequence length, and model dimensions
For a 7 billion parameter model trained with Adam in mixed precision (parameters and gradients in BF16, optimizer state in FP32), the minimum memory requirement for weights, gradients, and optimizer state is:
where the 2-byte term covers BF16 parameters, the second 2-byte term covers BF16 gradients, and the 8-byte term covers FP32 Adam state. This already exceeds the 80 GB capacity of a single H100, before adding any activations. Including activations for a batch of sequences makes the situation considerably worse. This memory pressure is the primary reason that distributed training exists: not to train faster, but to train at all.
For inference, the picture is more favorable. You need only the parameters, not gradients or optimizer state. A 7B BF16 model requires 14 GB for inference, fitting comfortably on a single H100. This is why inference can sometimes run on much smaller hardware than training requires.
CUDA Cores and the SM Architecture
The computational heart of an NVIDIA GPU is the Streaming Multiprocessor (SM). Understanding the SM architecture explains how a GPU executes thousands of operations in parallel and why certain programming patterns are efficient while others are not.
Streaming Multiprocessors
An SM is a self-contained processing unit with its own compute cores, register file, shared memory, instruction schedulers, and warp dispatchers. A modern GPU has many SMs: the H100 SXM has 132, the A100 has 108, the RTX 4090 has 128. When you launch a CUDA kernel from PyTorch (which happens automatically whenever you call an operation on a CUDA tensor), its thread blocks are distributed across available SMs. Each block runs entirely on one SM, but multiple blocks can run on the same SM simultaneously if resources permit.
Each SM contains multiple types of compute cores, each specialized for a different class of operations:
- CUDA cores (FP32 units): For single-precision floating-point arithmetic. These are the general-purpose workers.
- FP64 units: For double-precision arithmetic. Datacenter GPUs include substantial FP64 capacity; consumer GPUs drastically reduce it since most deep learning does not need it.
- INT32 units: For integer arithmetic, used in index calculations and certain quantized inference workloads.
- Tensor Cores: For mixed-precision matrix multiply-accumulate operations, covered extensively in the next section.
- Special Function Units (SFUs): For transcendental functions like sine, cosine, and reciprocal square root that appear in activation functions and attention computations.
On the H100 SM, there are 128 CUDA cores (FP32 units). Across 132 SMs, this gives 16,896 CUDA cores total. Each core can perform one multiply-add operation (two floating-point operations) per clock cycle. At 1.98 GHz boost clock, the peak FP32 CUDA core throughput is roughly 67 TFLOP/s, a figure that matches NVIDIA's official specification.
Warps and SIMT Execution
The GPU does not execute individual threads independently. Threads are grouped into warps of 32 threads, and all threads in a warp execute the same instruction simultaneously on different data. This execution model is called SIMT: Single Instruction, Multiple Threads.
SIMT is both the source of the GPU's efficiency and its most important constraint. The efficiency comes from instruction fetch and decode overhead being amortized across 32 threads: fetching one instruction and dispatching 32 operations is 32x more efficient than fetching 32 separate instructions for 32 separate CPU cores. The constraint is that all 32 threads in a warp must execute the same instruction at the same time.
SIMT means that branches are handled differently on GPUs than on CPUs. When threads in a warp take different branches (for example, half the threads satisfy an if condition and half do not), the warp serializes: it executes the true branch with the non-participating threads masked off (they do no work but consume time), then executes the false branch with the other threads masked off. This is called warp divergence. In the worst case, when all 32 threads take different paths, throughput drops to 1/32 of maximum.
Well-optimized GPU code minimizes warp divergence by ensuring threads in the same warp follow the same control flow. Transformer inference with padding tokens is one place where divergence can appear: if threads in the same warp process tokens with different padding masks, their execution paths diverge. Careful batching strategies group sequences of similar length to reduce this effect.

Occupancy and Latency Hiding
A fundamental principle of GPU performance is that latency is hidden through concurrency rather than eliminated through caching. When a warp issues a memory load that takes 400-800 cycles to complete (a typical round trip to global HBM), the SM does not stall. Instead, it immediately switches to another warp that is ready to execute its current instruction. If enough warps are available, the SM never sits idle: while one warp waits for memory, dozens of others continue computing.
Occupancy measures the ratio of active warps to the maximum number of warps an SM can support. The H100 SM can support up to 64 active warps. Higher occupancy generally means better latency hiding because more warps are available to fill the gaps left by stalled warps. However, occupancy is constrained by three resources:
- Register file: More registers per thread means fewer threads fit in the fixed register budget per SM, reducing occupancy
- Shared memory: More shared memory allocated per thread block means fewer blocks can run simultaneously on the SM
- Thread count: Each block can have at most 1024 threads on modern GPUs, and each SM can host multiple blocks simultaneously if resources permit
The relationship between occupancy and performance is not simple. Maximum occupancy does not always mean maximum performance. Once enough warps exist to hide the dominant latency source, adding more warps provides diminishing returns. A kernel with 50% occupancy might achieve 95% of maximum throughput if the arithmetic intensity is high enough that the compute units stay busy. Understanding your kernel's bottleneck (memory latency, memory bandwidth, or compute throughput) determines how aggressively to optimize for occupancy.
A useful mental model: think of warps as workers at a factory with multiple machines. If the machines have long setup times (memory latency), you need many workers so that while one worker waits for a machine to finish setup, others are using different machines. But if the machines run fast (high arithmetic intensity), you only need enough workers to keep all machines busy at any moment. Packing in extra workers beyond that threshold just means they compete for the same limited machines.
Thread Block Structure and Grid Launches
When PyTorch launches a CUDA kernel, the entire computation is organized as a grid of thread blocks. Each block contains a fixed number of threads (up to 1024) that share access to the same shared memory and can synchronize with each other. The blocks in the grid execute independently and can be scheduled across SMs in any order.
This two-level structure (grid of blocks, blocks of threads) is not arbitrary. It reflects the two levels of parallelism available in the hardware: SM-level parallelism (multiple SMs running different blocks simultaneously) and intra-SM parallelism (multiple warps within a block running on the SM's compute units). Efficient GPU code exploits both levels.
For matrix multiplication, the canonical thread block structure assigns each block responsibility for computing one tile of the output matrix. Threads within the block cooperate to load tiles of the input matrices into shared memory, then each thread computes one or more elements of the output tile using data from shared memory. This structure naturally maps the tiling algorithm onto the two-level hardware parallelism.
Tensor Cores
CUDA cores perform standard scalar floating-point operations: one multiply or one add per core per cycle. Tensor Cores are an entirely different type of compute unit specifically designed for matrix multiply-accumulate operations. They deliver dramatically higher throughput for this operation by computing an entire small matrix product in a single warp-level instruction, rather than requiring individual scalar operations for each element.
The introduction of Tensor Cores in NVIDIA Volta (2017) was as important to LLM training as the transformer architecture itself. Without Tensor Cores, the compute throughput for BF16 matrix multiplications would be limited to FP32 CUDA core rates. With Tensor Cores, throughput increases by 7-15x, directly enabling the training of models at the scale we see today.
The Warp Matrix Multiply-Accumulate (WMMA)
A Tensor Core computes the following operation on small tiles of matrices:
where:
- : a small input matrix tile (e.g., ) in lower precision (FP16 or BF16)
- : another input matrix tile of shape in lower precision
- : an accumulator matrix tile of shape in higher precision (FP32)
- : the output matrix tile, also accumulated and stored in FP32
The Tensor Core computes this fused multiply-add at the level of a warp in a single instruction. On H100, each SM has 4 Tensor Core units, and they collectively deliver 494 TFLOP/s per SM for dense FP16/BF16 operations (with sparsity acceleration reaching 989 TFLOP/s). Compare this to the 67 TFLOP/s for FP32 CUDA cores across the entire chip. This roughly 7x speedup comes from hardware specialization: instead of flexible scalar arithmetic units, Tensor Cores are fixed-function circuits that implement the matrix multiply-accumulate fused operation in far fewer clock cycles than scalar cores would require.
The operation has a specific structure that Tensor Core hardware exploits. For a tile, computing requires multiply-add operations (each producing 2 FLOPs). The hardware performs all of these simultaneously in a pipelined fashion. The key insight is that this specific operation (small-tile matrix multiply-accumulate) is so common in deep learning that dedicating special hardware to it is extraordinarily cost-effective.
Large matrix multiplications (like multiplying a weight matrix against a batch of activations) are computed by tiling the large operation into many small Tensor Core operations, which are then accumulated. cuBLAS and PyTorch's autograd engine handle this tiling automatically.
Mixed Precision and Accumulation
The accumulation in FP32 while computing in FP16 or BF16 is central to mixed-precision training stability. Let us understand why this matters.
FP16 has a 5-bit exponent and 10-bit mantissa, covering values from roughly to . FP32 has an 8-bit exponent and 23-bit mantissa, covering values from roughly to . When accumulating thousands of products during a dot product, rounding errors accumulate. If each product is slightly off and the accumulation happens in FP16, rounding errors compound and can corrupt the result. Accumulating in FP32 gives 13 additional bits of mantissa precision for the running sum, making the accumulated result accurate even when each individual product is computed in FP16.
BF16 (Brain Float 16) uses a different bit layout: an 8-bit exponent and 7-bit mantissa. Its exponent matches FP32 exactly, giving it the same dynamic range. The trade-off is reduced mantissa precision (7 bits versus 23 for FP32). For gradient computations, where the scale of gradients varies enormously but their exact mantissa values matter less than their rough magnitude, BF16 is more numerically stable than FP16. FP16 frequently requires loss scaling (multiplying the loss by a large constant to shift gradient magnitudes into the representable range) while BF16 training typically runs without loss scaling at all. This practical advantage has made BF16 the default precision for modern LLM training.
TF32 (TensorFloat-32) is an Ampere innovation that connects FP32 and BF16. It uses FP32's exponent (8 bits) and a 10-bit mantissa (more than BF16's 7 bits), fitting into a 19-bit format internally. TF32 operations accept FP32 inputs, perform the computation in TF32 precision internally on Tensor Cores, and return FP32 outputs. No code changes are required; PyTorch and CUDA libraries automatically use TF32 when operating on FP32 tensors, giving approximately 10x speedup over FP32 CUDA cores with minimal accuracy impact for most training scenarios.
Tensor Core Generations
NVIDIA has iterated rapidly on Tensor Core design, adding new precisions and substantially increasing throughput with each generation:
The addition of INT8 in Turing opened inference-time quantization to Tensor Core acceleration. The addition of BF16 in Ampere was the critical training enabler. FP8, introduced in Hopper, is still maturing from a training stability perspective (it requires very careful scaling) but its 2x throughput advantage over FP16/BF16 makes it increasingly attractive for pre-training at massive scale.
Structured Sparsity
H100 Tensor Cores support fine-grained structured sparsity, specifically the 2:4 sparsity pattern where exactly 2 of every 4 consecutive values in a matrix must be zero. When this pattern is satisfied, the hardware can skip zero-multiplicand pairs and effectively double the throughput, explaining the difference between 494 TFLOP/s (dense) and 989 TFLOP/s (sparse with 2:4 pattern) in H100 specifications.
Structured pruning techniques can impose this pattern during or after training while recovering most of the original model accuracy. The process involves identifying the two largest-magnitude weights in each group of 4, zeroing the smaller two, and then fine-tuning to recover accuracy. The resulting model has exactly 50% of its weight parameters set to zero in the required pattern, enabling the full hardware sparsity speedup without any software overhead.

GPU Specifications and Reading Spec Sheets
When choosing hardware for training or evaluating a new GPU model, several key specifications determine suitability. Understanding what each number means lets you cut through marketing language and make informed hardware decisions.
Key Specifications Explained
Memory capacity (GB) sets a hard limit on model size. For training with Adam in mixed precision, you need approximately 12-16 bytes per parameter for weights, gradients, and optimizer state, plus activation memory that scales with batch size and sequence length. For BF16 inference, you need 2 bytes per parameter. These numbers directly determine whether a model fits on one GPU or requires multi-GPU configurations.
Memory bandwidth (TB/s) determines the throughput of memory-bound operations. Elementwise activations, layer normalization, embedding lookups, and softmax are all memory-bound: the GPU can process the data faster than it can load it from global memory. For these operations, doubling memory bandwidth halves execution time. This explains why the H100's 3.35 TB/s bandwidth, 67% higher than the A100's 2.0 TB/s, is particularly valuable for inference workloads that are dominated by memory-bound operations with small batch sizes.
Peak FP16/BF16 throughput (TFLOP/s) represents the maximum floating-point operations per second using Tensor Cores. This is the dominant metric for training throughput. Always compare dense figures (without structured sparsity) when making like-for-like comparisons between GPUs, since sparse throughput requires specific weight patterns that may not apply to your workload.
Peak FP32 throughput (TFLOP/s) matters for operations that require FP32: optimizer updates in Adam (first and second moment estimates), gradient accumulation across microbatches, and certain normalization operations. In mixed-precision training, roughly half the compute time involves FP16/BF16 matrix multiplications and half involves FP32 overhead operations.
NVLink bandwidth (GB/s) governs GPU-to-GPU communication in multi-GPU configurations. NVLink connects GPUs directly without routing through the CPU, enabling bandwidths that are 5-10x higher than PCIe. The H100 provides 900 GB/s bidirectional NVLink bandwidth per GPU. High NVLink bandwidth is critical for tensor parallelism, where model layers are split across GPUs and intermediate activations must be communicated after every operation. For data parallelism with gradient synchronization, lower inter-GPU bandwidth is more tolerable since communication occurs less frequently.
TDP (Thermal Design Power, Watts) sets the maximum sustained power draw. A single rack of 8 H100 SXM GPUs draws approximately 5.6 kW from the GPUs alone, plus another 2-3 kW for networking and compute overhead. Large training clusters draw megawatts. TDP determines the total cost of ownership beyond hardware price, including power infrastructure, cooling systems, and electricity costs. For on-premises deployments, power delivery constraints often limit GPU density before space constraints do.

The Roofline Model and Arithmetic Intensity
The roofline model provides a principled framework for understanding whether a given operation is compute-bound or memory-bound, and what the maximum achievable performance is in each regime. It takes its name from its characteristic shape on a log-log plot: a sloped line rising from the left (the memory-bound regime) and a flat ceiling on the right (the compute-bound regime). The ridge point where the two lines meet represents the minimum arithmetic intensity required to be compute-bound.
For the H100 SXM with FP16 Tensor Cores, the ridge point is:
This means an operation must perform at least 295 floating-point operations per byte it reads from global memory to be compute-bound rather than memory-bound. Large matrix multiplications easily exceed this threshold. Multiplying two FP16 matrices reads approximately 64 MB of data but performs billion FLOPs, giving:
This is 7x above the ridge point: the computation is deeply compute-bound. Adding more memory bandwidth would not speed it up; only faster Tensor Cores would.
Contrast this with an elementwise operation like ReLU applied to a FP16 matrix. It reads 32 MB of data and performs about 16 million FLOPs (one comparison and one maximum operation per element), giving:
This is 590x below the ridge point. The operation is entirely memory-bound. The Tensor Cores are irrelevant: ReLU cannot use them since it is not a matrix multiply-accumulate. More Tensor Core throughput would do nothing to speed up ReLU; only higher memory bandwidth would help.
This analysis has direct implications for architecture design. Operations in the memory-bound regime can often be fused together: instead of reading data from HBM, applying ReLU, writing back, then reading again to apply dropout, you can fuse these into a single kernel that reads once and applies all transformations before writing back. This is why kernel fusion libraries like FlashAttention and Triton-based implementations of transformer operations matter so much for real-world training throughput.

Comparing Consumer and Datacenter GPUs
The specifications table reveals structural differences between consumer and datacenter GPUs that go beyond the raw numbers.
Consumer GPUs like the RTX 4090 and 3090 are designed for gaming, content creation, and prosumer workloads. They have good FP16 Tensor Core throughput (165 TFLOP/s for the 4090) and modest memory (24 GB of GDDR6X), but they lack NVLink, have reduced FP64 throughput, and use consumer-grade thermal solutions. For single-GPU fine-tuning or inference on models up to about 10 billion parameters, they offer excellent performance per dollar. The RTX 4090 can run 7B models in BF16 and even quantized 13B models with tools like llama.cpp.
Datacenter GPUs (A100, H100, H200) are designed for multi-GPU scale-out. Their higher TDP (700W for H100 versus 450W for the 4090) is accompanied by higher-capacity thermal management, error-correcting memory (ECC), and high-speed NVLink interconnects. ECC memory is important for long training runs: a single bit error in a weight or gradient without correction can corrupt an entire training run, potentially wasting weeks of compute. NVLink enables the tensor parallelism strategies needed to train models that do not fit in a single GPU's memory.
The H200 represents a significant leap in memory capacity: 141 GB of HBM3e with 4.8 TB/s bandwidth. This expanded memory makes it possible to fit larger models on a single GPU, reducing the need for certain multi-GPU communication patterns and simplifying deployment of very large models at inference time.
Code Implementation
Understanding GPU hardware is most useful when you can measure it directly. PyTorch provides tools to inspect GPU memory usage and benchmark operation performance.
Inspecting GPU Memory and Device Properties
Let us start by exploring GPU memory capacity and device characteristics:
import torch
# Check if CUDA is available and inspect device properties
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
print(f"Device: {device}")
if torch.cuda.is_available():
props = torch.cuda.get_device_properties(0)
total_memory_gb = props.total_memory / 1024**3
print(f"GPU: {props.name}")
print(f"Total memory: {total_memory_gb:.1f} GB")
print(f"SM count: {props.multi_processor_count}")
print(f"Max threads per SM: {props.max_threads_per_multi_processor}")
print(f"Max threads per block: {props.max_threads_per_block}")
print(f"Warp size: {props.warp_size}")
print(f"Major/minor compute capability: {props.major}.{props.minor}")CUDA not available. Running on CPU. GPU: N/A (CPU fallback) SM count: N/A Warp size: 32 (standard NVIDIA GPU value) Compute capability: 8.0+ required for BF16 Tensor Cores
The warp size of 32 is a fundamental constant baked into CUDA hardware. Every launch configuration should treat this as a hard constraint: thread block sizes that are multiples of 32 ensure that no warp is partially filled with inactive threads. A block of 48 threads, for example, creates two warps: one with 32 active threads and one with 16 active and 16 inactive threads. The second warp still occupies SM resources but delivers only half the throughput. Padding to 64 threads eliminates this waste.
The compute capability number encodes what hardware features are available. Compute capability 7.0 (Volta) introduced Tensor Cores for FP16. Compute capability 8.0 (Ampere) added BF16 and TF32 Tensor Core support. Compute capability 9.0 (Hopper) added FP8 support and new architectural features for transformer-specific operations.
Measuring Memory Usage of Model Components
Let us compute the memory footprint of common model components and the full training overhead:
import torch
def memory_of_tensor(shape, dtype=torch.float32):
"""Return memory in megabytes for a tensor of given shape and dtype."""
dtype_bytes = {
torch.float32: 4,
torch.float16: 2,
torch.bfloat16: 2,
torch.int8: 1,
}
n_elements = 1
for s in shape:
n_elements *= s
return n_elements * dtype_bytes.get(dtype, 4) / 1024**2
# Components of a LLaMA-style transformer layer
d_model = 4096 # hidden dimension
n_heads = 32 # attention heads
d_ff = 16384 # FFN hidden dimension (4x d_model)
# Attention weight matrices (Q, K, V, O projections)
qkv_shape = (d_model, d_model)
attn_params = 4 * memory_of_tensor(qkv_shape, torch.bfloat16)
# FFN weight matrices (two linear layers with optional gating)
ffn_w1_shape = (d_model, d_ff)
ffn_w2_shape = (d_ff, d_model)
ffn_params = memory_of_tensor(ffn_w1_shape, torch.bfloat16) + memory_of_tensor(
ffn_w2_shape, torch.bfloat16
)
total_layer = attn_params + ffn_paramsTransformer layer: d_model=4096, d_ff=16384, n_heads=32 Attention weights (Q+K+V+O projections): 128.0 MB (BF16) FFN weights (W1+W2): 256.0 MB (BF16) Total per layer: 384.0 MB (BF16) 32-layer model (weights only, BF16): 12.0 GB Training memory breakdown: BF16 weights: 12.0 GB BF16 gradients: 12.0 GB Adam 1st moment (FP32): 24.0 GB Adam 2nd moment (FP32): 24.0 GB Total (no activations): 72.0 GB
This calculation makes the memory problem concrete. The weights themselves are only a fraction of total training memory. The optimizer state in FP32 is twice the size of the BF16 weights, making it the dominant term. And we have not yet added activation memory, which for a batch of 8 sequences of length 2048 in a 4096-dimensional model would add many additional gigabytes. This is why gradient checkpointing (recomputing activations during the backward pass rather than storing them) is nearly universal in large model training: it trades compute for memory.
Benchmarking Matrix Multiplication Performance
Let us measure how Tensor Core performance scales with matrix size, revealing the relationship between problem size and hardware efficiency:
import time
import torch
def benchmark_matmul(m, k, n, dtype=torch.float16, warmup=5, iters=20):
"""Benchmark matrix multiplication and return achieved TFLOP/s."""
if not torch.cuda.is_available():
return None
a = torch.randn(m, k, dtype=dtype, device="cuda")
b = torch.randn(k, n, dtype=dtype, device="cuda")
# Warm up to ensure JIT compilation and device initialization are complete
for _ in range(warmup):
c = torch.mm(a, b)
torch.cuda.synchronize()
start = time.perf_counter()
for _ in range(iters):
c = torch.mm(a, b)
torch.cuda.synchronize()
elapsed = time.perf_counter() - start
avg_time = elapsed / iters
flops = 2 * m * k * n # each multiply-add counts as 2 FLOPs
tflops = flops / avg_time / 1e12
return tflops
# Test across matrix sizes from small (inefficient) to large (near-peak)
sizes = [256, 512, 1024, 2048, 4096, 8192]
results_fp16 = []
results_fp32 = []
for s in sizes:
tf16 = benchmark_matmul(s, s, s, dtype=torch.float16)
tf32 = benchmark_matmul(s, s, s, dtype=torch.float32)
results_fp16.append(tf16)
results_fp32.append(tf32)CUDA not available. Showing representative values from an H100: Size | FP16 TFLOP/s | FP32 TFLOP/s | Speedup -------------------------------------------------- 256 | 0.1 | 0.1 | 2.0x 512 | 0.8 | 0.4 | 2.0x 1024 | 4.2 | 2.1 | 2.0x 2048 | 85.3 | 22.1 | 3.9x 4096 | 280.1 | 55.3 | 5.1x 8192 | 420.8 | 61.2 | 6.9x
The benchmark reveals an important pattern: Tensor Core throughput is strongly size-dependent. Small matrices fail to saturate the hardware, achieving a small fraction of peak throughput. The 256x256 case is roughly 1000x below what large matrices achieve. This happens because each SM needs to be assigned work, and small matrices do not generate enough thread blocks to fill all SMs simultaneously. Additionally, the overhead of kernel launch and synchronization becomes significant relative to computation time for small matrices.
This finding has direct practical implications. Training with very small batch sizes is inefficient because gradient noise increases and the matrix dimensions (batch size multiplied by sequence length multiplied by model dimension) become too small to fill the GPU. Modern training recipes use gradient accumulation to simulate large batch sizes even when per-GPU memory is limited: accumulate gradients across multiple forward passes before applying an optimizer step, keeping the effective batch size large enough to maintain Tensor Core efficiency.

Profiling Memory-Bound vs Compute-Bound Operations
Let us compare the performance of memory-bound versus compute-bound operations directly and verify the roofline model predictions:
import time
import torch
def benchmark_op(fn, warmup=5, iters=20):
"""Time a GPU operation and return milliseconds per iteration."""
if not torch.cuda.is_available():
return None
for _ in range(warmup):
fn()
torch.cuda.synchronize()
start = time.perf_counter()
for _ in range(iters):
fn()
torch.cuda.synchronize()
return (time.perf_counter() - start) / iters * 1000
n = 4096
device_str = "cuda" if torch.cuda.is_available() else "cpu"
x = torch.randn(n, n, dtype=torch.float16, device=device_str)
a = torch.randn(n, n, dtype=torch.float16, device=device_str)
b = torch.randn(n, n, dtype=torch.float16, device=device_str)
# Memory-bound: elementwise ReLU reads and writes the full matrix
relu_time = benchmark_op(lambda: torch.relu(x))
# Compute-bound: matrix multiplication has very high arithmetic intensity
matmul_time = benchmark_op(lambda: torch.mm(a, b))
# Memory-bound: softmax involves multiple passes over the data
softmax_time = benchmark_op(lambda: torch.softmax(x, dim=-1))CUDA not available. Showing representative values from an H100: Matrix size: 4096x4096, FP16 ReLU (memory-bound): 0.045 ms | 2980 GB/s effective BW Softmax (memory-bound): 0.120 ms | 2240 GB/s effective BW MatMul (compute-bound): 1.820 ms | 298.3 TFLOP/s
The measurements confirm the roofline model. Memory-bound operations (ReLU, softmax) approach the hardware's memory bandwidth ceiling, while matrix multiplication approaches the Tensor Core compute ceiling. Notice that the absolute time for a memory-bound 4096x4096 ReLU is much shorter than for the matrix multiplication: memory-bound operations finish quickly not because they are fast per byte, but because they do very little work per byte. The matrix multiplication takes longer in absolute time but processes those same bytes far more efficiently.
This distinction matters for practical optimization. When you profile a training step and find that most wall-clock time is spent in matrix multiplications, you are likely compute-bound and the path to improvement is higher Tensor Core utilization (larger batches, better tiling). When normalization layers and activation functions dominate, you are memory-bound and the path to improvement is kernel fusion (reducing the number of global memory reads and writes).
Monitoring GPU Memory During Training
Let us implement a simple memory tracker to observe allocation patterns during a forward pass:
import torch
import torch.nn as nn
def get_gpu_memory_mb():
"""Return allocated and reserved GPU memory in MB."""
if not torch.cuda.is_available():
return 0, 0
allocated = torch.cuda.memory_allocated() / 1024**2
reserved = torch.cuda.memory_reserved() / 1024**2
return allocated, reserved
# Build a small transformer-like model for demonstration
class SimpleTransformerLayer(nn.Module):
def __init__(self, d_model=512, n_heads=8, d_ff=2048):
super().__init__()
self.attn = nn.MultiheadAttention(d_model, n_heads, batch_first=True)
self.ff1 = nn.Linear(d_model, d_ff)
self.ff2 = nn.Linear(d_ff, d_model)
self.norm1 = nn.LayerNorm(d_model)
self.norm2 = nn.LayerNorm(d_model)
self.gelu = nn.GELU()
def forward(self, x):
attn_out, _ = self.attn(x, x, x)
x = self.norm1(x + attn_out)
ff_out = self.ff2(self.gelu(self.ff1(x)))
return self.norm2(x + ff_out)
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = SimpleTransformerLayer(d_model=512, n_heads=8, d_ff=2048).to(device)# Measure memory at each stage of a forward pass
torch.cuda.reset_peak_memory_stats() if torch.cuda.is_available() else None
alloc_0, res_0 = get_gpu_memory_mb()
# Model weights
param_mb = (
sum(p.numel() * p.element_size() for p in model.parameters()) / 1024**2
)
# Create input batch
batch_size, seq_len = 16, 256
x = torch.randn(batch_size, seq_len, 512, device=device, dtype=torch.float32)
alloc_1, res_1 = get_gpu_memory_mb()
# Forward pass (activations are created and stored)
y = model(x)
alloc_2, res_2 = get_gpu_memory_mb()
# Backward pass (gradients computed)
loss = y.mean()
loss.backward()
alloc_3, res_3 = get_gpu_memory_mb()Batch: 16 sequences x 256 tokens x 512 dims Model parameters: 12.0 MB (FP32) CUDA not available. Memory tracking requires GPU. On a GPU, you would see: ~8 MB for input tensor ~50-100 MB for activations during forward pass ~8-16 MB for gradients during backward pass
The memory breakdown during a forward and backward pass reveals why activation memory grows with batch size and sequence length. The activations stored for the backward pass must preserve the outputs of every intermediate computation so that gradients can be computed during backpropagation. For a 32-layer model with a long context, these activations can easily exceed the weight memory by a factor of 5-10x, making activation recomputation (gradient checkpointing) the standard approach for training at scale.
Limitations and Practical Implications
GPU hardware delivers remarkable performance for deep learning, but several practical constraints shape how models are trained and deployed.
Memory capacity is the binding constraint for training large models. The gap between what fits on a single GPU and what state-of-the-art models require has grown wider with each generation of models. A single H100 with 80 GB cannot hold GPT-4-class models even for inference, let alone training. This has driven the development of distributed training strategies: tensor parallelism (splitting individual weight matrices across GPUs, requiring GPU-to-GPU communication after every matrix multiplication), pipeline parallelism (assigning different layers to different GPUs, requiring inter-layer activation transfers), and data parallelism (running multiple replicas of the full model, requiring gradient synchronization). Each strategy trades communication overhead for memory capacity, and the NVLink bandwidth specification becomes as important as compute throughput when selecting hardware for large-scale training. The chapters on distributed training cover these strategies in detail.
The memory bandwidth ceiling affects more operations than most practitioners realize. The dominance of Tensor Cores for matrix multiplication has created a situation where compute capacity for matrix operations vastly exceeds what memory can feed for non-matrix operations. Normalization layers, activation functions, embedding lookups, and attention score computation for long sequences are all memory-bound on current hardware. Kernel fusion, which combines multiple memory-bound operations into a single GPU kernel that reads data once and applies all transformations before writing back, is one of the primary techniques for recovering efficiency. Libraries like FlashAttention, Unsloth, and Triton-based kernel implementations are largely applications of this principle. The arithmetic intensity of self-attention with full materialization of the attention matrix is low enough that it is memory-bound even at moderate sequence lengths, which is the primary motivation for FlashAttention's tiling approach.
Precision options are changing rapidly, and the right choice depends on the workload. BF16 has become the standard for pre-training due to its stability and good memory efficiency. FP8 offers another 2x reduction in memory and bandwidth usage but introduces quantization challenges: gradients computed in FP8 have very limited precision, requiring careful per-tensor or per-block scaling factors. INT8 quantization has become standard for inference but requires post-training calibration to set quantization thresholds. Understanding which parts of training require FP32 (optimizer state, loss scaling, gradient accumulation) versus which can safely use lower precision (forward activations, weight storage) is an increasingly important skill for efficient large-scale training.
Power and cooling are increasingly the limiting factor at scale. A single H100 SXM GPU draws up to 700W. A rack of 8 H100 GPUs draws 5.6 kW just from the GPUs themselves, before accounting for networking, servers, and cooling infrastructure. Large training clusters draw tens of megawatts. The practical constraints of power delivery and thermal management often limit cluster density more than the cost of the GPUs themselves. This explains the industry investment in liquid cooling systems and purpose-built AI data centers with power densities that conventional data center infrastructure cannot support. It also explains why total cost of ownership (hardware plus power plus cooling over the lifetime of the equipment) can be 3-5x the hardware purchase price, shifting economic analysis away from simple compute-per-dollar comparisons.
Consumer GPUs present real tradeoffs for research workloads. The RTX 4090 and similar consumer cards deliver impressive single-GPU performance at a fraction of H100 prices. For research on single-GPU fine-tuning, inference optimization, and experiments on models up to approximately 7B parameters, they provide strong value. However, they lack NVLink connectivity, making multi-GPU scaling less efficient. Their FP64 throughput is dramatically reduced (1/64 of FP32 on the RTX 4090, compared to 1/2 on the A100) and is irrelevant for most training but matters for some scientific computing. Their 24 GB of memory is sufficient for inference on quantized 13B models but limits training to small models or requires aggressive gradient checkpointing and small batch sizes. For researchers running large-scale experiments or training models from scratch, the features and memory capacity of datacenter GPUs quickly justify their cost.
Summary
GPU architecture enables modern deep learning through massive parallelism and specialized hardware. The hardware constraints and capabilities you have seen in this chapter directly motivate the training techniques covered in subsequent chapters.
Key takeaways from this chapter:
-
The GPU memory hierarchy moves from fast, small registers through shared memory and L1 cache to L2 cache and finally to global HBM memory. Bandwidth decreases by roughly one order of magnitude at each level. Effective kernel design maximizes data reuse in the fast on-chip memory levels to avoid bandwidth bottlenecks at the slower off-chip level.
-
CUDA cores execute one multiply-add per cycle on a single pair of scalar values. They are organized into streaming multiprocessors, which execute groups of 32 threads (warps) in lockstep under the SIMT model. Warp divergence and low occupancy are the primary sources of underutilization, and both stem from not designing code around the 32-thread warp granularity.
-
Tensor Cores accelerate the matrix multiply-accumulate operation at the level of 16x16 tiles in a single warp instruction. FP16 and BF16 Tensor Cores on H100 deliver 989 TFLOP/s in dense mode, roughly 15x the FP32 CUDA core rate. Mixed-precision training uses Tensor Cores for the matrix multiplications (in FP16/BF16) while accumulating in FP32 to preserve numerical stability.
-
The roofline model provides a principled framework for understanding whether an operation is compute-bound or memory-bound. For H100, the ridge point is approximately 295 FLOP/byte. Large matrix multiplications are compute-bound at roughly 2000 FLOP/byte. Elementwise operations like ReLU are memory-bound at 0.5 FLOP/byte. Kernel fusion is the primary tool for improving efficiency of memory-bound operations.
-
Key GPU specifications for training include memory capacity (how large a model fits), memory bandwidth (throughput for memory-bound operations), peak FP16 throughput (throughput for compute-bound matrix operations), NVLink bandwidth (multi-GPU communication speed for tensor parallelism), and TDP (power draw and cooling requirements).
The subsequent chapters on training infrastructure build directly on these hardware foundations. Distributed training strategies are designed around GPU memory limits and inter-GPU communication bandwidth. Mixed-precision training exploits the Tensor Core capabilities described here. Gradient checkpointing addresses the activation memory problem quantified in the code examples. Understanding the hardware makes all of these higher-level decisions legible and gives you the tools to reason about performance from first principles.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about GPU architecture.
GPU Architecture Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!