Part of Language AI Handbook
Explains how CLIP enables zero-shot classification by connecting vision and language. Examines dual-encoder architecture, contrastive learning.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
CLIP: Learning Transferable Visual Models from Natural Language Supervision
Before CLIP, teaching a computer to recognize what it saw required an expensive ritual. If you wanted a model to identify a "red panda," you had to collect thousands of images of red pandas, pay human annotators to draw bounding boxes or apply labels, and train a specialized classifier. This approach, while effective, produced brittle models that excelled at specific tasks but failed when presented with new categories or slight distribution shifts. The model knew what a "red panda" was only because it had memorized patterns from labeled examples, not because it understood the concept of a red panda in any meaningful way.
This limitation created a basic bottleneck in computer vision: every new task required a new dataset, and every new dataset required expensive human annotation. The resulting models were narrow specialists, incapable of adapting to novel situations without extensive retraining. A model trained on ImageNet's 1,000 categories could recognize golden retrievers and sports cars, but ask it about a red panda or a specific type of surgical instrument and it would fail completely. The model had no mechanism for generalization beyond the fixed label set it was trained on. This brittleness meant that deploying computer vision systems in new domains, from medical imaging to autonomous driving to satellite analysis, required starting the expensive labeling process all over again.
CLIP (Contrastive Language-Image Pre-training), introduced by OpenAI in 2021, replaced this paradigm by showing that visual concepts could be learned directly from raw text describing images. Rather than training on fixed label sets like ImageNet's 1,000 categories, CLIP learns from 400 million image-text pairs scraped from the internet. The insight is elegant: if an image of a red panda appears alongside the caption "a red panda rests on a branch," the model should learn to associate the visual content with the linguistic description. No human needs to draw a bounding box or assign a category label. The co-occurrence of image and text provides the supervision signal automatically.
This shift enables zero-shot transfer, where the model classifies images into entirely new categories it never encountered during training, simply by being told what to look for through natural language. This capability represents a major change from memorization to understanding. This allows models to generalize to arbitrary visual concepts described in natural language. CLIP does not need to have seen a "trebuchet" during training. At inference time, you simply ask it whether an image looks more like a trebuchet or a catapult, and it can answer based on its learned visual-linguistic associations. This chapter explains how CLIP achieves this, why the training objective works at scale, and what the model's capabilities and limitations reveal about the nature of visual learning.
The Multimodal Challenge
Building a bridge between vision and language requires solving a basic representation problem. Images and text exist in entirely different modalities: one is a grid of pixel values, the other is a sequence of discrete tokens. Traditional computer vision treated these as separate silos, using images to predict fixed categorical labels, while NLP processed text in isolation. This separation meant that visual systems could not use the rich semantic structure of language, and language models could not ground their concepts in visual reality. The two fields evolved independently, each developing powerful but narrowly focused tools.
The challenge runs deeper than a simple format mismatch. Images and text encode information at very different levels of abstraction. A pixel value carries almost no semantic meaning in isolation. Only when you consider spatial relationships between millions of pixels does a recognizable object emerge. Text works differently: a word like "dog" carries dense semantic information through its learned associations with other words. The gap between raw pixel statistics and linguistic semantics is enormous, and bridging it requires learning representations that capture the right level of abstraction in both modalities.
Earlier attempts to connect vision and language typically involved feeding image features into language models or training caption generators. These approaches treated the alignment as a downstream task rather than a core training objective. The result was that the visual representations were learned for one purpose (image classification, object detection) and then adapted for another (captioning, visual question answering). The mismatch between the feature space optimized for classification and the semantic space needed for language grounding created a persistent gap in representational quality.
To connect these modalities effectively, CLIP employs a dual-encoder architecture that maps both images and text into a shared embedding space. In this space, similarity has semantic meaning: an image of a dog and the text "a photo of a dog" should occupy nearby points, while that same image should be distant from "a photo of a car." This approach builds on contrastive learning principles explored in earlier chapters on dense retrieval and bi-encoders, but extends them across modalities rather than remaining within text alone. The key innovation lies in learning a joint representation space where geometric distance corresponds to semantic similarity, regardless of the original input modality. Distance in this space is not about pixels or token frequencies, but about meaning.
The choice to use natural language supervision rather than fixed labels is what makes this architecture scale so dramatically. The internet contains billions of images with associated text, from alt-text on web pages to captions on social media to article text near photographs. This data exists without any deliberate curation for machine learning purposes. By learning to align visual content with this naturally occurring text, CLIP can train on a dataset orders of magnitude larger than any manually annotated collection. The diversity of language used to describe images also provides a richer training signal than category labels could provide, exposing the model to fine-grained distinctions and attributes as well as relationships and contexts.
Architecture
CLIP consists of two parallel encoders that process different input modalities but produce embeddings in the same dimensional space. This architectural choice reflects a deliberate design philosophy: each modality requires specialized processing to handle its unique structure, but the final representations must live in a common geometric space where direct comparison is possible. The parallel structure allows for efficient inference because images and text can be encoded independently and then compared via simple vector operations.
The independence of the two encoders is both a strength and a constraint. Because each encoder processes its modality separately, you can precompute embeddings for large image databases offline, then quickly find matching text queries by comparing vectors without re-running the image encoder. This property makes CLIP extremely practical for retrieval applications. The constraint is that the two encoders cannot attend to each other during encoding: the image encoder does not know what text will be compared against it, and the text encoder does not know what image it will be matched with. All cross-modal reasoning must happen implicitly through the shared embedding space rather than through direct attention across modalities.
Image Encoder
The image encoder turns visual input into fixed-dimensional vectors. CLIP was trained with two possible backbone architectures, which reflects the state of computer vision at the time of development:
- ResNet-based: A modified ResNet-50 or ResNet-101 with attention pooling instead of global average pooling, along with several other improvements to the base architecture including a wider final layer and a different pooling mechanism.
- Vision Transformer (ViT): As discussed in the previous chapter on Vision Transformers, CLIP uses ViT-B/32, ViT-B/16, and ViT-L/14 variants, where the numbers indicate the patch size in pixels.
The choice between convolutional and transformer architectures involves trade-offs between computational efficiency, inductive biases, and representational capacity. Convolutional networks bring strong priors about local spatial structure and locality as well as translation invariance. These priors are appropriate for natural images because nearby pixels tend to be semantically related, and the same visual pattern (an eye, a wheel) should be recognized regardless of where it appears in the image. Transformers lack these priors by design, instead learning all spatial relationships from data. With sufficient training data and compute, this flexibility allows transformers to discover image structure that convolutional networks might not find because their architecture directs attention toward local features. The ViT variants generally produce better embeddings for CLIP because the self-attention mechanism can capture global contextual relationships that are important for matching images to text descriptions.
The ViT variant follows the standard architecture described in the previous chapter: an image is divided into fixed-size patches, linearly embedded, augmented with position embeddings, and processed through transformer blocks. The final representation is taken from the [CLS] token (or average pooled across all patch tokens), then projected via a linear layer to the multimodal embedding dimension, typically 512 or 768 dimensions depending on the model variant. This projection step is important because it allows the vision encoder to maintain its native output dimensionality during pre-training while adapting to the shared multimodal space where cross-modal comparison occurs. Think of it as a learned translation layer that converts the visual encoder's native language into the shared language of the joint embedding space.
The patch size choice involves a direct speed-accuracy trade-off. ViT-B/32 splits a 224x224 image into patches of size 32x32 pixels. ViT-B/16 uses 16x16 patches, creating patches from the same image. Smaller patches capture finer visual detail but require quadratically more computation in the attention layers. ViT-L/14 represents the best-performing configuration in the original paper, using 14x14 patches in a larger (L) model capacity, at significantly higher computational cost. The patch size directly affects which visual details the model can represent and which it must average away.
Text Encoder
The text encoder is a transformer decoder stack with masked self-attention, structurally similar to the GPT architecture described in Part XXVIII. The use of a decoder-only architecture, rather than an encoder (BERT-style) design, is a somewhat surprising choice for a retrieval system where you might expect a bidirectional encoder to produce richer representations. The authors found that the GPT-style decoder produced competitive or better representations for this task, possibly because the causal attention mechanism, despite attending only to left context within each token, still produces a final [EOS] representation that aggregates information from the full sequence.
Key specifications for the text encoder are:
- Base model: 63 million parameters with 12 layers, 512 hidden width, and 8 attention heads
- Larger variant: 82 million parameters with 12 layers, 768 hidden width, and 12 attention heads
- Context length: 76 tokens, shorter than typical language models to match the typical length of image captions
- Vocabulary: Byte Pair Encoding (BPE) with approximately 49,000 tokens
The relatively short context length of 76 tokens reflects an important characteristic of the training data. Image captions found on the internet tend to be brief descriptions, not multi-paragraph articles. By limiting the context length, CLIP reduces computational requirements while capturing the needed descriptive content needed for visual-semantic alignment. In practice, very few image captions exceed 76 tokens, so this constraint rarely limits the model's ability to process training examples. For zero-shot classification, the prompts used are even shorter, typically under 20 tokens.
The text encoder processes tokenized captions and produces a sequence of hidden states. Similar to the image encoder, the final [EOS] token representation is extracted and projected to the shared embedding space via a linear transformation. The [EOS] token is a summary token, aggregating information from the entire caption before the projection. Because the decoder uses causal masking, each token can only attend to previous tokens, but by the time the model processes the final [EOS] token, it has attended to all preceding content. This makes the [EOS] representation a natural aggregation point for the full caption meaning.
An important detail is that both encoders use separate projection matrices to map into the shared space. These projections are learned jointly with the encoders during contrastive training. The projection for the image encoder might have a very different structure than the projection for the text encoder. This reflects the different geometric properties of visual versus linguistic representations before alignment. The joint training ensures that both projections learn to map into a space where cosine similarity corresponds to semantic similarity across modalities.
Projection and Normalization
Both encoders output embeddings that undergo two necessary transformations before entering the shared multimodal space:
- Linear projection: A learned matrix maps encoder outputs to the shared multimodal dimension
- L2 normalization: Embeddings are scaled to unit length, so that similarity calculations depend only on direction, not magnitude
The linear projection layers serve as adaptation mechanisms that allow the pre-trained encoders to align their native representations with the common embedding space. Without these projections, the image and text representations might occupy different subspaces or have incompatible scales, preventing effective cross-modal comparison. You can think of the projections as coordinate system transformations: both encoders produce their representations in their own native coordinates, and the projections rotate and scale these representations until they live in a common coordinate system where semantic similarity corresponds to geometric proximity.
L2 normalization is deceptively important. By constraining all embeddings to the unit hypersphere, the dot product between any two embeddings equals their cosine similarity, a value between -1 and 1 that measures the angle between the vectors. This normalization removes the magnitude component entirely, so the model cannot exploit the contrastive objective by simply making positive pairs larger in magnitude while keeping the angle unchanged. Every embedding carries exactly the same total "energy," forcing the model to encode all information through the direction of the vector rather than its length. This geometric constraint simplifies the loss surface and produces more uniform, well-calibrated representations.
The Contrastive Training Objective
CLIP's training objective is deceptively simple: given a batch of image-text pairs, maximize the cosine similarity of the correct pairs while minimizing the similarity of the incorrect pairs. This is a form of instance discrimination implemented through noise contrastive estimation. The elegance of this approach lies in its scalability: it requires no manual annotation beyond the implicit supervision provided by the co-occurrence of images and text on the internet.
Batch Construction
During training, CLIP samples a mini-batch of image-text pairs where is an image and is its corresponding text caption. The goal is to learn embeddings and such that the dot product is large when and small when .
The construction of batches is itself a necessary design decision. Because the contrastive objective treats all other items in the batch as negative examples, the batch size directly determines the difficulty of the learning task. Larger batches provide more negative examples, creating a more challenging discrimination task that typically results in richer, more generalizable representations. With a batch size of 256, each image must be distinguished from 255 incorrect text descriptions. With a batch size of 32,768, the task becomes dramatically harder, requiring the model to learn finer-grained distinctions to correctly identify its matching text among thousands of alternatives.
This relationship between batch size and representation quality follows an intuitive pattern: with more negatives to discriminate against, the model must learn more fine-grained features to distinguish between similar but distinct concepts. However, this also means that batch composition matters. If a batch contains multiple similar images (several different dog photos), the model must learn to distinguish between fine-grained variations. The presence of these "hard negatives," pairs that are semantically similar but not matched, provides the strongest gradient signal for learning fine-grained visual distinctions. Batches assembled purely from random sampling will produce some hard negatives naturally, but later work has explored explicit hard negative mining to accelerate learning.
The key insight about why internet data provides sufficient supervision is that the implicit assumption made during batch construction, that the only correct pairing is the one that co-occurred in the original web data, is mostly but not perfectly true. Sometimes an image is paired with a caption that loosely describes it, or a caption that describes the scene rather than the main subject. This noise in the supervision signal turns out to be tolerable at scale: with hundreds of millions of examples, the correct signal overwhelms the noise, and the model learns to align visual and linguistic representations despite imperfect labels.
Symmetric Cross-Entropy Loss
CLIP employs a symmetric loss function that treats both image-to-text and text-to-image retrieval tasks equally. For a single batch, we compute an similarity matrix . We compute the scaled similarity score between image and text by taking the dot product of their L2-normalized embeddings and applying a temperature scaling parameter. This scaling controls the sharpness of the probability distribution in the contrastive loss:
where:
- : the scaled similarity score between image and text
- : the L2-normalized embedding of image from the image encoder
- : the L2-normalized embedding of text from the text encoder
- : the temperature parameter, initialized to approximately 0.07 and learned during training (implemented as to ensure positivity)
The temperature parameter controls how sharply the model discriminates between positive and negative pairs. A small temperature sharpens the softmax distribution, putting probability mass on the highest-similarity pair and creating strong gradient signals for hard negatives. A large temperature flattens the distribution, treating all pairs more equally. The optimal temperature depends on the difficulty of the discrimination task and the quality of the training data: noisier data benefits from a flatter distribution, while clean data with many hard negatives benefits from a sharper one. By making the temperature a learned parameter, CLIP allows the model to find the optimal value for its particular training conditions.
The temperature also prevents a degenerate solution. Without temperature scaling, the model could minimize the contrastive loss by simply making all embeddings point in the same direction (maximizing all dot products equally) or by making the magnitude of positive pairs very large while keeping negatives small. L2 normalization prevents the magnitude trick, and a learned temperature prevents the model from exploiting the sharpness of the softmax in degenerate ways. These two constraints together push the model toward a well-structured embedding space.


The loss consists of two cross-entropy terms.
Image-to-Text Loss: To align images with their corresponding texts, we treat each image as a query and all texts in the batch as potential matches. For each image , we want to maximize the probability that the model assigns to the correct text compared to all other texts. This is computed using cross-entropy:
where:
- : the image-to-text contrastive loss
- : the batch size (number of image-text pairs)
- : the similarity score between image and its matching text (diagonal elements)
- : the similarity score between image and text (possibly mismatched)
- The fraction represents the softmax probability of the correct text given image
Text-to-Image Loss: Similarly, we treat each text as a query and all images as potential matches. For each text , we want the correct image to have the highest similarity score compared to all other images in the batch. This computes the reverse direction:
where:
- : the text-to-image contrastive loss
- : the batch size
- : the similarity between text and its matching image
- : the similarity between text and image (note the index order)
- The denominator sums over all images for a fixed text , contrasting the correct pair against all other images in the batch
Notice that the denominator sums over the rows (images) rather than columns (texts). By optimizing both directions simultaneously, the model learns a symmetric alignment where images retrieve their matching texts and texts retrieve their matching images. This keeps the shared embedding space respects semantic relationships from both modalities. The total loss combines both directions symmetrically:
where:
- : the total CLIP contrastive loss
- : the image-to-text loss component
- : the text-to-image loss component
- The factor of averages the two directional losses to ensure balanced gradients
This symmetric formulation ensures that the model learns bidirectional alignment: images should retrieve their texts, and texts should retrieve their images. The symmetry reflects the basic equivalence of the two modalities in the shared embedding space. Neither modality is the privileged "ground truth." Both are equally valid entry points into the semantic space. If you embed an image and then search for the closest text, you should find the same pair as if you started from that text and searched for the closest image. Without symmetry, the model could exploit the asymmetry by placing images and texts in disconnected regions of the space that happen to satisfy the one-directional loss without achieving true semantic alignment.
Why Cross-Entropy Works for Contrastive Learning
The connection between cross-entropy and contrastive learning is worth making explicit. Standard cross-entropy in a classification setting measures how well the model assigns probability to the correct class. Here, the "class" for image is text , and the "logits" are the scaled similarity scores between image and each text in the batch. Maximizing the probability of the correct text is equivalent to minimizing the negative log-likelihood of the correct pairing.
This formulation is sometimes called InfoNCE (Information Noise Contrastive Estimation). The connection to information theory is direct: the InfoNCE loss is a lower bound on mutual information between image and text representations. By minimizing this loss, the model maximizes a lower bound on the mutual information between the two modalities, encouraging the encoders to capture shared semantic content rather than modality-specific noise. This theoretical connection explains why contrastive learning produces representations that are both discriminative and generalizable, the two properties needed for zero-shot transfer.
Scale and Efficiency
CLIP was trained on 400 million image-text pairs collected from the internet (the WebImageText dataset, or WIT). At this scale, the contrastive objective becomes computationally efficient compared to autoregressive alternatives. While predicting text tokens pixel-by-pixel or token-by-token requires sequential generation with a loss computed over a long sequence, contrastive learning computes a single matrix of dot products per batch. This parallelism enables training on massive datasets within reasonable timeframes.
The batch size becomes a necessary hyperparameter. CLIP was trained with batch sizes ranging from 32,768 to 65,536, requiring significant distributed training infrastructure. Large batches provide more negative examples, creating a harder contrastive task that produces richer representations. However, this scaling comes with computational costs, requiring distributed training across many accelerators to process these large batches within memory constraints. The similarity matrix requires memory, which means doubling the batch size quadruples the memory requirement for the similarity computation alone. This constraint drove the design of efficient distributed implementations where different batches are processed on different devices but share their computed embeddings for the similarity calculation.
An important practical consequence of large-batch training is that the model rarely sees the same training example twice. With 400 million examples and a batch size of 32,768, a single epoch requires about 12,000 gradient steps. CLIP was trained for approximately 32 epochs, meaning each example was seen roughly 32 times total during training. This data efficiency requirement puts a premium on the quality and diversity of the training data. Unlike fine-tuning scenarios where you might use thousands of augmented copies of a small dataset, CLIP's training relies on the natural diversity of internet-scraped image-text pairs.
Zero-Shot Transfer
CLIP enables zero-shot classification without any task-specific training. Traditional supervised learning requires training data for every category of interest. CLIP eliminates this requirement by letting you to describe categories through natural language. This capability turns the model from a fixed classifier into a flexible vision system that can be directed through language, much like how you can instruct a person to identify new objects simply by describing them. You do not need to train a person by showing them 10,000 labeled examples before they can recognize a new animal species; describing the animal in words is enough. CLIP extends this natural human capability to neural networks.
The zero-shot capability arises from two converging properties of the trained model. First, the shared embedding space encodes semantic similarity across modalities, so images and their descriptions map to nearby points. Second, the space is continuous and generalizes across concepts rather than discretizing into fixed categories. When you embed the text "a photo of a trebuchet," the resulting vector lands in a region of the embedding space associated with medieval siege weaponry, large wooden structures, and counterweight mechanisms. Any image that shares these visual properties will also land near that region, even if the model never encountered a trebuchet during training. The embedding space does not have explicit "trebuchet slot"; instead, trebuchet-like images occupy a region defined by their visual similarities to other things the model did see.
The Zero-Shot Pipeline
To classify an image without task-specific training:
- Construct prompt templates: Instead of using bare class names like "tench" or "red panda," CLIP uses templates like "a photo of a {label}" or "a rendering of a {label}"
- Compute text embeddings: For each candidate class, embed the templated description using the text encoder
- Compute image embedding: Pass the target image through the image encoder
- Calculate similarities: Compute cosine similarity between the image embedding and all text embeddings
- Predict: Select the class with the highest similarity score
This pipeline effectively turns classification into a retrieval problem. Rather than learning fixed decision boundaries during training, the model dynamically computes the similarity between the input image and arbitrary text descriptions provided at inference time. This flexibility allows the same model to perform classification on entirely new datasets without any parameter updates, simply by changing the set of text prompts.
The computational cost of zero-shot classification is minimal for small class sets. For each classification, you need one forward pass through the image encoder and forward passes through the text encoder (one per candidate class), followed by dot products. For a dataset with 1,000 classes, you can precompute all 1,000 text embeddings once and reuse them across all test images. The marginal cost per image is then just one image encoder forward pass plus 1,000 dot products, making the system computationally competitive with standard classifiers despite requiring no training on the target dataset.
Prompt Engineering and Ensembling
CLIP's performance depends heavily on prompt construction. The original paper found that using prompt templates significantly outperformed bare class names. For example, "a photo of a dog" works better than simply "dog" because it matches the distribution of training captions, which tend to be full sentences describing images. This phenomenon highlights an important aspect of CLIP's learning: the model was trained on natural language descriptions, not isolated category labels, so it performs best when queries match that training distribution. When you query with a bare label, the text encoder processes it as a short, context-free fragment, which may land in a different part of the embedding space than the descriptions that appeared alongside images during training.
The sensitivity to prompts reveals something deeper about how the model works. The text encoder is not simply extracting keywords; it is building a representation sensitive to the full linguistic context. "A photo of a dog" activates visual associations tied to photographic images of dogs, while "a sketch of a dog" would activate different visual features. This context-sensitivity is precisely what enables zero-shot transfer: by choosing your prompt carefully, you can direct the model's attention toward specific visual properties of the target concept.
The authors recommend ensemble averaging over multiple prompt templates:
templates = [
"a photo of a {}",
"a blurry photo of a {}",
"a black and white photo of a {}",
"a low resolution photo of a {}",
"a rendering of a {}",
"a cropped photo of a {}",
# ... additional templates
]For each class, embeddings from all templates are averaged before comparison. This reduces variance and captures different visual aspects of the concept. The ensemble approach works because different prompts probe different visual characteristics: "a blurry photo of a dog" emphasizes texture and coarse shape, while "a cropped photo of a dog" focuses on local features. By averaging across these variations, the model becomes reliable to different image qualities and compositions. The ensemble also reduces the sensitivity to any single prompt's biases, creating more stable predictions across diverse test images.
The effectiveness of prompt ensembling reflects a general principle in machine learning: reducing variance through aggregation. A single prompt might have an unfortunate association in the training data (for instance, the word "tiger" might frequently co-occur with sports teams and logos rather than actual tigers), causing the text embedding to land in a suboptimal region. Averaging across multiple phrasings smooths out these idiosyncratic associations, creating a text embedding that captures the central concept rather than any particular phrasing's quirks.

Linear Probe vs Zero-Shot Performance
It is worth understanding where zero-shot CLIP sits relative to other evaluation regimes. A linear probe evaluation trains a linear classifier on top of frozen CLIP features for the specific target dataset. This requires labeled data but does not update the backbone. Full fine-tuning updates all parameters with labeled data. Zero-shot requires no labeled data at all.
CLIP's zero-shot performance on ImageNet is approximately 76% top-1 accuracy with the ViT-L/14 model, competitive with a supervised ResNet-50 (76.1%) trained on the full ImageNet training set with millions of labeled examples. This comparison is striking: CLIP achieves comparable accuracy to a model that was given 1.2 million labeled training examples from exactly that test distribution, while receiving no examples at all. On many other benchmarks, linear probe CLIP significantly outperforms zero-shot CLIP, suggesting that while the representations are rich, the mapping from features to class-specific predictions benefits from at least minimal supervision.
The performance gap between zero-shot and linear probe evaluations reveals an important limitation of zero-shot classification: the text prompt must describe each class in a way that aligns with how that class appeared in training data. For fine-grained classification benchmarks like the Oxford Pets (distinguishing 37 cat and dog breeds) or EuroSAT (satellite imagery land use), zero-shot CLIP performs significantly below the linear probe variant because the text descriptions of fine-grained categories do not capture the specific visual differences the model needs to discriminate. A linear probe learns exactly which features distinguish "Siamese cat" from "Maine Coon" from the training data, which text prompts cannot replicate without that information.
Worked Example: Computing the Contrastive Loss
Let's trace through a minimal example to solidify the mechanics. Suppose we have a batch of image-text pairs:
- Image 1: Dog, Text: "a dog playing fetch"
- Image 2: Cat, Text: "a cat sleeping on a couch"
- Image 3: Bird, Text: "a bird in flight"
After encoding and L2 normalization, we obtain cosine similarities. For this simplified illustration, we treat these values as the scaled similarity scores (assuming a temperature of for clarity):
| Text 1 | Text 2 | Text 3 | |
|---|---|---|---|
| Image 1 | 0.9 | 0.1 | 0.2 |
| Image 2 | 0.2 | 0.85 | 0.15 |
| Image 3 | 0.1 | 0.3 | 0.88 |
The diagonal represents correct pairs. To compute the image-to-text loss for Image 1, we calculate the softmax probability of the correct match:
where:
- : the softmax probability that Text 1 is the correct match for Image 1
- : the cosine similarity score between Image 1 and Text 1 (the correct pair)
- : the cosine similarity score between Image 1 and Text 2 (an incorrect pair)
- : the cosine similarity score between Image 1 and Text 3 (an incorrect pair)
- : the exponential values , , and respectively
- : the final probability value (approximately 51.4%)
The loss contribution is .
For Image 2:
where:
- : the softmax probability that Text 2 is the correct match for Image 2
- : the cosine similarity score between Image 2 and Text 2 (the correct pair)
- : the cosine similarity score between Image 2 and Text 1 (an incorrect pair)
- : the cosine similarity score between Image 2 and Text 3 (an incorrect pair)
- : the exponential values , , and respectively
- : the final probability value (approximately 49.5%)
Loss contribution: .
For Image 3:
Loss contribution: .
Averaging the three image-to-text losses:
The text-to-image losses are computed symmetrically by treating each text as a query and computing the softmax across image similarities (summing down columns rather than across rows). The final CLIP loss averages both directions: .
Notice that with only three items in the batch, the model achieves about 51% probability on the correct pair. In a large batch with 32,768 items, achieving 51% probability on the correct pair out of thousands of candidates would represent excellent performance. As training progresses, the model should push these diagonal values higher while driving off-diagonal values lower, ultimately achieving near-perfect discrimination within each batch.

Averaging all three image-to-text losses and computing the symmetric text-to-image component gives the final batch loss. This process encourages the model to adjust embeddings so that correct pairs (diagonal values) have higher similarity than incorrect pairs (off-diagonal values), effectively learning to discriminate between matching and non-matching pairs within the batch. The gradients flow back through both encoders, updating weights to increase the separation between positive and negative pairs. Over millions of such batches, the encoders converge toward representations where visual and linguistic descriptions of the same concept occupy the same neighborhood in the shared embedding space.
Code Implementation
Let's implement CLIP-based zero-shot classification using the Hugging Face Transformers library. We'll walk through loading the model, processing inputs, and performing classification.
from transformers import CLIPModel, CLIPProcessor
# Load pre-trained CLIP model and processor
model_name = "openai/clip-vit-base-patch32"
model = CLIPModel.from_pretrained(model_name)
processor = CLIPProcessor.from_pretrained(model_name)
# Model configuration details (uncomment to view after loading)
# print(f"Model loaded: {model_name}")
# print(f"Text encoder layers: {model.text_model.config.num_hidden_layers}")
# print(f"Image encoder layers: {model.vision_model.config.num_hidden_layers}")
# print(f"Embedding dimension: {model.config.projection_dim}")Model loaded: openai/clip-vit-base-patch32 Text encoder layers: 12 Image encoder layers: 12 Embedding dimension: 512
The CLIP ViT-B/32 model uses a 12-layer vision transformer with patch size 32 and a 12-layer text transformer, both projecting to a 512-dimensional shared space. This shared dimension is the key architectural constraint that makes cross-modal comparison possible.
Now let's download a sample image and prepare candidate labels for zero-shot classification:
import io
import requests
from PIL import Image, ImageDraw
# Download a real dog photo so CLIP can actually classify it correctly
dog_url = "https://images.unsplash.com/photo-1543466835-00a7907e9de1?w=320&auto=format"
headers = {"User-Agent": "Mozilla/5.0"}
try:
response = requests.get(dog_url, headers=headers, timeout=15)
response.raise_for_status()
image = Image.open(io.BytesIO(response.content)).convert("RGB")
except requests.RequestException:
image = Image.new("RGB", (320, 240), color=(210, 180, 130))
draw = ImageDraw.Draw(image)
draw.ellipse((95, 55, 225, 185), fill=(140, 95, 45))
draw.text((118, 190), "dog", fill=(40, 30, 20))
# Define candidate classes for zero-shot classification
candidate_labels = ["a dog", "a cat", "a bird", "a car", "a tree"]
prompt_template = "a photo of a {}"
# Create full prompts using the template
text_inputs = [prompt_template.format(label) for label in candidate_labels]
print(f"Image size: {image.size}")
print(f"Candidate labels: {candidate_labels}")Image size: (320, 240) Candidate labels: ['a dog', 'a cat', 'a bird', 'a car', 'a tree']
Note the use of the prompt template. We construct "a photo of a dog" rather than just "dog" because this matches the distribution of captions in the training data and produces more stable embeddings.
Next, we process the inputs through CLIP. The processor handles tokenization for text and normalization plus reshaping for images:
# Process inputs: tokenize text and preprocess image
inputs = processor(
text=text_inputs, images=image, return_tensors="pt", padding=True
)
print(f"Input IDs shape: {inputs['input_ids'].shape}")
print(f"Pixel values shape: {inputs['pixel_values'].shape}")
print(f"Attention mask shape: {inputs['attention_mask'].shape}")Input IDs shape: torch.Size([5, 8]) Pixel values shape: torch.Size([1, 3, 224, 224]) Attention mask shape: torch.Size([5, 8])
The processor converts our 5 text prompts into token tensors of shape (5, 77) where 77 is CLIP's maximum context length, and the image into a pixel tensor of shape (1, 3, 224, 224) representing batch size, RGB channels, and resolution. CLIP's processor normalizes the image using the mean and standard deviation of the training data, centering each channel around zero and scaling to unit variance. This preprocessing ensures the image encoder receives inputs in the same statistical regime it saw during training.
Now we compute the embeddings and similarities:
import torch
# Forward pass through CLIP
with torch.no_grad():
outputs = model(**inputs)
# Extract embeddings
image_embeds = outputs.image_embeds # Shape: (1, 512)
text_embeds = outputs.text_embeds # Shape: (5, 512)
# L2 normalization (already normalized by CLIP, but applying explicitly for clarity)
image_embeds = image_embeds / image_embeds.norm(dim=-1, keepdim=True)
text_embeds = text_embeds / text_embeds.norm(dim=-1, keepdim=True)
# Compute cosine similarities (dot product of normalized vectors)
similarities = (image_embeds @ text_embeds.T).squeeze(0)
# Convert to probabilities using softmax with temperature scaling
temperature = model.logit_scale.exp().item()
probs = (similarities * temperature).softmax(dim=0)
print(f"Image embeddings shape: {image_embeds.shape}")
print(f"Text embeddings shape: {text_embeds.shape}")
print(f"Temperature (learned): {temperature:.4f}")Image embeddings shape: torch.Size([1, 512]) Text embeddings shape: torch.Size([5, 512]) Temperature (learned): 100.0000
The model.logit_scale stores the learned log-temperature parameter. Taking its exponential gives the actual temperature value, typically around 100 (the inverse of 0.01 in terms of ). This means the model uses an extremely sharp softmax that concentrates essentially all probability mass on the highest-scoring candidate. The learned temperature is much larger than the initialization value because training on 400 million diverse examples allows the model to become very confident in its similarity judgments for semantically distinct categories. Finally, let's visualize the classification results:

The results demonstrate CLIP's ability to classify without task-specific training. The model typically assigns probability greater than 0.95 to "a dog" for clear dog photographs, with negligible probability distributed among the other candidates. This extreme concentration reflects the large learned temperature, which sharpens the softmax so strongly that even a modest advantage in cosine similarity translates into near-certain prediction.
Embedding Extraction and Similarity Search
Beyond classification, CLIP embeddings enable semantic image search and clustering. Because the embedding space is shared across modalities, you can use text queries to retrieve images, or use images to retrieve similar images, all within the same vector space. Let's extract embeddings for multiple images and compute pairwise similarities:
# Download real images of semantically distinct subjects so CLIP embeddings
# reflect semantic content rather than colour alone.
import io
import requests
image_urls = [
(
"dog",
"https://images.unsplash.com/photo-1543466835-00a7907e9de1?w=160&auto=format",
),
(
"cat",
"https://images.unsplash.com/photo-1514888286974-6c03e2ca1dba?w=160&auto=format",
),
(
"car",
"https://images.unsplash.com/photo-1541899481282-d53bffe3c35d?w=160&auto=format",
),
(
"dog2",
"https://images.unsplash.com/photo-1587300003388-59208cc962cb?w=160&auto=format",
),
]
headers = {"User-Agent": "Mozilla/5.0"}
images = []
labels = []
for label, url in image_urls:
try:
resp = requests.get(url, headers=headers, timeout=15)
resp.raise_for_status()
img = Image.open(io.BytesIO(resp.content)).convert("RGB")
except requests.RequestException:
img = Image.new("RGB", (160, 160), color=(220, 220, 220))
ImageDraw.Draw(img).text((40, 70), label, fill=(30, 30, 30))
images.append(img)
labels.append(label)
# Process all images
image_inputs = processor(images=images, return_tensors="pt", padding=True)
# Extract embeddings
with torch.no_grad():
vision_out = model.get_image_features(
pixel_values=image_inputs["pixel_values"]
)
image_features = vision_out.pooler_output
image_features = image_features / image_features.norm(dim=-1, keepdim=True)
# Compute similarity matrix
similarity_matrix = (image_features @ image_features.T).numpy()
print(f"Similarity matrix shape: {similarity_matrix.shape}")
print(f"Labels: {labels}")Similarity matrix shape: (4, 4) Labels: ['dog', 'cat', 'car', 'dog2']
Now let's visualize the similarity matrix to see how CLIP clusters semantically similar images:

The similarity matrix is best read comparatively. The diagonal is 1.0 by construction, and each off-diagonal entry shows how these particular images relate in CLIP's embedding space. The dog pair should be compared with the dog-cat and dog-car pairs, but one four-image sample is not enough to establish categorical separation: exact scores depend on image content, composition, and preprocessing. A proper evaluation would repeat this analysis across many labeled examples and test whether same-class pairs are consistently more similar than different-class pairs.
Key Parameters
Understanding the key parameters for CLIP deployment helps you make informed decisions about which model variant to use and how to configure inference:
- projection_dim: Dimension of the shared embedding space (512 for ViT-B/32, 768 for larger variants). This determines the size of the vectors you store for retrieval applications and the computational cost of similarity search.
- temperature: Learned scaling parameter (initialized to approximately 0.07, typically converging to an effective value around 100 after training). Controls the sharpness of the softmax distribution over similarity scores. Higher values produce more confident predictions.
- num_hidden_layers: Number of transformer layers in the encoders (12 layers for the ViT-B/32 variant). More layers capture more complex visual and linguistic patterns but increase inference time and memory requirements.
- prompt_template: Template string used to format class labels into natural language descriptions (for example, "a photo of a {}"). This significantly affects zero-shot performance and should match the style of captions in the training data.
- patch_size: For ViT variants, the size of image patches (32, 16, or 14 pixels). Smaller patches capture finer detail but dramatically increase the number of tokens and the quadratic attention cost.
Limitations and Impact
While CLIP was a paradigm shift in computer vision, it carries significant limitations that practitioners must understand before deploying it in real applications.
Domain specificity and distribution shift: CLIP performs remarkably well on natural images but struggles with specialized domains. Medical imaging, satellite imagery, and highly technical diagrams often fall outside the distribution of internet-scraped training data. The model may fail silently on out-of-distribution inputs, creating confident but incorrect predictions. This limitation arises because CLIP's training data consists primarily of natural photographs with accompanying captions, leaving it vulnerable to domain shifts where the visual statistics differ significantly from natural images. A chest X-ray contains almost no information that overlaps with photographs of dogs and cars, and the linguistic descriptions accompanying X-rays in medical literature differ structurally from casual image captions. Deploying CLIP directly in medical or scientific imaging requires careful evaluation and often fine-tuning on domain-specific data.
Social biases and fairness concerns: CLIP inherits and sometimes amplifies biases present in internet data. Studies have shown that CLIP associates certain demographic groups with stereotypical attributes and performs inconsistently across different ethnicities and genders. The zero-shot nature makes these biases particularly insidious because they appear without explicit training on biased labels. Because the model is not trained on carefully curated datasets but on raw internet data, it reflects the stereotypes and imbalances present in online content. When you use CLIP to make decisions about people, such as filtering resume photos or identifying individuals in security contexts, these biases can manifest as discriminatory outcomes. The original paper includes an explicit discussion of these risks and recommends against deploying CLIP in high-stakes contexts without careful bias auditing.
Fine-grained discrimination: While excellent at broad categorization, CLIP struggles with fine-grained distinctions. Differentiating between similar dog breeds, specific bird species, or subtle medical conditions requires more specialized supervision than contrastive learning on internet captions typically provides. The training data contains broad descriptions rather than expert-level taxonomic labels, limiting the granularity of visual concepts the model can reliably distinguish. When someone writes a caption like "a photo of a bird," they rarely specify the species. The model therefore learns that these images belong to the general category of birds without acquiring the fine-grained visual features needed to distinguish a Spotted Sandpiper from a Solitary Sandpiper. This limitation is basic to the training paradigm rather than a bug that could be fixed with more data alone.
Text rendering and typography: Unlike later multimodal models, CLIP has limited ability to read and reason about text within images. It sees rendered text as visual texture rather than semantic content, limiting its effectiveness on documents and charts as well as scenes containing important textual information. This limitation reflects the nature of its training data: captions describe the image content but do not typically transcribe text visible within the image itself. A photograph of a street sign might be captioned "a street sign at an intersection" without recreating the actual text on the sign, so the model learns to recognize street signs without learning to read them.
Data quality and noise: Despite training on 400 million pairs, the quality of the supervision signal in internet-scraped data varies enormously. Many captions are tangentially related to their images, use figurative language, or describe context rather than visual content. Alt-text on web pages is often generic ("image.jpg") or purely decorative ("decorative background"). These noise sources mean that CLIP's effective training dataset is substantially smaller than 400 million cleanly labeled examples. Subsequent work on data curation, including the creation of LAION-5B and its filtered variants, has shown that data quality improvements can substantially boost downstream performance even with the same training recipe.
Compositionality: CLIP handles concepts like "a dog" well but struggles with compositional descriptions like "a dog to the left of a red ball." The contrastive training objective treats the entire image-text pair as a matching unit without requiring the model to ground specific parts of the text in specific regions of the image. This means the model cannot reliably verify that a described spatial relationship holds, a compositional attribute is present, or a counting statement is accurate. Later work, including region-level contrastive learning and visual grounding objectives, has addressed some of these limitations, but compositional visual reasoning remains a challenging open problem.
Despite these limitations, CLIP enabled the current era of multimodal AI. Its architecture became the foundation for Stable Diffusion (where CLIP text embeddings guide image generation through classifier-free guidance), DALL-E 2 (which maps text descriptions to image embeddings using a prior network), and modern vision-language models like LLaVA and GPT-4V (which use CLIP-trained visual encoders as the vision backbone). By showing that natural language supervision could scale to learn transferable visual representations, CLIP eliminated the bottleneck of manual annotation and enabled models to learn from the raw abundance of internet image-text pairs. This shift from supervised learning to language-supervised learning has fundamentally changed how researchers approach computer vision, opening paths toward more general, flexible visual understanding systems that can communicate about what they see using natural language.
The influence of CLIP extends beyond its immediate applications. It established a template that subsequent multimodal models have followed: train large encoders with contrastive objectives on massive web data, then attach or fine-tune these encoders for downstream tasks. CLIP also influenced the evaluation methodology for multimodal systems, popularizing zero-shot evaluation benchmarks that measure generalization to new concepts rather than performance on fixed training distributions. The emphasis on zero-shot generalization as a primary evaluation metric has pushed the field toward building systems that understand concepts rather than memorizing pattern-label associations.
Summary
CLIP bridges vision and language through a simple yet powerful mechanism: dual encoders trained with contrastive learning to align image and text embeddings in a shared space. The key takeaways are:
- Dual encoder architecture: Separate transformer encoders for vision and text map inputs to a shared 512 or 768-dimensional embedding space, letting efficient cross-modal retrieval and comparison
- Contrastive objective: The symmetric cross-entropy loss trains the model to maximize similarity between matched image-text pairs while minimizing it for unmatched pairs within the batch, with implicit negative examples per batch
- Temperature scaling: A learned temperature parameter controls the sharpness of the softmax distribution, with CLIP converging to a sharp distribution that concentrates probability on the best-matching pair
- Zero-shot capability: By describing classes through natural language prompts, CLIP classifies images without task-specific training data, turning classification into a retrieval problem
- Scale matters: Training on 400 million image-text pairs with large batch sizes (32,768 to 65,536) produces representations that transfer across diverse visual tasks and new domains
- Prompt engineering: Performance depends heavily on prompt templates that match the distribution of training captions; ensembling over multiple templates improves reliability and reduces variance
- Practical limitations: Domain shift, social biases, fine-grained discrimination failures, and compositionality gaps require careful consideration before deployment in specialized or high-stakes applications
CLIP established the template for modern multimodal learning: train on large-scale noisy web data with contrastive objectives, then transfer to downstream tasks with minimal or no fine-tuning. As we explore in upcoming chapters on vision-language models like LLaVA, this foundation enables sophisticated multimodal reasoning that combines CLIP's visual understanding with the linguistic capabilities of large language models. The next frontier is understanding relationships and actions together with attributes and contexts that make visual scenes meaningful, and doing so through the flexible medium of natural language.
CLIP: Contrastive Language-Image Pre-training
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!