Skip to content
GenAI Learn/Diffusion Models & Multimodal Generation
Browsing as a guest. Sign in to save your progress and earn XP as you complete chapters.

Contrastive Embeddings: How CLIP Connects Images and Text

8 min read

You'll learn to

  • -Understand a joint embedding space as the mechanism that lets a dot product compare an image to a piece of text
  • -Explain the CLIP contrastive training objective and why it is symmetric
  • -Implement the CLIP loss from scratch and see why a mismatched pairing produces a much higher loss than a correct one

The GenAI Lab's Multi-Modal Retriever level uses CLIP to search images with text, and the Math module's embeddings chapter used cosine similarity to compare two vectors of the SAME type - but neither ever asked how an image and a caption end up in the same coordinate space in the first place, comparable by a plain dot product at all. That is CLIP's actual contribution, and it comes from a specific training objective: contrastive learning.

One Shared Space for Two Very Different Kinds of Data

An image encoder and a text encoder are two separate networks, with no reason a priori to produce comparable outputs. CLIP trains them together, on a large dataset of (image, caption) pairs, with one explicit goal: push each image's embedding close to its OWN caption's embedding, and away from every OTHER caption in the same training batch - simultaneously, in both directions.

The Contrastive Training Objective

With batch_size image/caption pairs, the similarity matrix (every image embedding dotted with every text embedding) has exactly one correct match per row: image i pairs with text i, on the diagonal. Treat each row, scaled by a learned temperature, as classification logits over "which caption matches this image" - the correct label is always the diagonal index - and that is just cross-entropy, the same loss function from Act 0.

CLIP Contrastive Loss
L = ½·[CE(scale·I·Tᵀ, diag) + CE(scale·T·Iᵀ, diag)]
I and T are the batch's (already-normalized) image and text embedding matrices. The loss is computed BOTH directions - image-to-text and text-to-image - and averaged, so the space does not end up lopsided in favor of just one direction.
The CLIP contrastive loss, from scratch
  • -The loss needs no manual labeling beyond "which caption came with which image" - which is exactly the metadata a normal (image, caption) dataset already has, no extra annotation required.
  • -Once trained, a plain dot product between the image encoder's output and the text encoder's output IS a meaningful similarity score - this is the mechanism that makes "search images by text description" and zero-shot image classification possible at all.
  • -The same idea trains embedding models for text-only search too (the vector databases from the RAG module rely on embeddings trained with a close cousin of this objective) - CLIP just runs it across two different encoders instead of one.

AI Lab's Level 48 implements this in numpy across a full batch at once, using matrix multiplication for the whole similarity matrix instead of a nested loop over pairs.

Interview Signal is part of Pro

See a real weak answer next to a real strong one for this exact topic.

Quiz is part of Pro

Test what you just read with a short quiz, and bank the XP.

Ready to Build This?

Implement it from scratch: Level 71 (Contrastive Embeddings / CLIP-Style Loss) in AI Lab.