Contrastive Embeddings: How CLIP Connects Images and Text
You'll learn to
- -Understand a joint embedding space as the mechanism that lets a dot product compare an image to a piece of text
- -Explain the CLIP contrastive training objective and why it is symmetric
- -Implement the CLIP loss from scratch and see why a mismatched pairing produces a much higher loss than a correct one
The GenAI Lab's Multi-Modal Retriever level uses CLIP to search images with text, and the Math module's embeddings chapter used cosine similarity to compare two vectors of the SAME type - but neither ever asked how an image and a caption end up in the same coordinate space in the first place, comparable by a plain dot product at all. That is CLIP's actual contribution, and it comes from a specific training objective: contrastive learning.
One Shared Space for Two Very Different Kinds of Data
An image encoder and a text encoder are two separate networks, with no reason a priori to produce comparable outputs. CLIP trains them together, on a large dataset of (image, caption) pairs, with one explicit goal: push each image's embedding close to its OWN caption's embedding, and away from every OTHER caption in the same training batch - simultaneously, in both directions.
The Contrastive Training Objective
With batch_size image/caption pairs, the similarity matrix (every image embedding dotted with every text embedding) has exactly one correct match per row: image i pairs with text i, on the diagonal. Treat each row, scaled by a learned temperature, as classification logits over "which caption matches this image" - the correct label is always the diagonal index - and that is just cross-entropy, the same loss function from Act 0.
- -The loss needs no manual labeling beyond "which caption came with which image" - which is exactly the metadata a normal (image, caption) dataset already has, no extra annotation required.
- -Once trained, a plain dot product between the image encoder's output and the text encoder's output IS a meaningful similarity score - this is the mechanism that makes "search images by text description" and zero-shot image classification possible at all.
- -The same idea trains embedding models for text-only search too (the vector databases from the RAG module rely on embeddings trained with a close cousin of this objective) - CLIP just runs it across two different encoders instead of one.
AI Lab's Level 48 implements this in numpy across a full batch at once, using matrix multiplication for the whole similarity matrix instead of a nested loop over pairs.
Interview Signal is part of Pro
See a real weak answer next to a real strong one for this exact topic.
Quiz is part of Pro
Test what you just read with a short quiz, and bank the XP.
Implement it from scratch: Level 71 (Contrastive Embeddings / CLIP-Style Loss) in AI Lab.