Multimodal AI
For years, AI models lived in silos: CNNs only looked at images, and Transformers only read text. Multimodal AI bridges these gaps, creating models that can natively understand text, images, video, and audio simultaneously—much like a human brain.
1. CLIP (Contrastive Language-Image Pretraining)
Developed by OpenAI, CLIP is arguably the most important multimodal model. It was trained on hundreds of millions of image-text pairs from the internet.
CLIP uses a Text Encoder and an Image Encoder. It is trained using a "Contrastive Loss"—it learns to make the embedding vector for an image (like a picture of a dog) perfectly match the embedding vector for its corresponding text ("a picture of a dog").
Why it matters: CLIP allows you to search through images using text prompts, and acts as the "vision" component for many other models (including Stable Diffusion and DALL-E).
2. Vision-Language Models (VLMs)
Models like GPT-4o and LLaVA are capable of taking an image as input alongside your text prompt. Under the hood, these models typically use a CLIP Vision Encoder to chop the image up into embedding tokens, and then simply feed those image tokens directly into a standard LLM alongside the text tokens!
Python Implementation: Zero-Shot Image Classification with CLIP
from PIL import Image
import requests
from transformers import CLIPProcessor, CLIPModel
# 1. Load the Model and Processor
model = CLIPModel.from_pretrained("openai/clip-vit-base-patch32")
processor = CLIPProcessor.from_pretrained("openai/clip-vit-base-patch32")
# 2. Get an image
url = "http://images.cocodataset.org/val2017/000000039769.jpg" # Two cats sleeping
image = Image.open(requests.get(url, stream=True).raw)
# 3. Define classes you want to search for
choices = ["a photo of a cat", "a photo of a dog", "a photo of a bird"]
# 4. Process and Predict
inputs = processor(text=choices, images=image, return_tensors="pt", padding=True)
outputs = model(**inputs)
logits_per_image = outputs.logits_per_image # image-text similarity score
probs = logits_per_image.softmax(dim=1) # convert to probabilities
for choice, prob in zip(choices, probs[0]):
print(f"{choice}: {prob.item()*100:.2f}%")