Skip to main content
Open In Colab Chapter 9 explores multimodal large language models that can process and understand both images and text, opening up new possibilities for AI applications.

Overview

Multimodal models combine vision and language understanding, enabling AI systems to perform tasks like image captioning, visual question answering, and image-text retrieval.

Key Topics Covered

  • CLIP (Contrastive Language-Image Pre-training)
  • Image and text embeddings
  • Similarity search across modalities
  • BLIP-2 for image captioning and VQA
  • Multimodal architectures

CLIP: Bridging Vision and Language

CLIP learns to align images and text in a shared embedding space, enabling zero-shot image classification and cross-modal retrieval.

Loading an Image

Setting Up CLIP

Creating Text Embeddings

Output:
Output:
Output:

Creating Image Embeddings

Output:
Output:
CLIP embeddings are 512-dimensional vectors that exist in the same space for both images and text.

Computing Similarity

Output:

Multi-Image Comparison

Compare multiple images with multiple captions to find the best matches.

Visualizing Similarities

The similarity matrix shows how well each image matches each caption:
Higher values indicate stronger semantic alignment between image and text.

Sentence-BERT for CLIP

Simplify CLIP usage with the Sentence Transformers library.

BLIP-2: Advanced Multimodal Understanding

BLIP-2 combines a vision encoder, Q-Former, and language model for sophisticated image understanding tasks.

Setting Up BLIP-2

Image Preprocessing

Output:

Text Preprocessing

Output:
Output:
BLIP-2 uses the GPT-2 tokenizer. The Ġ character represents spaces and is replaced with _ for clarity.

Use Case 1: Image Captioning

Generate descriptive captions for images automatically.
Output:

Rorschach Test Example

Output:
Image captioning can be used for accessibility (alt text), content moderation, search indexing, and automatic tagging.

Use Case 2: Visual Question Answering

Ask questions about images and get specific answers.
Output:

More VQA Examples

Output:
Output:

Architecture Overview

1

Vision Encoder

Processes images into visual features using a Vision Transformer (ViT).
2

Q-Former

Querying Transformer that bridges the vision and language modalities, extracting relevant visual information for the language model.
3

Language Model

Generates text based on visual features and optional text prompts. BLIP-2 uses OPT or FlanT5.

Key Takeaways

Shared Embedding Space

CLIP aligns images and text in a common space, enabling zero-shot classification and cross-modal retrieval.

Image Captioning

BLIP-2 can automatically generate descriptive captions for images without task-specific training.

Visual QA

Ask specific questions about image content and get accurate, grounded answers.

Multimodal Understanding

Modern architectures combine specialized vision and language components for sophisticated reasoning.

Applications

  • Accessibility: Generate alt text for images automatically
  • Content Moderation: Detect inappropriate visual content
  • E-commerce: Search products by image or description
  • Education: Interactive visual learning with VQA
  • Healthcare: Medical image analysis with natural language queries
  • Robotics: Enable robots to understand and describe their visual environment

Model Comparison

Further Resources