Overview
This chapter takes you inside the transformer architecture to understand how Large Language Models actually work. You’ll learn about the model’s internal layers, how it processes embeddings, the attention mechanism, and key optimizations like KV caching that make text generation efficient.This chapter requires a GPU for running examples. In Google Colab, go to Runtime > Change runtime type > Hardware accelerator > GPU > GPU type > T4.
Learning Objectives
By the end of this chapter, you will:- Understand the transformer architecture and its components
- Know how to inspect model layers and parameters
- Learn how the language model head produces token probabilities
- Understand key-value caching and why it matters
- Be able to analyze model outputs and internal states
Setting Up
Install the required dependencies:Loading the Model
Let’s load Phi-3 and examine its architecture:Model Architecture Overview
Let’s inspect the model’s structure:Understanding the Architecture
1
Embedding Layer
Converts token IDs (32,064 vocab) into 3,072-dimensional vectors
2
Transformer Layers (32 layers)
Each layer contains:
- Self-attention: Captures relationships between tokens
- MLP (Feed-forward): Processes each position independently
- Layer normalization: Stabilizes training
3
Language Model Head
Projects the final hidden state (3,072 dims) back to vocabulary size (32,064) to produce logits
Key Components Explained
The Attention Mechanism
The attention mechanism allows each token to “look at” other tokens in the sequence:Query, Key, Value (QKV) Projections
Query, Key, Value (QKV) Projections
- Query (Q): What I’m looking for
- Key (K): What I have to offer
- Value (V): The actual information I contain
Output Projection
Output Projection
Rotary Positional Embeddings
Rotary Positional Embeddings
The Feed-Forward Network (MLP)
Each transformer layer includes a position-wise feed-forward network:Expansion
First layer expands from 3,072 to 16,384 dimensions (5.3x larger!)
Non-linearity
SiLU (Swish) activation adds non-linear transformations
Compression
Second layer compresses back to 3,072 dimensions
Knowledge Storage
These large intermediate dimensions store factual knowledge
The Inputs and Outputs of the Model
Let’s see what the model actually processes:Examining Model Internals
Let’s process a simple prompt and inspect the internal representations:Understanding the Shapes
From Logits to Tokens
The final step is selecting the next token from the probability distribution:Visualization of the Process
Optimizing Generation with KV Caching
Text generation requires generating one token at a time. Without optimization, this would be extremely slow!The Problem
- Token 1: Process all input tokens
- Token 2: Process input + token 1
- Token 3: Process input + token 1 + token 2
- …
- Token 100: Process input + 99 generated tokens
The Solution: KV Caching
The key insight: attention Keys and Values for previous tokens never change. We can cache them!1
First Token
Compute K and V for all input tokens, cache them
2
Subsequent Tokens
Only compute K and V for the new token, reuse cached values
3
Massive Speedup
Eliminate redundant computation
Benchmarking the Difference
With caching enabled:KV caching provides a 3.3x speedup! This optimization is critical for making LLMs practical for real-time applications.
Memory Considerations
While KV caching speeds up generation, it requires memory:Model Parameters
Let’s count the parameters in Phi-3:Embedding Layer
Embedding Layer
Each Transformer Layer
Each Transformer Layer
- Attention QKV: 3,072 × 9,216 = 28.3M
- Attention output: 3,072 × 3,072 = 9.4M
- MLP gate/up: 3,072 × 16,384 = 50.3M
- MLP down: 8,192 × 3,072 = 25.2M
All 32 Layers
All 32 Layers
LM Head
LM Head
Grand Total
Grand Total
Understanding Attention Patterns
Different layers learn different patterns:Early Layers
Focus on syntax and local patterns (nearby words)
Middle Layers
Capture semantic relationships and facts
Late Layers
Handle high-level reasoning and task-specific patterns
Final Layer
Prepares information for token prediction
Advanced Topics
Temperature and Sampling
Whendo_sample=True, the model doesn’t just pick the highest probability token:
Batch Processing
The model can process multiple sequences simultaneously:Practical Applications
Understanding the internals enables advanced techniques:Prompt Engineering
Knowing attention mechanisms helps craft better prompts
Fine-tuning
Target specific layers for parameter-efficient training
Inference Optimization
Optimize KV cache and batch sizes
Model Compression
Prune layers or reduce dimensions strategically
Performance Tips
1
Use Flash Attention
Install flash-attention for 2-3x speedup in attention computation
2
Enable torch.compile()
PyTorch 2.0+ can compile the model for faster execution
3
Optimize Batch Size
Larger batches improve GPU utilization but require more memory
4
Use Mixed Precision
FP16 or BF16 reduces memory and increases speed with minimal quality loss
Next Steps
Now that you understand how LLMs work internally, you’re ready to apply them to real-world tasks!Chapter 4: Text Classification
Learn how to use LLMs for classification tasks
Chapter 5: Text Clustering
Explore unsupervised learning with embeddings
