Overview
This chapter explores how to fine-tune BERT and other representation models for classification tasks. We’ll cover supervised fine-tuning, layer freezing strategies, and parameter-efficient approaches.Use a GPU for fine-tuning. In Google Colab, select Runtime > Change runtime type > Hardware accelerator > GPU > GPU type > T4.
When to Fine-Tune BERT
Fine-tune BERT models when:- You have labeled classification data (100+ examples minimum)
- You need task-specific predictions (sentiment, topic, intent)
- Pre-trained models need domain adaptation
- You want better performance than zero-shot approaches
Dataset Preparation
Loading the Data
We’ll use the Rotten Tomatoes dataset for sentiment classification:Model and Tokenizer Setup
Tokenization
Supervised Fine-Tuning
Define Metrics
Training Configuration
Training and Evaluation
Layer Freezing Strategies
Understanding BERT Layers
Freeze All Except Classification Head
Freezing layers provides:
- Faster training: 15s vs 62s (4x speedup)
- Lower F1: 0.638 vs 0.857
- Less overfitting: Useful with limited data
Freeze Lower Layers (0-5)
A balanced approach that preserves pre-trained features while allowing task adaptation:Fine-Tuning Strategy Comparison
BERT Architecture Layers
1
Embeddings
Word, position, and token type embeddings (indices 0-4)
2
Transformer Layers 0-5
Lower layers capture syntax and basic semantics (indices 5-100)
3
Transformer Layers 6-11
Upper layers capture task-specific patterns (indices 101-196)
4
Pooler and Classifier
Task-specific output layers (indices 197-200)
Hyperparameter Recommendations
Learning Rate
Batch Size
Epochs
Training Loss Progression
Typical training loss over 1 epoch (534 steps):Model Inspection
Examine trainable parameters:Best Practices
1
Start with Pre-trained Models
Always use models pre-trained on large corpora (BERT, RoBERTa, DistilBERT)
2
Match Tokenizer and Model
Ensure tokenizer corresponds to the model architecture:
bert-base-uncased→ lowercase, no accent marksbert-base-cased→ preserves case and accents
3
Monitor for Overfitting
4
Use Mixed Precision
Common Issues and Solutions
Out of Memory (OOM)
Poor Performance
- Increase training epochs
- Reduce frozen layers
- Try different learning rates (1e-5, 3e-5, 5e-5)
- Check for data quality issues
Slow Training
- Enable fp16 mixed precision
- Freeze more layers
- Increase batch size if memory allows
- Use gradient accumulation
Next Steps
- Experiment with RoBERTa and DeBERTa for better performance
- Try few-shot learning with SetFit
- Implement LoRA for parameter-efficient fine-tuning
- Use different classification heads (pooling strategies)
