Skip to main content
Open In Colab

Overview

This chapter explores how to fine-tune BERT and other representation models for classification tasks. We’ll cover supervised fine-tuning, layer freezing strategies, and parameter-efficient approaches.
Use a GPU for fine-tuning. In Google Colab, select Runtime > Change runtime type > Hardware accelerator > GPU > GPU type > T4.

When to Fine-Tune BERT

Fine-tune BERT models when:
  • You have labeled classification data (100+ examples minimum)
  • You need task-specific predictions (sentiment, topic, intent)
  • Pre-trained models need domain adaptation
  • You want better performance than zero-shot approaches

Dataset Preparation

Loading the Data

We’ll use the Rotten Tomatoes dataset for sentiment classification:
Dataset statistics:

Model and Tokenizer Setup

Tokenization

Supervised Fine-Tuning

Define Metrics

Training Configuration

Training and Evaluation

Training results (full fine-tuning):
Evaluation results:

Layer Freezing Strategies

Understanding BERT Layers

BERT-base has 12 transformer layers. Freezing lower layers preserves general language understanding while fine-tuning upper layers for your task.

Freeze All Except Classification Head

Results (frozen layers):
Freezing layers provides:
  • Faster training: 15s vs 62s (4x speedup)
  • Lower F1: 0.638 vs 0.857
  • Less overfitting: Useful with limited data

Freeze Lower Layers (0-5)

A balanced approach that preserves pre-trained features while allowing task adaptation:
Results (partial freezing):

Fine-Tuning Strategy Comparison

BERT Architecture Layers

1

Embeddings

Word, position, and token type embeddings (indices 0-4)
2

Transformer Layers 0-5

Lower layers capture syntax and basic semantics (indices 5-100)
3

Transformer Layers 6-11

Upper layers capture task-specific patterns (indices 101-196)
4

Pooler and Classifier

Task-specific output layers (indices 197-200)

Hyperparameter Recommendations

Learning Rate

Batch Size

Epochs

BERT models overfit quickly. Recommendations:
  • Large datasets (>10k): 2-3 epochs
  • Medium datasets (1k-10k): 3-4 epochs
  • Small datasets (<1k): 4-6 epochs with validation monitoring

Training Loss Progression

Typical training loss over 1 epoch (534 steps):

Model Inspection

Examine trainable parameters:

Best Practices

1

Start with Pre-trained Models

Always use models pre-trained on large corpora (BERT, RoBERTa, DistilBERT)
2

Match Tokenizer and Model

Ensure tokenizer corresponds to the model architecture:
  • bert-base-uncased → lowercase, no accent marks
  • bert-base-cased → preserves case and accents
3

Monitor for Overfitting

4

Use Mixed Precision

Common Issues and Solutions

Out of Memory (OOM)

Poor Performance

  • Increase training epochs
  • Reduce frozen layers
  • Try different learning rates (1e-5, 3e-5, 5e-5)
  • Check for data quality issues

Slow Training

  • Enable fp16 mixed precision
  • Freeze more layers
  • Increase batch size if memory allows
  • Use gradient accumulation

Next Steps

  • Experiment with RoBERTa and DeBERTa for better performance
  • Try few-shot learning with SetFit
  • Implement LoRA for parameter-efficient fine-tuning
  • Use different classification heads (pooling strategies)