Skip to main content
Open In Colab

Overview

This chapter explores a two-step approach for fine-tuning generative LLMs:
  1. Supervised Fine-Tuning (SFT): Teach the model to follow instructions
  2. Direct Preference Optimization (DPO): Align outputs with human preferences
We’ll use QLoRA for memory-efficient training on consumer GPUs.
Use a GPU with at least 15GB VRAM. In Google Colab, select Runtime > Change runtime type > Hardware accelerator > GPU > GPU type > T4.

When to Fine-Tune Generation Models

Fine-tune generative models when:
  • You need specific output formats or styles
  • Domain-specific knowledge is required
  • You want to align with specific preferences or guidelines
  • Base models don’t follow instructions well enough
  • You have quality instruction-response pairs

Step 1: Supervised Fine-Tuning (SFT)

Data Preprocessing

Format data using the model’s chat template:
Formatted example:

Model Setup with QLoRA

QLoRA combines 4-bit quantization with LoRA for efficient fine-tuning on limited hardware.

LoRA Configuration

Training Configuration

SFT Training

Training progress (375 steps):
Training metrics:

Step 2: Preference Tuning (DPO)

DPO Dataset Preparation

Dataset statistics:

Load SFT Model

DPO Training Configuration

DPO training progress (200 steps):
DPO loss starts higher than SFT and decreases more gradually as the model learns preferences.

Merge Adapters

Inference

Sample output:

Two-Step Training Comparison

1

Supervised Fine-Tuning

Purpose: Learn instruction following and task completionData: Instruction-response pairsLoss: Cross-entropy on next token predictionResult: Model can follow instructions but may not align with preferences
2

Direct Preference Optimization

Purpose: Align with human preferences and quality standardsData: Prompt with chosen/rejected response pairsLoss: DPO loss encouraging chosen over rejectedResult: Model generates preferred outputs matching human judgment

QLoRA Benefits

Hyperparameters Guide

SFT Hyperparameters

DPO Hyperparameters

Best Practices

1

Data Quality

  • Use diverse, high-quality instruction data
  • Ensure chat template consistency
  • Filter low-quality responses
  • Balance different task types
2

QLoRA Configuration

3

Training Stability

  • Monitor loss curves for smooth descent
  • Use gradient checkpointing for memory
  • Enable fp16 mixed precision
  • Start with lower learning rates if unstable
4

Evaluation

  • Test on held-out examples
  • Compare base vs SFT vs DPO outputs
  • Evaluate instruction following quality
  • Check for catastrophic forgetting
Avoid these common mistakes:
  • Overfitting: Monitor validation loss, stop early
  • Wrong template: Use exact chat format from pre-training
  • High learning rate: Can destabilize the model
  • Insufficient warmup: Causes training instability

Memory Optimization

For Limited VRAM

For Faster Training

Training Time Estimates

SFT (3,000 examples on T4):
DPO (5,922 examples on T4):

Next Steps

  • Experiment with larger models (7B, 13B parameters)
  • Try different LoRA ranks (8, 16, 32, 64)
  • Implement PPO for more complex preference learning
  • Use RLHF for human-in-the-loop refinement
  • Evaluate with LM Eval Harness or MT-Bench